<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>L.W. (2006). The situated nature of concepts. American Journal of
Psychology</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Knowledge Extraction on Multidimensional Concepts: Corpus Pattern Analysis (CPA) and 1 Concordances</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pilar León Araúz</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arianne Reimerink</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pamela Faber</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Translation and Interpreting, University of Granada</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2006</year>
      </pub-date>
      <fpage>327</fpage>
      <lpage>332</lpage>
      <abstract>
        <p>Multidimensionality of concepts in multidisciplinary domains is a problem terminographers have to deal with. We apply Corpus Pattern Analysis (CPA; Pustejovsky, Hanks, &amp; Rumshisky, 2004) to extract conceptual dimensions according to context. The dynamic nature of these concepts is exemplified with the case study of SAND. On the other hand, knowledge patterns (KPs) often convey different conceptual relations and are therefore polysemic structures. The development of pattern-based constraints can help to disambiguate them and at the same time avoid conceptual noise, which would be a first step towards the systematization of automatic knowledge extraction. Two KPs are analyzed in detail: rang* from, which conveys the conceptual relation is_a, and the polysemic KP formed by. Mots clés: Extraction de connaissances, Multidimensionalité, Corpus Pattern Analysis, Patron de connaissance.</p>
      </abstract>
      <kwd-group>
        <kwd>Knowledge extraction</kwd>
        <kwd>Multidimensionality</kwd>
        <kwd>Corpus Pattern Analysis</kwd>
        <kwd>Knowledge pattern</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Résumé: La multidimensionalité conceptuelle dans les domaines
multidisciplinaires est un problème auquel les terminographes doivent faire
face. On a appliqué le Corpus Pattern Analysis (CPA ; Pustejovsky et al.,
2004) afin d’extraire les dimensions conceptuelles selon le contexte. La nature
dynamique qui caractérise les concepts est illustrée avec l’exemple de SAND.
Par ailleurs, les patrons de connaissance (KPs) expriment très souvent de
différentes relations conceptuelles, étant ainsi des structures polysémiques.
Dans ce sens, le développement de contraintes basées sur les KPs peut être une
bonne approche pour contourner le problème de la polysémie et en même temps
éviter le bruit conceptuel dans les concordances. Cela représenterait un premier
pas vers la systématisation de l’extraction automatique de connaissances. Dans
cet article, on présente l’approche suivi dans l’analyse de deux KPs: rang*
from, exprimant la relation is_a, et formed by, un des KPs les plus
polysémiques.
1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>According to Kageura (1997: 120) the characteristics of a concept are frequently
specified from different points of view or facets (function, material, shape...) and the
set of characteristics that constitutes a concept is normally multidimensional.
Moreover, when concepts are represented in different contexts (i.e. specialized
subdomains) certain facets or dimensions become more or less salient. Corpus
linguistics provides the clues to distinguish which dimensions are activated in each
case and a sound methodology must be applied to study this contextual
multidimensionality in a consistent way.</p>
      <p>
        Corpus Pattern Analysis
        <xref ref-type="bibr" rid="ref2">(CPA; Pustejovsky, Hanks, &amp; Rumshisky, 2004; Hanks &amp;
Pustejovsky, 2005)</xref>
        investigates syntagmatic criteria for distinguishing different
meanings of a polysemous predicate. The procedure consists of three components: (1)
the manual discovery of selection context patterns for specific verbs; (2) the automatic
recognition of instances of the identified patterns; (3) the automatic acquisition of
patterns for unanalyzed cases (Rumshisky et al., 2006: 329).
      </p>
      <p>
        We apply the CPA approach in a slightly different way. We analyse concordances
starting with a direct search of specialized terms. After that, they are classified
according to the dimensions they show, where different knowledge patterns (KPs;
        <xref ref-type="bibr" rid="ref4">Barrière, 2004</xref>
        ) can be associated to different conceptual relations. Then we analyse
multidimensionality according to context. In this sense, Rumshisky et al. (2006)
indicate several ways in which different context dimensions expressed in the selection
context patterns can affect the semantic interpretation of a predicate. According to
them, the most frequent source of meaning differentiation of verbs lies in contrasting
the argument types filling each argument slot (idem, 330). However, here CPA is
applied to identify how context can affect the relational behaviour of a concept. In our
experience, the most frequent source of context differentiation of a concept’s
behaviour lies in contrasting the specific values filling each dimension.
      </p>
      <p>On the other hand, a pattern-based search is conducted in order to find new
relations among other concepts. The problem is that knowledge patterns are very often
polysemic structures that convey different conceptual relations. As a result, the
development of pattern-based constraints can help to disambiguate them and at the
same time avoid conceptual noise.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Conceptual dimensions</title>
      <p>Contextual multidimensionality is derived from the situated nature of concepts. It
occurs especially in the case of concepts with a low degree of specificity. We call
them versatile concepts because they are involved in a myriad of events and they are
not always related to the same concepts or through the same relations, especially in
interdisciplinary domains.</p>
      <p>For example, the concept SAND is generally (or prototypically) defined as a kind of
sediment located in the sea, rivers or soil layers. However, in real texts, SAND
activates many other dimensions. In a more general domain, such as GEOLOGY, the
concept is linked to others through: type, as a kind of SEDIMENT; attribute, related to
grain size as a classification parameter; and material, linking the concept to the
natural elements of which it is part (VALLEY, SOIL, AQUIFER, DESERT, etc):</p>
      <p>However, in a COASTAL PROCESS domain (see Fig. 2), salient dimensions become:
material, although values (natural elements) are restricted to coastal ones (SAND
BARRIER, SAND BERM, SAND SPIT, BEACH, etc.); and patient, where the concept is
involved in certain natural processes (WAVE ACTION, STORMS, LONGSHORE CURRENT,
DEPOSITION). If context is again restricted to the COASTAL DEFENCE domain (see Fig.
3), dimensions are still the same, but values are focalized to artificial elements
(FENCE, BERM, DIKE) or processes (TRAPPING, PUMPING, DUMPING). Furthermore,
there is a new dimension, highlighting the functional nature of the concept in this
specific context (SAND protects DUNE-BLUFFS, SAND BODY is used for BEACH
NOURISHMENT, etc.). These three domains form a hierarchy (GEOLOGY COASTAL
PROCESS COASTAL DEFENCE), but in a completely different domain, changes are
more remarkable:</p>
      <p>In the WATER TREATMENT domain, a new dimension is found, where SAND is
linked to a particular instrument used in water treatment plants. The functional
dimension now has a different value (FILTRATION) and patient and material are no
longer representative conceptual dimensions.</p>
      <p>Yeh and Barsalou (2006) claim that when situations are incorporated into a
cognitive task, processing becomes more tractable than when situations are ignored,
and the same can be applied to knowledge acquisition processes. As a result, a more
believable representational system should account for re-conceptualization according
to the situated nature of concepts.</p>
      <p>We are developing context-based conceptual networks by dividing the
interdisciplinary field of the environment2 in different contextual domains, but first we
need to know which concepts are activated in each situation and how to extract this
information. This is where CPA can help to accomplish our aim. In our approach,
patterns expressing contextual information delimit conceptual dimensions, which at
the same time delimit domain membership.</p>
      <p>First of all, if a concept only activates a particular dimension in just one domain,
the identification of any KP linking the concept to that dimension will be enough to
associate the concept to a concrete domain. In the case of SAND, if a KP expresses the
instrument dimension, the concept will be automatically assigned to the WATER
TREATMENT domain.</p>
      <p>Nevertheless, domain disambiguation is not that easy when different domains can
activate the same dimensions and, as a result, KPs are usually the same. For example,
if SAND is found next to KPs like consist of, comprising, formed from, containing, or
composed of, the material dimension ascribes concepts to three possible domains:
GEOLOGY, COASTAL PROCESSES and COASTAL DEFENCE. Sometimes, certain KPs
express a particular dimension in just one domain, such as made of, where SAND is
always related to COASTAL DEFENCE because the pattern needs the activation of an
artificial concept. However, most of the time, this is not the case, and domain
disambiguation requires a second step based on the kind of values associated to each
dimension. At this stage, semantic annotation seems the only way to differentiate</p>
      <sec id="sec-3-1">
        <title>2 http://manila.ugr.es/visual/</title>
        <p>domain membership. As mentioned above, in the GEOLOGY domain materials must be
natural elements found in nature, in the COASTAL PROCESS domain, materials are
restricted to those found in the coastal area, and in the COASTAL DEFENCE domain
materials are no longer natural elements. Consequently, annotation should be
conceptoriented, differentiating all concepts in the hierarchy and assigning each of them to
particular contexts.
3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Knowledge patterns</title>
      <p>Many KPs can be found in the manual identification process. However, automatic
extraction needs a certain level of reliability to be effective. In table 1 we show some
of the most reliable patterns for seven conceptual relations in our specialized domain:
Relation
Is_a
Part_of
Made_of
Located_at
Result_of
Has_function
Effected_by</p>
      <p>Most patterns are general language expressions and can be applied to many other
domains. Domain-specific patterns are generally more reliable, such as built of/from
and constructed of. However, even reliable ones show a certain degree of conceptual
noise and polysemy. Thus it is necessary to discover certain rule constraints according
to each KP’s specific needs. In the next sections we will deal with the KPs rang* from
and formed by.
3.1</p>
      <sec id="sec-4-1">
        <title>KP rang* from</title>
        <p>One of the most reliable KPs that activates the relation is_a in the environmental
domain is rang* from. However, when rang* from is followed by a number (see Fig.
7), the relation that is expressed is always of magnitude. The same goes for the
combination with a number written in full, or an adverb or an article and a number.
For example, measurable concepts such as GROWTH TIMES and RECHARGE EVENTS are
defined by certain time and amount combinations. Therefore, if our aim is to analyze
the is_a relation, there should also be restrictions on words expressing duration such
as minute, hour, month, day, week, etc. (see Fig. 5).</p>
        <p>Furthermore, if the terms are too far away from the KP, it is very hard to extract
useful information from the concordance (see Fig. 6). This noise could be partly
solved with the implementation of a candidate term extractor. This way, terms which
did not fall under the consideration of specialized terms would be excluded from the
extraction process from the beginning. Anaphora would still be a problem, but the
validation of conceptual propositions would be at least more efficient.</p>
        <p>The knowledge pattern rang* from has proven to be a very reliable pattern if the
above-mentioned restrictions are taken into account. It is not only very informative for
the extraction of hyponymic relations, but also for relations of coordination, as the
expression requires at least two coordinate concepts:</p>
        <p>From figure 7, the following information can be extracted: FINE SAND and PEBBLE
is_a BEACH SEDIMENTS; LONG BEACH and STRAIGHT BEACH is_a BEACH; SAND, MUD,</p>
        <sec id="sec-4-1-1">
          <title>CLAY and BOULDER is_a SEDIMENT.</title>
          <p>3.2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>KP formed by</title>
        <p>As in the case of rang* from, certain constraints must be applied in order to avoid
noise. However, the main problem of the KP formed by is polysemy. Concordances in
figure 8 show the way formed by works in the three different dimensions it can
express, although the result dimension prevails over part and material (see Fig. 8).</p>
        <p>The disambiguation of this polysemic KP requires different steps. If the KP is
followed by a verb, it is definitely related to the result dimension. Instead, if the KP is
followed by a noun, it can be linked to any of the three dimensions. Then the
difference lies in two factors: if the noun is a process concept type, the concept still
falls into the result dimension; if the noun is an object concept type, the dimension can
be either part or material, but if the noun is uncountable, it will always refer to the
material dimension, whereas countable nouns will always link wholes with parts.</p>
        <p>In order to construct a knowledge resource based on the real behaviour of concepts
in all the possible contexts of a domain, all the dimensions of multidimensional
concepts must be taken into account. On the other hand, we have shown how
constraints can be defined for KPs to facilitate knowledge extraction. Needless to say,
there is still much work to be done before achieving an automatic knowledge
extraction system. For example, even if all KPs were tightly constrained, we are still
analysing terms, and synonymic values may seem different concepts. That could be
solved if knowledge representation and extraction processes were complementary. In
this way, an ontological system could inform the extraction system and semantic
annotation could be based on an already defined hierarchy. In any case, manual
validation would still be necessary.
context. Proceedings of the 20th international conference on Computational Linguistics.
RUMSHISKY, A., HANKS, P., HAVASI, C. &amp; PUSTEJOVSKY, J. (2006). Constructing a
Corpusbased Ontology using Model Bias. Proceedings of FLAIRS 2006, Melbourne Beach, Florida,</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Geneva</surname>
          </string-name>
          , p.
          <fpage>924</fpage>
          -
          <lpage>931</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>HANKS</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>PUSTEJOVSKY</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>A Pattern dictionary for natural language Processing</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>Revue Française de linguistique appliquée 10: 2</source>
          , p.
          <fpage>63</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>BARRIÈRE</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>“BUILDING A CONCEPT HIERARCHY FROM CORPUS ANALYSIS”</article-title>
          .
          <string-name>
            <surname>Terminology</surname>
            <given-names>PUSTEJOVSKY</given-names>
          </string-name>
          , J.,
          <string-name>
            <given-names>P. HANKS</given-names>
            , &amp;
            <surname>RUMSHISKY</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2004</year>
          ).
          <source>Automated induction of sense in</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>