<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extracting Concrete Entities through Spatial Relations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Olga Acosta</string-name>
          <email>oacostal@uc.cl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>César Aguilar</string-name>
          <email>caguilara@uc.cl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Facultad de Letras, Pontificia Universidad Católica de Chile</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper focuses on the automated extraction of concrete entities from a specialized-domain corpus. Then, in a bootstrapping phase, the candidates are used to extract new candidates. Concrete entities are automatically identified by a set of spatial features. In a spatial scene something is located by virtue of the spatial properties associated with a reference object. The axial properties are represented by place adverbs. Additionally, for identifying referent objects in a sentence we consider syntactical patterns extracted by chunking. In order to reduce noise in results, we take into account a corpus comparison approach and linguist heuristics. Results show high precision in candidates with high weights.</p>
      </abstract>
      <kwd-group>
        <kwd>Concrete entities</kwd>
        <kwd>lexical relation</kwd>
        <kwd>information extraction</kwd>
        <kwd>term extraction</kwd>
        <kwd>axial properties</kwd>
        <kwd>nominalization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In recent years, the automatic mining of relevant knowledge in the biomedical domain
has become in an interesting research area, particularly in tasks related to the
generation of taxonomies and ontologies
        <xref ref-type="bibr" rid="ref18 ref24">(Smith and Kumar, 2004)</xref>
        . This kind of tasks
require the design and implementation of efficient information extraction (IE) methods,
capable of identifying and extracting textual patterns that contain such relevant
knowledge.
      </p>
      <p>Therefore, in this work we propose a methodology for the automatic extraction of
concrete entities implicit in medical documents. Then, in a bootstrapping phase, these
candidates are used for extracting a larger set of new candidates.</p>
      <p>Linguistically speaking, a main concern is those noun phrases (NP) whose
modifiers are relational adjectives and where the noun head is a concrete entity, because
relational adjectives introduce semantic features which describe specific properties
such as formal, constitutive, telic and agentive qualities (Fábregas, 2007). The
identification of this type of NP contributes to delimit the number of possible semantic
relations. For testing our method, we work with a corpus of medical texts in Spanish.</p>
      <p>We organize our paper as follows: in section 2 we define what a concrete entity is,
taking into account the description proposed by Fellbaum (1998) for classifying
names in WordNet. Then, in section 3, we show a brief explanation about the
representation of space in natural language, according to a cognitive framework. In section
4, we describe the most common deverbal nominalizations in specialized texts. In
section 5 we explain the relation noun + relational adjective in order to delineate a set
of linguistic heuristics useful for filtering non-relevant adjectives. In section 6 we
describe our methodology. In section 7 we offer a description of preliminary results.
Finally, in section 8, we give our conclusions.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>Concrete entities</title>
      <p>
        We understand all that exists in the world as a concrete entity which something can be
predicated (in Aristotle’s categories: substance). For example, concrete entities can be
artifactual categories like vehicles, clothing and weapons, or natural kinds like birds,
fruits and vegetables
        <xref ref-type="bibr" rid="ref17 ref20 ref21">(Landau and Jackendoff, 1993; Murphy, 2002)</xref>
        . This is in line
with 8 of the 25 main categories considered in the WordNet hierarchy for nouns
denoting tangible things: {animal, fauna}, {artifact}, {body}, {food}, {natural object},
{person, human being}, {plant, flora}, {substance}. From our point of view these
categories can be collapsed in artifactual and natural kinds.
3.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Space in language and cognition</title>
      <p>Levinson (2004) points out that the spatial thinking is a crucial feature in our lives:
we constantly consult our spatial memories in events such as finding our way across
town, giving route directions, searching for lost keys, and so on. This importance is
mirrored in real discourse where knowledge about formal, agentive, constitutive and
telic features, as well as spatial features, are found in specialized domains.</p>
      <p>There are three frames of reference lexicalized in language: intrinsic, relative
and absolute frame. Intrinsic frame involves an object-centred coordinate system,
where the coordinates are determined by the “inherent features”, sidedness or facets
of the objet to be used as the ground (i.e., he’s in front of the house). Relative frame of
reference presupposes a viewpoint where a perceiver is located, a figure and ground
distinct from the viewpoint. Thus, it offers a triangulation of three points, and utilizes
coordinates fixed on viewpoint to assign directions to figure and ground (i.e., the ball
is to the left of the tree). Finally, absolute frame refers to the fixed direction provided
by gravity (i.e., he’s north of the house).
3.1.</p>
      <sec id="sec-3-1">
        <title>Work related</title>
        <p>Mani et al. (2010) focused on the problem of extracting information about places,
considering both absolute and relative references. Their goal was on grounding such
references to precise positions that can be characterized in terms of geo-coordinates.
These authors use a supervised approach to mark up PLACE tags in documents.
SpatialML is an annotation scheme derived from this work and which has been applied to
annotated corpora in English and Mandarin Chinese. An automatic tagger for
SpatialML extents scores 86.9 F-measure, which is a reasonable performance. On the
other hand, Clementini et al. (1997) propose a unified framework for the qualitative
representation of positional information in a two-dimensional space in order to
perform spatial reasoning. The orientation and distance relations for objects modeled as
points can determine positional information. The implicit characteristics of an object
are its topology and its extension, while, with respect to other objects, topological,
orientation, and distance relations have to be considered.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Axial properties</title>
        <p>Evans (2007) explains that a spatial scene is a linguistic unit containing information
based on our spatial experience. This space is structured according to four parameters:
a figure (or trajector), a referent object (that is, a landmark), a region and —in certain
cases— a secondary reference object. These two reference objects configure a
reference frame. We can understand this configuration by considering the following
example: un auto está estacionado detrás de la escuela (Eng.: “a car is parked behind the
school”). In this sentence, un auto is the figure and la escuela is the referent object.
The region is established by the combination of the adverb detrás1 which sketches a
spatial relation with the referent object. This relation encodes the location of the
figure.</p>
        <p>Moreover, Evans (2007) points out the existence of axial properties, that is, a set of
spatial features associated to a specific referent object. Considering again the sentence
a car is parked near to the school, we can identify the location of the car searching
for it in the region near to the school. Therefore, this search can be performed because
the referent object (the school) has a set of axial divisions: front, back and side areas.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Axial properties and place adverbs</title>
        <p>
          Axial properties are linguistically represented by place adverbs. In this experiment we
only consider adverbs functioning in Spanish with preposition de
          <xref ref-type="bibr" rid="ref2">(Acosta and
Aguilar, 2015)</xref>
          :
        </p>
        <p>Enfrente, delante (Engl. In front to/of); Detrás, atrás (Engl. Behind);
sobre, encima (Engl. On); abajo, debajo (Engl. under); dentro, adentro
(Engl. In/inside); fuera, afuera (Engl. Out/outside); arriba (Engl. Above/
over).</p>
        <p>Additionally, we use some synonymous nouns such as exterior (outside) and
interior (in), as well as side nouns synonymous with the dimensions left and right.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Nominalization</title>
      <p>According to Martin (1993: 203-220) and Vivanco (2006), from a linguistic
perspective, the discourse neutrality in science and technology is presented by means of
im1 In English, behind is a preposition. In contrast, in Spanish is an adverb.
personation: missing second person, low presence of first person, abundance of
impersonal verbs and passive voice, as well as nominalizations hiding actions made by
the subject. These nominalizations are used by scientists to support their arguments,
coining new terms by means of nouns and summarizing information previously
provided in a text.</p>
      <p>In line with the frequent use of nominalization in specialized texts, in the case of
Spanish, Cademártori, Parodi and Venegas (2006) show data concerning the use of
deverbal nominalizations in three domains: commercial, maritime and industrial. The
most used suffixes for constructing nouns are: -ción, -miento, -sión, and -dor.
5.</p>
    </sec>
    <sec id="sec-5">
      <title>Adjectives-Noun modifiers</title>
      <p>
        An adjective is a grammatical category whose function is to modify nouns
        <xref ref-type="bibr" rid="ref11">(Demonte,
1999)</xref>
        . There are two kinds of adjectives: descriptive and relational adjectives. The
descriptive adjectives refer to constitutive features of the modified noun characterized
by means of a single physical property: color, form, character, predisposition, sound,
and so on, e.g., el libro azul (Eng.: “the blue book”). On the other hand, relational
adjectives assign a set of properties, i.e., all the characteristics jointly defining names
as sea: puerto marítimo (Eng.: “maritime port”). In terminology, relational adjectives
represent an important element for building specialized terms. For example, inguinal
hernia, venereal disease and others are considered terms in medicine as opposed to
NPs with more contextual interpretations like rare hernia, serious disease, and
critical disorder.
5.1.
      </p>
      <sec id="sec-5-1">
        <title>Identifying syntactically non-relevant adjectives</title>
        <p>
          If we consider the internal structure of adjectives, we can identify two types:
permanent and episodic adjectives
          <xref ref-type="bibr" rid="ref11">(Demonte, 1999)</xref>
          . The first kind of adjectives represents
stable situations, permanent properties characterizing individuals. These adjectives
a r e l o c a t e d o u t s i d e o f a n y s p a t i a l o r t e m p o r a l r e s t r i c t i o n ( i . e . ,
psicópata/“psychopath”). On the other hand, episodic adjectives refer to transient
situations or properties implying change and with time-space limitations.
        </p>
        <p>
          Almost all descriptive adjectives derived of participles belong to this latter class as
well all adjectival participles (i.e., harto/“jaded”). Spanish is one of the few languages
that in its syntax represent this difference in the meaning of adjectives. In many
languages this difference is only recognizable through interpretation. In Spanish,
individual properties can be predicated with the verb ser, and episodic properties with the
verb estar, which is an essential test to recognize what class an adjective belongs to.
In this sense, with the goal of identifying and extracting non-relevant adjectives, we
propose extracting adjectives predicated with the verb estar
          <xref ref-type="bibr" rid="ref1">(Acosta, Aguilar and
Sierra, 2013)</xref>
          .
        </p>
        <p>Another linguistic heuristic for identifying descriptive adjectives is that only these
kinds of adjectives accept degree adverbs or are part of comparative constructions,
e.g., muy alto/“very high”, Juan es más alto que Pedro/“John is taller than Peter”.
Finally, only descriptive adjectives can precede a noun because —in Spanish—
relational adjectives are always postposed (e.g., la antigua casa/“the old house”).
5.2.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Types of relational adjectives</title>
        <p>According to Bosque (1993) relational adjectives such as salivary in the noun phrase
salivary gland belong to a kind of relational adjectives which do not occupy positions
in the argument structure of the predicate, but they denote entities which establish a
specific relation with the head noun. Bosque refers to these relational adjectives as
classification relational adjectives, while the term thematic relational adjectives is
left for the other group, e.g., the case of renal infection, where infection is derived
from a verb.
6.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Methodology</title>
      <p>In this paper we propose a methodology for extracting concrete entities from a
specialized domain corpus with part-of-speech tags.
6.1.</p>
      <sec id="sec-6-1">
        <title>Part-of-Speech Tagging</title>
        <p>
          Part-of-Speech (POS) tagging is the process of assigning a grammatical category to
each word in a corpus. The most common taggers used for Spanish are TreeTagger
          <xref ref-type="bibr" rid="ref23">(Schmid, 1994)</xref>
          and FreeLing2
          <xref ref-type="bibr" rid="ref7">(Carreras et al., 2004)</xref>
          . In this experiment, we use
FreeLing because it is more precise than TreeTagger for tagging texts in Spanish. The
following example shows a sentence in Spanish tagged with the FreeLing
tagger:
el/DA tipo/NC más/RG común/AQ de/SP lesión/NC ocurrir/VM cuando/CS
algo/PI irritar/VM el/DA superficie/NC externo/AQ del/PDEL ojo/NC
6.2.
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>Chunking</title>
        <p>Chunking is the process of identifying and classifying segments of a sentence by
grouping the major parts-of-speech that form basic non-recursive phrases.</p>
        <p>
          In this work, we concern the automated extraction of concrete entities. Concrete
entities relevant to a domain are terms and the most productive patterns of terms
consist of a noun and zero or more adjectives
          <xref ref-type="bibr" rid="ref26">(Vivaldi, 2001)</xref>
          . Using FreeLing tags, these
patterns can be represented as a regular expression in a single pattern:
&lt;NC&gt;&lt;AQ&gt;*
The above regular expression is considered in the first phase of extraction of
candidates.
2 FreeLing based on the tags of the EAGLES group.
        </p>
        <p>Concrete entities can be located in spatial scenes as figures or reference objects. In
this experiment, only reference objects are extracted with their axial properties that
can be linguistically represented as:</p>
        <p>&lt;RG|NC&gt;&lt;PDEL&gt;&lt;DA&gt;?&lt;NC&gt;&lt;AQ&gt;*
The regular expressions used to extract non-relevant adjectives according to the
linguistic heuristics mentioned in section 5.1 are:
&lt;RG&gt;&lt;AQ&gt;
&lt;VAE&gt;&lt;AQ&gt;
&lt; D.*|P.*|F.* |S.*&gt;&lt;AQ&gt;&lt;NOUN&gt;
Where RG, AQ and VAE as tagged with FreeLing, correspond to adverbs, adjectives
and the verb estar, respectively. Tags &lt;D.*|P.*|F.*|S.*&gt; correspond to determinants,
pronouns, punctuation signs and prepositions. The expression &lt;D.*|P.*|F.*|S.*&gt; is a
restriction to reduce noise, since elements wrongly tagged by FreeLing as adjectives
are extracted without this restriction.
6.3.</p>
      </sec>
      <sec id="sec-6-3">
        <title>Bootstrapping phase</title>
        <p>We use the candidates to concrete entities obtained in the first step as seeds for
extracting more candidates. On the one hand, we assume that coordinating phrases
where a good candidate occurs have a high probability of containing other good
candidates for a concrete entity:</p>
        <p>&lt;NC&gt;&lt;AQ&gt;*&lt;CC&gt;&lt;NC&gt;&lt;AQ&gt;*
Where &lt;CC&gt; tag corresponds to the disjunction (i.e.: kidney or liver) and conjunction
(i.e.: kidney and liver).</p>
        <p>On the other hand, noun phrases with at least an adjective take advantage of the
noun head of candidates for a concrete entity for finding more specific candidates
(i.e., artery-femoral artery):</p>
        <p>&lt;NC&gt;&lt;AQ&gt;+
6.4.</p>
      </sec>
      <sec id="sec-6-4">
        <title>Reducing noise</title>
        <p>We sought to remove non-relevant words from noun phrases before ranking
candidates for concrete entities. After the chunking phase, noise was reduced by removing
non-relevant open-class words. One of our goals consists of building this stopword
list as automatically as possible.</p>
        <p>Since concrete entities are terms in the domain, a list of non-relevant words from
the domain (i.e., stopword list) can be used to refine the terminology obtained from an
automatic process. We considered a list constructed with high frequency words in a
reference corpus to have drawbacks because, apart from the selection by occurrence
frequency (in the domain corpus, words with high frequency can be terms), human
supervision is required in order to determine whether a word is relevant to the
domain.</p>
        <p>
          Given the above, we consider that linguistic heuristics operating in a specific
language can be taken into account in order to automate the selection of non-relevant
words. One of the disadvantages, however, is that this leads to language dependence.
For the case of adjectives, in Spanish, characteristic features have been proposed in
order to distinguish between descriptive and relational adjectives as mentioned in
section 5. On the other hand, with a corpus comparison approach, we obtain both
nouns and adjectives where the relative frequency in a reference corpus is greater or
equal than in the domain corpus. These words can be used as part of the stopword list.
Additionally, we take into account empirical evidence concerning the use of deverbal
nominalizations in specialized discourse
          <xref ref-type="bibr" rid="ref8">(Cadermártori, Parodi and Venegas, 2006)</xref>
          for removing phrases where noun heads are indicative of actions, events and states but
not concrete entities (in a NP with a noun head of this type, a thematic relational
adjective is found). In this sense, suffixes as –ción, -miento, and –sión were used for
filtering out noun phrases. Finally, a short list with the more frequent non-relevant
nouns operating as noun heads in phrases: form, type, kind, cause, effect and so on,
were considered for removing noun phrases.
        </p>
        <p>Adjectives from the reference corpus can be used as a fixed-size list where
nonrelevant adjectives automatically extracted from the domain can be added. These can
be obtained taking into account the three heuristics mentioned in section 5.1. Then,
these adjectives can be manually reviewed in order to determine their relevance to
any specialized knowledge domain (i.e., adjectives as relevant, important, necessary,
appropriate, and so on can be considered for the stopword list). This is a fixed-size list
and can be the base-list where non-relevant adjectives automatically extracted from
the domain can be added.
6.5.</p>
      </sec>
      <sec id="sec-6-5">
        <title>Ranking words</title>
        <p>
          We evaluate termhood of simple words by means of rank difference
          <xref ref-type="bibr" rid="ref9">(Kit and Liu,
2008)</xref>
          between two different corpora as in the formula (1). Given the syntactical
pattern used for terms in this study, we take into account only nouns and adjectives in
both corpora because they are the kind of words most used for building terms:
(1)
Where fdom and Ndom correspond to the absolute occurrence frequency of wi and the
size of the domain corpus, respectively. Similarly, fref and Nref correspond to absolute
occurrence frequency of wi and the size of the reference corpus.
        </p>
        <p>Kit and Liu (2008) only focus on extracting single-word term candidates, so they
only weigh words occurring in both the domain and the general corpus. In our
experiment we also consider words that only occur in the domain corpus. We assumed that
the reference corpus is large enough to filter out non-relevant words, hence words
only occurring in the domain corpus have a higher probability of being relevant and
the word’s frequency reflects its importance:
We consider that the larger the reference corpus, the higher the exhaustivity3 of open
class words of general usage, as well as a higher probability that specialty terms occur
at least one time (the reference corpus was collected from an online newspaper where
news about science and technology are published too), so that we would expect a
higher precision in ranking.
6.6.</p>
      </sec>
      <sec id="sec-6-6">
        <title>Ranking multi-word term candidates</title>
        <p>Formally, if a candidate noun phrase (np) has a length of n words, w1 w2 …wn, where
n&gt;1, then the ranking of the candidate np is the sum of the frequency of np as a whole
plus the weights of all the individual words wi:
(2)
(3)
7.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Results</title>
      <p>This section presents the results of our experiment considering a subset of 1,200,000
tokens of the MedLineplus corpus.
7.1.</p>
      <sec id="sec-7-1">
        <title>Sources of textual information</title>
      </sec>
      <sec id="sec-7-2">
        <title>Domain corpus</title>
        <p>The source of textual information is constituted by a set of documents of the medical
domain, basically human body diseases and related topics (surgeries, treatments, and
so on). These documents were collected from MedlinePlus in Spanish.</p>
        <p>The size of the corpus is 1.2 million tokens, but we carried out our experiment with
a subset of 200,000 words in order to determine manually the number of concrete
entities present in the results. As an ongoing work, we are manually determining how
many concrete entities are present in the complete corpus. We chose a medical
domain due to the availability of textual resources in digital format. Finally, we assume
that the choice of domain does not suppose a very strong constraint for generalizing
the results to other domains.</p>
      </sec>
      <sec id="sec-7-3">
        <title>Reference corpus</title>
        <p>
          With the goal of ranking words relevant to the domain by means of their relative
frequency ratio, a large reference corpus was collected from an online newspaper4 with
new articles from 2014 (the size of corpus is about 5 million tokens). URLs from the
3 Exhaustivity of a document description is the coverage it provides for the main topics of the
document. So, if we add new vocabulary terms to a document, the exhaustivity of the
document description increases
          <xref ref-type="bibr" rid="ref3">(Baeza and Ribeiro, 2011)</xref>
          .
4 www.lajornada.com.mx. Mexican newspaper with information available online.
main heads were automatically extracted using the Python library BeautifulSoup5.
Then, this set of URLs was introduced in WebBootCat, a search tool of Sketch
Engine6, in order to automatically collect the textual information from each WEB page.
The description of the structure of the reference corpus is showed in table 1.
The programming language used in order to automate all tasks required was Python
version 3.4 as well as the NLTK module version 3.0
          <xref ref-type="bibr" rid="ref5">(Bird, Klein and Loper, 2009)</xref>
          .
Additionally, the POS tagger used in this experiment was FreeLing which is included
in Sketch Engine.
5
6 https://the.sketchengine.co.uk
        </p>
        <sec id="sec-7-3-1">
          <title>Analysis of results</title>
          <p>The first phase of extraction of candidates to concrete entity without filters achieves a
global precision of 56%. The tables 2 and 3 show precision with different thresholds
of candidates starting with the better ranked candidates. With the stopword list built as
mentioned in section 6.4, we achieve a global precision of 76%. Global precision with
a stopword list reflects an improvement of 20%, but a significant loss of 17% of true
candidates. As can be seen from these tables, the ranking of words and noun phrases
is useful for sorting results from the most relevant to the least relevant results.</p>
        </sec>
      </sec>
      <sec id="sec-7-4">
        <title>Bootstrapping phase</title>
        <p>The bootstrapping phase taking into account coordinating phrases achieves a set of
1248 candidates, of which 262 are new true candidates. The global precision with this
second phase is of 47%, with a precision by thresholds as shown in table 3. The
advantage of this phrase structure is that single-word candidates can be extracted.</p>
        <p>On the other hand, the bootstrapping phase considering noun phrases achieves a
set of 2796 candidates, of which 1534 are good candidates. The global precision of
this phase is of 55%, with a precision by thresholds as shown is table 3. One
disadvantage of this structure is that only candidates with at least one adjective can be
selected.</p>
        <p>Table 3 shows a better performance with noun phases. The identification of the
concrete entities present in corpus is an ongoing task that will let us evaluate in terms
of recall too.</p>
        <sec id="sec-7-4-1">
          <title>7.4. Discussion</title>
          <p>The candidates in a bootstrapping phase give us insight about the kind of semantic
relations implicit in noun phrases of the type &lt;NC&gt;&lt;AQ&gt;. Given the phase of
reduction of non-relevant adjectives, we have a great deal of relational adjectives where it
is possible to find different relations. For example, salivary gland has implicit a telic
relation. On the other hand, testicular gland has a part-whole or locative relation.
Finally, meibomian gland may be considered as a specific type of gland.</p>
          <p>
            With respect to the extraction of lexical relations, specifically
hyponymy-hypernymy relations
            <xref ref-type="bibr" rid="ref16 ref27">(Hearst, 1992; Wilks, Slator and Guthrie, 1995; Pantel and
Pennacchiotti, 2006)</xref>
            , as well as meronymy relations
            <xref ref-type="bibr" rid="ref15 ref4">(Berland and Charniak, 1999; Girju,
Badulescu and Moldovan, 2006)</xref>
            , these works are based on patterns where two terms
are located in the context of a sentence: the hand has fingers, the dog is an animal,
and so on, but there are few jobs working with noun phrases, which we consider it is
very important because we could consider a noun phrase as salivary gland as an
hyponym of gland, but it is clear that if we dig a little deeper that the semantic relation
implicit is telic.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Conclusions</title>
      <p>We discussed a methodology for extracting concrete entities in the medical domain.
Concrete entities have been studied since Aristotle’s works, particularly in his
biological and zoological descriptions. According to Aristotle’s categories (the first
category), many things can be predicated of substances. We assume that substances are
concrete entities, with a more extended meaning, i.e.: the eight tangible categories
formulated by Fellbaum for WordNet (1998). Thus, we consider that the automated
identification and extraction of this kind of information is an important advance in further
NLP tasks.</p>
      <p>Cognitive abilities as the spatial knowledge and his representation in natural
language are important for our extraction methodology. We observe that spatial
descriptions are frequent in specialized discourses. Additionally, we propose a further step of
bootstrapping in order to find a great number of candidates for concrete entities.
Candidates with a concrete entity as a noun head and a relational adjective show semantic
relations as part-whole, locative, agentive and telic, which can be interpreted, at first,
as hyponymy/hyperonymy relations.</p>
      <p>On the other hand, to assign relevance to words is an important step for ranking
candidates, according to our exposed results. In this sense, as ongoing work, we are
collecting more information about science and technology at the same electronic
journal in order to improve the results in the ranking process.</p>
      <p>Finally, it is necessary to mention that POST taggers as FreeLing and TreeTagger
fail in the task of identifying nouns, adjectives and verbs closely related with the
domain. This failure has a negative impact on the results. We believe it is important to
face this problem in future extraction tasks.</p>
      <sec id="sec-8-1">
        <title>Acknowledgments</title>
        <p>This paper has been supported by the National Commission for Scientific and
Technological Research (CONICYT) of Chile, Project Numbers: 3140332 and
11130565.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Acosta</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aguilar</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Sierra</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <article-title>Using Relational Adjectives for Extracting Hyponyms from Medical Texts</article-title>
          . In A. Lieto &amp; M. Cruciani (eds.),
          <source>Proceedings of the First International Workshop on Artificial Intelligence and Cognition (AIC</source>
          <year>2013</year>
          ),
          <source>CEUR Workshop Proceedings</source>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>44</lpage>
          .Torino, Italy. (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Acosta</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Aguilar</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Extraction of Concrete Entities and Part-Whole Relations</article-title>
          . In B. Sharp &amp; R. Delmonte (eds.),
          <source>Natural Language Processing and Cognitive Science. Proceedings</source>
          <year>2014</year>
          , pp.
          <fpage>89</fpage>
          -
          <lpage>100</lpage>
          . Berlin, De Gruyter (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Baeza</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Riveira</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Modern Information</surname>
            <given-names>Retrieval</given-names>
          </string-name>
          , 2nd ed. New York, Addison Wesley (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Berland</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Charniak</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <article-title>Finding parts in very large corpora</article-title>
          .
          <source>In Proceedings of the 37th Annual Meeting of the Association for Computational Linguistics</source>
          , pp.
          <fpage>57</fpage>
          -
          <lpage>64</lpage>
          . College Park, Maryland, USA, ACL Publications (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Loper</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Natural Language Processing with Python</surname>
          </string-name>
          , Sebastropol, Cal.,
          <string-name>
            <surname>O'Reilly</surname>
          </string-name>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bosque</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>Sobre las diferencias entre los adjetivos relacionales y los calificativos</article-title>
          . Revista Argentina de Lingüística,
          <source>No. 9</source>
          , pp.
          <fpage>10</fpage>
          -
          <lpage>48</lpage>
          (
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Carreras</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Chao</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Padró</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Padró</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>FreeLing: An Open-Source Suite of Language Analyzers</article-title>
          . In M.T. Lino et al. (eds.)
          <source>Proceedings of the 4th International Conference on Language Resources and Evaluation LREC</source>
          <year>2004</year>
          , pp.
          <fpage>239</fpage>
          -
          <lpage>242</lpage>
          . Lisbon, Portugal, ELRA Publications (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Cademártori</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parodi</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Venegas</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>El discurso escrito y especializado: caracterización y funciones de las nominalizaciones en los manuales técnicos</article-title>
          , Literatura y Lingüística,
          <source>No. 17</source>
          , pp.
          <fpage>243</fpage>
          -
          <lpage>265</lpage>
          (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Chunyu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <article-title>Measuring mono-word termhood by rank difference via corpus comparison</article-title>
          .
          <source>Terminology</source>
          ,
          <volume>14</volume>
          (
          <issue>2</issue>
          ),
          <fpage>204</fpage>
          -
          <lpage>229</lpage>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Clementini</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Di</surname>
            <given-names>Felice</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            , &amp;
            <surname>Hernández</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Qualitative representation of positional information</article-title>
          .
          <source>Artificial intelligence</source>
          ,
          <volume>95</volume>
          (
          <issue>2</issue>
          ),
          <fpage>317</fpage>
          -
          <lpage>356</lpage>
          (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Demonte</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>El adjetivo</article-title>
          . Clases y usos.
          <article-title>La posición del adjetivo en el sintagma nominal</article-title>
          . In I. Bosque &amp; V. Demonte (eds.),
          <source>Gramática descriptiva de la lengua española</source>
          , Vol.
          <volume>1</volume>
          ,
          <issue>Cap</issue>
          . 3, pp.
          <fpage>129</fpage>
          -
          <lpage>215</lpage>
          . Madrid, Espasa-Calpe (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>A Glossary of Cognitive Linguistics, Edinburgh</article-title>
          ,
          <string-name>
            <surname>UK</surname>
          </string-name>
          , Edinburgh University Press (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Fábregas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>The internal syntactic structure of relational adjectives</article-title>
          ,
          <source>Probus</source>
          ,
          <volume>19</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>36</lpage>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>WordNet: An Electronic Lexical Database</article-title>
          , Cambridge, Mass., MIT Press (
          <year>1998</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Girju</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Badulescu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Moldovan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>Automatic discovery of part-whole relations</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>32</volume>
          (
          <issue>1</issue>
          ),
          <fpage>83</fpage>
          -
          <lpage>135</lpage>
          (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Hearst</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Automatic Acquisition of Hyponyms from Large Text Corpora</article-title>
          .
          <source>In Proceedings of the Fourteenth International Conference on Computational Linguistics</source>
          , pp.
          <fpage>539</fpage>
          -
          <lpage>545</lpage>
          , Nantes, France.
          <source>ACL Publications</source>
          (
          <year>1992</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Landau</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Jackendoff</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>What and where in spatial language and spatial cognition, Behavioral and</article-title>
          brain sciences,
          <volume>16</volume>
          (
          <issue>02</issue>
          ),
          <fpage>255</fpage>
          -
          <lpage>265</lpage>
          (
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Levinson</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Space in Language and Cognition: Explorations in Cognitive Diversity</article-title>
          , Cambridge, UK, Cambridge University Press (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Mani</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doran</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hitzeman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quimby</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Richer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Clancy</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>SpatialML: annotation scheme, resources, and evaluation</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>44</volume>
          (
          <issue>3</issue>
          ),
          <fpage>263</fpage>
          -
          <lpage>280</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>James R</given-names>
          </string-name>
          .
          <article-title>Technicality and abstraction: Language for the creation of specialized texts</article-title>
          . In M.
          <string-name>
            <surname>A.K. Halliday</surname>
          </string-name>
          &amp;
          <string-name>
            <surname>James R. Martin</surname>
          </string-name>
          .
          <article-title>Writing science: Literacy and discursive power</article-title>
          , pp.
          <fpage>203</fpage>
          -
          <lpage>220</lpage>
          , London, The Falmer Press (
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Murphy</surname>
          </string-name>
          , G.
          <source>The Big Book of Concepts</source>
          . Cambridge, Mass., MIT Press (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Pustejovsky</surname>
          </string-name>
          .
          <source>J. The generative lexicon</source>
          , Cambridge, Mass., MIT Press (
          <year>1996</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <article-title>Probabilistic part-of-speech tagging using decision trees</article-title>
          .
          <source>In Proceedings of the International Conference on New Methods in Language Processing</source>
          , Vol.
          <volume>12</volume>
          , pp.
          <fpage>44</fpage>
          -
          <lpage>49</lpage>
          . Manchester, UK (
          <year>1994</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Controlled vocabularies in bioinformatics: a case study in the gene ontology</article-title>
          ,
          <source>Drug Discovery Today: BIOSILICO</source>
          ,
          <volume>2</volume>
          (
          <issue>6</issue>
          ),
          <fpage>246</fpage>
          -
          <lpage>252</lpage>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Vivanco</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>El español de la ciencia y la tecnología</article-title>
          , Madrid, Arco Libros (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Vivaldi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>Extracción de Candidatos a Término mediante combinación de estrategias heterogéneas</article-title>
          .
          <source>PhD Dissertation</source>
          . Barcelona, Universidad Politècnica de Catalunya (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Wilks</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Slator</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Guthrie</surname>
            ,
            <given-names>L. Electric</given-names>
          </string-name>
          <string-name>
            <surname>Words</surname>
          </string-name>
          , Cambridge, Mass., MIT Press (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Winston</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaffin</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Herrmann</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>A taxonomy of part-whole relations</article-title>
          ,
          <source>Cognitive science 11</source>
          (
          <issue>4</issue>
          ),
          <fpage>417</fpage>
          -
          <lpage>444</lpage>
          (
          <year>1987</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>