<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extraction of Definitional Contexts from Restricted Domains by Measuring Synthetic Judgements and Word Relevance</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>César Aguilar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olga Acosta</string-name>
          <email>olgalimx@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Pontificia Universidad Católica de Chile</institution>
          ,
          <addr-line>Santiago de</addr-line>
          <country country="CL">Chile</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Pontificia Universidad Católica de Chile</institution>
          ,
          <addr-line>Santiago de</addr-line>
          <country country="CL">Chile</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>161</fpage>
      <lpage>166</lpage>
      <abstract>
        <p>In this article we present an ongoing work for extracting conceptual information from specialized-domain texts. Concepts are forms of dividing the world in classes and they are the fundamental pieces for constructing ontologies. In this sense, ontology learning is the (semi-) automatic support for constructing an ontology. Input data are required for the ontology learning and this data are the basic source from which to learn the relevant concepts for a domain, their definitions as well the relations holding between them. With this necessity in mind, we propose here a methodology that takes into account the level of synthetic judgements and word relevance in a sentence in order to filter out and rank sentences. Sentences with high relevance and low level of synthetic judgements should have at least a predicative verb characteristic of analytical definitions for being good candidates.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Concepts are one of the most fundamental pieces
of the cognition: humans daily use concepts for
interacting with others and the world. According
to Smith (1988), concepts mirror the way that we
divide the world into classes, and much of what
we learn, communicate, and reason involves
relations among these classes. Additionally,
        <xref ref-type="bibr" rid="ref9">Rosch
(1978)</xref>
        argues that concepts promote the
cognitive economy because the human beings attempt
to gain as much information as possible about its
environment while minimizing cognitive effort
and resources.
      </p>
      <p>Currently, due to the accelerated growth of
digital information on the Web and other media as
well the urgent necessity of obtaining relevant
information in a fast and efficient way from
these huge text sources, automated methods or
approaches have been developed. For instance,
in Maedche and Staab (2004) define ontology
learning as a number of complementary
disciplines that feed on different types of unstructured
and semi-structured data in order to support a
semi-automatic ontology engineering process. In
line with this, Cimiano (2006) describes various
sub-processes for constructing an ontology from
texts where the concept extraction is an
important phase. So, the ontology learning needs
input data from which to learn the relevant
concepts for a given domain.</p>
      <p>According to these ideas, in this paper we sketch
a methodology for recognizing candidates to
analytical definitional contexts, according to the
work developed by Sierra et al. (2008). We
organize our work as follows: in section 2 we
present general information about analytical
definitions and the automated extraction of
conceptual information. In section 3 we describe the
function of adjectives as modifiers of a noun as
well the distinction among descriptive and
relational adjectives and the relation of descriptive
adjectives with synthetic judgements in an
attributive form. In section 4 we summarize the
methodology proposed. In section 5 we show
some preliminary results. Finally, in section 6 we
present the future work.</p>
    </sec>
    <sec id="sec-2">
      <title>Conceptual information</title>
      <p>We consider as conceptual information the
information expressed by specialized definitions,
particularly in analytical definitions constituted
by Genus Term and Differentia, following the
criteria formulated by Smith (2004). In fact, this
author considers that information expressed by
these kinds of definitions is relevant to create
ontologies based in lexical relations, specifically
hyponymy/hypernymy and meronymy/holonymy
relations. Smith argues that these relations, from
a philosophical point of view, are basic and
universal.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Analytical Definitions</title>
      <p>An analytical definition is a formula for
describing a concept, denoted by a linguistic tag, in
terms of a superordinate concept (Genus Term),
and a differentia distinguishing the concept
defined from others with the same Genus Term.
For example, the next definition provides a
description of the concept lightning conductor
using one of the most common verbs (i.e., to be)
for introducing a definition. In this case, the
genus is the concept device while the differentia
describes the function of the lightning conductor:
[Lightning conductor Term] is a [device Genus Term]
[that allows to protect the electrical systems
against surges of atmospheric origin Differentia].
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Definitional contexts</title>
      <p>Sierra et al., (2008) proposed a based-pattern
method for extracting terms and definitions in
Spanish. This relevant information is expressed
in textual fragments called definitional contexts
(or DCs) and are constituted by: a term, a
definition, and linguistic or metalinguistic forms, such
as verb phrases, typographical markers and/or
pragmatic patterns, for example:</p>
      <p>The primary energy, in general terms, is
defined as an energetic resource that has not been
affected for any transformation, with the
exception of its extraction.</p>
      <p>We can see here a DC sequence formed by the
term primary energy, the definition that
resource that… and the verb pattern is defined as,
as well other characteristic units such as the
pragmatic pattern in general terms and the
typographical marker (bold font) that in this case
emphasizes the presence of the term.</p>
      <p>For achieving this objective, the authors employ
verb patterns operating as connectors between
terms and definitions. Such patterns
syntactically are predicative phrases (or PrP), configured
around a verb that operates as a head of this PrP
(e.g., to be, to characterize, to conceive, to
consider, to describe, to define, to understand, to
know, to refer, to denominate, to call, to name).</p>
    </sec>
    <sec id="sec-5">
      <title>Adjectives</title>
      <p>Based on Demonte (1999), adjectives are
syntactic units modifying the noun’s meaning and
associating it with one or various attributes. There
are two kinds of adjectives which assign
properties to nouns: descriptive and relational
adjectives. On the one hand, descriptive adjectives
refer to constitutive features of the modified
noun. These features are exhibited or
characterized by means of a single physical property:
color, form, character, predisposition, sound, and so
on: la silla verde (e.g., the green chair). On the
other hand, relational adjectives assign a set of
properties, i.e., all the characteristics jointly
defining names as sea: puerto marítimo (e.g.,
maritime port). In terminology, relational adjectives
represent an important element for building
specialized terms, e.g.: inguinal hernia, venereal
disease, psychological disorder and others are
considered terms in medicine. In contrast, rare
hernia, serious disease and critical disorder
seem more descriptive judgments and closely
related with a specific context.
3.1</p>
    </sec>
    <sec id="sec-6">
      <title>Syntactical Identification</title>
    </sec>
    <sec id="sec-7">
      <title>Relevant Adjectives of Non</title>
      <p>In line with what was just mentioned, if we
consider the internal structure of adjectives, two
kinds of adjectives can be identified: permanent
and episodic adjectives (Demonte, 1999). The
first kinds of adjectives represent stable
situations, permanent properties characterizing
individuals. These adjectives are located outside of
any spatial or temporal restriction (i.e.,
psicópata- psychopath). On the other hand,
episodic adjectives refer to transient situations or
properties implying change and with time-space
limitations. Almost all descriptive adjectives
derived of participles belong to this latter class as
well all adjectival participles (i.e., harto-jaded,
limpio-clean). Spanish is one of the few
languages that in syntax represent this difference in
the meaning of adjectives. In many languages
this difference is only recognizable through
interpretation. In Spanish, individual properties can
be predicated with the verb ser, and episodic
properties with the verb estar.</p>
      <p>Another linguistic heuristics for identifying
descriptive adjectives is that only these kinds of
adjectives accept degree adverbs, and they can be
part of comparative constructions, for example,
muy alto (Eng.: very high). Finally, only
descriptive adjectives can precede a noun because
in Spanish relational adjectives are always
postposed, i.e.: la antigua casa (Eng.: the old
house).
3.2</p>
    </sec>
    <sec id="sec-8">
      <title>Synthetic Judgements and Descriptive</title>
    </sec>
    <sec id="sec-9">
      <title>Adjectives</title>
      <p>According to Kant (2013), analytic sentences are
those whose truth seems to be knowable by
knowing the meanings of the constituent words
alone (e.g., gynecologists are doctors), unlike the
more usual synthetic ones (e.g., gynecologists
are rich), whose truth is knowable by both
knowing the meaning of the words and
something about the world.</p>
      <p>We believe that synthetic judgements in an
attributive position (e.g., rich gynecologists) are
common in non-relevant sentences in specialized
domains. This kind of judgements can be
recognized from the descriptive adjectives obtained by
linguistic heuristics mentioned in section 3.1.</p>
    </sec>
    <sec id="sec-10">
      <title>Methodology</title>
      <p>
        We present here our methodology for extracting
conceptual information from a medical domain
corpus. The input data consist of a corpus with
POS tagged with FreeLing
        <xref ref-type="bibr" rid="ref3">(Carreras et al.,
2004)</xref>
        .
4.1
      </p>
    </sec>
    <sec id="sec-11">
      <title>Sentence Segmentation</title>
      <p>The heuristics assumed here in order to segment
our corpus by sentences take into account that a
sentence must be separated by a point, to have at
least a main verb, and the number of words must
be greater than 10 words because the most short
DC would have a single word term, the most
long predicative verb-is defined as, a possible
article preceding genus, genus term and, in this
case, some arbitrary limit of words for the
differentia).
4.2</p>
    </sec>
    <sec id="sec-12">
      <title>Filtering out Sentences by Predicative</title>
    </sec>
    <sec id="sec-13">
      <title>Verbs</title>
      <p>The set of sentences obtained by the above step
are filtered out by considering predicative verbs
mentioned in section 2.2, that is, if there is at
least a predicative verb; then it is a good
candidate to DC. For the case of to be, if it is the first
word of the sentence, then it is discarded.
4.3</p>
    </sec>
    <sec id="sec-14">
      <title>Chunking</title>
      <p>
        We have used the library of Natural Language
NLTK
        <xref ref-type="bibr" rid="ref2">(Bird, Klein and Loper, 2009)</xref>
        in the
Python language, for implementing a chunker in
order to extract descriptive adjectives with
heuristics described in section 3.1.
      </p>
      <p>In this work, we propose a phase of
quantification of synthetic judgments in candidate
sentences as a further filter of non-relevant
sentences. We assumed here that synthetic
judgments are descriptive adjectives in an attributive
position (e.g., rare syndrome). So, the higher
amount of synthetic judgments in a sentence, the
more likely sentence is non-relevant. We
considered the set of descriptive adjectives obtained by
heuristics as a mechanism for this quantification
of syntheticity.</p>
      <p>Acosta, Aguilar and Sierra (2013) point out
relational adjectives have a higher probability of
being part of terms. The heuristics considered in
this experiment are:
&lt;RG&gt;&lt;AQ&gt;
&lt;VAE&gt;&lt;AQ&gt;
&lt;D.*|P.*|F.*|S.*&gt;&lt;AQ&gt;&lt;NC&gt;
Where RG, AQ and VAE as tagged with
FreeLing, correspond to adverbs, adjectives and
the verb estar, respectively. The tags
&lt;D.*|P.*|F.*|S.*&gt; correspond to determinants,
pronouns, punctuation signs and prepositions.
The expression &lt;D.*|P.*|F.*|S.*&gt; is a
restriction to reduce noise, since elements wrongly
tagged by FreeLing as adjectives are extracted
without this restriction.
4.4</p>
    </sec>
    <sec id="sec-15">
      <title>Weighting Words</title>
      <p>
        We evaluated relevance of simple words by
means of a corpus comparison approach by
applying the relative frequency ratio
        <xref ref-type="bibr" rid="ref8">(Manning and
Schütze, 1999)</xref>
        between two different corpora as
in (1). Given that the syntactical pattern of most
common terms in Spanish is &lt;NC&gt;&lt;AQ&gt;
(Vilvaldi, 2004), we take into account only
nouns and adjectives in both corpora:
(1)
Where , correspond to the absolute
occurrence frequency of wi and the size of the domain
corpus, respectively. Similarly, ,
correspond to absolute occurrence frequency of wi
and the size of the reference corpus. The measure
in (1) is only calculated for wi’s, where
. Otherwise, wi can be used as part of a
list of non-relevant words for purposes of
quantifying non-relevance in sentences. On the other
hand, words only occurring in domain are
weighted as in (2). We assume that the reference
corpus is large enough for filter out non-relevant
words, hence words only occurring in the
domain corpus will have a higher probability of
being relevant so that the word’s frequency can
reflect its importance:
(2)
4.5
      </p>
    </sec>
    <sec id="sec-16">
      <title>Relevance of Sentences</title>
      <p>The ranking of sentences is done by adding up
the individual ranks of words present in the
sentence. Formally, if s (that is, a sentence) has a
length of n words, w1 w2 …wn, where n&gt;10, then
the ranking of the candidate s is the sum of the
weights of all the individual words wi W, where
W are all of the relevant words weighted as
mentioned in section 4.4. In contrast, if wi W, then
its weight is zero.</p>
    </sec>
    <sec id="sec-17">
      <title>Preliminary Results</title>
      <p>Considering descriptive adjectives automatically
extracted by heuristics for quantifying
syntheticity, the first results show to be a good filter in
order to remove non-relevant fragments by
setting thresholds related with the number of
descriptive adjectives in sentences. At the same
time, the ranking of words achieves to sort
sentences according to its relevance for the domain.
Additionally, given that only sentences with
predicative verbs are considered, a subset of the
better ranked sentences are analytical DCs.
If we take into account words where relative
frequency in reference is greater or equal than in
domain (given its higher occurrence in reference
than in domain, we assume they are non-relevant
words) as part of this list for removing
nonrelevant sentences by setting thresholds (here,
nouns and adjectives are included) improve
significantly the results.</p>
    </sec>
    <sec id="sec-18">
      <title>Future results</title>
      <p>In a future phase of this experiment, we will
implement a syntactic phase in order to remove
more non-relevant sentences. For instance,
sentences with to be verb are the most common
sentences and which produce so much noise in
results. Given this, we consider that a syntactic
phase capable to assure the occurrence of
specific syntactic structures will be an important
advance in order to perform a better filtering.
On the other hand, we will continue with the
recollection of more information for increasing
the sections of science and technology in our
reference corpus, in order to improve the word
weighing and the calculation of relevance
sentences.</p>
    </sec>
    <sec id="sec-19">
      <title>Acknowledgments</title>
      <p>This paper has been supported by the National
Commission for Scientific and Technological Research
(CONICYT) of Chile, Project Numbers: 3140332 and
11130565.
semantic relation extraction. Terminology, 14(1):
74-98.</p>
      <p>Barry Smith. 2004. Beyond concepts: ontology as
reality representation. In Formal Ontologies in
Information Systems, ed. by Achille Varzi and Laure
Vieu, pp. 73-84., IOS Press, Amsterdam.</p>
      <p>Edward Smith. 1988. Concepts and Thought. In
Psychology of human thought, ed. by Robert J.
Sternberg, pp. 19-49. Cambridge University Press,
Cambridge, UK.</p>
      <p>Jorge Vivaldi. 2004. Extracción de candidatos a
términos mediante la combinación de estrategias
heterogéneas. Ph. D. Dissertation. IULA -UPF,
Barcelona.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Olga</given-names>
            <surname>Acosta</surname>
          </string-name>
          , Gerardo Sierra and
          <string-name>
            <given-names>César</given-names>
            <surname>Aguilar</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Extraction of Definitional Contexts using Lexical Relations</article-title>
          .
          <source>International Journal of Computer Applications</source>
          ,
          <volume>34</volume>
          (
          <issue>6</issue>
          ):
          <fpage>46</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Steven</given-names>
            <surname>Bird</surname>
          </string-name>
          , Ewan Klein and
          <string-name>
            <given-names>Edward</given-names>
            <surname>Loper</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <string-name>
            <given-names>Natural</given-names>
            <surname>Language Processing whit Python. O'Reilly</surname>
          </string-name>
          , Sebastropol, Cal.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Carreras</surname>
          </string-name>
          , Isaac Chao, Lluís Padró, and
          <string-name>
            <given-names>Muntsa</given-names>
            <surname>Padró</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>FreeLing: An Open-Source Suite of Language Analyzers</article-title>
          .
          <source>In Proceedings of the 4th International Conference on Language Resources and Evaluation LREC</source>
          <year>2004</year>
          , ed. by Maria Teresa Lino et al., pp.
          <fpage>239</fpage>
          -
          <lpage>242</lpage>
          . ELRA Publications, Lisbon, Portugal.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Cimiano</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Ontology Learning and Population from Text</article-title>
          . Springer, Berlin.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Violeta</given-names>
            <surname>Demonte</surname>
          </string-name>
          .
          <article-title>El adjetivo</article-title>
          . Clases y usos.
          <article-title>La posición del adjetivo en el sintagma nominal</article-title>
          . In Gramática descriptiva de la lengua española, ed.
          <source>by Ignacio Bosque and Violeta Demonte</source>
          . Vol.
          <volume>1</volume>
          ,
          <issue>Ch</issue>
          . 3, pp.
          <fpage>129</fpage>
          -
          <lpage>215</lpage>
          . Espasa-Calpe, Madrid.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Immanuel</given-names>
            <surname>Kant</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Crítica de la razón pura, edited and traslated to Spanish by Pedro Ribas</article-title>
          . Taurus, Madrid.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Maedche</surname>
          </string-name>
          and
          <string-name>
            <given-names>Steffen</given-names>
            <surname>Staab</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Ontology Learning</article-title>
          . In Handbook on Ontologies, ed.
          <source>by Steffen Staab and Rudi Studer</source>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>190</lpage>
          . Springer, Berlin.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          and
          <string-name>
            <given-names>Hinrich</given-names>
            <surname>Schütze</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Foundations of Statistical Natural Language Processing</article-title>
          . MIT Press, Cambridge, Mass.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Rosch</surname>
          </string-name>
          .
          <year>1978</year>
          .
          <article-title>Principles of categorization</article-title>
          . In Cognition and Categorization, ed.
          <source>by Elinor Rosh and Barbara Lloyd</source>
          , pp.
          <fpage>27</fpage>
          -
          <lpage>48</lpage>
          . Lawrence Erlbaum Associates, Hillsdale, NJ.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Gerardo</given-names>
            <surname>Sierra</surname>
          </string-name>
          , Rodrigo Alarcón, César Aguilar and
          <string-name>
            <given-names>Carme</given-names>
            <surname>Bach</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Definitional verbal patterns for</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>