<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extracting knowledge-rich information from definitions. A corpus-based approach to building a conceptual-based terminological resource</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Margarida Ramos</string-name>
          <email>mvramos@fcsh.unl.pt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rute Costa</string-name>
          <email>rute.costa@fcsh.unl.pt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>NOVA CLUNL, Centro de Linguística da Universidade NOVA de Lisboa, Avenida de Berna 26-C</institution>
          ,
          <addr-line>1069-061 Lisboa</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2004</year>
      </pub-date>
      <abstract>
        <p>This paper aims to describe a text-mining approach on a domain corpus (cork) within the theoretical framework of the dual dimension of terminology to create a terminological dictionary and correlate it with an ontology. We will make some considerations on (i) domain specificities; (ii) lexical markers; (iii) automatic corpus processing using Sketch Engine; (iv) representation of lexical networks using CmapTools; and (v) representation of the concept system using Protégé. The goal of the ontology is to logically support the coherence and quality of the natural language definitions contained in the terminological resource.</p>
      </abstract>
      <kwd-group>
        <kwd>1 terminology</kwd>
        <kwd>definition</kwd>
        <kwd>domain-ontology</kwd>
        <kwd>knowledge-rich information</kwd>
        <kwd>domain terminological dictionary</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>This paper aims to demonstrate a method for developing a terminological dictionary based on a
domain ontology. To this end, we will describe the methods used to capture specialised lexical and
conceptual knowledge from the corpus and use it to develop a dedicated ontology. The terminological
resource will consist of a linguistic description of the specialised concepts, based on the formal
definitions of the concepts that make up the cork ontology, the OntoCork [11].</p>
      <p>
        The method used in this paper is corpus-driven. The corpus was compiled based on rigorous criteria
specific to terminological work [10], where the specialised context of text production is a key-element.
In this sense, the corpus is composed of technical explanatory and normative (standards) texts. For
corpus analysis, we used Sketch Engine2 to find and systematise lexical-semantic relationships. During
the corpus analysis process, we found two types of relevant knowledge-rich information [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]: definitions
and definitional contexts. Definitions are one of the components of the glossary’s microstructure that
can be found at the end of the normative texts. The purpose of these definitions is to achieve a consensus
among the members of the cork community. On the other hand, the definitional contexts are integral
parts of the texts and have relevant specialised lexical-semantic markers in their structure.
      </p>
      <p>Our method encompasses two stages:
(i) From the linguistic analysis of the lexical markers, and the corresponding lexical-semantic
relations observed between the terms, we systematise the results into lexical maps using CmapTools3.</p>
      <p>(ii) Based on the previous stage, we proceed to the conceptual analysis and subsequent formal
representation. The conceptual analysis grounds the identification of conceptual relations obtained by
interpreting the lexical-semantic relations observed between two terms. To infer conceptual relations –
such as the associative type – and to identify characteristics that will help us in the process of building
concept systems, we resort to deductive mechanisms employing the Aristotelian formula: X=Y+DC to
build OntoCork.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Domain corpus: cork</title>
      <p>
        The Cork Corpus was built up from texts produced within the cork industry. The internal and
external criteria [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] used to build our specific-domain corpus are systematised in Table 1.
      </p>
      <sec id="sec-2-1">
        <title>Error! Reference source not found. 1</title>
      </sec>
      <sec id="sec-2-2">
        <title>Internal and external criteria of the cork corpus</title>
        <sec id="sec-2-2-1">
          <title>Criteria</title>
          <p>Degree of specialisation
Source validation
Type
Content adequacy
Synchronism (≤ 10 years)</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>Purpose/description</title>
          <p>Produced by experts and semi-experts
Entities recognised as an authority
Technical-explanatory; normative
On cork/Cork stopper</p>
          <p>Given the fast evolution of technology
The corpus comprises 98 texts written in European Portuguese (see Figure 1).
corpora collection
academic articles</p>
          <p>16%</p>
          <p>These texts were produced by experts from different organisations and in different domains related
to the cork industry. The texts were collected according to the following criteria: (1) texts produced by
and for the scientific community in the domain of cork; (2) texts produced by experts for quasi-experts;
and (3) texts produced by experts for non-experts.</p>
          <p>Considering the 98 documents of the corpus, we have obtained the quantitative data shown in Table</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Terminological data extraction</title>
      <p>2.</p>
      <sec id="sec-3-1">
        <title>Error! Reference source not found. 2</title>
      </sec>
      <sec id="sec-3-2">
        <title>Quantitative data of the corpus</title>
        <p>Tokens
Words
Sentences</p>
        <sec id="sec-3-2-1">
          <title>Total number</title>
          <p>1,712,652
1,217,968
48,031</p>
          <p>
            For the corpus exploration and linguistic data analysis, we mainly focused on 43 texts produced in
two communicative settings, namely (i) expert–semi-expert, and (ii)
expert–quasiexperts/professionals, while the remaining 55 texts were used as a reference corpus [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ] so that we could
compare a given terminological data extraction (see Figure 2). The corpus was processed using Sketch
Engine, with which we compiled, annotated, and queried the corpus employing advanced searches in
Corpus Query Language (CQL) format, where regular expressions (regex) are applied.
Scientific : expert-expert
Regulatory: semi-expert - expert
Marketing : semi-expert - non-expert
Narrative-Informative : semi-expert - non-expert
Economics : expert- semi-expert
Technical-explanatory &amp; normative : expert
quasi-expert / professional
6
          </p>
          <p>Communicative setting of text production
16
29
4
43
16
27
n=98
Error! Reference source not found. 2: Subcorpus under focus (43 texts) based on the communicative
setting of text production</p>
          <p>
            Among the results presented in Table 2, the most frequent terms are “cortiça” [cork]4 and “rolha”
[stopper]. Given the high frequency of these two terms, we analysed the contexts in which they occur
in the subcorpus (43 texts) using the Word Sketch function as a first option, with which we identified
some candidate terms such as “ROLHA COLMATADA” [colmated stopper] (in capital letters). We then
moved on to simple queries (concordances) to search for polylexical terms containing adjectives in their
pattern. Once the most common morphosyntactic structures of terms were identified, we decided to
improve our search for terms and definitions employing advanced queries, namely through regex, so
that we could capture knowledge-rich contexts (KRC) [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ], e.g., definitions (definitions found in context)
and definitional contexts (contexts explaining what the concept is; thus valuable for understanding
and/or elaborating proper definitions).
3.1.
          </p>
          <p>Exploring the corpus with text mining methods</p>
          <p>Based on the patterns we have identified within definitional contexts, we explored the subcorpus
with advanced queries using regexes. For this paper, we will highlight two specific regexes that proved
productive in isolating lexical relations between terms, but also in finding definitional contexts where
the generic term is expanded in its syntax (see Table 3).</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Error! Reference source not found. 3</title>
      </sec>
      <sec id="sec-3-4">
        <title>Linguistic expressions commonly used by experts</title>
        <p>Definitional contexts (pt)
(1) Rolha que foi submetida a um tratamento químico
com o objectivo de desinfectar e/ou homogeneizar
a cor e/ou branquear.
(2) Rolha cuja superfície lateral foi submetida a uma
operação de abrasão para a tornar cilíndrica ou
diminuir o seu diâmetro.</p>
        <p>Literal translation into English
en: Stopper that was submitted to chemical treatment with
the aim of disinfecting and/or homogenising the colour and /
or bleaching
en: Stopper whose side surface was submitted to an abrasion
operation to make it cylindrical or to reduce its diameter.]</p>
        <p>
          The first regex has the following structure: "rolha"[tag="V.P.*SF"], whose formulation aims to
match patterns as ONLY forms of “rolha” [stopper] followed by ANY past participle ONLY in the singular
and feminine inflection. For the elaboration of this regex, we considered the linguistic expressions used
repeatedly by the experts, such as the past participle co-occurring with a term (see Table 4). The
outcomes of this query, namely 69 hits, delivered the most productive patterns for identifying lexical
markers, such as “x foi submetida a y” [x was submitted to y], as well as terms whose morphosyntactic
structures fall under our search patterns, such as [Noun + Past Participle], e.g., “x acabada” [finished
X] or “X terminada” [finalised X] where X is a term and Y corresponds to a structure that has proved
4 Our translation
to be rich in knowledge information, i.e. information provided by the experts that allows us to perceive
their conceptualisations [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
        <p>Considering the satisfactory results obtained, we decided to expand its formulation (regex 2):
"rolha"[(tag="D.*"|(tag="S.*")]?[tag="A.*"]?"cortiça"?[]{0,4}"rolha"[]{0,4}[tag="V.
P.*SF"]. In this case, we want to match a context in which the terms “rolha” [stopper] and “cortiça”
[cork] may co-occur with either adjectives, past participles, or found duplicated, in addition to the
functional forms. Out of the 55 hits matched by this regex, 48 were either a description or a definition.</p>
        <p>From the whole set of descriptions or definitions semi-automatically extracted from the Cork
Corpus, we decided to select ten (10) definitions for linguistic and conceptual analysis (see Table 4).</p>
      </sec>
      <sec id="sec-3-5">
        <title>Error! Reference source not found. 4</title>
        <p>Ten (10) definitions to organise a typology of cork stoppers
#
1
2
3
4
5
6
7
8
10 definitions (literal translations from pt)</p>
        <p>10 definitions (pt) extracted from the Cork Corpus
stopper rolha
Product obtained from natural cork and / or Produto obtido da cortiça natural e / ou de cortiça
agglomerated cork, consisting of one or more aglomerada, constituído por uma ou mais peças,
pieces, intended to seal bottles or other containers destinado a vedar garrafas ou outros recipientes e a
and to preserve their contents. (5.1 - NORM) preservar o seu conteúdo. (5.1 - NORM)
STOPPER ROLHA
piece of cork, usually cylindrical, conical or prismatic peça de cortiça, em geral cilíndrica, troncocónica ou
quadrangular, sometimes with rounded or prismática quadrangular, por vezes de arestas
chamfered lateral edges, consisting of one or laterais boleadas ou chanfradas, constituída por um
several glued elements and intended to seal the ou vários elementos colados e destinada a vedar os
containers or contribute to their water tightness. recipientes ou a contribuir para a sua
(7.8 – TECH) estanquicidade (7.8 – TECH)
natural cork stopper rolha de cortiça natural
Stopper consisting entirely of natural cork Rolha totalmente constituída por cortiça natural.
Note: Natural cork stoppers that have been Nota: As rolhas naturais que tenham sido
submitted to the sealing operation (see 6.5.5) are submetidas à operação de colmatagem (ver 6.5.5)
commonly referred to as colmated natural stoppers. são comummente designadas por rolhas naturais
(5.5 – NORM) colmatadas. (5.5 – NORM)
colmated natural cork stopper rolha de cortiça natural colmatada
The colmated natural cork stopper is a stopper A rolha de cortiça natural colmatada é uma rolha
made of natural cork in which its lenticels are filled feita de cortiça natural em que são obturadas as
with a mixture of glues and cork powder from the suas lenticelas com uma mistura de colas e pó de
dimensional finishing processes of natural cork cortiça proveniente dos acabamentos dimensionais
stoppers. (6.1 – REP) das rolhas de cortiça natural. (6.1 – REP)
agglomerated cork stopper rolha de cortiça aglomerada
Stopper obtained by the agglutination of cork Rolha obtida pela aglutinação de granulado de
granules with a size between 0,25 mm and 8 mm, cortiça com dimensão compreendida entre 0,25mm
with addition of binders, by means of extrusion or e 8mm, com adição de ligantes, através de extrusão
moulding and composed of at least 51% by weight ou moldagem e composta, pelo menos, por 51 % de
of cork granules. (5.5 – NORM) granulado de cortiça, em peso. (5.5 – NORM)
agglomerated stopper: rolha aglomerada:
piece of agglomerated cork, obtained by extrusion peça de cortiça aglomerada, obtida por extrusão ou
or moulding (3.1 – STUD) moldagem (3.1 – STUD)
“SnoNnt+f.oB”ndp.id:sspitIksenoskrpstufhposoieersfrmdnd.aee(tsd5ui.gbr5nay–alactNioobOrnokR,dgM“ylnuo)”efidangdtogiclooamnteeesortarhtebeodntuchmoerbnkedarsn.d raedRNomioo“sllhnctbhaoa”oa:sdsnNfuoi+osetnrcsismlotitzasoaapddddoaeeossspc.i.ogo(nrr5tau.i5ççmaã–oncN,ao“OtrnupR”roMainld)cdeoicclaaodrotoinçsaúnmaugmeloroomudeeeramda
technical stopper rolha técnica
Technical stoppers are composed of a very dense As rolhas técnicas são constituídas por um corpo de
body of agglomerated cork with disks of natural cortiça aglomerada, muito denso, com discos de
cork glued to one end - or to both ends. Technical cortiça natural colados no seu topo – ou em ambos
stoppers with one disk on each end are called 1+1 os topos. As rolhas técnicas com um disco em cada
technical stoppers; those with two disks of natural topo são designadas rolhas técnicas 1+1. Com dois
cork on each end are called 2+2 technical stopper; discos de cortiça natural em cada topo chamam-se
and those with two disks glued at only one of the
ends are called 2+0 technical stoppers. (6.1 – REP)
rounded stopper
Stopper whose edges of one or two ends were
rounded by abrasion. (5.5 – NORM)
marked stopper
Stopper whose lateral surface or ends were marked
in ink or by fire (7.6 – TECH)
rolhas técnicas 2+2, e com dois discos em apenas
um dos topos chamam-se rolhas técnicas 2+0. (6.1 –
REP)
rolha boleada
Rolha cujas arestas de um ou dois topos foram
arredondadas, por abrasão. (5.5 – NORM)
ROLHA MARCADA
Rolha cuja superfície lateral ou topos foram
marcados a tinta ou a fogo. (7.6 – TECH)</p>
        <p>For this paper, we will consider only one definition, namely &lt;Rolha de cortiça natural&gt; [natural cork
stopper] (see line 3 in Table 4), to demonstrate our linguistic and conceptual analysis. However, instead
of using the definitional statement written in Portuguese, we have decided to use its literal translation
into English for clarity.</p>
      </sec>
      <sec id="sec-3-6">
        <title>Error! Reference source not found. 5</title>
      </sec>
      <sec id="sec-3-7">
        <title>Linguistic analysis of the definition of &lt;Natural cork stopper&gt;</title>
        <sec id="sec-3-7-1">
          <title>Concept</title>
          <p>&lt;Natural cork stopper&gt;</p>
        </sec>
        <sec id="sec-3-7-2">
          <title>Definition in context</title>
          <p>stopper consisting entirely of natural cork
Note: Natural cork stoppers that have been submitted to the sealing operation (see 6.5.5) are commonly
referred to as colmated natural stoppers</p>
          <p>
            In the second sentence – inserted as a note in the definition – another piece of information is obtained
from the analysis of the statement “natural cork stoppers that have been submitted to sealing operation”.
Here, the lexical marker is “submitted to” [submetidas à] and relates the term “natural cork stopper”
[rolha natural] to the term “sealing operation” [operação de colmatagem]. The term “sealing operation”
– which indicates an operation/activity – is related by the lexical marker “submitted to” [submetidas à]
to the term “natural cork stopper” – which we already know to be an object. The interpretation of their
meanings allows us to infer that the lexical-semantic relation established is meronymy, subtype
[ACTIVITY-FEATURE] [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ] (see Map 1 for the former, and Map 1.1 for the latter, in Figure 3).
          </p>
        </sec>
      </sec>
      <sec id="sec-3-8">
        <title>Error! Reference source not found.3: Lexical Map 1 and Lexical Map 1.1</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. The conceptual analysis</title>
      <p>The conceptual analysis corresponds to the second stage of the analysis of the definition in focus.
The differential characteristics found in this definition are expressed by /natural cork/, /natural/,
/colmated/ and /sealing operation/. The observations of this analysis are systematised in Table 6 and are
based on the lexical markers found in the linguistic analysis of the definition. At the same time, based
on the linguistic interpretation of the data, we extrapolated to conceptual relation identifiers.</p>
      <sec id="sec-4-1">
        <title>Error! Reference source not found. 6</title>
      </sec>
      <sec id="sec-4-2">
        <title>The conceptual analysis of the definition of &lt;Natural cork stopper&gt;</title>
        <p>As systematised in Table 6, we propose three conceptual relation identifiers, namely, (1)
has_substance, (2) is_a, and (3) has_process.</p>
        <p>
          (1) has_substance is expressed by the lexical marker “consisting entirely of” [totalmente
constituída por], which refers to the substance of the object. As we know from the linguistic analysis,
the term “natural cork” points to the notion of substance, a material that a given object can be made of.
Since &lt;Stopper&gt; is an object made of a substance, we propose the conceptual relation identifier
has_substance to represent such a semantic relation. This semantic relation mirrors a pragmatic
association - e.g., a thematic connection through virtue or experience, or a dependency between
concepts established by the proximity of time and space [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] - in which a &lt;Stopper&gt; is a [PRODUCT]
obtained from a substance, more specifically a [RAW MATERIAL]. From the interpretation of this
information, we assume that an associative conceptual relation is in place, subtype PRODUCT – RAW
MATERIAL, in which stopper points to the meaning of PRODUCT, and natural cork points to the meaning
of RAW MATERIAL. This interpretation can be represented as follows: [stopper] PRODUCT has_substance
[natural cork] RAW MATERIAL.
        </p>
        <p>
          The dichotomy PRODUCT – RAW MATERIAL has twofold importance at this point of the conceptual
analysis: on the one hand, it underpins the subtype of the associative relation, while on the other hand,
it is included in the Aristotelian formula [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ],[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] known as X = Y + DC, where X=specific concept;
Y=genus; and DC=differential characteristics. The purpose of using such a formula is to identify, for
7
Transcription in
        </p>
        <p>X=Y+DC
natural cork</p>
        <p>stopper
[SPECIES] =</p>
        <p>stopper
[GENUS] + DC ?
natural cork
stopper
[SPECIES]
= stopper
[GENUS]
+ natural cork</p>
        <p>[DC]
natural cork
[GENUS] = cork
[GENUS] +
natural [DC]
? [SPECIES] =
natural cork
stopper
[GENUS] +</p>
        <p>sealing
operation [DC]</p>
        <p>colmated
natural stopper
[SPECIES] =
natural cork
stopper
[GENUS]
+ colmated [DC]</p>
        <p>Differential
characteristi</p>
        <p>cs
/natural
cork/
/natural/
/sealing
operation/
/colmated/
the task of concept modelling, the characteristics stated in the definition under analysis. In order to use
such a formula, one must first identify two concepts: the specific concept and its genus. [We will
develop this further in the paper].</p>
        <p>(2) is_a relation: &lt;Natural Cork Stopper&gt; is the subordinate concept, which we have labelled
[SPECIES], and &lt;Stopper&gt; is the superordinate concept, which we have labelled [GENUS]. This
assumption can be represented as: [natural cork stopper] SPECIES is_a [ stopper] GENUS. Once the genus
and the species have been identified, we can then insert these two elements in the formula X SPECIES =
Y GENUS + DC, where: X = natural cork stopper; Y = stopper. Differential characteristics are inferred in
a second stage: considering that [stopper] PRODUCT has_substance [natural cork] RAW MATERIAL, we can
conclude that X [natural cork stopper] = Y [stopper] + DC [natural cork]. The first statement of the
definition conveys the information represented by the first interpretation above, with the dichotomy
[SPECIES-GENUS], which can be represented in the form of a conceptual map (see Figure 4). Conceptual
map 1 is built by applying a differentiae dichotomy in which the differential characteristic /natural cork/
underlies one of the subdivision criteria5.</p>
        <p>Error! Reference source not found.4: Conceptual Map 1 - two composition types of
&lt;Natural_cork_stopper&gt; in CmapTools
Conceptual Map 1 (Figure 4) is the conceptual representation of the first statement of the definition,
from which we have inferred that a &lt;Natural cork stopper&gt; is_a &lt;Cork stopper&gt;. Two axes of analysis
are considered in this map: Substance and Parts (the ‘Parts’ axis was inferred from Definition 1; see
Table 4). The conceptual information represented here, namely the axes of analysis Substance and Parts
– whose underlying characteristics are /natural cork/, /mono piece/ and /multi piece/ – will be some of
the coordinates for the elaboration of the formal description of the concept NaturalCorkStopper in
5 According to (ISO/FDIS 1087), the “subdivision criterion [is the] type of characteristic according to which a superordinate concept is
divided into subordinated concepts.” (2019 (E), p. 5).
Protégé. Finally, &lt;Multi_piece_natural_cork_stopper&gt; will help us to formally describe types of
&lt;Stoppers&gt; composed of several Parts – not only made of &lt;Natural_cork&gt;, but also of
&lt;Agglomerated_cork&gt; and &lt;Mixed_cork&gt;. Here, the characteristics fall under the axis of analysis
‘Parts’ and are the coordinates for modelling multi-part concepts.</p>
        <p>(3) has_process: Following the same method, the analysis of the note from which we obtained
the information: &lt;Natural cork stopper&gt; is submitted to /sealing operation/, was represented in a second
map (Figure 5). This piece of information grounds the conceptual relation identifier we have named as
has_process.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Error! Reference source not found.5:</title>
        <p>&lt;Mono_piece_natural_cork_stopper_with_sealing_operation&gt;</p>
      </sec>
      <sec id="sec-4-4">
        <title>Conceptual map of</title>
        <p>Conceptual Map 2 (Figure 5) is the representation of the two sentences of the definition in focus.
Therefore, three axes of analysis are now considered: Substance, Parts, and Finishing Processes, to
which the characteristics /with sealing operation/ and /without sealing operation/, were added. As
represented in Conceptual Map 2, the characteristics /with sealing operation/ and /without sealing
operation/ led us to a different level of concept representation, i.e., the concept
&lt;Mono_piece_natural_cork_stopper_with_sealing_operation&gt;, verbally designated by “colmated cork
stopper”, is a specialisation of &lt;Mono_piece_natural_cork_stopper&gt;, in turn, verbally designated by
“natural cork stopper”. Therefore, these two concepts should not be treated at the same level, nor should
they be defined in the same definitional context, either in natural language or in (semi)formal languages.</p>
        <p>The conceptual relations we have inferred from the analysis of the lexical markers observed in the
first five definitions (see Table 4), is summarised in Table 7.</p>
      </sec>
      <sec id="sec-4-5">
        <title>Error! Reference source not found. 7</title>
      </sec>
      <sec id="sec-4-6">
        <title>Overview of the conceptual relations inferred from lexical markers</title>
        <p>‘is a’
‘commonly
referred as’
‘is a’</p>
        <p>Conceptual relation
identifier
is_a
is_a
is_a
‘intended to’</p>
        <p>has_function
PARTITIVE
PARTITIVE
stopper [SPECIES] = product [GENUS] + several
pieces [PARTS=DC]
stopper [SPECIES] = piece of cork [GENUS] + one
element [PARTS=DC]
stopper [SPECIES] = piece of cork [GENUS] + several
elements [PARTS=DC]</p>
        <p>
          As shown in Table 7, differential characteristics (DC) can be any characteristic in a given definition
according to the formula of an intensional definition [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], so that, depending on what is added to the
intension of the [GENUS], the understanding of the concept’s place in the concept system is provided.
The same happens with the associative relation, although with several other axes of analysis involved.
Here, DC share semantic labels in a more productive variety, namely [SUBSTANCE]; [FUNCTION];
[PROCESS] and [SHAPE], given the prolific semantic relations identified between concepts.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Building the ontology</title>
      <p>
        For the task of building OntoCork, we used the editor Protégé [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. OntoCork is an ontology in which
the concepts of the domain of cork are systematised through logical constructs. The descriptive domain
properties (conceptual relations) elaborated to develop the ontology ground on the five axes of analysis
that we have previously retained, namely Part, Substance, Shape, Finishing Process, and Function, as
systematised in Table 8.
      </p>
      <sec id="sec-5-1">
        <title>Error! Reference source not found. 8</title>
      </sec>
      <sec id="sec-5-2">
        <title>Five core conceptual relations of the ontology</title>
        <p>Axis of analysis
FUNCTION
SUBSTANCE
PARTS
FINISHING
PROCESS
SHAPE</p>
        <p>Format in
Protégé
hasFunction
IsMadeOf
hasStructure
hasProcess
hasShape</p>
        <p>Type of conceptual relation
associative relation, subtype [OBJECT-FUNCTION]
associative relation, covering both subtypes [RAW MATERIAL – PRODUCT] and
[MATTER/SUBSTANCE – PROPERTY]
partitive relation [PART-WHOLE]
associative relation, within the subtype [PROCESS-RESULT]
associative relation, subtype [OBJECT-SHAPE]</p>
        <p>For this paper, we will present the description of the characteristics that build up the formal definition
of &lt;Natural cork stopper&gt;, a closure with a body-structure of 1 &lt;Part&gt; submitted to &lt;Sealing process&gt;,
in addition to the classification provided by the reasoner HermiT 6 as a &lt;Semi-finished&gt; object (see
Figure 6).
6 HermiT – a plugin reasoner of Protégé (http://www.hermit-reasoner.com/)</p>
      </sec>
      <sec id="sec-5-3">
        <title>Error! Reference source not found.6:</title>
      </sec>
      <sec id="sec-5-4">
        <title>ColmatedMonoPieceNaturalCorkStopper, in Protégé</title>
      </sec>
      <sec id="sec-5-5">
        <title>Concept description of</title>
        <p>Figure 7 is the ontological representation of ColmatedMonoPieceNaturalCorkStopper, in
Ontograf7, where we can observe several concepts systematised, either vertically: in a hierarchical
dependency, or horizontally: in a pragmatic (associative) dependency, according to the differential
characteristics. For clarity, we have decided to elide the visualisation of the associative relations
between concepts that are not in focus in the following lines.</p>
      </sec>
      <sec id="sec-5-6">
        <title>Error! Reference source not found.7:</title>
      </sec>
      <sec id="sec-5-7">
        <title>ColmatedMonoPieceNaturalCorkStopper, in Ontograf</title>
      </sec>
      <sec id="sec-5-8">
        <title>Ontological representation of</title>
        <p>As illustrated in Figure 7, the ColmatedMonoPieceNaturalCorkStopper is a specification of
MonoPieceNaturalCorkStopperWithFinishingProcess. The subsumption relation is represented by
vertical blue arcs, and the associative relations are represented by horizontal dashed lines. The concepts
7 https://protegewiki.stanford.edu/wiki/OntoGraf
ColmatedMonoPieceNaturalCorkStopper and LenticelsColmation are linked by the associative
relation, subtype [PROCESS-RESULT]: hasLenticelsColmationOperation. This conceptual relation is based on
the differential characteristic /with sealing operation/, which was drawn from the analysis of the
definition of &lt;Natural cork stopper&gt;. Thus, hasLenticelsColmationOperation is the associative relation that
induces the specification of MonoPieceNaturalCorkStopperWithFinishingProcess by differentia.
Finally, it is also possible to see a hierarchical representation of FinishingProcesses, in which the
involved operation of the concept we have just described is assigned as the most specific concept of
this hierarchy. The interpretation of this subsumption is: LenticelsColmation is a kind of
QualityTreatment, which is a kind of SurfaceTreatment, which in turn is a kind of
SemifinishingProcess, all of these are kinds of FinishingProcesses.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>With this research, we wanted to explain the method used to build an ontology from human analyses
of linguistic data. Linguistic and conceptual levels of analysis are to be analysed in relation to one
another but as distinct phenomena. Texts are vehicles for knowledge transfer. Analysing texts to extract
the characteristics of concepts, linguistically expressed by lexical markers pointing to lexical-semantic
relations, allowed us to effectively capture the conceptual relations that are specific to the domain
through the formula X=Y+DC. As demonstrated in this study, we were able to propose a preliminary
conceptual organisation of the subject field. We have bridged three main aspects in our study: (i) the
classical aspects of the Aristotelian logic; (ii) the methodology of our terminological work – where
characteristics play a fundamental role in the analysis or the drafting of intensional definitions; and (iii)
the formal definitions, for which we have used Protégé and the inherent Web Ontology Language
(OWL) [12] to formally describe the concepts of the domain in order to relate them via abstract syntaxes
and thus achieve formal reasoning, as concepts are consistently defined in a ‘reason-able’ ontology.</p>
      <p>In future work, we intend to model the conceptual and the linguistic information contained in the
resources we have developed, namely the ontology, the corpus, and a glossary (in progress) developed
with Lexonomy8, as linked data with the use of interoperable Linked Open Vocabularies9.</p>
    </sec>
    <sec id="sec-7">
      <title>7. References</title>
      <p>8 https://github.com/elexis-eu/lexonomy
9 https://lov.linkeddata.es/dataset/lov/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Atkins</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clear</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ostler</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>1992</year>
          ).
          <article-title>Corpus Design Criteria</article-title>
          .
          <source>Literary and Linguistic Computing</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          . doi:
          <volume>10</volume>
          .1093/llc/7.1.
          <fpage>1</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hardie</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>McEnery</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>A Glossary of Corpus Linguistics</article-title>
          . Edinburgh: Edinburgh University Press.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>ISO</surname>
          </string-name>
          <year>704</year>
          .
          <article-title>(</article-title>
          <year>2009</year>
          ).
          <article-title>Travail terminologique - Principes et méthodes</article-title>
          .
          <source>NF ISO 704, 1er tirage</source>
          <year>2009</year>
          - 12
          <string-name>
            <surname>-P. La Plaine</surname>
          </string-name>
          Saint-Denis: Association Française de Normalisation.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4] ISO/FDIS 1087.
          <article-title>(2019 (E))</article-title>
          .
          <article-title>Terminology work and terminology science - Vocabulary</article-title>
          . Suisse: ISO.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L</given-names>
            <surname>'Homme</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. C.</surname>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>La Terminologie: principes et techniques - Paramètres</article-title>
          . Montréal, Canadá: Les presses de l'Université de Montréal.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6] Meyer,
          <string-name>
            <surname>I.</surname>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Extracting Knowledge-Rich contexts for terminography: a conceptual and methodological framework</article-title>
          . In D. Bourigault,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jacquemin</surname>
          </string-name>
          , &amp;
          <string-name>
            <surname>M.-C. L'Homme</surname>
          </string-name>
          (Eds.),
          <source>Recent Advances in Computational Terminology</source>
          (Vol.
          <volume>2</volume>
          , pp.
          <fpage>279</fpage>
          -
          <lpage>302</lpage>
          ). Amsterdam / Philadelphia: John Benjamins B.V.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Musen</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>The Protégé project: A look back and a look forward</article-title>
          .
          <source>doi:10.1145/2557001</source>
          .25757003
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Pearson</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>Terms in context</article-title>
          . Amsterdam: John Benjamins B.V.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Pottier</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>1992</year>
          ).
          <article-title>Théorie et analyse en Linguitique (2</article-title>
          , corrigée ed.). Paris: HACHETTE, Supérieur.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>