<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>CLiC-it</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Towards a Multi-Level Annotation Format for the Interoperability of Automatic Term Extraction Corpora</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicola Cirillo</string-name>
          <email>nicirillo@unisa.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniela Vellutino</string-name>
          <email>dvellutino@unisa.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Terminology, Automatic Term Extraction, Linguistic Linked Data, OntoLex-Lemon</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Salerno</institution>
          ,
          <addr-line>132 Via Giovanni Paolo II, Fisciano (SA), 84084</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>9</volume>
      <abstract>
        <p>English. The main corpora used as benchmarks in Automatic Term Extraction are represented in diferent formats. Unfortunately, none of these formats covers the wide range of linguistic phenomena related to terminology. To address this issue, we propose to encode Automatic Term Extraction corpora in RDF using the OntoLex-Lemon and the NLP Interchange Format ontologies. Furthermore, we developed a small Italian corpus on waste management legislation to provide an example of the proposed formalization.</p>
      </abstract>
      <kwd-group>
        <kwd>Extraction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Automatic Term Extraction - ATE is an NLP task that
involves recognizing terms in specialized corpora. As
with most NLP tasks, ATE research benefits from
annotated corpora that are employed as training data and
evaluation benchmarks. Nevertheless, existing term
annotation schemata are far from capturing the complex
organization that characterizes the terminology of
specialized languages [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. Most ATE studies, with a few
exceptions [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4, 5, 6</xref>
        ], overlook the complex organization of
terms in specialized languages assuming that the terms
contained in a corpus belong to a single domain. At best,
they draw a diference between
domain terms (that
belong to the investigated domain) and out-of-domain terms
(that belong to diferent domains). Unfortunately, this
assumption is too simplistic since every specialized corpus
contains terms from diferent subject fields. Moreover,
in the interest of reusability, researchers who use
terminology corpora in their work must be able to define the
subject fields of interest according to their needs.
      </p>
      <p>
        Furthermore, ATE corpora do not adhere to standard
formats used to encode terminological data like TermBase
eXchange, an ISO standard [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and OntoLex-Lemon, a
W3C standard. This lack of standardization poses
interoperability issues and hinders the evaluation of ATE
tools. The adoption of a standard format will provide at
      </p>
      <p>CEUR</p>
      <p>Workshop Proceedings (CEUR-WS.org)
nEvelop-O
(D. Vellutino)
CEUR
htp:/ceur-ws.org
ISN1613-073
least three main benefits:
• It will grant the interoperability of termbases.</p>
      <p>Therefore if a term is already present in an
existing termbase, it could be imported.
• It will grant the interoperability of corpora,
meaning that multiple corpora could be combined to
cover diferent languages and subject fields.
• It will ease the efort made to evaluate ATE tools
from both sides developers and users.
available on GitHub3
In this paper, we propose a custom form design of
multilevel annotation to formalize ATE corpora in RDF format
by using the OntoLex-Lemon1 and the NLP Interchange
Format - NIF2 ontologies to represent termbases and
corpora, respectively. Moreover, we develop a small
annotated corpus to provide a proof-of-concept. The corpus
and the code employed in its formalization are publicly</p>
      <p>The remainder of this paper is organized as follows.
Section 2 lays out the main feature of terms. Section 3
gives an overview of the main ATE corpora. Section 4
illustrates the proposed formalization schema. In Section
5 we describe the corpus annotation experiment. Finally,</p>
      <sec id="sec-2-1">
        <title>Section 6 provides conclusions.</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Features of terms</title>
      <p>
        According to ISO, a term is a ”designation that represents
a general concept by linguistic means” [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Therefore,
1https://www.w3.org/community/ontolex/wiki/Final_Model_
Specification
      </p>
      <sec id="sec-3-1">
        <title>2https://nif.readthedocs.io/en/latest/</title>
      </sec>
      <sec id="sec-3-2">
        <title>3https://github.com/nicolaCirillo/lod4term</title>
        <p>corpus</p>
        <p>GENIA
ACL-RD TEC</p>
        <p>ACTER
discontinuous terms
yes
no
yes
nested terms
yes
no
only in the termbase
corpus format</p>
        <p>XML
XML; vert</p>
        <p>TSV
termbase format
none
TSV
TSV
terms have linguistic and conceptual features and a sound
annotation schema must account for both. The most
relevant are illustrated above.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Related Work</title>
      <p>With regards to the format of ATE benchmark corpora,
there is no agreed standard. The most popular corpora
are encoded in diferent formats as summarized in Table
1.</p>
      <p>Nested terms A term is nested when it is contained
into another (longer) term. For example, the term
competent authority of dispatch contains both the terms
competent authority and dispatch, joined by the preposition
of.</p>
      <p>
        The GENIA corpus [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is composed of 2000 English
abstracts taken from the MEDLINE database. It is focused
on the biology domain, specifically on transcription
factors in human blood cells. The corpus is encoded in XML
Discontinuous terms A term is discontinuous when with each occurrence of a term being enclosed in the
there is unrelated linguistic material between its words. &lt;term&gt; tag. Discontinuous and nested terms are allowed.
From the conceptual perspective, each term is an instance
Sometimes, discontinuous terms are also nested. For
of a class defined in the GENIA ontology (e.g. the term
example, the term prevention of pollution is discontinuous
ifbroblastic tumour is an instance of the Tissue class).
when it appears inside the term integrated prevention and Being constituted of 300 abstracts from the ACL
Ancontrol of pollution. thology Reference Corpus, the ACL RD-TEC corpus [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
has been developed with the intent of providing a term
Term variants A term variant is a term that expresses extraction corpus on which computational linguists are
the same concept as other terms. For example, the terms experts themselves. It is available in XML and in a
vertiair pollution and atmospheric pollution are term variants cal format (i.e. one token per line). Discontinuous and
because they both refer to the ”contamination of the in- nested terms are not allowed. From the conceptual
perdoor or outdoor environment by any chemical, physical spective terms are categorized following the guidelines
or biological agent that modifies the natural characteris- (e.g. technology and method, tool and library, language
tics of the atmosphere”.4 Acronym and abbreviations resource, etc.).
are specific kinds of term variants. Resolving abbrevia- The ACTER corpus [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is composed of multiple
subtions is one of the goals of the Simple Text Track at CLEF corpora covering four diferent subject fields and three
2023.5 languages. ACTER is specifically made to test ATE tools
on diferent topics and languages while retaining a
conTerminology layer A terminology layer is a set of sistent annotation and, thus, comparable results. It is
terms belonging to a given subject field. For example, the available in TSV (one token per line) with IOB (Inside,
terminology layer of waste management comprises terms Outside, Beginning) or IO (Inside, Outside) tags. The list
such as incineration plant, separate collection, and landfill . of terms found in the corpus is also made available.
DisSome ATE techniques focus on isolating terminology continuous and nested terms are allowed but the latter
layers [
        <xref ref-type="bibr" rid="ref5 ref6 ref9">5, 6, 9</xref>
        ] are not represented in the IOB and IO formats. From the
conceptual perspective terms are classified according to
Translation equivalent A translation equivalent is a a domain-independent annotation schema composed of
term of a natural language that denotes the same con- four labels: specific term , common term, out-of-domain
cept as another term of another natural language. For term, and not term. Moreover, it distinguishes terms from
example, the term autorità competente di spedizione is Named Entities.
the Italian equivalent of the English term competent
authority of dispatch. Finding translation equivalents from
comparable corpora is an ATE subtask [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <sec id="sec-4-1">
        <title>4https://iate.europa.eu/entry/result/3567909/en 5http://simpletext-project.com/2023/clef/tasks</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Proposed Format</title>
      <p>the Dublin Core ontology. Moreover, to grant
interoperability, we propose to employ DBPedia categories [17] to
An ATE corpus has two main components: a termbase represent subject fields. In this way, a term belonging to
and the actual corpus. The termbase contains the list a given subject field (e.g. waste management),
automatiof unique terms (types) that appear in the corpus and cally belongs also to the subject fields it descends from
provides information for each of them. Conversely, the (e.g. waste; sustainability and environmental
managecorpus contains contextualized instances (tokens) of the ment; economy and the environment), in a multi-level
terms in the termbase. We propose to formalize both re- fashion.
sources using formats based on RDF/OWL on account of
their interoperability. Besides, linked data formats have 4.2. Corpus Representation
already improved the benchmarking of Named Entity
Recognition [13].</p>
      <p>To formalize the corpus, we propose to use the NIF and
the POWLA8 ontologies. NIF is based on RDF/OWL
4.1. Termbase Representation and has been developed to achieve interoperability
between NLP tools, language resources, and annotations.</p>
      <p>We propose to use the OntoLex-Lemon ontology to en- It provides multiple benefits. First of all, it does not rely
code the termbase, for various reasons. First of all, strictly on tokenization like TSV formats (see Section 3).
OntoLex-Lemon is already a standard among terminol- It provides support for terminology annotation and has
ogists [14, 15, 16]. In addition, it can represent many already been used for this purpose within the FREME
linguistic and conceptual information that are of interest project [18]. The only drawback of NIF is that it
canto ATE and its subtasks. not represent discontinuous terms. To this end, we use</p>
      <p>In OntoLex-Lemon, there are three main entities: en- POWLA nodes to join NIF strings, as suggested in [19].
tries, concepts, and senses. Entries are instances of the Then, we link each POWLA node to the corresponding
LexicalEntry class. They are linguistic units with one LexicalSense in the termbase to produce unambiguous
or more forms. For example, the term heap is an entry annotations (see Appendix A).
with a singular form (heap) and a plural form (heaps).</p>
      <p>Concepts are instances of the LexicalConcept class.</p>
      <p>They represent units of thought. For example, the con- 5. Example Corpus
cept heap is defined as ”engineered facility for the
deposit of solid waste on the surface”.6 Finally, senses are In order to test our approach and provide a
proof-ofinstances of the LexicalSense class. They are entry- concept, we run an annotation experiment on a
Euroconcept pairs. For example, the entry heap has mul- pean directive, namely the Italian version of the Directive
tiple senses one of which couples this entry with the 2006/21/EC of the European Parliament and of the Council
concept defined above. Entries are related to senses of 15 March 2006 on the management of waste from
extracthrough the sense property and concepts through the tive industries and amending Directive 2004/35/EC (26,882
lexicalicalizedSense property. Moreover, entries can tokens).
also be directly linked to concepts through the evokes
property. 5.1. Annotation</p>
      <p>Furthermore, Ontolex-Lemon allows the
representation of many linguistic and conceptual features that are
of interest to ATE and its subtasks (see Section 2). Term
variants are easy to identify because they are entries
referring to the same concept. However, OntoLex-Lemon
also allows to directly link entries and senses and
speciifes their relation via the LexInfo 7 ontology. Namely, the
synonym property of LexInfo links two senses with the
same meaning, abbreviationFor links an abbreviation
to its full form, and translation links two terms that
are translations of each other. OntoLex-Lemon can also
represent nested terms by means of the subterm property
of its decomposition module. Lastly, terminology layers
can be handled by assigning senses and concepts to the
respective subject field through the subject property of
Two non-expert annotators carried out the annotation.</p>
      <p>They were instructed to identify terms in the corpus and
classify them according to the subject field (i.e. law, EU
law, waste management, waste management law,
environment, other ). Particular attention has been paid to
the identification of nested and discontinuous terms (see
Section 2). After the annotation phase, we asked
annotators to revise the list of unique terms they found (i.e.
the termbase) to delete incorrect ones and revise nested
terms. Finally, we kept in the corpus only the
annotations of terms that were in the revised termbases and
standardized the annotation of nested terms. Namely,
we removed their manual annotations and automatically
tagged them according to the subterms provided in the
revised termbase, thus ensuring consistency.
6https://iate.europa.eu/entry/result/3504812/en
7https://lexinfo.net/</p>
      <sec id="sec-5-1">
        <title>8https://github.com/acoli-repo/powla</title>
        <p>before revision
after revision
subject fields (Fleiss’ k)</p>
        <p>
          To estimate the inter-annotator agreement on term
identification, we computed the F-score measure, similar
to [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], for both the corpus and the termbase, before and
after the revision process (see Table 2). Moreover, to
estimate the agreement on subject fields, we computed
the Fleiss’ k only on terms identified by both annotators.
        </p>
        <p>Inter-annotator agreement scores confirm the benefits
of the revision process. Even though the agreement on
the termbase shows only a little improvement after the
revision (+0.034), the efect on the corpus is much more
relevant (+0.149) as a result of the standardization of
nested terms.</p>
        <p>In the final dataset, we joined the annotations of both
annotators and linked the resulting termbase to IATE9 by
associating each concept with the respective IATE entry
when it exists.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions and Future Work</title>
      <p>The lack of standardization in the representation of ATE
corpora constitutes a bottleneck for the evaluation of ATE
tools. To address this issue, we proposed an RDF-based
formalization that employs OntoLex-Lemon to represent
termbases and NIF to represent corpora. We showed that
these formats are able to represent the wide range of
linguistic and conceptual phenomena that characterize
terminology. In addition, we developed a small corpus
about waste management legislation in order to provide
an example of the proposed formalization.</p>
      <p>In future, we plan to convert the major ATE corpora
into the proposed format, to further improve ATE
standardization. Moreover, we intend to increase the size and
quality of the small ATE corpus we developed.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>The authors contributed to this paper as follows.
sections 1, 2, and 6 are attributed to Daniela Vellutino while
sections 3, 4, and 5 are attributed to Nicola Cirillo.</p>
      <sec id="sec-7-1">
        <title>9https://iate.europa.eu/home</title>
        <p>tec 2.0: A language resource for evaluating term
extraction and entity recognition methods,
European Language Resources Association (ELRA),
2016, pp. 1862–1868. URL: https://aclanthology.org/</p>
        <p>L16-1294.
[13] M. Röder, R. Usbeck, A.-C. N. Ngomo, Gerbil –
benchmarking named entity recognition and
linking consistently, Semantic Web 9 (2018) 605–625.</p>
        <p>doi:10.3233/SW-170286.
[14] P. Martín-Chozas, T. Declerck, Representing
multilingual terminologies with ontolex-lemon, in:
Proceedings of the 1st International Conference on
Multilingual Digital Terminology Today (MDTT
2022), CEUR Workshop Proceedings, 2022.
[15] S. Piccini, F. Vezzani, A. Bellandi, Tbx and ‘lemon’:</p>
        <p>What perspectives in terminology?, Digital
Scholarship in the Humanities 38 (2023) i61–i72. URL:
https://doi.org/10.1093/llc/fqad025. doi:10.1093/
llc/fqad025.
[16] M. Fiorelli, A. Stellato, T. Lorenzetti, A. Turbati,</p>
        <p>P. Schmitz, E. Francesconi, N. Hajlaoui, B. Batouche,
Towards ontolex-lemon editing in vocbench 3,
AIDAinformazioni, Rivista di scienze
dell’informazione (2018).
[17] J. Lehmann, R. Isele, M. Jakob, A. Jentzsch, D.
Kontokostas, P. N. Mendes, S. Hellmann, M. Morsey,
P. V. Kleef, S. Auer, C. Bizer, Dbpedia - a
largescale, multilingual knowledge base extracted from
wikipedia, Semantic Web 6 (2015) 167–195. doi:10.</p>
        <p>3233/SW-140134.
[18] M. Dojchinovski, F. Sasaki, T. Gornostaja, S.
Hellmann, E. Mannens, F. Salliau, M. Osella, P. Ritchie,
G. Stoitsis, K. Koidl, M. Ackermann, N. Chakraborty,
Freme: Multilingual semantic enrichment with
linked data and language technologies, volume 8,
European Language Resources Association (ELRA),
2016, pp. 4180–4183. URL: https://aclanthology.org/</p>
        <p>L16-1660.
[19] P. Cimiano, C. Chiarcos, J. P. McCrae, J. Gracia,</p>
        <p>Linguistic Linked Data Representation,
Generation and Applications, 1 ed., Springer Cham, 2020.</p>
        <p>doi:10.1007/978-3-030-30225-2.</p>
        <p>A. Example of RDF files
@ p r e f i x dbc : &lt; h t t p s : / / dbpedia . org / page / Category : &gt; .
@ p r e f i x d c t : &lt; h t t p : / / p u r l . org / dc / terms / &gt; .
@ p r e f i x decomp : &lt; h t t p : / / www. w3 . org / ns / lemon / decomp# &gt; .
@ p r e f i x l e x i n f o : &lt; h t t p : / / www. l e x i n f o . n e t / o n t o l o g y / 2 . 0 / l e x i n f o # &gt; .
@ p r e f i x o n t o l e x : &lt; h t t p : / / www. w3 . org / ns / lemon / o n t o l e x # &gt; .
@ p r e f i x termbase : &lt; h t t p : / / example . com / termbase / &gt; .
termbase : e n t r y _ r i f i u t o _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e a o n t o l e x : M u l t i w o r d E x p r e s s i o n ;
decomp : subterm termbase : e n t r y _ i n d u s t r i a _ e s t r a t t i v a ,</p>
        <p>termbase : e n t r y _ r i f i u t o ;
o n t o l e x : c a n o n i c a l F o r m termbase : f o r m _ r i f i u t o _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e ;
o n t o l e x : otherForm termbase : f o r m _ r i f i u t i _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e ;
o n t o l e x : s e n s e termbase : r i f i u t o _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e _ s e n s e 1
o n t o l e x : e v o k e s termbase : c o n c e p t _ r i f i u t o _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e .
termbase : f o r m _ r i f i u t o _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e a o n t o l e x : Form ;
l e x i n f o : gender l e x i n f o : m a s c u l i n e ;
l e x i n f o : number l e x i n f o : s i n g u l a r ;
o n t o l e x : w r i t t e n R e p ” r i f i u t o d e l l e i n d u s t r i e e s t r a t t i v e ” @it .
termbase : f o r m _ r i f i u t i _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e a o n t o l e x : Form ;
l e x i n f o : gender l e x i n f o : m a s c u l i n e ;
l e x i n f o : number l e x i n f o : p l u r a l ;
o n t o l e x : w r i t t e n R e p ” r i f i u t i d e l l e i n d u s t r i e e s t r a t t i v e ” @it .
termbase : r i f i u t o _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e _ s e n s e 1 a o n t o l e x : L e x i c a l S e n s e ;
o n t o l e x : i s L e x i c a l i z e d S e n s e O f termbase : c o n c e p t _ r i f i u t o _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e ;
o n t o l e x : i s S e n s e O f termbase : e n t r y _ r i f i u t o _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e ;
d c t : s u b j e c t dbc : Waste_management ;
l e x i n f o : synonym termbase : r i f i u t o _ d e r i v a n t e _ d a l l e _ i n d u s t r i e _ e s t r a t t i v e _ s e n s e 1 ,
termbase : r i f i u t o _ g e n e r a t o _ d a l l e _ i n d u s t r i e _ e s t r a t t i v e _ s e n s e 1 ,
termbase : r i f i u t o _ p r o d o t t o _ d a l l e _ i n d u s t r i e _ e s t r a t t i v e _ s e n s e 1 ,
termbase : r i f i u t o _ p r o v e n i e n t e _ d a l l e _ i n d u s t r i e _ e s t r a t t i v e _ s e n s e 1 .
_ : terms a powla : Root .
_ : term1 a powla : Node ;
powla : s t r i n g ” r i f i u t i ” ;
powla : h a s P a r e n t _ : terms ;
i t s r d f : term ” yes ” ;
i t s r d f : t e r m I n f o R e f termbase : r i f i u t o _ s e n s e 1 .
_ : term2 a powla : Node ;
powla : s t r i n g ” i n d u s t r i e e s t r a t t i v e ” ;
powla : h a s P a r e n t _ : terms ;
i t s r d f : term ” yes ” ;
i t s r d f : t e r m I n f o R e f termbase : i n d u s t r i a _ e s t r a t t i v a _ s e n s e 1 .
_ : term3 a powla : Node ;
powla : s t r i n g ” r i f i u t i d e l l e i n d u s t r i e e s t r a t t i v e ” ;
powla : h a s P a r e n t _ : terms ;
i t s r d f : term ” yes ” ;
i t s r d f : t e r m I n f o R e f termbase : r i f i u t o _ d e l l e _ i n d u s t r i e _ e s t r a t t i v e _ s e n s e 1 .
c o r p u s : doc1 a n i f : Context ,
n i f : O f f s e t B a s e d S t r i n g ;
n i f : i s S t r i n g ” D i r e t t i v a 2 0 0 6 / 2 1 / CE d e l Parlamento europeo e d e l C o n s i g l i o . . . ” .
c o r p u s : doc1 # o f f s e t _ 1 0 5 _ 1 1 2 a n i f : O f f s e t B a s e d S t r i n g ,
n i f : Word ,
powla : Node ;
n i f : anchorOf ” r i f i u t i ” ;
n i f : r e f e r e n c e C o n t e x t c o r p u s : doc1 ;
powla : h a s P a r e n t _ : term1 ,</p>
        <p>_ : term3 .
c o r p u s : doc1 # o f f s e t _ 1 1 3 _ 1 1 8 a n i f : O f f s e t B a s e d S t r i n g ,
n i f : Word ,
powla : Node ;
n i f : anchorOf ” d e l l e ” ;
n i f : r e f e r e n c e C o n t e x t c o r p u s : doc1 ;
powla : h a s P a r e n t _ : term3 ;
powla : n e x t c o r p u s : doc1 # o f f s e t _ 1 1 3 _ 1 1 8 .
c o r p u s : doc1 # o f f s e t _ 1 1 3 _ 1 1 8 a n i f : O f f s e t B a s e d S t r i n g ,
n i f : Word ,
powla : Node ;
n i f : anchorOf ” i n d u s t r i e ” ;
n i f : r e f e r e n c e C o n t e x t c o r p u s : doc1 ;
powla : h a s P a r e n t _ : term2 ,</p>
        <p>_ : term3 ;
powla : n e x t c o r p u s : doc1 # o f f s e t _ 1 2 9 _ 1 3 9 .
c o r p u s : doc1 # o f f s e t _ 1 2 9 _ 1 3 9 a n i f : O f f s e t B a s e d S t r i n g ,
n i f : Word ,
powla : Node ;
n i f : anchorOf ” e s t r a t t i v e ” ;
n i f : r e f e r e n c e C o n t e x t c o r p u s : doc1 ;
powla : h a s P a r e n t _ : term2 ,</p>
        <p>_ : term3 .</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Vellutino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Maslias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rossi</surname>
          </string-name>
          ,
          <article-title>Verso l'interoperabilità semantica di iate. studio preliminare per il dominio “gestione dei rifiuti urbani”, Terminologie specialistiche e difusione dei saperi (</article-title>
          <year>2016</year>
          )
          <fpage>1</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Vellutino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Maslias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rossi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mangiacapre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Montoro</surname>
          </string-name>
          ,
          <article-title>Verso l'interoperabilità semantica di iate. studio preliminare sul lessico dei fondi strutturali e d'investimento europei, Diversité et Identité Culturelle en Europe/Diversitate si Identitate Culturala in Europa (</article-title>
          <year>2016</year>
          )
          <fpage>1</fpage>
          -
          <lpage>254</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Vellutino</surname>
          </string-name>
          ,
          <string-name>
            <surname>L'</surname>
          </string-name>
          <article-title>italiano istituzionale per la comunicazione pubblica</article-title>
          ,
          <source>il Mulino</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lenci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Montemagni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Pirrelli</surname>
          </string-name>
          , G. Venturi,
          <article-title>Ontology learning from italian legal texts</article-title>
          , in: Law,
          <article-title>Ontologies and the Semantic Web</article-title>
          , IOS Press,
          <year>2009</year>
          , pp.
          <fpage>75</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Bonin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Montemagni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Venturi</surname>
          </string-name>
          ,
          <article-title>A contrastive approach to multi-word extraction from domain-specific corpora</article-title>
          ,
          <source>in: Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)</source>
          ,
          <source>European Language Resources Association (ELRA)</source>
          , Valletta, Malta,
          <year>2010</year>
          . URL: http://www.lrec-conf. org/proceedings/lrec2010/pdf/553_Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Drouin</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-C. L'Homme</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Robichaud</surname>
          </string-name>
          ,
          <article-title>Lexical profiling of environmental corpora</article-title>
          ,
          <source>in: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ),
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>ISO</surname>
          </string-name>
          , ISO
          <volume>30042</volume>
          :
          <fpage>2019</fpage>
          -
          <article-title>Management of terminology resources - TermBase eXchange (TBX</article-title>
          ),
          <source>International Organization for Standardization</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>ISO</surname>
          </string-name>
          , ISO
          <volume>1087</volume>
          :
          <fpage>2019</fpage>
          -
          <article-title>Terminology work and terminology science</article-title>
          ,
          <source>International Organization for Standardization</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N.</given-names>
            <surname>Cirillo</surname>
          </string-name>
          ,
          <article-title>Isolating terminology layers in complex linguistic environments: a study about waste management (short paper)</article-title>
          ,
          <source>in: Proceedings of the 2nd International Conference on Multilingual Digital Terminology Today (MDTT</source>
          <year>2023</year>
          ), volume
          <volume>3427</volume>
          , CEUR Workshop Proceedings,
          <year>2023</year>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3427</volume>
          /short3.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rigouts Terryn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Hoste</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Lefever,</surname>
          </string-name>
          <article-title>In no uncertain terms: a dataset for monolingual and multilingual automatic term extraction from comparable corpora</article-title>
          ,
          <source>Language Resources and Evaluation</source>
          <volume>54</volume>
          (
          <year>2020</year>
          )
          <fpage>385</fpage>
          -
          <lpage>418</lpage>
          . doi:
          <volume>10</volume>
          .1007/ s10579- 019- 09453- 9.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ohta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tateisi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tsujii</surname>
          </string-name>
          ,
          <article-title>Genia corpus - a semantically annotated corpus for biotextmining</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>19</volume>
          (
          <year>2003</year>
          ). doi:
          <volume>10</volume>
          .1093/ bioinformatics/btg1023.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B.</given-names>
            <surname>Qasemizadeh</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.-K. Schumann</surname>
          </string-name>
          , The acl rd-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>