<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Computational Lexicon for Italian: building the morphological layer by harmonizing and merging existing resources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Flavia Sciolette</string-name>
          <email>flavia.sciolette@ilc.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simone Marchi</string-name>
          <email>simone.marchi@ilc.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emiliano Giovannetti</string-name>
          <email>emiliano.giovannetti@ilc.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Istituto di Linguistica Computazionale “Antonio Zampolli” (CNR-ILC), Area della Ricerca del CNR di Pisa</institution>
          ,
          <addr-line>Via G. Moruzzi, 1, 56124 Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>computational lexicon, lexical resources, morphology, morphological harmonization</p>
      <p>© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License will populate the morphological layer of the
computatasks can take advantage of lexical resources, for exam- tion of a new computational lexicon for the Italian
lanCLiC-it 2023: 9th Italian Conference on Computational Linguistics, sion of the morphological layer, carried out through the</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        A significant number of digital lexical resources are
available for many languages. In CLARIN Virtual Language
Observatory (VLO)1, a search for “lexicalResource” of
Italian provides 52 results. Two resources appear in
several versions and updates: Parole-Simple-Clips (PSC)2
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], a multilayered lexicon, and ItalWordNet3. The most
part of the results includes monolingual and multilingual
domain terminologies. Amongst the notable resources
are worth mentioning Italian Function Words (IFWs)4
and Italian Content Words (ICWs)5, two lists in JSON
Lines format developed for supporting POS tagging and
syntactic parsing of Italian. In fact, a number of NLP
ple sentiment analysis [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] but also “semantic role labeling,
verb sense disambiguation and ontology mapping” [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>However, ICWs includes hundreds of thousands of
forms generated automatically and not manually revised,
∗Corresponding author.
†These authors contributed equally.
nEvelop-O
LGOBE</p>
      <p>0000-0002-7998-9768 (F. Sciolette); 0000-0003-4320-6466
(S. Marchi); 0000-0002-0716-1160 (E. Giovannetti)
CEUR
Workshop
Proce dings
htp:/ceur-ws.org
ISN1613-073</p>
      <p>CEUR</p>
      <p>Workshop Proceedings (CEUR-WS.org)
3https://dspace-clarin-it.ilc.cnr.it/repository/xmlui/handle/20.500.
11752/ILC-977
4https://lindat.mff.cuni.cz/repository/xmlui/handle/11372/LRT-2893
5https://lindat.mff.cuni.cz/repository/xmlui/handle/11372/LRT-2894
CEUR</p>
      <p>ceur-ws.org
which, despite being morphologically correct, have very
low, if not zero, usage frequency.</p>
      <p>
        Although not listed in Clarin’s VLO, we also mention
SIMPLELex-it, since built similarly to our lexicon by
combining various existing resources [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>Even in lexical resources which have been manually
developed and revised, however, the linguistic coverage
of entries can pose problems, both in terms of lexical
coverage and content of entries. Hence, integrating
information from diferent sources, as we did in this work,
can be efective in filling the gaps, though it can present
several challenges in terms of harmonization of distinct
formats and models.</p>
      <sec id="sec-2-1">
        <title>We here describe the first steps towards the construc</title>
        <p>
          guage that we called CompL-it. We started from the
enrichment of an existing resource, LexicO6 [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], a
computational lexicon which, in turn, was derived from the
already cited PSC lexicon.
        </p>
        <p>
          In particular, this first phase was focused on the
expanrated values (or TSV)8.
integration of two other resources: a list of lemmatized
forms generated by the morphological analyzer MAGIC7
[
          <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
          ], and a set of Italian treebanks. The obtained
resource, constituted of nearly 800 thousand forms, was
made available in a CoNLL-like format as a tabular
sepa
        </p>
        <p>
          This core of forms, lemmas, and morphological traits
tional lexicon CompL-it under construction, which will
later be released in the form of Linguistic Linked Open
6https://dspace-clarin-it.ilc.cnr.it/repository/xmlui/handle/20.500.
7https://dspace-clarin-it.ilc.cnr.it/repository/xmlui/handle/20.500.
11752/ILC-1002
8https://github.com/klab-ilc-cnr/CompL-it_morphological_layer
Data (LLOD) (see Section 5). For this reason, it was cho- the following treebanks: i) ISDT [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]; ii) VIT - Venice
Italsen not to update the relational database of LexicO, but ian Treebank [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]; iii) TUT [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]; iv) ParlaMint-It, based
to use the CoNLL-like format as a temporary data repre- on ParlaMint-It corpus10 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
sentation format.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. The building of the</title>
      <p>morphological layer</p>
    </sec>
    <sec id="sec-4">
      <title>2. The sources</title>
      <p>The sources we considered for building the
morphological layer of CompL-it difer from each other for model, 3.1. Harmonization
vocabularies, and aims; for this reason, it was first neces- Morphological data are represented in the considered
sary to carry out a harmonization process to make the resources in diferent ways. The vocabulary labels of each
resources comparable to each other, as shown in Sec- resource was mapped into LexInfo11, the data category
tion 3. ontology for OntoLex-Lemon model12, de facto standard</p>
      <p>
        Regarding the choice of sources, we opted to include for representing lexical resources in the Semantic Web.
only resources for which manual revision was docu- In the case of M-GLF and LexicO, it involved the direct
mented. In this sense, we chose not to delve into the conversion of their custom tagsets - specific for Italian
data at this initial stage. Corrective actions, aimed at into the nomenclature of LexInfo.
preventing the generation and propagation of errors, fo- The LexicO and M-GLF vocabularies also follow a
difcused on issues that could be resolved through automatic ferent theoretical approach compared to the UD used
processes and were independent of data evaluation, such in treebanks. In the first two cases, the vocabulary is
as redundancy or the comparison of entries to assess designed for lexical resources, and the POS tags are
fineinformation richness (see 3.3). grained, often finding a direct counterpart in LexInfo, as
LexInfo serves as an ontology for this type of resource. In
2.1. LexicO the case of UD, used for corpus annotation, word
descriptions are assigned a “universal” POS tag, further specified
LexicO is available on CLARIN as a relational database, by features defined in the Universal Features vocabulary.
and shares the same linguistic model of PSC, which is The cases addressed in the mapping can be classified
based on the theory of Generative Lexicon by James into three types: i) perfect correspondence; in these cases,
Pustejovsky [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. LexicO contains four layers of linguistic the value was directly converted into the LexInfo
vocabinformation: morphology, syntax, semantics, and phonol- ulary; ii) correspondence of POS in combination with
anogy. other value; in these cases, the mapping associated a
LexInfo label with a combination of POS and a morphological
2.2. MAGIC feature, as seen in the case of demonstratives in UD; iii)
correspondence not present in LexInfo; in this case, a new
class was formalized and linked to OLIA13. The tables for
mapping has been made available on GitHub14.
      </p>
      <p>MAGIC is a morphological analyzer which includes three
modules: a lexicon compiler for Italian, the
morphological analyzer itself, and the morphological generator.</p>
      <p>With an ad hoc script, we extracted all forms generated
by the morphological analyzer. The generated output
consists of a series of linguistic objects called “words”,
for each of which lemmas, morphosyntactic types, and
features are specified. This resource was made available
on CLARIN as “MAGIC - Generated Lemmatized Forms”
(M-GLF).</p>
      <sec id="sec-4-1">
        <title>2.3. Universal Dependencies treebanks</title>
        <p>Treebanks are collected and listed in the Universal
Dependencies (UD) repository9. We excluded non-manually
revised treebanks from the selection. Additionally, we
excluded treebanks aimed at representing specific case
studies that could introduce sparsely attested forms into
the lexicon or introduce excessive “noise”. We considered</p>
        <sec id="sec-4-1-1">
          <title>9https://github.com/UniversalDependencies</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Conversion to CoNLL-like format</title>
        <p>Once the vocabularies were harmonized, each resource
was converted into a file in CoNLL-like format.</p>
        <p>The choice of this format is primarily due to two
reasons: i) UD treebanks are already in tabular format
(CoNLL is a TSV); ii) the information from LexicO and
M-GLF does not have a specific output format (the former
is stored in a relational database, while the latter is in a
textual format that does not adhere to any standard) and
can be easily transformed into a TSV format.
10https://www.clarin.eu/parlamint
11https://github.com/ontolex/lexinfo
12https://www.w3.org/2016/05/ontolex/
13https://github.com/acoli-repo/olia
14https://github.com/klab-ilc-cnr/Tables-for-mapping-of-Italian-Lexicon-ComplIt</p>
        <p>In a first phase, from each of the aforementioned re- Similarly, Figure 2 shows the distribution of lemmas
sources, a list of forms with lemmatization and morpho- per resource.
logical traits was extracted. Subsequently, the obtained
lists were converted in distinct CoNLL-like files, using
ad hoc developed Perl scripts. In this phase, the tagsets
were also converted according to the mappings in LexInfo
mentioned in the previous section.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Merging</title>
        <p>The merging process of the three resources represented
by the CoNLL-like files was divided into two phases: i)
initially, two resources were compared and combined
in a partial merge; ii) subsequently, the third resource
was added to the comparison to obtain the final output.
The algorithm compared two entries at a time. If two
entries were equal in terms of form, POS, and lemma,
then their morphological features were compared. If the
features of the first entry constituted a subset of those
belonging to the second entry, the latter was considered
for the final output, being richer in linguistic information.
The algorithm, developed in Java, was made available on
GitHub15.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Evaluation</title>
      <p>The resulting output consists of 790,758 forms associated
with 102,000 lemmas and the relative traits.</p>
      <p>In CompL-it, each form is associated with the following
data: lemma, part of speech (POS), and morphological
features specific to the considered POS. Table 1 shows an
example of form for the lemma “gatto” (cat).</p>
      <p>Despite the evident larger size of M-GLF compared to
LexicO, it is important to specify that the choice to use
this latter as the reference base was mainly qualitative,
in particular for its multilevel structure16, which will be
exploited for enriching CompL-it in subsequent works
(Section 5).</p>
      <p>In Table 2, the number of forms per POS in LexicO is
compared to the final CompL-it resource, along with the
respective percentage increase.</p>
      <p>It is worth noting the significant increase in values,
particularly for adjectives and adverbs, which have a
lower coverage in LexicO17.</p>
      <p>To conclude this section, we provide in Table 3 a
quantitative comparison between CompL-it and some of the
lexical resources mentioned in the introduction,
specifically PSC, ItalWordNet, and SIMPLELex-IT. We excluded
ICWs from the comparison due to the mentioned issue
of overgenerating forms (which doesn’t align well with
the need to represent lexically precise data) and IFWs, as
it contains many multiword entries that we have chosen
to exclude from our lexicon at this time.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusions and future works</title>
      <p>In this article, we documented a first step towards
building a new computational lexicon of Italian. A set of
approximately 800 thousand lemmatized forms with
morphological features was created, through the integration
of existing resources. In the next phases, this lexical core
will be converted as a LLOD based on the OntoLex-lemon
model, making the resulting lexicon more easily
shareable, interoperable, and compliant with Semantic Web
standards. Additionally, new linguistic layers will be
added, starting from semantics, by using the information
already available in LexicO and by integrating data from
WordNets for Italian.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was conducted in the context of the TALMUD
project and the scientific cooperation between S.c.a r.l.
PTTB and CNR-ILC.</p>
      <sec id="sec-7-1">
        <title>Making Italian Parliamentary Records Machine</title>
        <p>Actionable: the Construction of the ParlaMint-IT
Corpus, in: Proceedings of the Workshop
ParlaCLARIN III within the 13th Language Resources
and Evaluation Conference, 2022, pp. 117–124. URL:
https://aclanthology.org/2022.parlaclarin-1.17/.
[13] N. Ruimy, M. Monachini, R. Distante, E. Guazzini,
S. Molino, M. Ulivieri, N. Calzolari, A. Zampolli,
Clips, a multi-level Italian computational lexicon:
A glimpse to data, in: Proceedings of the Third
International Conference on Language Resources
and Evaluation (LREC02), 2002.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lenci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Busa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Calzolari</surname>
          </string-name>
          , E. Gola,
          <string-name>
            <given-names>M.</given-names>
            <surname>Monachini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ogonowski</surname>
          </string-name>
          , I. Peters,
          <string-name>
            <given-names>W.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ruimy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zampolli</surname>
          </string-name>
          , SIMPLE: A
          <article-title>General Framework for the Development of Multilingual Lexicons</article-title>
          ,
          <source>International Journal of Lexicography</source>
          <volume>13</volume>
          (
          <year>2000</year>
          )
          <fpage>249</fpage>
          -
          <lpage>263</lpage>
          . doi:
          <volume>10</volume>
          .1093/ijl/13.4. 249.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T. N.</given-names>
            <surname>Prakash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Aloysius</surname>
          </string-name>
          ,
          <source>Textual Sentiment Analysis Using Lexicon Based Approaches, Annals of the Romanian Society for Cell Biology</source>
          (
          <year>2021</year>
          )
          <fpage>9878</fpage>
          -
          <lpage>85</lpage>
          . URL: http://annalsofrscb.ro/index. php/journal/article/view/3734.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Brown</surname>
          </string-name>
          , J. Windisch,
          <string-name>
            <given-names>G.</given-names>
            <surname>Kazeminejad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zaenen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pustejovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmer</surname>
          </string-name>
          ,
          <article-title>Semantic Representations for NLP Using VerbNet and the Generative Lexicon</article-title>
          ,
          <source>Frontiers in Artificial Intelligence</source>
          <volume>5</volume>
          (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .3389/frai.
          <year>2022</year>
          .
          <volume>821697</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mazzei</surname>
          </string-name>
          ,
          <article-title>Building a computational lexicon by using SQL</article-title>
          ,
          <source>in: Proceedings of the Third Italian Conference on Computational Linguistics CLiCit</source>
          <year>2016</year>
          :
          <article-title>5-6 December 2016</article-title>
          , Napoli,
          <year>2016</year>
          . doi:
          <volume>10</volume>
          . 4000/books.aaccademia.
          <year>1808</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Sciolette</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Giovannetti</surname>
          </string-name>
          , S. Marchi,
          <article-title>LexicO: an Italian Computational Lexicon derived from ParoleSimple-Clips, Umanistica Digitale 7 (</article-title>
          <year>2023</year>
          )
          <fpage>169</fpage>
          -
          <lpage>193</lpage>
          . doi:
          <volume>10</volume>
          .6092/issn.2532-
          <issue>8816</issue>
          /15176.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Battista</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Pirrelli</surname>
          </string-name>
          ,
          <article-title>Una Piattaforma di Morfologia Computazionale per l'Analisi e la Generazione delle Parole Italiane</article-title>
          ,
          <source>Technical Report, ILC-CNR Technical Report</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V.</given-names>
            <surname>Pirrelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Battista</surname>
          </string-name>
          ,
          <article-title>The Paradigmatic Dimension of Stem Allomorphy in Italian Verb Inflection</article-title>
          ,
          <source>Rivista di Linguistica</source>
          <volume>12</volume>
          (
          <year>2000</year>
          )
          <fpage>307</fpage>
          -
          <lpage>379</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pustejovsky</surname>
          </string-name>
          , The Generative Lexicon, MIT Press, Cambridge, MA,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Simi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Montemagni</surname>
          </string-name>
          , Less is More?
          <article-title>Towards a Reduced Inventory of Categories for Training a Parser for the Italian Stanford Dependencies</article-title>
          ,
          <source>in: Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>83</fpage>
          -
          <lpage>90</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Delmonte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bristot</surname>
          </string-name>
          , S. Tonelli, VIT - Venice Italian Treebank:
          <article-title>Syntactic and Quantitative Features</article-title>
          ,
          <source>in: Proceedings of the Sixth International Workshop on Treebanks and Linguistic Theories</source>
          , volume
          <volume>1</volume>
          ,
          <year>2007</year>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , C. Bosco, PartTUT: The Turin University Parallel Treebank, in: Harmonization and
          <article-title>development of resources and tools for Italian Natural Language Processing within the PARLI project</article-title>
          ,
          <source>LNCS</source>
          , Springer Verlag,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Agnoloni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bartolini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Frontini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Montemagni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Marchetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Quochi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ruisi</surname>
          </string-name>
          , G. Venturi,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>