<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multi-Word Expressions in spoken language: PoliSdict</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Università di Salerno</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Salerno</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Network Contacts</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Molfetta</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Consiglio Nazionale delle Ricerche</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>English. The term multiword expressions (MWEs) is referred-to a group of words with a unitary meaning, not inferred from that of the words that compose it, both in current use and in technical-specialized languages. In this paper, we describe PoliSdict an Italian electronic dictionary composed of multi-word expressions (MWEs) automatically extracted from a multimodal corpus grounded on political speech language, currently being developed at the "Maurice Gross" Laboratory of the Department of Political Sciences, Social and Communication of the University of Salerno, thanks to a loan from the company Network Contacts. We introduce the methodology of creation and the first results of a systematic analysis which considered terminological labels, frequency labels, recurring syntactic patterns, further proposing an associated ontology.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. Con il termine polirematica si fa
generalmente riferimento ad un gruppo di parole con
significato unitario, non desumibile da quello
delle parole che lo compongono, sia nell’uso
corrente sia in linguaggi tecnico-specialistici. In
questo contributo viene presentato PoliSdict un
dizionario elettronico in lingua italiana composto
da espressioni polirematiche occorrenti nel
parlato spontaneo estratte a partire da un corpus
multimodale di dominio politico in lingua
italiana in corso di ampliamento presso il Laboratorio
“Maurice Gross” del Dipartimento di Scienze
Politiche, Sociali e della Comunicazione
dell’Università degli Studi di Salerno, grazie a
un finanziamento della società Network Contacts.
Viene presentata la metodologia di creazione ed i
primi risultati di un'analisi sistematica che ha
considerato etichette terminologiche, marche
d'uso e pattern ricorrenti, proponendo infine
un’ontologia associata.
1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        The term multi-word expressions (MWEs)
includes a wide range of constructions such as
noun compounds, adverbials, binomials, verb
particles constructions, collocations, and idioms
        <xref ref-type="bibr" rid="ref35">(Vietri, 2014)</xref>
        .
        <xref ref-type="bibr" rid="ref6">D'Agostino &amp; Elia (1998)</xref>
        consider MWUs part of a continuum in which
combinations can vary from a high degree of
variability of co-occurrence of words
(combinations with free distribution), to the
absence of variability of co-occurrence1. They
identify four different types of combinations of
phrases or sentences, namely (i) with a high
degree of variability of co-occurrence among
words; (ii) with a limited degree of variability of
co-occurrence among words; (iii) with no or
almost no variability of co-occurrence among
words; (iv) with no variability of co-occurrence
among words. The essential role played by
MWEs in Natural Language Processing (NLP)
and linguistic analysis in general has been long
recognised, as confirmed by then numerous
dedicated workshops and special issues of
journals discussing this subject in recent years
        <xref ref-type="bibr" rid="ref16 ref5">(CSL, 2005; JLRE, 2009)</xref>
        , and this appears more
clear if we consider as the detection of MWEs
represents a real issue in several NLP tasks such
as semantic parsing and machine translation
(Fellbaum, 2011). According to
        <xref ref-type="bibr" rid="ref4">Chiari (2012)</xref>
        regarding the Italian language a line of great
1 Concerning compositionality, the study of
        <xref ref-type="bibr" rid="ref23">Nunberg et al.
(1994)</xref>
        is noteworthy. This study undermines the issue of
compositionality, as widely emphasized in
        <xref ref-type="bibr" rid="ref35">Vietri (2014)</xref>
        .
interest is represented by the works of Annibale
Elia and Simonetta Vietri
        <xref ref-type="bibr" rid="ref28 ref33 ref34 ref6">(Elia, D'Agostino et al
1985, Vietri 1986, D'Agostino and Elia 1998,
Vietri 2004)</xref>
        . Finally the discussion concerning
the MWEs in Italian lexicography has been
systematized in the GRADIT
        <xref ref-type="bibr" rid="ref10">(De Mauro 1999)</xref>
        which records 132.000 different MWEs, whose
collection was coordinated by Annibale Elia at
the Department of Communication Sciences of
the University of Salerno. This research is part of
the larger project BIG 4 M.A.S.S. conducted by
the company Network Contacts2 in collaboration
with the Department of Social Politics and
Communication, which received funding to
develop semantic and syntactic modules of
Italian.
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>Related work</title>
      <p>
        In the last twenty years or so MWEs have been
an increasingly important concern for NLP.
MWEs have been studied for decades in
phraseology under the term phraseological unit.
But in the early 1990s, MWEs received
increasing attention in corpus-based
computational linguistics and NLP. Early
influential work on MWEs includes
        <xref ref-type="bibr" rid="ref29">Smadja
(1993)</xref>
        ,
        <xref ref-type="bibr" rid="ref7">Dagan and Church (1994)</xref>
        ,
        <xref ref-type="bibr" rid="ref38">Wu (1997)</xref>
        ,
        <xref ref-type="bibr" rid="ref8">Daille (1995)</xref>
        ,
        <xref ref-type="bibr" rid="ref37">Wermter and Chen (1997)</xref>
        ,
        <xref ref-type="bibr" rid="ref19">McEnery et al. (1997)</xref>
        , and
        <xref ref-type="bibr" rid="ref20">Michiels and Dufour
(1998)</xref>
        . These studies address the automatic
treatment of MWEs and their applications in
practical NLP and information systems. An
important research contribution is the Multiword
Expression Project carried out at Stanford
University, which began in 2001 to investigate
means to encode a variety of MWEs in precision
grammars 3 . Other major work has been
conducted at Lancaster University, which
resulted in a large collection of semantically
annotated English, Finnish and Russian MWE
dictionary resources for a semantic annotation
tool
        <xref ref-type="bibr" rid="ref22 ref26 ref28">(Rayson et al. 2004; Lo¨fberg et al. 2005;
Piao et al. 2005; Mudraya et al. 2006)</xref>
        . Since
then, many advances have been made, either
looking at MWEs in general
        <xref ref-type="bibr" rid="ref36 ref39">(Zhang et al., 2006;
Villavicencio et al., 2007)</xref>
        , or focusing on
2 Network Contacts, is one of the national leader players in
the areas of BPO (business process outsourcing), CRM
(customer relationship management), Digital Interaction and
Call&amp;Contact Center services. Over the years, it has built
numerous partnership with some of the most recognized
national academic players, such as the University of
Salerno, so as to face stimulating research challenges in the fields
of Artificial Intelligence and Natural Language Processing.
3 For more information cfr. http://mwe.stanford.edu
specific MWE types, such as collocations
        <xref ref-type="bibr" rid="ref24">(Pearce, 2002)</xref>
        , phrasal verbs (Baldwin, 2005;
Ramisch et al., 2008) or compound nouns (Keller
et al., 2002). A popular type-independent
alternative to MWE identification is to use
statistical AMs
        <xref ref-type="bibr" rid="ref36 ref39">(Evert and Krenn, 2005; Zhang et
al., 2006; Villavicencio et al., 2007)</xref>
        . Concerned
MWE identification and extraction from
monolingual corpora,
        <xref ref-type="bibr" rid="ref17 ref17 ref18 ref18">Kim and Baldwin (2006)</xref>
        proposed a method for automatically identifying
English verb particle constructions (VPCs),
Pecina (2009) reported an evaluation of a set of
lexical association measures based on the Prague
Dependency Treebank and the Czech National
Corpus,
        <xref ref-type="bibr" rid="ref31">Strik et al. (2010)</xref>
        investigated the
possible ways of automatically identifying Dutch
MWEs in speech corpora. Related to lexical
representation of MWEs in a lexicon and a
syntactic treebank,
        <xref ref-type="bibr" rid="ref12">Gregoire (2010</xref>
        ) discusses the
design and implementation of a Dutch Electronic
Lexicon of Multiword Expressions (DuELME),
which contains over 5,000 Dutch multiword
expressions.
        <xref ref-type="bibr" rid="ref2">Bejcˇek and Stranak (2010</xref>
        ) describe
the annotation of multiword expressions found
within the Prague Dependency Treebank. In
NLP, MWEs in spoken language have been
studied in the field of automatic speech
recognition, generally with the aim of
establishing to what extent modeling such
expressions can help reducing word error rate
        <xref ref-type="bibr" rid="ref30">(Strik and Cucchiarini 1999)</xref>
        . So a review of
related work about MWEs highlights the lack of
electronic dictionaries of Italian MWEs for
spoken language, hence the idea of creating an
ad hoc dictionary starting from a resource of
political domain. That being said, it should be
specified here that this study represents an initial
experiment on a relatively small sample, since a
larger balanced corpus would be necessary for a
broader coverage. Political discourse offers
interesting cues for analysis and experimentation
(Frank, 1996; Dixon, 2002; Callander &amp; Wilkie,
2007; Osborne, 2014). In recent years, political
speech has earned much attention (Guerini et al.,
2008; 2013; Esposito et al., 2015) for purposes,
ranging from analysis of communication
strategies (Muelle, 1973; Wilson, 1990; Wilson,
2011), persuasive Natural Language Processing,
politicians’ rhetoric (Stover &amp; Ibroscheva, 2017)
and virality of information diffusion (Caliandro
&amp; Balina, 2015). Regarding MWs resources for
Italian we may mention recent contributions such
as PANACEA (Platform for Automatic,
Normalized Annotation and Cost-Effective Acquisition
of Language Resources for Human Language
Techologies) that includes Italian word n-grams
and Italian word/tag/lemma n-grams in the
"Labour" (LAB) domain
        <xref ref-type="bibr" rid="ref3">(Bel at al., 2012)</xref>
        and also
PARSEME-IT Corpus, an annotated Corpus of
Verbal Multiword Expressions in Italian
        <xref ref-type="bibr" rid="ref21">(Monti
et al., 2017)</xref>
        .
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>PoliSdict</title>
      <p>
        According to
        <xref ref-type="bibr" rid="ref15">Gross (1999)</xref>
        the lexicographic data
available in machine-readable format are printed
dictionaries, electronic dictionaries and corpora.
In particular dictionaries are built for being used
by programs, with their content made of
alphanumerical codes which represent the
grammatical data that can be reasonably
formalized at this moment in time. The creation
and management of the electronic dictionary of
MWEs in Italian spoken language took place
through four main steps:
• lexical acquisition from corpus
• lexicon-based identification of MWEs
• information extraction
• identification of most recurrent PoS
patterns
The first step concerns the lexical acquisition.
We automatically extract MWEs starting from
PoliModalCorpus
        <xref ref-type="bibr" rid="ref32">(Trotta et al., 2018)</xref>
        , a political
domain corpus for Italian language currently
composed of transcriptions4 of 59 face-to-face
interviews (14:00:00 hours) held during the
political talk show “In mezz'ora in più” (from 24
September 2017 to 14 January 2018) and 18
speeches (7:02:39 hours) held during the election
campaign for regional elections (from December
24th 2014 to March 4th 2015) by the then
candidate Vincenzo De Luca5. The dimension of
the individual corporus is indicated below (Tab.
1).
      </p>
      <sec id="sec-4-1">
        <title>Type</title>
      </sec>
      <sec id="sec-4-2">
        <title>PoliModalCorpus 11,231</title>
      </sec>
      <sec id="sec-4-3">
        <title>De Luca Corpus 7,225</title>
        <p>Total 18,456
Table 1 - Corpus statistics overview</p>
      </sec>
      <sec id="sec-4-4">
        <title>Token</title>
        <p>158,543
56,672
215,251</p>
        <p>TTR
0.07
0.12
0.08
4 Using a semi-supervised speech-to-text methodology
(Google API + manual transcription).
5 It should be specified here that our is an initial experiment
on a relatively small sample, since a larger balanced corpus
would be necessary for a broader coverage.</p>
        <p>
          In a second step – exploiting the theoretical
backgroung offered by the Lexicon-Grammar6
framework - we identified the MWEs by
processing the corpus in Nooj7 (Elia et al., 2010)
and using the Compound-Word Electronic
Dictionaries (DELAC-DELACF)
          <xref ref-type="bibr" rid="ref9">(De Bueriis &amp;
Elia, 2008)</xref>
          , which includes compound words and
sequences formed by two or more words which
jointly construct single units of meaning, thanks
to which it was also possible to attribute a
terminological label to each identified MWEs. It
has to be noticed that in this step our efforts
focused on the extraction of nominal compounds,
leaving the extraction and integration of
adverbial and adjectival compounds for future
research. In a third phase the extracted MWEs
were manually verified using the GRADIT (De
Mauro, 2000). This operation has allowed us to
identify 356 MWEs compared to 882 identified
by DELAC-DELACF and to attribute to each
compound expression the respective frequency
label documented by the GRADIT. In a fourth
phase a structural analysis of the extracted
MWEs was carried out and the most recurring
part of speech patterns were identified. Therefore
the terminological labels 8 are distributed as
follows: &lt;econ&gt; 112, &lt;fig&gt; 37, &lt;dige&gt; 36,
&lt;pol&gt; 21, &lt;med&gt; 179. Even though we extracted
the MWEs from interviews of political kind, the
MWEs tagged with the &lt;pol&gt; (political) labels
are only 21. Following the most recurrent
frequency label we found were: TS10 (167) (i.e.
abuso di ufficio), CO 11 (136) (i.e. arredo
urbano), CO - TS (30) (i.e. istituto di credito).
The methodological approach of the
Lexicongrammar has also restricted the taxonomic
6 Gross (1975) shows that every verb has a unique behavior,
characterized by different properties and constraints. In
general, no ether verb has an identical syntactic paradigm.
Consequently, the properties of each verbal construction
must be represented in a lexicon-grammar.
7 NooJ is a knowledge-based NLP tool based on huge
handcrafted linguistic resources, i.e. Dictionaries, derivational
grammars.
          <xref ref-type="bibr" rid="ref35">(Vietri, 2014)</xref>
          .
8 Being an essentially terminological dictionary,
DELACDELACF assigns one or more terminology labels to each
single entry, based on the areas of knowledge in which a
specific compound has been attested. Currently the domains
are 173 and the most populated is that of medicine.
9 The terminological labels with a frequency lower than 17
are not mentioned.
10 Technical-specialist use (107,194 words have this
acronym and are known above all in relation to specific contexts
of science or technology, eg amicina).
11 Common use (as many as 47.060 words are used and
understood and understood, regardless of profession or
origin, to anyone with a higher level of education, eg
allusivo).
analysis of compound polysematic words today
they are naturally combined with the notion of
compound nouns set by Gross and which can be
described as “the sequence of their grammatical
categories, in the same way as for adverbs”
          <xref ref-type="bibr" rid="ref14">(Gross, 1986)</xref>
          . Starting from this point of view,
we may indicate how the most recurring patterns
in our dictionary were respectively: N + A - valid
for 218 words (like lavori forzati ecc), N di N
(82) (i.e. economia di scala), N + N (30) (i.e.
estratto conto), N prep N (22) (i.e. ministero del
lavoro), N a N (2) (i.e. corpo a corpo), N da N
(2) (i.e. macchina da guerra). Notice that, since
in this study we are dealing with nominal MWEs
the syntactic head of the compounds is always
represented by the name in patterns like N + A
and A + N, N + N, while in more complex
patterns, as N a N and the like, we found
controversial the identification of a single word
as syntactic head. Since our primary interest was
to identify and systematically arrange the
extracted knowledge from a lexicographic point
of view, we decided to deepen the syntactic
analysis (which is to say the explicitation of the
syntactic heads and the syntactic category of
each MWE) during research steps to be included
in near future research. Starting from the
information extracted so far we have then created
an electronic dictionary where to each MWE are
associated information about gender and number,
part of speech pattern, frequency labels, and
terminological label. The dictionary was created
using the XML as markup language following
the TEI standard12 and adding the tags &lt;mark&gt;
in order to include the frequency tags indicated
by the GRADIT and &lt;label&gt; to indicate the
knowledge domain in which the word is attested,
indicated to the DELAC-DELACF dictionaries).
The choice of exploiting this markup language is
motivated by its extreme generalization and
flexibility
          <xref ref-type="bibr" rid="ref27">(Pierazzo, 2005)</xref>
          and in order to
represent the MWEs in a common format and to
enable linkage (Calzolari et al., 2002). The
adopted formalism uses the following tags:
● &lt;entry&gt;: contains a single structured
entry in any kind of lexical resource,
such as a dictionary or lexicon
● &lt;form&gt;: (form information group)
groups all the information on the written
12 P5: Guidelines for Electronic Text Encoding and
Interchange, Version 3.4.0. Last updated on 23rd July 2018,
revision 1fa0b54.
        </p>
        <p>and spoken forms of one headword
● &lt;gramGrp&gt;: (grammatical information
group) groups morpho-syntactic
information about a lexical item, e.g.
pos, gen, number
● &lt;mark&gt;: frequency label from GRADIT
● &lt;label&gt;: terminological label from</p>
        <p>DELAC-DELACF
The dictionary therefore appears as follows:
&lt;entry&gt;
&lt;form&gt;
&lt;orth&gt;abuso d'ufficio&lt;/orth&gt;
&lt;type&gt;multiword expression&lt;/type&gt;
&lt;/form&gt;
&lt;gramGrp&gt;
&lt;gram type= "pos"&gt;NdiN&lt;/gram&gt;
&lt;gram type="gen"&gt;m&lt;/gram&gt;
&lt;gram type="num"&gt;s&lt;/gram&gt;
&lt;/gramGrp&gt;
&lt;mark&gt;TS&lt;/mark&gt;
&lt;label&gt;dige&lt;/label&gt;
&lt;/entry&gt;
&lt;entry&gt;
&lt;form&gt;
&lt;orth&gt;agente atmosferico&lt;/orth&gt;
&lt;type&gt;multiword expression&lt;/type&gt;
&lt;/form&gt;
&lt;gramGrp&gt;
&lt;gram type= "pos"&gt;NA&lt;/gram&gt;
&lt;gram type="gen"&gt;m&lt;/gram&gt;
&lt;gram type="num"&gt;s&lt;/gram&gt;
&lt;/gramGrp&gt;
&lt;mark&gt;TS&lt;/mark&gt;
&lt;label&gt;meteor&lt;/label&gt;
&lt;/entry&gt;
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Ontologic expansion dictionary of the xml</title>
      <p>Following the creation of the dictionary we also
decided to organize the knowledge retrieved
from the exploited datasets as an ontological
dictionary which is actually under construction
and that will be freely avilable under Creative
Commons License (CC+BY-NC-ND). The
choice to build such a linguistic resource is
grounded on the idea that a formal representation
of the MWEs may not only help software agents
in the automatic recognition of compound words
within written/oral texts, but can still enhance the
resolution of referential expression such as
Primo Ministro, Santo Padre and the like, which
is to say of those frozen expressions that bear
pragmatic references pointing to subject/object
that are likely to change over medium/short
periods of time. In order to perform a deeper
pragmatic disambiguation of MWEs we
exploited the descriptive capability of the
Ontology Web Language (OWL), a standard
markup language provided by the World Wide
Web (W3C) Consortium for the formalization of
vocabularies of terms covering specific domains
of knowledge. Following the W3C guidelines we
shaped the electronic dictionary so that to each
MWE a set of description classes and linking
relationship are attached, according to the
lexicon-grammar analysis previously performed
and transposed into the ontology. Here is an
example of the metadata scheme provided for the
compound expression campagna elettorale:
●
●
●
●
●
●</p>
      <sec id="sec-5-1">
        <title>Class “DELAC-DELACF Label”:</title>
        <p>&lt;pol&gt; (politic)</p>
      </sec>
      <sec id="sec-5-2">
        <title>Class “GRADIT” Label: CO</title>
        <p>(Common)</p>
      </sec>
      <sec id="sec-5-3">
        <title>Class “Syntactic Pattern”: N(oun) +</title>
        <p>A(djective)</p>
      </sec>
      <sec id="sec-5-4">
        <title>Data property “Corpus frequency”: 52</title>
      </sec>
      <sec id="sec-5-5">
        <title>Data property “Occurrence”:</title>
        <p>Berlusconi comincia la sua campagna
elettorale andando in Tunisia a
commemorare Craxi, che ne pensa di
questa decisione?</p>
      </sec>
      <sec id="sec-5-6">
        <title>Data property “DBpedia redirection link”:</title>
        <p>http://it.dbpedia.org/resource/Campagna
_elettorale/html
As we can notice the first three classes plus the
first two data properties directly derive from the
linguistic analysis and their ontological
formalisation may serve as powerful search
filters in case of description logic queries
submitted over the electronic dictionary. To what
concerns the DBpedia redirection link property
class, this derives from the Italian section of
DBpedia project (Auer et al., 2007) and will
serve as core mechanism for the pragmatic
resolution of the compound expression. It should
be further noticed that the mapping effort
between the extracted MWEs and DBpedia
virtually put the work in progress ontology on
the fifth and last level of Berner Lee’s Open Data
scale, which is to say on the level reserved for
web semantic compliant resources additionally
providing redirection links to other web datasets
for the contextualisation of the described
knowledge, following the initial proposal of
(Bizer et al., 2008 ).
5</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Future work</title>
      <p>In this work we described the initial steps for the
development and formalization of PoliSdict, an
electronic dictionary of spoken language MWEs.
We illustrated the methdology used to build the
resource and the preliminary results that we
obtained from a systematic analysis. For what is
related to future research we consider necessary
exploiting standard association measures (like
mutual information or log-likelihood ratio) to get
an index of cohesion within the identified
expressions and compare the use and
collocations of MWEs between corpora of
written and spoken language in order to
understand which of them are the most used.
Considering this study as an initial experiment
on a relatively small sample, a larger balanced
corpus would be necessary for a broader
coverage, therefore we intend to proceed with
the expansion of the corpus and the associated
dictionary. Following we will make the
described resources freely accessible by means
of graphical interface, so as to offer the
possibility to browse and explore data, also
allowing the free use of the source codes for
research purposes under Creative Commons
License (CC+BY-NC-ND.
6</p>
    </sec>
    <sec id="sec-7">
      <title>Aknowledgments</title>
      <p>We would like to thank Network Contacts s.r.l.
for their willingess to help us with valuable
research insights and for the support during the
writing of this paper. We would also like to
thank the anonymous reviewers for their helpful
suggestions.
Proceedings of the 16th Annual Conference of the
European Association for Machine Translation,
Trento, Italy.
dei
testi:
Ramisch, C., Schreiner, P., Idiart, M., &amp;
Villavicencio, A. (2008, June). An evaluation of
methods for the extraction of multiword
expressions. In Proceedings of the LREC
Workshop-Towards a Shared Task for Multiword
Expressions (MWE 2008) (pp. 50-53).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Baldwin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Villavicencio</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2002</year>
          ,
          <article-title>August). Extracting the unextractable: A case study on verbparticles</article-title>
          .
          <source>In proceedings of the 6th conference on Natural language learning-</source>
          Volume
          <volume>20</volume>
          (pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Bejček</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Straňák</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Annotation of multiword expressions in the Prague dependency treebank</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>44</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>7</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Bel</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poch</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Toral</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>PANACEA (Platform for Automatic, Normalised Annotation and Cost-Effective Acquisition of Language Resources for Human Language Technologies)</article-title>
          . In Calzolari, N.,
          <string-name>
            <surname>Fillmore</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grishman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ide</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lenci</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>MacLeod</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Zampolli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2002</year>
          , May).
          <article-title>Towards Best Practice for Multiword Expressions in Computational Lexicons</article-title>
          .
          <source>In LREC.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Chiari</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Collocazioni e polirematiche nel lessico musicale italiano. Lingua, letteratura e cultura italiana"</article-title>
          .
          <source>Atti del convegno Internazionale</source>
          ,
          <volume>50</volume>
          ,
          <fpage>165</fpage>
          -
          <lpage>190</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>CSL.</surname>
          </string-name>
          <year>2005</year>
          .
          <article-title>Special issue on Multiword Expressions of Computer Speech</article-title>
          &amp; Language, volume
          <volume>19</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>D'Agostino</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Elia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>Il significato delle frasi: un continuum dalle frasi semplici alle forme polirematiche</article-title>
          .
          <source>AA</source>
          . VV,
          <article-title>Ai limiti del linguaggio</article-title>
          .
          <source>Bari: Laterza</source>
          ,
          <fpage>287</fpage>
          -
          <lpage>310</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Dagan</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Church</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>1994</year>
          ,
          <article-title>October)</article-title>
          .
          <article-title>Termight: Identifying and translating technical terminology</article-title>
          .
          <source>In Proceedings of the fourth conference on Applied natural language processing</source>
          (pp.
          <fpage>34</fpage>
          -
          <lpage>40</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Daille</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>Combined approach for terminology extraction: lexical statistics and linguistic filtering</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>De Bueriis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Elia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Lessici elettronici e descrizioni lessicali, sintattiche</article-title>
          , morfologiche ed ortografiche.
          <source>Plectica</source>
          , Salerno.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>De Mauro</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>1999</year>
          ).
          <source>Gradit. Torino: UTET</source>
          , 1.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Maienborn</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>von Heusinger</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Portner</surname>
            ,
            <given-names>P</given-names>
          </string-name>
          . (Eds.). (
          <year>2011</year>
          ).
          <article-title>Semantics: An international handbook of natural language meaning</article-title>
          (Vol.
          <volume>1</volume>
          ). Walter de Gruyter.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Grégoire</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>DuELME: a Dutch electronic lexicon of multiword expressions</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>44</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>23</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Gross</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Thématisation des compléments circonstanciels</article-title>
          .
          <source>In Le poids des mots. Hommage à Alicja Kacprzak;. Wydawnictwo Uniwersytetu Łódzkiego.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Gross</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>1986</year>
          ,
          <article-title>August)</article-title>
          .
          <article-title>Lexicon-grammar: the representation of compound words</article-title>
          .
          <source>In Proceedings of the 11th conference on Computational linguistics</source>
          (pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Gross</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>1999</year>
          ).
          <article-title>A bootstrap method for constructing local grammars</article-title>
          .
          <source>In Proceedings of the Symposium on Contemporary Mathematics</source>
          (pp.
          <fpage>229</fpage>
          -
          <lpage>250</lpage>
          ). University of Belgrad.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>JLRE.</surname>
          </string-name>
          <year>2009</year>
          .
          <article-title>Special issue on Multiword Expressions of the Journal of Language Resources and Evaluation</article-title>
          , volume to appear.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S. N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Baldwin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2006</year>
          , April).
          <article-title>Automatic identification of English verb particle constructions using linguistic features</article-title>
          .
          <source>In Proceedings of the Third ACL-SIGSEM Workshop on Prepositions</source>
          (pp.
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S. N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Baldwin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2006</year>
          , April).
          <article-title>Automatic identification of English verb particle constructions using linguistic features</article-title>
          .
          <source>In Proceedings of the Third ACL-SIGSEM Workshop on Prepositions</source>
          (pp.
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>McEnery</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Langé</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oakes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Véronis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>1997</year>
          ).
          <article-title>The exploitation of multilingual annotated corpora for term extraction</article-title>
          .
          <source>Corpus annotation--- linguistic information from computer text corpora</source>
          ,
          <fpage>220</fpage>
          -
          <lpage>230</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Michiels</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Dufour</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>DEFI, a tool for automatic multi-word unit recognition, meaning assignment and translation selection</article-title>
          .
          <source>In Proceedings of the first international conference on language resources &amp; evaluation</source>
          (pp.
          <fpage>1179</fpage>
          -
          <lpage>1186</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Monti J.</given-names>
            ,
            <surname>di Buono</surname>
          </string-name>
          <string-name>
            <given-names>M.P.</given-names>
            ,
            <surname>Sangati</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          (
          <year>2017</year>
          )
          <article-title>PARSEME-It Corpus An annotated Corpus of Verbal Multiword Expressions in Italian</article-title>
          .
          <source>In: CLIC-It 2017 Proceedings - Rome 11-13 December</source>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Mudraya</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Babych</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rayson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp; Wilson,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Developing a Russian semantic tagger for automatic semantic annotation</article-title>
          .
          <source>Corpus Linguistics</source>
          <year>2006</year>
          ,
          <fpage>290</fpage>
          -
          <lpage>297</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Nunberg</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sag</surname>
            ,
            <given-names>I. A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Wasow</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>1994</year>
          ).
          <source>Idioms. Language</source>
          ,
          <volume>70</volume>
          (
          <issue>3</issue>
          ),
          <fpage>491</fpage>
          -
          <lpage>538</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Pearce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2002</year>
          , May).
          <article-title>A Comparative Evaluation of Collocation Extraction Techniques</article-title>
          .
          <source>In LREC.</source>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Lexical association measures and collocation extraction</article-title>
          .
          <source>Language resources and evaluation</source>
          ,
          <volume>44</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>137</fpage>
          -
          <lpage>158</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Piao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Archer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mudraya</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rayson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garside</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McEnery</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <article-title>&amp;</article-title>
          <string-name>
            <surname>Wilson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>A large semantic lexicon for corpus annotation</article-title>
          .
          <source>Corpus Linguistics</source>
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Pierazzo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>La codifica un'introduzione. Carocci editore</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Rayson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Archer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>McEnery</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>The UCREL semantic analysis system</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Smadja</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>1993</year>
          ).
          <article-title>Retrieving collocations from text: Xtract</article-title>
          . Computational linguistics,
          <volume>19</volume>
          (
          <issue>1</issue>
          ),
          <fpage>143</fpage>
          -
          <lpage>177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Strik</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Cucchiarini</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>1999</year>
          ).
          <article-title>Modeling pronunciation variation for ASR: A survey of the literature</article-title>
          .
          <source>Speech Communication</source>
          ,
          <volume>29</volume>
          (
          <issue>2-4</issue>
          ),
          <fpage>225</fpage>
          -
          <lpage>246</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Strik</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hulsbosch</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Cucchiarini</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Analyzing and identifying multiword expressions in spoken language</article-title>
          .
          <source>Language resources and evaluation</source>
          ,
          <volume>44</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>41</fpage>
          -
          <lpage>58</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>Trotta</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albanese</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <article-title>Polimodalcorpus: verso la costruzione del primo corpus multimodale di dominio politico in italiano; Proceedings of the XXVIII Ass</article-title>
          .I.Term International Conference, Salerno,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>Vietri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Lessico-grammatica dell'italiano</article-title>
          . Metodi, descrizioni e applicazioni (p.
          <fpage>304</fpage>
          ). UTET Università.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <surname>Vietri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>1985</year>
          ).
          <article-title>Lessico e sintassi delle espressioni idiomatiche: una tipologia tassonomica dell'italiano</article-title>
          . Liguori.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>Vietri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Idiomatic constructions in Italian: a lexicon-grammar approach</article-title>
          (Vol.
          <volume>31</volume>
          ). John Benjamins Publishing Company.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <surname>Villavicencio</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kordoni</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Idiart</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ramisch</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Validation and evaluation of automatically acquired multiword expressions for grammar engineering</article-title>
          .
          <source>In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL).</source>
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>Wermter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>1997</year>
          ).
          <article-title>Cautious steps towards hybrid connectionist bilingual phrase alignment</article-title>
          .
          <source>In Recent Advances in Natural Language Processing</source>
          (Vol.
          <volume>97</volume>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>1997</year>
          ).
          <article-title>Stochastic inversion transduction grammars and bilingual parsing of parallel corpora</article-title>
          .
          <source>Computational linguistics</source>
          ,
          <volume>23</volume>
          (
          <issue>3</issue>
          ),
          <fpage>377</fpage>
          -
          <lpage>403</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kordoni</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villavicencio</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Idiart</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2006</year>
          ,
          <article-title>July)</article-title>
          .
          <article-title>Automated multiword expression prediction for grammar engineering</article-title>
          .
          <source>In Proceedings of the workshop on multiword expressions: Identifying and exploiting underlying properties</source>
          (pp.
          <fpage>36</fpage>
          -
          <lpage>44</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>