<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Language resources for Italian: towards the development of a corpus of annotated Italian multiword expressions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shiva Taslimipoor</string-name>
          <email>shiva.taslimi@wlv.ac.uk</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anna Desantis, Manuela Cherchi</string-name>
          <email>annadesantis_91@libero.it</email>
          <email>annadesantis_91@libero.it, manuealacherchi82@gmail.com</email>
          <email>manuealacherchi82@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ruslan Mitkov</string-name>
          <email>r.mitkov@wlv.ac.uk</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Johanna Monti</string-name>
          <email>jmonti@unior.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>"L'Orientale" University of Naples</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Sassari</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Wolverhampton</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2010</year>
      </pub-date>
      <abstract>
        <p>English. This paper describes the first resource annotated for multiword expressions (MWEs) in Italian. Two versions of this dataset have been prepared: the first with a fast markup list of out-of-context MWEs, and the second with an in-context annotation, where the MWEs are entered with their contexts. The paper also discusses annotation issues and reports the inter-annotator agreement for both types of annotations. Finally, the results of the first exploitation of the new resource, namely the automatic extraction of Italian MWEs, are presented.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Italiano. Questo contributo descrive
la prima risorsa italiana annotatata con
polirematiche. Sono state preparate due
versioni del dataset: la prima con una
lista di polirematiche senza contesto, e
la seconda con annotazione in contesto.
Il contributo discute le problematiche
emerse durante l’annotazione e riporta
il grado di accordo tra annotatori per
entrambi i tipi di annotazione. Infine
vengono presentati i risultati del primo
impiego della nuova risorsa, ovvero
l’estrazione automatica di polirematiche
per l’italiano.
ever, despite being desiderata for linguistic
analysis and language learning, as well as for
training and evaluation of NLP tasks such as term
extraction (and Machine Translation in multilingual
scenarios), resources annotated with MWEs are a
scarce commodity
        <xref ref-type="bibr" rid="ref21 ref22">(Schneider et al., 2014b)</xref>
        . The
need for such types of resources is even greater for
Italian which does not benefit from the variety and
volume of resources as does English.
      </p>
      <p>This paper outlines the development of a new
language resource for Italian, namely a corpus
annotated with Italian MWEs of a particular class:
verb-noun expressions such as fare riferimento,
dare luogo and prendere atto. Such
collocations are reported to be the most frequent class of
MWEs and of high practical importance both for
automatic translation and language learning. To
the best of our knowledge, this is the first resource
of this kind in Italian.</p>
      <p>
        The development of this corpus is part of a
multilingual project addressing the challenge of
computational treatment of MWEs. It covers English,
Spanish, Italian and French and its goal is to
develop a knowledge-poor methodology for
automatically identifying MWEs and retrieving their
translations
        <xref ref-type="bibr" rid="ref25">(Taslimipoor et al., 2016)</xref>
        for any pair
of languages. The developed methodology will
be used for Machine Translation and
multilingual dictionary compilation, and also in
computeraided tools to support the work of language
learners and translators.
      </p>
      <p>Two versions of the above resource have been
produced. The first version consists of lists
of MWEs annotated out-of-context with a view
to performing fast evaluation of the developed
methodology (out-of-context mark-up). The
second version consists of annotated MWEs along
with their concordances (in-context annotation).
The latter type of annotation is time-consuming,
but provides the contexts for the MWEs annotated.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Annotation of MWEs: out-of-context mark-up and in-context annotation</title>
      <p>
        After more than two decades of computational
studies on MWEs, the lack of a proper gold
standard is still an issue. Lexical resources like
dictionaries have limited coverage of these
expressions
        <xref ref-type="bibr" rid="ref11">(Losnegaard et al., 2016)</xref>
        and there is also
no proper tagged corpus of MWEs in any language
        <xref ref-type="bibr" rid="ref21 ref22">(Schneider et al., 2014b)</xref>
        .
      </p>
      <p>
        Most previous studies on the computational
treatment of MWEs have focused on extracting
types (rather than tokens)1 of MWEs from corpora
        <xref ref-type="bibr" rid="ref16 ref17 ref18 ref20 ref23 ref26 ref4">(Ramisch et al., 2010; Villavicencio et al., 2007;
Rondon et al., 2015; Salehi and Cook, 2013)</xref>
        . The
widely-used toolboxes of MWEToolkit
        <xref ref-type="bibr" rid="ref17 ref4">(Ramisch
et al., 2010)</xref>
        or Xtract
        <xref ref-type="bibr" rid="ref24">(Smadja, 1993)</xref>
        extract
expressions if their statistical occurrences represent
the likelihood of them being MWEs. The
evaluation for the type-based extraction of MWEs has
been mostly performed against a dictionary
        <xref ref-type="bibr" rid="ref4">(de
Caseli et al., 2010)</xref>
        , lexicon
        <xref ref-type="bibr" rid="ref16 ref20 ref23">(Pichotta and
DeNero, 2013)</xref>
        or list of human-annotated expressions
        <xref ref-type="bibr" rid="ref26">(Villavicencio et al., 2007)</xref>
        . However, there are
some examples like the expression have a baby,
which in exactly the same form and structure,
might be an MWE (meaning to give birth ) in some
contexts and a literal expression in others.
      </p>
      <p>As for the automatic identification of the tokens
of MWEs, Fazly et al. (2009) make use of both
linguistic properties and the local context, in
determining the class of an MWE token. They
report an unsupervised approach to identifying
idiomatic and literal usages of an expression in
context. Their method is evaluated on a very small
sample of expressions in a small portion of the
British National Corpus (BNC), which were
annotated by humans. Schneider et al. (2014a)
developed a supervised model whose purpose is to
identify MWEs in context. Their methodology results
in a corpus of automatically annotated MWEs. It
is not clear, however, if the methodology is able
to tag one specific expression as an MWE in one
context and non-MWE in another. The PARSEME
shared task2 is also devoted to annotating verbal
1Type refers to the canonical form of an expression, while
token refers to each instance (usage) of the expression in any
morphological form in text.</p>
      <p>2http://typo.uni-konstanz.de/parseme/index.php/2-general/
142-parseme-shared-task-on-automatic-detection-of-verbal-mwes
MWEs in several languages. The shared task,
while having interesting discussions on the area,
has embarked upon the labour-intensive
annotation of verbal MWEs.</p>
      <p>
        Since there is no list of verb-noun MWEs in
Italian, we first automatically compile a list of
such expressions, to be annotated by human
experts. This is based on previous attempts at
extracting a lexicon of MWEs (as in
        <xref ref-type="bibr" rid="ref27">(Villavicencio,
2005)</xref>
        ). Annotators are not provided with any
context and hence the task is more feasible in terms
of time. Human annotators are asked to label the
expressions as MWEs only if they have sufficient
degrees of idiomaticity. In other words, a Verb +
Noun MWEs does not convey literal meaning in
that the verb is delexicalised.
      </p>
      <p>
        However, we believe that idiomaticity is not a
binary property; rather it is known to fall on a
continuum from completely semantically transparent,
or literal, to entirely opaque, or idiomatic
        <xref ref-type="bibr" rid="ref7">(Fazly et al., 2009)</xref>
        . This makes the task of
out-ofcontext marking-up of the expression more
challenging for annotators, since they have to pick a
value according to all the possible contexts of a
target expression. This ambiguity and the fact that
there are many expressions that in some contexts
are MWEs and in some contexts not, prompted us
to initiate a subsequent annotation where MWEs
are tagged in their contexts. The idea is to
extract the concordances around all the occurrences
of a Verb + Noun expression and provide
annotators with these concordances in order to be able
to decide the degree of idiomaticity of the specific
verb-noun expression. We compare the reliability
of the in-context and out-of-context annotations by
way of the agreement between annotators.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Experimental expressions</title>
        <p>
          Highly polysemous verbs, such as give and take
in English and fare and dare in Italian widely
participate in Verb+Noun MWEs, in which they
contribute a broad range of figurative meanings that
must be recognised
          <xref ref-type="bibr" rid="ref6">(Fazly et al., 2007)</xref>
          . We
focus on four mostly frequent Italian verbs: fare,
dare, prendere and trovare. We extract all the
occurrences of these verbs when followed by any
noun, from the itWaC corpus
          <xref ref-type="bibr" rid="ref3">(Baroni and
Kilgarriff, 2006)</xref>
          , using SketchEngine
          <xref ref-type="bibr" rid="ref9">(Kilgarriff et al.,
2004)</xref>
          . For the first experiment all the Verb+Noun
types are extracted when the verb is lemmatised;
and for the second experiment all the
concordances of these verbs when followed by a noun
are generated.
The extraction of Verb+Noun candidates of the
four verbs in focus and the removal of the
expressions with frequencies lower than 20, results in a
dataset of 3; 375 expressions. Two native
speakers annotated every candidate expression with 1
for an MWE if the expression was idiomatic and
with 0 for a non-MWE if the expression was
literal. We have also defined the tag 2 for the
expressions that in some contexts behave as MWEs
and in others do not, e.g. dare frutti, which has
a literal usage that means to produce fruits but in
some contexts means to produce results and is an
MWE in these contexts. While this out-of-context
‘fast track’ annotation procedure saves time and
yields a long list of marked-up expressions,
annotators often feel uncomfortable due to the lack
of context. The information about the agreements
between annotators in terms of Kappa is shown
in Table 2 and is compared with the in-context
annotation of MWEs as explained in Section 2.3.
2.3
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Annotating Verb+Noun(s) in context</title>
        <p>We design an annotation task, in which we provide
a sample of all usages of any type of Verb+Noun
expression to be annotated. For this purpose, we
employ the SketchEngine to list all the
concordances of each verb when it is followed by a noun.
Concordances include the verb in focus with
almost ten words before and ten words after that.
The SketchEngine reports only 100; 000
concordances for each query. Among them, we filter out
the concordances that include Verb+Noun
expressions with frequencies lower than 50 and we
randomly select 10% of the concordances for each
verb. As a result, there are 30; 094 concordances
to be annotated. The two annotators annotate all
usages of Verb+Noun expressions in these
concordances, considering the context that the expression
occurred in, marking up MWEs with 1 and
expressions which are not MWEs, with 0. Table 1
reports on the details of annotation tasks and Table
2 shows the agreement details for them.
2.4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Discussion</title>
        <p>As seen in Table 2, the inter-annotator agreement
is significantly higher when annotating the
expressions in context. One of the main causes of
disagreements in out-of-context annotation is
concerned with abstract nouns. The annotation of
expressions composed of a verb followed by a noun
with an abstract meaning is a more complicated
process as the candidate expression may carry a
figurative meaning. Each annotator uses their
intuition to annotate them and it leads to random
tags for these expression (e.g. fare notizia, dare
identità, prendere possesso) when they are
out-ofcontext. However, in the case of in-context
annotation, concordances composed of abstract nouns
have been annotated in the majority of cases with
1 by both annotators.</p>
        <p>In-context annotation is also very helpful for
annotating expressions with both idiomatic and
literal meanings. An interesting observation,
reported in Table 3, is related to the number of
expressions that are detected with the two different
usages of idiomatic and non-idiomatic, in context.</p>
        <p>As can be seen in Table 3,3 among the 1; 649
types of expressions in concordances, 530 (32%)
of them could be MWEs in some context and
nonMWEs in others (context-depending), according
to the first annotator. This annotator has annotated
only 3% of the expressions with tag ‘2’ without
context.</p>
        <p>3Note that the numbers in Table 3 cannot be interpreted
to validate agreement between annotators, i.e. no conclusion
about agreement can be derived from 3.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>First use of the MWE resource: comparative evaluation of the automatic extraction of Italian MWEs</title>
      <p>In our multilingual project (see Section 1) we
regard the automatic translation of MWEs as a
twostage process. The first stage is the extraction of
MWEs in each of the languages; the second stage
is a matching procedure for the extracted MWEs in
each language which proposes translation
equivalents. In this study the extraction of MWEs is
based on statistical association measures (AMs).</p>
      <p>
        These measures have been proposed to
determine the degree of compositionality, and fixedness
of expressions. The more compositional or fixed
expressions are, the more likely it is that they are
MWEs
        <xref ref-type="bibr" rid="ref2 ref5">(Evert, 2008; Bannard, 2007)</xref>
        . According
to Evert (2008), there is no ideal association
measure for all purposes. We aim to evaluate AMs
as a baseline approach against the annotated data
which we prepared. We focus on a selection of
five AMs which have been more widely discussed
to be the best measures to identify MWEs. These
are: MI3
        <xref ref-type="bibr" rid="ref15">(Oakes, 1998)</xref>
        , log-likelihood
(Dunning, 1993), T-score
        <xref ref-type="bibr" rid="ref10">(Krenn and Evert, 2001)</xref>
        ,
logDice (Rychlý, 2008) and Salience
        <xref ref-type="bibr" rid="ref9">(Kilgarriff et al.,
2004)</xref>
        all as defined in SketchEngine. We compare
the performance of these AMs and also frequency
of occurrence (Freq) as the sixth measure to rank
the candidate MWEs. We evaluate the effect of
these measures in ranking MWEs on both kinds of
datasets.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Experiments on type-based extraction of</title>
      </sec>
      <sec id="sec-3-2">
        <title>MWEs</title>
        <p>In the first experiment, the list of all extracted Verb
+ Noun combinations (as explained in Section 2.1)
are ranked according to the above measures that
are computed from itWaC as a reference corpus.
To perform the evaluation against the list of
annotated expressions, we process all 2,415
expressions for which the annotators agreed on tags 0
or 1. After ranking the expressions by the
measures, we examine the retrieval performance of
each measure by computing the 11-point
Interpolated Average Precision (11-p IAP). This reflects
the goodness of a measure in ranking the relevant
items (here, MWEs) before the irrelevant ones. To
this end, the interpolated precision at the 11
recall values of 0, 10%, ..., 100% is calculated. As
detailed in Manning et al. (2008), the
interpolated precision at a certain recall level, r, is defined
as the highest precision found for any recall level
r0 r. The average of these 11 points is reported
as 11-p IAP in Table 4.</p>
        <p>As can be seen in Table 4, the selected
association measures generally perform with similar
performance in ranking this type of MWEs, with
M I3 performing slightly better than others.
In the second experiment, we seek to establish
the effect of these measures on identifying the
usages of MWEs in our dataset of in-context
annotations. We set a threshold for each score
that we have computed for Verb+Noun
expression types. By setting thresholds we compute the
classification accuracy of the measures to
identify MWEs among the usages of Verb+Noun
expressions in a corpus. Specifically, each candidate
of a Verb+Noun in the concordances is
automatically tagged as an MWE if its lemmatised form
has a score higher than the threshold, and as a
nonMWE, otherwise. For each measure, we compute
the arithmetic mean (average) of all the values of
that measure for all expressions, and set the
resulted average value as a threshold.</p>
        <p>The accuracies of classifying the candidate
Verb+Noun expressions are computed based on
the human annotations of the concordances and
are shown in Table 5. The classification
accuracies of AMs are also very close to each other (see
Table 5); however, this time Log-likelihood and
F req fare slightly better than others in classifying
tokens of Verb+Noun expressions.
Our new resource of concordances contains
useful linguistic information related to usages of
expressions and as such important features can be
extracted from the resource to help identifying
MWEs. One of these features can be obtained
from the statistics of different possible inflections
of the verb component of an expression. Based on
the premise of the fixedness of MWEs, we expect
that the verb component of a verb-noun MWE
occurs only in a limited number of inflections. We
implement this feature by dividing the frequency
of occurrences of each expression by the number
of inflections that the verb component occurs in.
Note that to count the number of different
inflections of the verb component, we rely on the
subcorpus of concordances that we gathered.</p>
        <p>We evaluate this approach only on 1,077
expressions that occur in concordances. We rank
the expressions according to this newly computed
score and we call this score, which depends on the
inflection varieties, INF-VAR. For all verbs, the
INF-VAR performs comparably to Frequency in
ranking MWEs higher than non-MWEs, but for
the verb trovare, we obtain better 11-p IAP using
this score than by using Frequency (see Table 6).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and future work</title>
      <p>In this paper, we outline our work towards a
goldstandard dataset which is tagged with Italian
verbnoun MWEs along with their contexts. We show
the reliability of this dataset by its considerable
inter-annotator agreement compared to the
moderate inter-annotator agreement on annotated
verbnoun expressions presented without context. We
also report the results of automatic extraction of
MWEs using this dataset as a gold-standard. One
of the advantages of this dataset is that it includes
both 0-tagged and 1-tagged tokens of expressions
and it can be used for classification and other
statistical NLP approaches. In future work, we are
interested in extracting context features from
concordances in this resource to automatically
recognise and classify the expressions that are MWEs in
some contexts but not MWEs in others.</p>
      <p>Ted Dunning. 1993. Accurate methods for the
statistics of surprise and coincidence.
COMPUTATIONAL LINGUISTICS, 19(1):61–74.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Timothy</given-names>
            <surname>Baldwin</surname>
          </string-name>
          and Su Nam Kim.
          <year>2010</year>
          .
          <article-title>Multiword expressions</article-title>
          .
          <source>In Handbook of Natural Language Processing</source>
          , second edition., pages
          <fpage>267</fpage>
          -
          <lpage>292</lpage>
          . CRC Press.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Colin</given-names>
            <surname>Bannard</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>A measure of syntactic flexibility for automatically identifying multiword expressions in corpora</article-title>
          .
          <source>In Proceedings of the Workshop on a Broader Perspective on Multiword Expressions</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          and
          <string-name>
            <given-names>Adam</given-names>
            <surname>Kilgarriff</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Large linguistically-processed web corpora for multiple languages</article-title>
          .
          <source>In Proceedings of the Eleventh Conference of the European Chapter of the Association for Computational Linguistics: Demonstrations, EACL '06</source>
          , pages
          <fpage>87</fpage>
          -
          <lpage>90</lpage>
          , Stroudsburg, PA, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Helena Medeiros de Caseli</surname>
          </string-name>
          , Carlos Ramisch,
          <article-title>Maria das Graças Volpe Nunes, and</article-title>
          <string-name>
            <given-names>Aline</given-names>
            <surname>Villavicencio</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Alignment-based extraction of multiword expressions</article-title>
          .
          <source>Language resources and evaluation</source>
          ,
          <volume>44</volume>
          (
          <issue>1- 2</issue>
          ):
          <fpage>59</fpage>
          -
          <lpage>77</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Evert</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Corpora and collocations</article-title>
          .
          <source>In Corpus Linguistics. An International Handbook</source>
          , volume
          <volume>2</volume>
          , pages
          <fpage>1212</fpage>
          -
          <lpage>1248</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Afsaneh</given-names>
            <surname>Fazly</surname>
          </string-name>
          , Suzanne Stevenson, and Ryan North.
          <year>2007</year>
          .
          <article-title>Automatically learning semantic knowledge about multiword predicates</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>41</volume>
          (
          <issue>1</issue>
          ):
          <fpage>61</fpage>
          -
          <lpage>89</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Afsaneh</given-names>
            <surname>Fazly</surname>
          </string-name>
          , Paul Cook, and
          <string-name>
            <given-names>Suzanne</given-names>
            <surname>Stevenson</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Unsupervised type and token identification of idiomatic expressions</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>35</volume>
          (
          <issue>1</issue>
          ):
          <fpage>61</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Sylviane</given-names>
            <surname>Granger</surname>
          </string-name>
          and
          <string-name>
            <given-names>Fanny</given-names>
            <surname>Meunier</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Phraseology: an interdisciplinary perspective</article-title>
          . John Benjamins Publishing Company.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Adam</given-names>
            <surname>Kilgarriff</surname>
          </string-name>
          , Pavel Rychlý, Pavel Smrz, and
          <string-name>
            <given-names>David</given-names>
            <surname>Tugwell</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>The sketch engine</article-title>
          .
          <source>In EURALEX 2004</source>
          , pages
          <fpage>105</fpage>
          -
          <lpage>116</lpage>
          , Lorient, France.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Brigitte</given-names>
            <surname>Krenn</surname>
          </string-name>
          and
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Evert</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Can we do better than frequency? a case study on extracting pp-verb collocations</article-title>
          .
          <source>Proceedings of the ACL Workshop on Collocations</source>
          , pages
          <fpage>39</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Gyri</given-names>
            <surname>Smørdal</surname>
          </string-name>
          <string-name>
            <surname>Losnegaard</surname>
          </string-name>
          , Federico Sangati, Carla Parra Escartín, Agata Savary, Sascha Bargmann, and
          <string-name>
            <given-names>Johanna</given-names>
            <surname>Monti</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Parseme survey on mwe resources</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ), Paris, France.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Christopher D Manning</surname>
            ,
            <given-names>Prabhakar</given-names>
          </string-name>
          <string-name>
            <surname>Raghavan</surname>
            , and
            <given-names>Hinrich</given-names>
          </string-name>
          <string-name>
            <surname>Schütze</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Introduction to information retrieval</article-title>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Johanna</given-names>
            <surname>Monti</surname>
          </string-name>
          and
          <string-name>
            <given-names>Amalia</given-names>
            <surname>Todirascu</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Multiword units translation evaluation in machine translation: another pain in the neck</article-title>
          ?
          <source>In Proceedings of MUMTTT workshop</source>
          , Corpas Pastor
          <string-name>
            <given-names>G</given-names>
            ,
            <surname>Monti</surname>
          </string-name>
          <string-name>
            <given-names>J</given-names>
            ,
            <surname>Mitkov</surname>
          </string-name>
          <string-name>
            <surname>R</surname>
          </string-name>
          , Seretan V (eds) (
          <year>2015</year>
          ),
          <article-title>Multi-word Units in Machine Translation</article-title>
          and
          <string-name>
            <given-names>Translation</given-names>
            <surname>Technology</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Johanna</given-names>
            <surname>Monti</surname>
          </string-name>
          , Ruslan Mitkov, Gloria Corpas Pastor, and
          <string-name>
            <given-names>Violeta</given-names>
            <surname>Seretan</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Multi-word units in machine translation and translation technologies</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Michael P.</given-names>
            <surname>Oakes</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Statistics for Corpus Linguistics</article-title>
          . Edinburgh: Edinburgh University Press.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Karl</given-names>
            <surname>Pichotta and John DeNero</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Identifying phrasal verbs using many bilingual corpora</article-title>
          .
          <source>In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP</source>
          <year>2013</year>
          ), Seattle, WA, October.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Carlos</given-names>
            <surname>Ramisch</surname>
          </string-name>
          , Aline Villavicencio, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Boitet</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>mwetoolkit: a Framework for Multiword Expression Identification</article-title>
          .
          <source>In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2010</year>
          ), Valetta, Malta, May. European Language Resources Association.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Alexandre</given-names>
            <surname>Rondon</surname>
          </string-name>
          , Helena Caseli, and
          <string-name>
            <given-names>Carlos</given-names>
            <surname>Ramisch</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Never-ending multiword expressions learning</article-title>
          .
          <source>In Proceedings of the 11th Workshop on Multiword Expressions</source>
          , pages
          <fpage>45</fpage>
          -
          <lpage>53</lpage>
          , Denver, Colorado, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Rychlý</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>A lexicographer-friendly association score</article-title>
          .
          <source>In RASLAN 2008</source>
          , pages
          <fpage>6</fpage>
          -
          <lpage>9</lpage>
          , Brno. Masarykova Univerzita.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Bahar</given-names>
            <surname>Salehi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Cook</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Predicting the compositionality of multiword expressions using translations in multiple languages</article-title>
          .
          <source>Second Joint Conference on Lexical and Computational Semantics (* SEM)</source>
          ,
          <volume>1</volume>
          :
          <fpage>266</fpage>
          -
          <lpage>275</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Nathan</given-names>
            <surname>Schneider</surname>
          </string-name>
          , Emily Danchik, Chris Dyer, and
          <string-name>
            <surname>Noah</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
          </string-name>
          . 2014a.
          <article-title>Discriminative lexical semantic segmentation with gaps: Running the MWE gamut</article-title>
          .
          <source>TACL</source>
          ,
          <volume>2</volume>
          :
          <fpage>193</fpage>
          -
          <lpage>206</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Nathan</given-names>
            <surname>Schneider</surname>
          </string-name>
          , Spencer Onuffer, Nora Kazour, Emily Danchik, Michael T. Mordowanec, Henrietta Conrad, and
          <string-name>
            <surname>Noah</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
          </string-name>
          . 2014b.
          <article-title>Comprehensive annotation of multiword expressions in a social web corpus</article-title>
          .
          <source>In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)</source>
          , pages
          <fpage>455</fpage>
          -
          <lpage>461</lpage>
          , Reykjavik, Iceland.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Violeta</given-names>
            <surname>Seretan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Eric</given-names>
            <surname>Wehrli</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Syntactic concordancing and multi-word expression detection</article-title>
          .
          <source>International Journal of Data Mining, Modelling and Management</source>
          ,
          <volume>5</volume>
          (
          <issue>2</issue>
          ):
          <fpage>158</fpage>
          -
          <lpage>181</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Frank</given-names>
            <surname>Smadja</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Retrieving collocations from text: Xtract</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>19</volume>
          :
          <fpage>143</fpage>
          -
          <lpage>177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Shiva</given-names>
            <surname>Taslimipoor</surname>
          </string-name>
          , Ruslan Mitkov, Gloria Corpas Pastor, and
          <string-name>
            <given-names>Afsaneh</given-names>
            <surname>Fazly</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Bilingual contexts from comparable corpora to mine for translations of collocations</article-title>
          .
          <source>In Proceedings of the 17th International Conference on Intelligent Text Processing and Computational Linguistics</source>
          ,
          <source>CICLing'16</source>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Aline</given-names>
            <surname>Villavicencio</surname>
          </string-name>
          , Valia Kordoni, Yi Zhang, Marco Idiart, and
          <string-name>
            <given-names>Carlos</given-names>
            <surname>Ramisch</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Validation and evaluation of automatically acquired multiword expressions for grammar engineering</article-title>
          . In EMNLPCoNLL, pages
          <fpage>1034</fpage>
          -
          <lpage>1043</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Aline</given-names>
            <surname>Villavicencio</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>The availability of verbparticle constructions in lexical resources: How much is enough? Computer Speech</article-title>
          &amp; Language,
          <volume>19</volume>
          (
          <issue>4</issue>
          ):
          <fpage>415</fpage>
          -
          <lpage>432</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>