<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Advances in Multiword Expression Identification for the Italian language: The PARSEME shared task edition 1.1</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Johanna Monti</string-name>
          <email>jmonti@unior.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silvio Ricardo Cordeiro</string-name>
          <email>silvioricardoc@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carlos Ramisch</string-name>
          <email>carlos.ramisch@lis-lab.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Federico Sangati</string-name>
          <email>fsangati@unior.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Agata Savary</string-name>
          <email>agata.savary@univ-tours.fr</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Veronika Vincze</string-name>
          <email>vinczev@inf.u-szeged.hu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aix Marseille Univ</institution>
          ,
          <addr-line>CNRS, LIS, Marseille</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>MTA-SZTE Research Group on Artificial Intelligence</institution>
          ,
          <country country="HU">Hungary</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University L'Orientale</institution>
          ,
          <addr-line>Naples</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Tours</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. This contribution describes the results of the second edition of the shared task on automatic identification of verbal multiword expressions, organized as part of the LAW-MWE-CxG 2018 workshop, co-located with COLING 2018, concerning both the PARSEME-IT corpus and the systems that took part in the task for the Italian language. The paper will focus on the main advances in comparison to the first edition of the task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Multiword expressions (MWEs) are a particularly
challenging linguistic phenomenon to be handled
by NLP tools. In recent years, there has been a
growing interest in MWEs since the possible
improvements of their computational treatment may
help overcome one of the main shortcomings of
many NLP applications, from Text Analytics to
Machine Translation. Recent contributions to this
topic, such as Mitkov et al. (2018) and Constant
et al. (2017) have highlighted the difficulties that
this complex phenomenon, halfway between
lexicon and syntax, characterized by idiosyncrasy on
various levels, poses to NLP tasks.</p>
      <p>This contribution will focus on the advances in
the identification of verbal multiword expressions
(VMWEs) for the Italian language. In Section 2
we discuss related work. In Section 3 we give an
overview of the PARSEME shared task. In Section
4 we present the resources developed for the
Italian language, namely the guidelines and the
corpus. Section 5 is devoted to the annotation
process and the inter-annotator agreement. Section 6
briefly describes the thirteen systems that took part
in the shared task and the results obtained. Finally,
we discuss conclusions and future work (Section
7).
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        MWEs have been the focus of the PARSEME
COST Action, which enabled the organization of
an international and highly multilingual research
community
        <xref ref-type="bibr" rid="ref15">(Savary et al., 2015)</xref>
        . This
community launched in 2017 the first edition of the
PARSEME shared task on automatic
identification of verbal MWEs, aimed at developing
universal terminologies, guidelines and
methodologies for 18 languages, including the Italian
language
        <xref ref-type="bibr" rid="ref1 ref14">(Savary et al., 2017)</xref>
        . The task was
colocated with the 13th Workshop on Multiword
Expressions (MWE 2017), which took place
during the European Chapter of the Association for
Computational Linguistics (EACL 2017). The
main outcomes for the Italian language were the
PARSEME-IT Corpus, a 427-thousand-word
annotated corpus of verbal MWEs in Italian
        <xref ref-type="bibr" rid="ref1 ref11">(Monti
et al., 2017)</xref>
        and the participation of four
systems1, namely TRANSITION, a transition-based
dependency parsing system
        <xref ref-type="bibr" rid="ref1">(Al Saied et al., 2017)</xref>
        ,
SZEGED based on the POS and dependency
modules of the Bohnet parser
        <xref ref-type="bibr" rid="ref1 ref17">(Simko´ et al., 2017)</xref>
        ,
ADAPT
        <xref ref-type="bibr" rid="ref1 ref9">(Maldonado et al., 2017)</xref>
        and RACAI
        <xref ref-type="bibr" rid="ref1 ref5">(Boros¸ et al., 2017)</xref>
        , both based on sequence
la1http://multiword.sourceforge.net/
sharedtaskresults2017
beling with CRFs. Concerning the identification
of verbal MWEs some further recent contributions
specifically focusing on the Italian language are:
A supervised token-based identification
approach to Italian Verb+Noun expressions that
belong to the category of complex
predicates
        <xref ref-type="bibr" rid="ref1 ref19">(Taslimipoor et al., 2017)</xref>
        . The
approach investigates the inclusion of
concordance as part of the feature set used in
supervised classification of MWEs in detecting
literal and idiomatic usages of expressions.
All concordances of the verbs fare (‘to do/ to
make’), dare (‘to give’), prendere (‘to take’)
and trovare (‘to find’) followed by any noun,
taken from the itWaC corpus
        <xref ref-type="bibr" rid="ref2">(Baroni and
Kilgarriff, 2006)</xref>
        using SketchEngine (Kilgarriff
et al., 2004) are considered.
      </p>
      <p>
        A neural network trained to classify and rank
idiomatic expressions under constraints of
data scarcity
        <xref ref-type="bibr" rid="ref1 ref4">(Bizzoni et al., 2017)</xref>
        .
      </p>
      <p>With reference to corpora annotated with VMWEs
for the Italian language and in comparison with the
state of the art described in Monti et al. (2017),
there are no further resources available so far. At
the time of writing, therefore, the PARSEME-IT
VMWE corpus still represents the first sample of
a corpus which includes several types of VMWEs,
specifically developed to foster NLP applications.
The corpus is freely available, with the latest
version (1.1) representing an enhanced corpus with
some substantial changes in comparison with
version 1.0 (cf. Section 4).
3</p>
    </sec>
    <sec id="sec-3">
      <title>The PARSEME shared task</title>
      <p>The second edition of the PARSEME shared task
on automatic identification of verbal multiword
expressions (VMWEs) was organized as part of
the LAW-MWE-CxG 2018 workshop co-located
with COLING 2018 (Santa Fe, USA)2 and aimed
at identifying verbal MWEs in running texts.
According to the rules set forth in the shared task,
system results could be submitted in two tracks:
CLOSED TRACK: Systems using only the
provided training/development data - VMWE
annotations + morpho-syntactic data (if any)
- to learn VMWE identification models
and/or rules.</p>
      <p>2https:http://multiword.sourceforge.
net/lawmwecxg2018
OPEN TRACK: Systems using or not the
provided training/development data, plus any
additional resources deemed useful (MWE
lexicons, symbolic grammars, wordnets, raw
corpora, word embeddings, language
models trained on external data, etc.). This track
includes notably purely symbolic and
rulebased systems.</p>
      <p>The PARSEME members elaborated for each
language i) annotation guidelines based on annotation
experiments ii) corpora in which VMWEs are
annotated according to the guidelines. Corpora were
split in training, development and tests corpora for
each language. Manually annotated training and
development corpora were made available to the
participants in advance, in order to allow them to
train their systems and to tune/optimize the
systems’ parameters. Raw (unannotated) test corpora
were used as input to the systems during the
evaluation phase. The contribution of the
PARSEMEIT research group3 to the shared task is described
in the next section.</p>
    </sec>
    <sec id="sec-4">
      <title>4 Italian resources for the shared task</title>
      <p>The PARSEME-IT research group contributed to
the edition 1.1 of the shared task with the
development of specific guidelines for the Italian language
and with the annotation of the Italian corpus with
over 3,700 VMWEs.
4.1</p>
      <sec id="sec-4-1">
        <title>The shared task guidelines</title>
        <p>
          The 2018 edition of the shared task relied on
enhanced and revised guidelines
          <xref ref-type="bibr" rid="ref12">(Ramisch et al.,
2018)</xref>
          . The guidelines4 are provided with Italian
examples for each category of VMWE.
        </p>
        <p>The guidelines include two universal categories,
i.e. valid for all languages participating in the task:</p>
      </sec>
      <sec id="sec-4-2">
        <title>Light-verb constructions (LVCs) with two</title>
        <p>subcategories: LVCs in which the verb is
semantically totally bleached (LVC.full) like
in fare un discorso (‘to give a speech’), and
LVCs in which the verb adds a causative
meaning to the noun (LVC.cause) like in dare
il mal di testa (‘to give a headache’);</p>
      </sec>
      <sec id="sec-4-3">
        <title>Verbal idioms (VIDs) like gettare le perle ai</title>
        <p>porci (‘to throw pearls before swine’).</p>
        <p>3https://sites.google.com/view/
parseme-it/home</p>
        <p>4http://parsemefr.lif.univ-mrs.fr/
parseme-st-guidelines/1.1/
Three quasi-universal categories, valid for some
language groups or languages but non-existent or
very exceptional in others are:</p>
        <p>Inherently reflexive verbs (IRV) which are
those reflexive verbal constructions which
(a) never occur without the clitic e.g.
suicidarsi (‘to suicide’), or when (b) the IRV
and non-reflexive versions have clearly
different senses or subcategorization frames e.g.
riferirsi (‘to refer’) opposed to riferire (‘to
report / to tell’);</p>
      </sec>
      <sec id="sec-4-4">
        <title>Verb-particle constructions (VPC) with</title>
        <p>two subcategories: fully non-compositional
VPCs (VPC.full), in which the particle
totally changes the meaning of the verb, like
buttare giu` (‘to swallow’) and semi
noncompositional VPCs (VPC.semi), in which
the particle adds a partly predictable but
nonspatial meaning to the verb like in andare
avanti (‘to proceed’);</p>
      </sec>
      <sec id="sec-4-5">
        <title>Multi-verb constructions (MVC) com</title>
        <p>posed by a sequence of two adjacent verbs
like in lasciar perdere (‘to give up’).</p>
        <p>An optional experimental category (if admitted
by the given language, as is the case for Italian) is
considered in a post-annotation step:</p>
        <p>Inherently adpositional verbs (IAVs),
which consist of a verb or VMWE and an
idiomatic selected preposition or
postposition that is either always required or, if
absent, changes the meaning of the verb
significantly, like in confidare su (‘to trust
on’).</p>
        <p>Finally, a language-specific category was
introduced for the Italian language:</p>
        <p>Inherently clitic verbs (LS.ICV) formed by
a full verb combined with one or more
nonreflexive clitics that represent the
pronominalization of one or more complements
(CLI). LS.ICV is annotated when (a) the verb
never occurs without one non-reflexive clitic,
like in entrarci (‘to be relevant to
something’), or (b) when the LS.ICV and the
nonclitic versions have clearly different senses
or subcategorization frames like in prenderle
(‘to be beaten’) vs prendere (‘to take’).
4.2</p>
      </sec>
      <sec id="sec-4-6">
        <title>The PARSEME-IT corpus</title>
        <p>
          The PARSEME-IT VMWE corpus version 1.1 is
an updated version of the corpus used for edition
1.0 of the shared task. It is based on a selection
of texts from the PAIS A` corpus of web texts
          <xref ref-type="bibr" rid="ref8">(Lyding et al., 2014)</xref>
          , including Wikibooks, Wikinews,
Wikiversity, and blog services. The
PARSEMEIT VMWE corpus was updated in edition 1.1
according to the new guidelines described in the
previous section. Table 4.2 summarizes the size of
the corpus developed for the Italian language and
presents the distribution of the annotated VMWEs
per category.
        </p>
        <p>The training, development and test data are
available in the LINDAT/Clarin repository5, and
all VMWE annotations are available under
Creative Commons licenses (see README.md files
for details). The released corpus’ format is based
on an extension of the widely-used CoNLL-U file
format.6
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Annotation process</title>
      <p>
        The annotation was manually performed in
running texts using the FoLiA linguistic annotation
tool7
        <xref ref-type="bibr" rid="ref20">(van Gompel and Reynaert, 2013)</xref>
        by six
Italian native speakers with a background in
linguistics, using a specific decision tree for the Italian
language for joint VMWE identification and
classification.8
      </p>
      <p>In order to allow the annotation of IAVs, a new
pre-processing step was introduced to split
compound prepositions such as della (‘of the’) into two
tokens. This step was necessary to annotate only
lexicalised components of the IAV, as in portare
alla disperazione, where only the verb and the
preposition a should be annotated, without the
article la.</p>
      <p>Once the annotation was completed, in order to
reduce noise and to increase the consistency of the
annotations, we applied the consistency checking
tool developed for edition 1.0 (Savary et al.,
forthcoming). The tool groups all annotations of the
same VMWE, making it possible to spot
annotation inconsistencies very easily.</p>
      <p>5http://hdl.handle.net/11372/LRT-2842
6http://multiword.sf.net/cupt-format
7http://mwe.phil.hhu.de/
8http://parsemefr.lif.univ-mrs.fr/
parseme-st-guidelines/1.1/?page=itdectree</p>
      <sec id="sec-5-1">
        <title>5.1 Inter-annotator agreement</title>
        <p>A small portion of the corpus consisting in 1,000
sentences was double-annotated. In
comparison with the previous edition, the inter-annotator
agreement shown in Table 2 increased, although it
is still not optimal.9 The improvement is probably
due to the fact that, this time, the group was based
in one place with the exception of one annotator,
and several meetings took place prior to the
annotation phase in order to discuss the new guidelines.</p>
        <p>The two annotators involved in the IAA task
annotated 191 VMWEs with no disagreement, but
there were several problems, which led to 44 cases
of partial disagreement and 250 cases of total
disagreement:</p>
        <p>PARTIAL MATCHES LABELED, (25 cases)
in which there is at least one token of the
VMWE in common between two annotators
and the labels assigned are the same. The
disagreement mainly concerns the lexicalized
elements as part of the VMWE, as in the case
of the VID porre in cattiva luce (‘make look
bad’). Annotators disagreed, indeed, about
considering the adjective cattiva (‘bad’) as
9As mentioned in Ramisch et al. (2018), the estimation of
chance agreement in span and cat is slightly different
between 2017 and 2018, therefore these results are not directly
comparable.
part of the VID.</p>
        <p>EXACT MATCHES UNLABELED, (18 cases) in
which the annotators agreed on the
lexicalized components of the VMWE to be
annotated but not the label. This type of
disagreement is mainly related to fine-grained
categories such as LVC.cause and LVC.full as
in the case of dare . . . segnale (to give . . .
a signal) or VPC.full and VPC.semi as for
mettere insieme (‘to put together’)
PARTIAL MATCHES UNLABELED, (1 case)
in which there is at least one token of the
VMWE in common between two annotators
but the labels assigned are different, such as
in buttar-si in la calca (‘to join the crowd’)
classified as VID by the first annotator and
buttar-si (‘to throw oneself’) classified as
IRV by the second one in the following
sentence: [. . . ] attendendo il venerd`ı sera per
buttarsi nella calca del divertimento [. . . ].
(‘waiting for the Friday evening to join the
crowd for entertainment’)</p>
        <p>ANNOTATIONS CARRIED OUT ONLY BY
ONE OF THE ANNOTATORS: This is the
category which collects the most numerous
examples of disagremeent between annotators:
106 VMWE were annotated only by
annotator 1 and 144 by annotator 2.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>The systems and the results of the shared task for the Italian language</title>
      <p>
        Whereas only four systems took part in edition 1.0
of the shared task for the Italian language, in
edition 1.1, fourteen systems took on this challenge.
The system that took part in the PARSEME shared
task are listed in Table 3: 12 took part in the closed
track and two in the open one. The two systems
that took part in the open track reported the
resources that were used, namely SHOMA used
pretrained wikipedia word embeddings
        <xref ref-type="bibr" rid="ref18 ref3">(Taslimipoor
and Rohanian, 2018)</xref>
        , while Deep-BGT
        <xref ref-type="bibr" rid="ref3">(Berk
et al., 2018)</xref>
        relied on the BIO tagging scheme
and its variants
        <xref ref-type="bibr" rid="ref16">(Schneider et al., 2014)</xref>
        to
introduce additional tags to encode gappy
(discontinuous) VMWEs. A distinctive characteristic of the
systems of edition 1.1 is that most of them
(GBDNER-resplit and GBD-NER-standard, TRAPACC,
and TRAPACC-S, SHOMA, Deep-BGT) use
neural networks, while the rest of the systems adopt
other approaches: CRF-DepTree-categs and
CRFSeq-nocategs are based on a tree-structured CRF,
MWETreeC and TRAVERSAL on syntactic trees
and parsing methods, Polirem-basic and
Poliremrich on statistical methods and association
measures, and finally varIDE uses a Naive Bayes
classifier. The systems were ranked according
two types of evaluation measures
        <xref ref-type="bibr" rid="ref12">(Ramisch et al.,
2018)</xref>
        : a strict per-VMWE score (in which each
VMWE in gold is either deemed predicted or not,
in a binary fashion) and a fuzzy per-token score
(which takes partial matches into account). For
each of these two, precision (P), recall (R) and
F1-scores (F) were calculated. Table 3 shows the
ranking of the systems which participated in the
shared task for the Italian language. The
systems with highest MWE-based Rank for Italian
have F1 scores that are mostly comparable to the
scores obtained in the General ranking of all
languages (e.g. TRAVERSAL had a General F1 of
54.0 vs Italian F1 of 49.2, being ranked first in
both cases). Nevertheless, the Italian scores are
consistently lower than the ones in the General
ranking, even if only by a moderate margin,
suggesting that Italian VMWEs in this specific corpus
might be particularly harder to identify. One of the
outliers in the table is MWETreeC, which predicts
much fewer VMWEs than in the annotated
corpora. This turned out to be true for other languages
as well. The few VMWEs that were predicted only
obtained partial matches, which explains why its
MWE-based score was 0. Another clear outlier is
Polirem-basic. Both Polirem-basic and
Poliremrich had predictions for Italian, French and
Portuguese. Their scores are somewhat comparable
in the three languages, suggesting that the lower
scores are a characteristic of the system and not
some artifact of the Italian corpus.
      </p>
      <p>
        TRASVERSAL
        <xref ref-type="bibr" rid="ref21">(Waszczuk, 2018)</xref>
        was the best
performing system in the closed track, while
SHOMA
        <xref ref-type="bibr" rid="ref18 ref3">(Taslimipoor and Rohanian, 2018)</xref>
        performed best in the open one. As shown in
Figure 1, comparing the MWE-based F1 scores for
each label for the two best performing systems,
TRASVERSAL obtained overall better results for
almost all VMWEs categories with the exception
of VID and MVC, for which SHOMA showed a
better performance.
Having presented the results of the PARSEME
shared task edition 1.1, the paper described the
advances achieved in this last edition in
comparison with the previous one, but also highlighted
that there is room for further improvements. We
are working on some critical areas which emerged
during the annotation task in particular with
reference to some borderline cases and the refinement
of the guidelines. Future work will focus on
maintaining and increasing the quality and the size of
the corpus but also on extending the shared task to
other MWE categories, such as nominal MWEs.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>Our thanks go to the Italian annotators Valeria
Caruso, Maria Pia di Buono, Antonio Pascucci,
Annalisa Raffone, Anna Riccio for their
contributions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Hazem</given-names>
            <surname>Al</surname>
          </string-name>
          <string-name>
            <surname>Saied</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Matthieu</given-names>
            <surname>Constant</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Marie</given-names>
            <surname>Candito</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>The ATILF-LLF system for PARSEME shared task: a transition-based verbal multiword expression tagger</article-title>
          .
          <source>In Proceedings of the 13th Workshop on Multiword Expressions (MWE</source>
          <year>2017</year>
          ), pages
          <fpage>127</fpage>
          -
          <lpage>132</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          and
          <string-name>
            <given-names>Adam</given-names>
            <surname>Kilgarriff</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Large linguistically-processed web corpora for multiple languages</article-title>
          .
          <source>In Proceedings of the Eleventh Conference of the European Chapter of the Association for Computational Linguistics: Posters &amp; Demonstrations</source>
          , pages
          <fpage>87</fpage>
          -
          <lpage>90</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>Go¨zde Berk, Berna Erden</article-title>
          , and Tunga Gu¨ngo¨r.
          <year>2018</year>
          .
          <article-title>Deep-bgt at parseme shared task 2018: Bidirectional lstm-crf model for verbal multiword expression identification</article-title>
          .
          <source>In Proceedings of the Joint Workshop on Linguistic Annotation</source>
          , Multiword Expressions and
          <string-name>
            <surname>Constructions (LAW-MWE-CxG-</surname>
          </string-name>
          2018), pages
          <fpage>248</fpage>
          -
          <lpage>253</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Yuri</given-names>
            <surname>Bizzoni</surname>
          </string-name>
          , Marco S. G. Senaldi, and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Deep-learning the Ropes: Modeling Idiomaticity with Neural Networks</article-title>
          .
          <source>In Proceedings of the Fourth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2017</year>
          ), Rome, Italy,
          <source>December 11-13</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Tiberiu</surname>
            <given-names>Boros¸</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sonia</surname>
            <given-names>Pipa</given-names>
          </string-name>
          , Verginica Barbu Mititelu, and Dan Tufis¸.
          <year>2017</year>
          .
          <article-title>A data-driven approach to verbal multiword expression detection. PARSEME Shared Task system description paper</article-title>
          .
          <source>In Proceedings of the 13th Workshop on Multiword Expressions (MWE</source>
          <year>2017</year>
          ), pages
          <fpage>121</fpage>
          -
          <lpage>126</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Mathieu</given-names>
            <surname>Constant</surname>
          </string-name>
          , Gu¨ls¸en Eryig˘it, Johanna Monti, Lonneke van der Plas, Carlos Ramisch,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Rosner</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Amalia</given-names>
            <surname>Todirascu</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <source>Multiword Expression Processing: A Survey. Computational Linguistics</source>
          ,
          <volume>43</volume>
          (
          <issue>4</issue>
          ):
          <fpage>837</fpage>
          -
          <lpage>892</lpage>
          . URL https://doi.org/10.1162/ COLI_a_
          <fpage>00302</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>David</given-names>
            <surname>Tugwell</surname>
          </string-name>
          .
          <year>2004</year>
          . Itri-04
          <article-title>-08 the sketch engine</article-title>
          .
          <source>Information Technology</source>
          ,
          <volume>105</volume>
          :
          <fpage>116</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Verena</given-names>
            <surname>Lyding</surname>
          </string-name>
          , Egon Stemle, Claudia Borghetti, Marco Brunello, Sara Castagnoli, Felice Dell'Orletta, Henrik Dittmann, Alessandro Lenci, and
          <string-name>
            <given-names>Vito</given-names>
            <surname>Pirrelli</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The PAISA` Corpus of Italian Web Texts</article-title>
          .
          <source>In Proceedings of the 9th Web as Corpus Workshop (WaC-9)</source>
          , pages
          <fpage>36</fpage>
          -
          <lpage>43</lpage>
          . Association for Computational Linguistics, Gothenburg, Sweden. URL http://www.aclweb.org/ anthology/W14-0406.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Alfredo</given-names>
            <surname>Maldonado</surname>
          </string-name>
          , Lifeng Han, Erwan Moreau, Ashjan Alsulaimani, Koel Chowdhury, Carl Vogel, and Qun Liu.
          <year>2017</year>
          .
          <article-title>Detection of Verbal Multi-Word Expressions via Conditional Random Fields with Syntactic Dependency Features and Semantic Re-Ranking</article-title>
          .
          <source>In Proceedings of the 13th Workshop on Multiword Expressions (MWE</source>
          <year>2017</year>
          ), pages
          <fpage>114</fpage>
          -
          <lpage>120</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Mitkov</surname>
          </string-name>
          , Johanna Monti, Gloria Corpas Pastor, and
          <string-name>
            <given-names>Violeta</given-names>
            <surname>Seretan</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Multiword units in machine translation and translation technology</article-title>
          , volume
          <volume>341</volume>
          . John Benjamins Publishing Company.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Johanna</given-names>
            <surname>Monti</surname>
          </string-name>
          , Maria Pia di Buono, and
          <string-name>
            <given-names>Federico</given-names>
            <surname>Sangati</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>PARSEME-IT Corpus</article-title>
          .
          <article-title>An annotated Corpus of Verbal Multiword Expressions in Italian</article-title>
          .
          <source>In Fourth Italian Conference on Computational Linguistics-CLiC-it</source>
          <year>2017</year>
          , pages
          <fpage>228</fpage>
          -
          <lpage>233</lpage>
          . Accademia University Press.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Carlos</given-names>
            <surname>Ramisch</surname>
          </string-name>
          , Silvio Ricardo Cordeiro, Agata Savary, Veronika Vincze, Verginica Barbu Mititelu, Archna Bhatia, Maja Buljan, Marie Candito, Polona Gantar, Voula Giouli, Tunga Gu¨ngo¨r, Abdelati Hawwari, Uxoa In˜urrieta, Jolanta Kovalevskaite˙,
          <string-name>
            <surname>Simon</surname>
            <given-names>Krek</given-names>
          </string-name>
          , Timm Lichte, Chaya Liebeskind, Johanna Monti, Carla Parra Escart´ın,
          <string-name>
            <surname>Behrang</surname>
            <given-names>QasemiZadeh</given-names>
          </string-name>
          , Renata Ramisch, Nathan Schneider, Ivelina Stoyanova, Ashwini Vaidya, and
          <string-name>
            <given-names>Abigail</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Edition 1.1 of the PARSEME Shared Task on Automatic Identification of Verbal Multiword Expressions</article-title>
          . In the Joint Workshop on Linguistic Annotation,
          <article-title>Multiword Expressions and Constructions (LAW-MWE-CxG2018)</article-title>
          , pages
          <fpage>222</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <article-title>forthcoming. PARSEME multilingual corpus of verbal multiword expressions</article-title>
          . In Stella Markantonatou, Carlos Ramisch, Agata Savary, and Veronika Vincze, editors,
          <article-title>Multiword expressions at length and in depth. Extended papers from the MWE 2017 workshop</article-title>
          . Language Science Press, Berlin, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Agata</given-names>
            <surname>Savary</surname>
          </string-name>
          , Carlos Ramisch, Silvio Cordeiro, Federico Sangati, Veronika Vincze,
          <string-name>
            <surname>Behrang</surname>
            <given-names>QasemiZadeh</given-names>
          </string-name>
          , Marie Candito, Fabienne Cap, Voula Giouli, Ivelina Stoyanova, and
          <string-name>
            <given-names>Antoine</given-names>
            <surname>Doucet</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>The parseme shared task on automatic identification of verbal multiword expressions</article-title>
          .
          <source>In Proceedings of the 13th Workshop on Multiword Expressions (MWE</source>
          <year>2017</year>
          ), pages
          <fpage>31</fpage>
          -
          <lpage>47</lpage>
          . Association for Computational Linguistics, Valencia, Spain. URL http://www. aclweb.org/anthology/W17-1704.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Agata</given-names>
            <surname>Savary</surname>
          </string-name>
          , Manfred Sailer, Yannick Parmentier, Michael Rosner, Victoria Rose´n, Adam Przepio´rkowski, Cvetana Krstev, Veronika Vincze, Beata Wo´jtowicz, Gyri Smørdal Losnegaard, Carla Parra Escart´ın, Jakub Waszczuk, Mathieu Constant, Petya Osenova, and
          <string-name>
            <given-names>Federico</given-names>
            <surname>Sangati</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>PARSEME - PARSing and Multiword Expressions within a European multilingual network</article-title>
          .
          <source>In 7th Language</source>
          &amp; Technology Conference:
          <article-title>Human Language Technologies as a Challenge for Computer Science and Linguistics (LTC</article-title>
          <year>2015</year>
          ). Poznan´, Poland. URL https://hal.archivesouvertes.fr/hal-01223349.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Nathan</given-names>
            <surname>Schneider</surname>
          </string-name>
          , Emily Danchik,
          <source>Chris Dyer, and Noah A Smith</source>
          .
          <year>2014</year>
          .
          <article-title>Discriminative lexical semantic segmentation with gaps: running the mwe gamut</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          ,
          <volume>2</volume>
          :
          <fpage>193</fpage>
          -
          <lpage>206</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Katalin</given-names>
            <surname>Ilona</surname>
          </string-name>
          <string-name>
            <surname>Simko´</surname>
          </string-name>
          ,
          <source>Vikto´ria Kova´cs, and Veronika Vincze</source>
          .
          <year>2017</year>
          .
          <article-title>USzeged: Identifying Verbal Multiword Expressions with POS Tagging and Parsing Techniques</article-title>
          .
          <source>In Proceedings of the 13th Workshop on Multiword Expressions (MWE</source>
          <year>2017</year>
          ), pages
          <fpage>48</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Shiva</given-names>
            <surname>Taslimipoor</surname>
          </string-name>
          and
          <string-name>
            <given-names>Omid</given-names>
            <surname>Rohanian</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Shoma at parseme shared task on automatic identification of vmwes: Neural multiword expression tagging with high generalisation</article-title>
          . arXiv preprint arXiv:
          <year>1809</year>
          .03056.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Shiva</given-names>
            <surname>Taslimipoor</surname>
          </string-name>
          , Omid Rohanian, Ruslan Mitkov, and
          <string-name>
            <given-names>Afsaneh</given-names>
            <surname>Fazly</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Investigating the Opacity of Verb-Noun Multiword Expression Usages in Context</article-title>
          .
          <source>In Proceedings of the 13th Workshop on Multiword Expressions (MWE</source>
          <year>2017</year>
          ), pages
          <fpage>133</fpage>
          -
          <lpage>138</lpage>
          .
          <article-title>Association for Computational Linguistics</article-title>
          . URL http: //aclweb.org/anthology/W17-1718.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Maarten van Gompel</surname>
            and
            <given-names>Martin</given-names>
          </string-name>
          <string-name>
            <surname>Reynaert</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>FoLiA: A practical XML Format for Linguistic Annotation-a descriptive and comparative study</article-title>
          .
          <source>Computational Linguistics in the Netherlands Journal</source>
          ,
          <volume>3</volume>
          :
          <fpage>63</fpage>
          -
          <lpage>81</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Jakub</given-names>
            <surname>Waszczuk</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Traversal at parseme shared task 2018: Identification of verbal multiword expressions using a discriminative treestructured model</article-title>
          .
          <source>In Proceedings of the Joint Workshop on Linguistic Annotation</source>
          , Multiword Expressions and
          <string-name>
            <surname>Constructions (LAW-MWECxG-</surname>
          </string-name>
          2018), pages
          <fpage>275</fpage>
          -
          <lpage>282</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>