<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards an Italian Learner Treebank in Universal Dependencies</article-title>
      </title-group>
      <abstract>
        <p>In this paper we describe the preliminary work on a novel treebank which includes texts written by learners of Italian drawn from the VALICO corpus. Data processing mostly involved the application of Universal Dependencies formalism and error annotation. First, we parsed the texts on UDPipe trained on the existent Italian UD treebanks, then we manually corrected them. The particular focus of this paper is on a one-hundred-sentence sample of the collection, used as a case study to define an annotation scheme for identifying the linguistic phenomena characterizing learners' interlanguage.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        The increasing interest in Learner Corpora
(henceforth LC) is twofold motivated. On the one hand,
LC are an especially valuable source of
knowledge for interlanguage varieties. They allow
indepth comparisons of non-native varieties,
helping to elucidate the properties of the
interlanguage developed by learners with different mother
tongues and learning levels. For this reason, LC
are important resources enabling data-driven
studies exploited within several research areas, such
as Second Language Acquisition, Foreign
Language Teaching, Contrastive Interlanguage
Analysis, Computer-aided Error Analysis,
ComputerAssisted Language Learning and L2
Lexicography (e.g.
        <xref ref-type="bibr" rid="ref15 ref26 ref31 ref37 ref38">(Pravec, 2002; Granger, 2008; McEnery
and Xiao, 2011)</xref>
        ). On the other hand, LC have
raised considerable computational interest, which
is closely related to their usefulness in tasks
such as Native Language Identification (Jarvis
      </p>
      <p>
        Copyright c 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
and Paquot, 2015; Malmasi, 2016),
GrammaticalError Detection and Correction
        <xref ref-type="bibr" rid="ref19 ref27">(Leacock et al.,
2015; Ng et al., 2014)</xref>
        , and Automated Essay
Scoring
        <xref ref-type="bibr" rid="ref16">(Higgins et al., 2015)</xref>
        .
      </p>
      <p>
        In this paper we describe the development of a
novel learner Italian treebank, i.e. VALICO-UD,
in which Universal Dependencies (UD)
formalism is tied to error annotation. The considerations
of the annotation process, carried out on a set of
one hundred sentences selected from a subcorpus
of VALICO1 (see Table 1)
        <xref ref-type="bibr" rid="ref20 ref7">(Corino and Marello,
2017)</xref>
        , allowed us to test a pilot scheme which
pinpoints some of the features of L2 Italian.
      </p>
      <p>This paper is organized as follows: in Section 2
we provide an overview of LC, focusing on
Italian resources in particular; in Section 3 we present
the data and the error annotation of VALICO-UD;
in Section 4 we offer some examples of how we
applied literal annotation to the learner sentences
(LS) and, finally, in Section 6 we present
conclusion and future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        LC, also called interlanguage or L2 corpora, are
collections of data produced by foreign or
second language learners
        <xref ref-type="bibr" rid="ref15">(Granger, 2008)</xref>
        . Most LC
projects were launched in the nineties and focused
mainly on learner English
        <xref ref-type="bibr" rid="ref25">(Tono, 2003)</xref>
        , but
recently we have witnessed an increasing interest
in LC for other target languages. This has
contributed to the establishment of learner corpus
research
        <xref ref-type="bibr" rid="ref25">(Tono, 2003)</xref>
        .
      </p>
      <p>
        LC can be enriched with Part of Speech (PoS)
tagging, syntactic, semantic, discourse structure
and error-tagging (with explicit or implicit target
hypotheses2) annotation
        <xref ref-type="bibr" rid="ref12">(Garside et al., 1997)</xref>
        . To
provide linguistic annotation, NLP tools are
often used
        <xref ref-type="bibr" rid="ref17">(Huang et al., 2018)</xref>
        and combined with
      </p>
      <sec id="sec-2-1">
        <title>1http://www.valico.org/</title>
        <p>
          2A reconstructed LS on which error identification is based
          <xref ref-type="bibr" rid="ref33">(Reznicek et al., 2013)</xref>
          .
human post-editing in order to overcome issues
arising from the failures of the automatic
analysis
          <xref ref-type="bibr" rid="ref13 ref14 ref9">(Geertzen et al., 2013; Granger et al., 2009;
Dahlmeier et al., 2013)</xref>
          .
        </p>
        <p>
          Among the 14 learner Italian corpora registered
in the Learner Corpora around the World list3,
the majority are in the form of plain texts, or they
only annotate PoS (COLI, LOCCLI and CAIL24,
and VALICO), while only MERLIN
          <xref ref-type="bibr" rid="ref5">(Boyd et al.,
2014)</xref>
          annotates syntax and errors (with explicit
target hypotheses).
        </p>
        <p>
          Although MERLIN contains 816 texts written
in non-native Italian
          <xref ref-type="bibr" rid="ref5">(Boyd et al., 2014)</xref>
          , they are
not balanced for learners’ mother tongue and are
not annotated using a standard annotation for
syntax, which would allow comparisons with other
resources. To fill this gap, we decided to develop
VALICO-UD, a L1-balanced resource developed
within the UD formalism, thus providing a greater
potential for contrastive analysis. Indeed, a
UDannotated LC can be compared with other LC
(therefore different interlanguages) or also with
native corpora of the L1 involved. For all these
reasons, we decided to develop this new learner
Italian treebank within the UD formalism.
References were the English and Chinese experiences,
respectively the English Second Language (ESL)
          <xref ref-type="bibr" rid="ref2">(Berzak et al., 2016)</xref>
          and the Chinese Foreign
Language (CFL)
          <xref ref-type="bibr" rid="ref20">(Lee et al., 2017)</xref>
          treebanks.
        </p>
        <p>
          The scholars involved in the annotation of the
ESL and CFL treebanks decided to follow a
wellestablished line of work, for which learner
language analysis is centered upon morpho-syntactic
surface evidence. This is motivated by various
studies, e.g.
          <xref ref-type="bibr" rid="ref22 ref32">(D´ıaz-Negrillo et al., 2010; Ragheb
and Dickinson, 2012)</xref>
          , in which the difference
between morphological and distributional PoS is
stressed. We decided to follow this line of research
annotating discrepancies between morphological
and distributional PoS, as described in the next
sections. However, in lieu of carrying out manual
annotation from scratch, such as in the ESL, we
combined automatic annotation and manual
postediting (as shown in the next section).
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data and annotation</title>
      <p>
        The data of VALICO-UD are drawn from the
VALICO corpus
        <xref ref-type="bibr" rid="ref20 ref7">(Corino and Marello, 2017)</xref>
        , a
3https://uclouvain.be/en/researchinstitutes/ilc/cecl/learner-corpora-around-the-world.html.
      </p>
      <p>4COLI, LOCCLI and CAIL2 are developed at Universita`
per Stranieri di Perugia and coordinated by Stefania Spina.
collection of non-native Italian texts elicited by
comic strips proposed to the learners. It consists of
a selection of narrative and descriptive texts
providing a large variety of structures beyond simple
presentative/existential constructions.</p>
      <p>The portion of VALICO that we selected for the
treebank is made up of 237 texts (2,261 LS)
organized in four sections as shown in Table 1.</p>
      <p>L1</p>
      <sec id="sec-3-1">
        <title>English (EN)</title>
        <p>French (FR)
German (DE)</p>
        <p>Spanish (ES)
EN+FR+DE+ES
# Texts</p>
        <p># LS Tokens</p>
        <p>
          Although the unpredictability and variation of
a learner product, in terms of vocabulary,
morphology and syntax, makes parsing a LC an
especially challenging task
          <xref ref-type="bibr" rid="ref22 ref23 ref36 ref8">(Corino and Russo, 2016;
D´ıaz-Negrillo et al., 2010)</xref>
          , it is highly
recommendable for smoothly retrieving interlanguage
features. Due to this peculiarity of interlanguage,
keeping separated the LS from its specifically built
target hypothesis (TH) is highly recommended
          <xref ref-type="bibr" rid="ref21">(Lu¨deling et al., 2005)</xref>
          .
        </p>
        <p>
          Our annotation scheme for learner Italian uses
the inventory of the Italian UD PoS tags and
dependency relations
          <xref ref-type="bibr" rid="ref3 ref4">(Bosco et al., 2013; Bosco et
al., 2014)</xref>
          and the related guidelines. In addition,
we tried to follow as much as possible the ESL
treebank to have comparable resources.
        </p>
        <p>
          First, we trained UDPipe
          <xref ref-type="bibr" rid="ref36">(Straka et al., 2016)</xref>
          on the Italian UD corpora, which include
standard texts, ISDT
          <xref ref-type="bibr" rid="ref4">(Bosco et al., 2014)</xref>
          , and Twitter
posts, POSTWITA-UD
          <xref ref-type="bibr" rid="ref34">(Sanguinetti et al., 2018)</xref>
          .
Second, we automatically parsed VALICO-UD.
Third, we manually corrected the treebank. This
step is currently ongoing and we envision the
treebank to be released in the UD repository in a few
months.
        </p>
        <p>For each sentence in VALICO-UD we provide
two distinct versions both annotated in UD and
tied to an error encoding system (see Section 3.1):
one version for the LS and the other for its TH.
The latter will differ from the former only when
some errors occur. As a trial for this scheme, we
selected one hundred sentences (i.e. sample set)
containing each at least one error to be annotated.</p>
        <sec id="sec-3-1-1">
          <title>3.1 Error Annotation</title>
          <p>
            In writing the TH we decided to adhere as much as
possible to the LS and to focus on linguistic
correctness (e.g. grammaticality) rather than
linguistic appropriateness (e.g. register)
            <xref ref-type="bibr" rid="ref33">(Reznicek et al.,
2013)</xref>
            5. For this reason, sometimes we sacrificed
naturalness for the sake of adherence to the LS.
This principle was applied also to lexical errors
requiring replacement. For instance, in Figure 1, the
term “rubadore” in the LS was replaced with
“rubatore” and not with its more common synonym
“ladro”, thief.6 With this principle in mind, we
decided to correct words if they are not present
neither in the VINCA corpus7 (the reference corpus
specifically compiled for VALICO and containing
texts based on the same comic strips but written by
Italian native speakers) nor in our reference
dictionary, Il Nuovo Vocabolario di Base della Lingua
Italiana
            <xref ref-type="bibr" rid="ref10">(De Mauro, 2016)</xref>
            . In fact, the VINCA
corpus is quite small and the language used sounds
quite unnatural though being produced by
speakers whose mother tongue is namely Italian (see
Corino and Marello (2017, p. 12)).
          </p>
          <p>
            Once the target hypotheses are written, we
applied to them a coding system based on Nicholls
(2003), which was used also in the NUCLE
            <xref ref-type="bibr" rid="ref9">(Dahlmeier et al., 2013)</xref>
            and FCE
            <xref ref-type="bibr" rid="ref38">(Yannakoudakis
et al., 2011)</xref>
            corpora. Our system follows
Nicholls’s same principle: “the first letter
repre5In the future we plan to provide a second TH, focusing
on linguistic appropriateness.
          </p>
          <p>6Although “rubadore” is reported and marked as obsolete
in the Italian Dictionary Olivetti, “rubatore” is the variant
reported in De Mauro (2016), our reference dictionary.
7http://www.valico.org/vinca.html
sents the general type of error (e.g. wrong form,
omission), while the second letter identifies the
word class of the required word”.</p>
          <p>
            To provide a finer-grained description of errors,
we used a large variety of letters in the first and
second position (e.g. I: inflection, X: auxiliary)
and a third letter which encodes information about
some grammatical features (e.g. T: tense, M:
mood, G: gender)
            <xref ref-type="bibr" rid="ref35">(Simone, 2008, pp. 303–346)</xref>
            and other phenomena involved (e.g.
capitalization, language transfer and government). Finally,
Nicholls included a catch-all code (CE: complex
error) to cover complex, multiple errors. In our
sample set, we did not use it because we managed
to describe all errors encountered using nested
XML tags. However, we do not exclude that,
applying the error codes to the whole corpus, we
might find particularly complex errors which need
to be marked using this code.
          </p>
          <p>Figure 1 shows an annotation example of a LS
along with its corresponding TH in the typical
CoNLL-U format and with the resource-specific
fields used to encode the error information. The
sent id field contains the identification code of
the sentence: in the example, NameSurname001
(anonymized here) indicates the unique identifier
of the text and refers to the transcribers name and
surname; the following two-digit number, 35 in
the example, indicates the position of the sentence
in the text; finally, LS or TH indicates learner
sentence and target hypothesis, respectively. The text
field contains the uncoded sentence (which can be
the learner sentence or the target hypothesis). The
err field contains the error annotation based on
the coding scheme introduced above. The foreign
field includes the index and the PoS of the words
which are considered errors due to language
transfer. The context field contains the index and the
PoS of the words which need replacement due to
wrong context-bound lexical choices8. Finally, in
line with the ESL, we used the segment field when
a sentence was wrongly divided and the typo field
to indicate PoS distributional-morphological
discrepancies.</p>
          <p>
            In the error-annotated sentence (the “err” field
mentioned above), we report the wrong form(s)
inside the hii h/ii tag and the corrected form(s)
inside the hci h/ci tag. Figure 3 shows three
examples of nested tag and two examples of cascade
errors (i.e. an error which is due to the correction
of another token)
            <xref ref-type="bibr" rid="ref1 ref14">(Andorno and Rastelli, 2009,
p. 52)</xref>
            . The hMAXi h/MAXi tag at the beginning
of the sentence, for example, indicates a missing
existential-construction pronoun, i.e. “Sono”
(are) instead of “Ci sono” (there are). After
the insertion of the missing pronoun “Ci”, the
capital “S” in “Sono” needs to be changed into
a lowercase “s”: this is a case in which we have
a cascade capitalization error and we mark it
adding a hashtag after the normal error code, as
in hSVS#i h/SVS#i. Another cascade error is
found in the next nested tag: we have an Inflection
Determiner Gender error which is caused by
the correction of the expression “tanti cofferi”,
involving a determiner and a noun (“cofferi” is a
8Only those choices in which there is no mismatch
between distributional and morphological PoS are registered in
this field.
          </p>
          <p>German word adapted to Italian and meaning
luggages); thus, we have a cascade hIDG#i h/IDG#i
tag which embeds a hFNLi h/FNLi tag (Form
Noun Language transfer). The next three
tags, hMARi h/MARi, hSARi h/SARi and
hSVi h/SVi, indicate Missing pronoun (A)
Relative (“che”, that), Spelling pronoun Relative
(“ce” instead of “che”) and Spelling Verb errors
(“qurda” instead of “guarda”, look), respectively.
There is, finally, another example of nested tag
involving an Inflection Determiner Gender and an
Unneccessary preposiTion errors; this has been
used to indicate the multiple-step shift from the
LS “sulle” (on the Fem Pl) to its TH counterpart
“i” (the Masc Pl): the shift involved a change
in the gender of the article (from feminine to
masculine) and the drop of the preposition “su”
(on), mistakenly used in the LS.</p>
          <p>In order to ensure consistency across different
annotators, the error annotation guidelines
provide a hierarchical order to be applied when
dealing with nested tags. We organized the errors
in a pyramid with at the bottom mechanical
errors (i.e. tokenization, capitalization, spelling and
punctuation) and, proceeding towards the apex,
morphological (derivation and inflection), lexical
(form and replace), and syntactic (missing,
unnecessary and word order) errors. For example,
following this hierarchical order, mechanical
errors should be corrected before a syntactic error.
However, cascade errors make an exception and
change the correction order, as we seen in Figure
3 in which we have a cascade capitalization error
(SVS#) caused by a missing pronoun error (MAX)</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Error category</title>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Derivation</title>
        <p>Form
Inflection
Spelling
Word segmentation
Word order
Missing word
Unnecessary word
Replace word</p>
      </sec>
      <sec id="sec-3-3">
        <title>Total D F I</title>
        <p>S
T
W
M
U
R
–
and a cascade inflection error (IDG#) due to a
lexical error (FNL).</p>
        <p>In the LS sample set, containing 1,860 tokens,
we marked 496 errors (which represent 26,66% of
the LS sample set tokens) distributed as shown in
Table 2.</p>
        <p>Tag
# occ
In this Section we describe how we applied literal
annotation to the (morpho-)syntactic structure of
the LS in particular, relying on the Universal
Dependencies scheme.</p>
        <sec id="sec-3-3-1">
          <title>Literal Annotation</title>
          <p>We annotated UD PoS and relations sticking as</p>
          <p>much as possible to the literal reading of the
learner sentence, thereby creating a treebank in
line with the two existing learner treebanks in the
UD framework (ESL and CFL).</p>
          <p>Argument Structure: When some extraneous or
unnecessary prepositions occur, we annotate the
dependencies accordingly. Figure 2 shows a LS in
which the verb “guardare”, look, is used as an
intransitive verb, thus we annotate its direct object
as an oblique9.</p>
          <p>Missing or Unnecessary Words: We annotate
literally when there are missing or unnecessary
words. In the example in Figure 2 the clitic
pronoun “ci” is missing , thus we treated “sono” as a
copular verb. There are other cases in which the
clitic pronoun “ci” is mistakenly combined with
the verb to be forming an existential clause, and
consequently causing a distributional mismatch
(e.g. LS: “[...] non ci era pericoloso o violento”,
TH: “[...] non era pericoloso o violento”10). In
these cases we mark in the “typo” field the
morphological PoS and in the PoS column the
distributional PoS, cf. Figure 1.</p>
          <p>Extraneous Word Forms: When the learner
misuses existent word forms, we annotate them
literally. In Figure 4, the learner used a gerund,
“leggendo” (reading), instead of the infinitive “ a
9In all the examples SE stands for spelling error, REFL
for reflexive pronoun, PP for past participle, GE for gerund
and Impf for imperfect tense.</p>
          <p>10LS: “[...] not there it-be Impf dargerous or violent”, TH:
“[...] not it-be Impf dangerous or violent”.
leggere” (to read). We then labeled it as an
adverbial clause in the LS (Figure 4) and as an open
clausal complement in the TH (Figure 5).</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>Exceptions to Literal Annotation</title>
          <p>Spelling: Some examples of spelling errors are
presented in Figure 2. We lemmatize and
PoStag them referring to their correct versions,
similarly to Andorno and Rastelli (2009, p. 58). Thus,
“ce” was treated as “che”, which,11, and “qurda”
as “guarda” look.</p>
          <p>Word Formation: We do not treat literally valid
words that are contextually implausible. We
consider them differently depending on the PoS of the
intended word: if the intended word has the same
PoS we signal it in the “context” field (e.g. LS:
“[...] salvando una ragazza indefessa”, TH: “[...]
salvando una ragazza indifesa”12), if it is different
in the “typo” field (cf. Figure 1).</p>
          <p>Nonexistent Words: In cases in which the learner
wrote a word which does not exist in Italian and
it is arguably a foreign word, we signal it in the
“foreign” field13. In the example in Figure 1 the
word “cara” (i.e. an adjective translatable into
beloved) is arguably a transfer from the Spanish
noun meaning face. In this case we lemmatize it
with the correct lemma of “cara”. In addition, in
the “typo” field we mark the occurring mismatch
between distributional and morphological PoS.
Word Tokenization: If one word is mistakenly
segmented into two, we use the “goeswith”
relation, as germane to UD annotation guidelines14. If
two words are mistakenly segmented into one, we
use X as PoS and decide the relation on a
caseby-case basis. For example in LS: “[...] butta tutto
perterra”, TH: “[...] butta tutto per terra”15 we
assigned to “perterra” PoS ‘X’ and dependency
relation ‘obl’.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5 Inter-Annotator Agreement</title>
      <p>As stated above, the complete manual revision of
the treebank is still in progress; however, with
the aim of assessing the annotation quality of this
preliminary sample set, as well as the quality of
the annotation guidelines (especially the ones
con11When “ce” is used instead of “c’e`”, there is, we treat it as
a single token and mark it as root, in line with what we would
have done if it were “c’e`”.</p>
      <p>12LS: “[...] saving a untiring girl”, TH: “[...] saving a
vulnerable girl”.</p>
      <p>
        13The lemma will be its Italian (quasi-)equivalent.
14https://universaldependencies.org/u/overview/typos.html
15[...] he-throw everything on the ground.
cerning the LS section) both LS and TH
sections were annotated by two independent
annotators. The inter-annotator agreement was then
computed, considering two measures in
particular: UAS (Unlabeled Attachment Score) and
LAS (Labeled Attachment Score) for the
assignment of both parent node and dependency relation,
and the Cohen’s kappa coefficient
        <xref ref-type="bibr" rid="ref6">(Cohen, 1960)</xref>
        for dependency relations only (similarly to Lynn
(2016)). UAS and LAS were computed with the
script provided in the second CoNLL shared task
on multilingual parsing
        <xref ref-type="bibr" rid="ref39">(Zeman et al., 2018)</xref>
        16.
The results are reported in Table 3, and though
showing slightly higher results for the TH set,
overall they are very close across the sets.
Especially as regards the LS section, this is evidence of
the guidelines clarity and of the annotators’
consistency, even when dealing with non-canonical
syntactic structures.
      </p>
      <p>set
LS
TH</p>
      <p>UAS
In this paper we introduced VALICO-UD and
proposed an annotation scheme suitable for texts of
learner Italian encompassing both UD and error
annotation. Our scheme follows the principle of
“literal annotation” and takes PoS and dependency
morphological-distributional mismatches into
account. Our error tag set seems adequate to
bookmark errors, providing also a fine-grained
description of some of them.</p>
      <p>
        There are a number of possible applications for
the monolingual parallel treebank proposed in this
paper. In the near future, we plan to apply the tree
edit distance to LS and TH to measure linguistic
competence. Recently, the tree edit distance has
been applied to various tasks
        <xref ref-type="bibr" rid="ref11 ref29 ref37">(Emms, 2008;
Tsarfaty et al., 2011; Plank et al., 2015)</xref>
        , and a study
has formalized the notion of syntactic
anisomorphism
        <xref ref-type="bibr" rid="ref30">(Ponti et al., 2018)</xref>
        . We aim to explore a
correlation between these notions and the linguistic
competence to describe the achievements of
foreign language learners.
      </p>
      <p>16http://universaldependencies.org/conll18/evaluation.html</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Cecilia</given-names>
            <surname>Maria</surname>
          </string-name>
          Andorno and
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Rastelli</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Un'annotazione orientata alla ricerca acquisizionale</article-title>
          .
          <source>In Cecilia Maria Andorno and Stefano Rastelli</source>
          , editors,
          <article-title>Corpora di italiano L2: tecnologie, metodi, spunti teorici</article-title>
          , pages
          <fpage>49</fpage>
          -
          <lpage>70</lpage>
          . Guerra.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Yevgeni</given-names>
            <surname>Berzak</surname>
          </string-name>
          , Jessica Kenney, Carolyn Spadine, Jing Xian Wang, Lucia Lam, Keiko Sophie Mori,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Garza</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Boris</given-names>
            <surname>Katz</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Universal Dependencies for Learner English</article-title>
          .
          <source>In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>737</fpage>
          -
          <lpage>746</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          , Montemagni Simonetta, and
          <string-name>
            <given-names>Simi</given-names>
            <surname>Maria</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Converting Italian Treebanks: Towards an Italian Stanford Dependency Treebank</article-title>
          .
          <source>In Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse</source>
          , pages
          <fpage>61</fpage>
          -
          <lpage>69</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          , Felice Dell'Orletta, Simonetta Montemagni, Manuela Sanguinetti, and
          <string-name>
            <given-names>Maria</given-names>
            <surname>Simi</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The EVALITA 2014 Dependency Parsing Task</article-title>
          .
          <source>In Proceedings of EVALITA</source>
          <year>2014</year>
          <article-title>Evaluation of NLP and Speech Tools for Italian</article-title>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Adriane</given-names>
            <surname>Boyd</surname>
          </string-name>
          , Jirka Hana, Lionel Nicolas, Detmar Meurers, Katrin Wisniewski, Andrea Abel, Karin Scho¨ne, Barbora Stindlova´, and
          <string-name>
            <given-names>Chiara</given-names>
            <surname>Vettori</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The MERLIN corpus: Learner Language and the CEFR</article-title>
          .
          <source>In Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>1281</fpage>
          -
          <lpage>1288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <year>1960</year>
          .
          <article-title>A Coefficient of Agreement for Nominal Scales</article-title>
          .
          <source>Educational and Psychological Measurement</source>
          ,
          <volume>20</volume>
          (
          <issue>1</issue>
          ):
          <fpage>37</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Corino</surname>
          </string-name>
          and
          <string-name>
            <given-names>Carla</given-names>
            <surname>Marello</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Italiano di stranieri. I corpora VALICO e VINCA</article-title>
          . Guerra.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Corino</surname>
          </string-name>
          and
          <string-name>
            <given-names>Claudio</given-names>
            <surname>Russo</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Parsing di Corpora di Apprendenti di Italiano: un Primo Studio su VALICO</article-title>
          .
          <source>In Proceedings of the 3rd Italian Conference on Computational Linguistics</source>
          , CLiC-it
          <year>2016</year>
          , pages
          <fpage>105</fpage>
          -
          <lpage>110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Dahlmeier</surname>
          </string-name>
          , Hwee Tou Ng, and Siew Mei Wu.
          <year>2013</year>
          .
          <article-title>Building a large annotated corpus of learner English: The NUS corpus of learner English</article-title>
          .
          <source>In Proceedings of the Eighth Workshop on Innovative Use of NLP for Building Educational Applications</source>
          , pages
          <fpage>22</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Tullio De Mauro</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Il Nuovo Vocabolario di Base della Lingua Italiana</article-title>
          . Internazionale, http://www.internazionale.it/opinione/tullio-demauro/
          <year>2016</year>
          /12/23/il-nuovo
          <article-title>-vocabolario-di-basedella-lingua-italiana.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Martin</given-names>
            <surname>Emms</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Tree Distance and Some Other Variants of Evalb</article-title>
          .
          <source>In Proceedings of the Sixth International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>1373</fpage>
          -
          <lpage>1379</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Roger</given-names>
            <surname>Garside</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Geoffrey N.</given-names>
            <surname>Leech</surname>
          </string-name>
          , and
          <string-name>
            <surname>Tony McEnery</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Corpus Annotation: Linguistic Information from Computer Text Corpora</article-title>
          . Taylor &amp; Francis.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Jeroen</given-names>
            <surname>Geertzen</surname>
          </string-name>
          , Theodora Alexopoulou, and
          <string-name>
            <given-names>Anna</given-names>
            <surname>Korhonen</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Automatic Linguistic Annotation of Large Scale L2 Databases: The EF-Cambridge Open Language Database (EFCAMDAT)</article-title>
          .
          <source>In Proceedings of the 31st Second Language Research Forum</source>
          , pages
          <fpage>240</fpage>
          -
          <lpage>254</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Sylviane</given-names>
            <surname>Granger</surname>
          </string-name>
          , Estelle Dagneaux, Fanny Meunier, and
          <string-name>
            <given-names>Magali</given-names>
            <surname>Paquot</surname>
          </string-name>
          .
          <year>2009</year>
          . International Corpus of Learner English. Louvain University Press.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Sylviane</given-names>
            <surname>Granger</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Learner Corpora</article-title>
          . In Anke Lu¨deling and Merja Kyto¨, editors,
          <source>Corpus Linguistics</source>
          , volume
          <volume>1</volume>
          , pages
          <fpage>259</fpage>
          -
          <lpage>275</lpage>
          . Walter de Gruyter.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Derrick</given-names>
            <surname>Higgins</surname>
          </string-name>
          , Chaitanya Ramineni, and
          <string-name>
            <given-names>Klaus</given-names>
            <surname>Zechner</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Learner Corpora and Automated Scoring</article-title>
          . In Sylviane Granger, Gae¨tanelle Gilquin, and Fanny Meunier, editors,
          <source>The Cambridge Handbook of Learner Corpus Research</source>
          , pages
          <fpage>587</fpage>
          -
          <lpage>604</lpage>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Yan</given-names>
            <surname>Huang</surname>
          </string-name>
          , Akira Murakami, Theodora Alexopoulou, and
          <string-name>
            <given-names>Anna</given-names>
            <surname>Korhonen</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Dependency Parsing of Learner English</article-title>
          .
          <source>International Journal of Corpus Linguistics</source>
          ,
          <volume>23</volume>
          (
          <issue>1</issue>
          ):
          <fpage>28</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Scott</given-names>
            <surname>Jarvis</surname>
          </string-name>
          and
          <string-name>
            <given-names>Magali</given-names>
            <surname>Paquot</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Learner Corpora and Native Language Identification</article-title>
          . In Sylviane Granger, Gae¨tanelle Gilquin, and Fanny Meunier, editors,
          <source>The Cambridge Handbook of Learner Corpus Research</source>
          , pages
          <fpage>605</fpage>
          -
          <lpage>628</lpage>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Claudia</given-names>
            <surname>Leacock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Chodorow</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Joel</given-names>
            <surname>Tetrault</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Automatic Grammar- and Spell-Checking for Language Learners</article-title>
          . In Sylviane Granger, Gae¨tanelle Gilquin, and Fanny Meunier, editors,
          <source>The Cambridge Handbook of Learner Corpus Research</source>
          , pages
          <fpage>567</fpage>
          -
          <lpage>586</lpage>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>John</given-names>
            <surname>Lee</surname>
          </string-name>
          , Herman Leung, and
          <string-name>
            <given-names>Keying</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Towards Universal Dependencies for Learner Chinese</article-title>
          .
          <source>In Proceedings of the NoDaLiDa 2017 Workshop on Universal Dependencies (UDW</source>
          <year>2017</year>
          ), pages
          <fpage>67</fpage>
          -
          <lpage>71</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Anke</given-names>
            <surname>Lu</surname>
          </string-name>
          ¨deling, Maik Walter, Emil Kroymann, and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Adolphs</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Multi-level error annotation in learner corpora</article-title>
          .
          <source>In Proceedings of Corpus Linguistics</source>
          <year>2005</year>
          , volume
          <volume>1</volume>
          , pages
          <fpage>14</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Ana D´</surname>
            ıaz-Negrillo, Detmar Meurers, Salvador Valera, and
            <given-names>Holger</given-names>
          </string-name>
          <string-name>
            <surname>Wunsch</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Towards Interlanguage POS Annotation for Effective Learner Corpora in SLA and FLT</article-title>
          .
          <source>Language Forum</source>
          ,
          <volume>36</volume>
          (
          <issue>1-2</issue>
          ):
          <fpage>139</fpage>
          -
          <lpage>154</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Teresa</given-names>
            <surname>Lynn</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Irish Dependency Treebanking and Parsing</article-title>
          .
          <source>Ph.D. thesis</source>
          , Dublin City University, Ireland and Macquarie University, Sydney, Australia.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Shervin</given-names>
            <surname>Malmasi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Native Language Identification: explorations and applications</article-title>
          .
          <source>Ph.D. thesis</source>
          , Macquarie University, Sydney, Australia.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Yukio</given-names>
            <surname>Tono</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Learner corpora: design, development and applications</article-title>
          .
          <source>In Proceedings of the Corpus Linguistics</source>
          <year>2003</year>
          conference, pages
          <fpage>800</fpage>
          -
          <lpage>809</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Tony</given-names>
            <surname>McEnery and Richard Xiao</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>What corpora can offer in language teaching and learning</article-title>
          . In Eli Hinkel, editor,
          <source>Handbook of Research in Second Language Teaching and Learning</source>
          , volume
          <volume>2</volume>
          , pages
          <fpage>364</fpage>
          -
          <lpage>380</lpage>
          . Routledge.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Hwee</given-names>
            <surname>Tou</surname>
          </string-name>
          <string-name>
            <surname>Ng</surname>
          </string-name>
          , Siew Mei Wu, Ted Briscoe, Christian Hadiwinoto, Raymond Hendy Susanto, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Bryant</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The CoNLL-2014 shared task on grammatical error correction</article-title>
          .
          <source>In Proceedings of the Eighteenth Conference on Computational Natural Language Learning: Shared Task</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Diane</given-names>
            <surname>Nicholls</surname>
          </string-name>
          .
          <year>2003</year>
          . The Cambridge Learner Corpus:
          <article-title>Error coding and analysis for lexicography and ELT</article-title>
          .
          <source>In Proceedings of the Corpus Linguistics 2003 Conference</source>
          , volume
          <volume>16</volume>
          , pages
          <fpage>572</fpage>
          -
          <lpage>581</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Plank</surname>
          </string-name>
          ,
          <article-title>He´ctor Mart´ınez Alonso, Zˇeljko Agic´</article-title>
          ,
          <string-name>
            <given-names>Danijela</given-names>
            <surname>Merkler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Anders</given-names>
            <surname>Søgaard</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Do dependency parsing metrics correlate with human judgments?</article-title>
          <source>In Proceedings of the Nineteenth Conference on Computational Natural Language Learning</source>
          , pages
          <fpage>315</fpage>
          -
          <lpage>320</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>Edoardo</given-names>
            <surname>Maria</surname>
          </string-name>
          <string-name>
            <surname>Ponti</surname>
          </string-name>
          , Roi Reichart, Anna Korhonen, and Ivan Vulic´.
          <year>2018</year>
          .
          <article-title>Isomorphic transfer of syntactic structures in cross-lingual NLP</article-title>
          .
          <source>In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>1531</fpage>
          -
          <lpage>1542</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <given-names>Norma A.</given-names>
            <surname>Pravec</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Survey of Learner Corpora</article-title>
          .
          <source>ICAME journal</source>
          ,
          <volume>26</volume>
          (
          <issue>1</issue>
          ):
          <fpage>8</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <given-names>Marwa</given-names>
            <surname>Ragheb</surname>
          </string-name>
          and
          <string-name>
            <given-names>Markus</given-names>
            <surname>Dickinson</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Defining syntax for learner language annotation</article-title>
          .
          <source>In Proceedings of COLING 2012: Posters</source>
          , pages
          <fpage>965</fpage>
          -
          <lpage>974</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <given-names>Marc</given-names>
            <surname>Reznicek</surname>
          </string-name>
          , Anke Lu¨deling, and
          <string-name>
            <given-names>Hagen</given-names>
            <surname>Hirschmann</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Competing target hypotheses in the falko corpus</article-title>
          . In Ana D´
          <article-title>ıaz-</article-title>
          <string-name>
            <surname>Negrillo</surname>
            ,
            <given-names>Nicolas</given-names>
          </string-name>
          <string-name>
            <surname>Ballier</surname>
          </string-name>
          , and Paul Thompson, editors,
          <source>Automatic treatment and analysis of learner corpus data</source>
          , volume
          <volume>59</volume>
          , pages
          <fpage>101</fpage>
          -
          <lpage>123</lpage>
          . John Benjamins Publishing Company.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Cristina Bosco, Alberto Lavelli, Alessandro Mazzei, Oronzo Antonelli, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Tamburini</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>PoSTWITA-UD: an Italian Twitter Treebank in Universal Dependencies</article-title>
          .
          <source>In Proceedings of the Eleventh International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>1768</fpage>
          -
          <lpage>1775</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <given-names>Raffaele</given-names>
            <surname>Simone</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Fondamenti di linguistica</article-title>
          .
          <source>Laterza.</source>
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <given-names>Milan</given-names>
            <surname>Straka</surname>
          </string-name>
          , Jan Hajic, and Jana Strakova´.
          <year>2016</year>
          .
          <article-title>Udpipe: Trainable pipeline for processing conll-u files performing tokenization, morphological analysis, pos tagging and parsing</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>88</fpage>
          -
          <lpage>99</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <given-names>Reut</given-names>
            <surname>Tsarfaty</surname>
          </string-name>
          , Joakim Nivre, and
          <string-name>
            <given-names>Evelina</given-names>
            <surname>Andersson</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Evaluating dependency parsing: Robust and heuristics-free cross-annotation evaluation</article-title>
          .
          <source>In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>385</fpage>
          -
          <lpage>396</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <given-names>Helen</given-names>
            <surname>Yannakoudakis</surname>
          </string-name>
          , Ted Briscoe, and
          <string-name>
            <given-names>Ben</given-names>
            <surname>Medlock</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>A new dataset and method for automatically grading esol texts</article-title>
          .
          <source>In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies</source>
          , pages
          <fpage>180</fpage>
          -
          <lpage>189</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Zeman</surname>
          </string-name>
          , Jan Hajicˇ, Martin Popel, Martin Potthast, Milan Straka, Filip Ginter, Joakim Nivre, and
          <string-name>
            <given-names>Slav</given-names>
            <surname>Petrov</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>CoNLL 2018 shared task: Multilingual parsing from raw text to universal dependencies</article-title>
          .
          <source>In Proceedings of the CoNLL</source>
          <year>2018</year>
          <article-title>Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies</article-title>
          , pages
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>