<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Quale testo e` scritto meglio? A Study on Italian Native Speakers' Perception of Writing Quality</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aldo Cerulli</string-name>
          <email>a.cerulli1@studenti.unipi.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dominique Brunato⋄</string-name>
          <email>dominique.brunato@ilc.cnr.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felice Dell'Orletta⋄</string-name>
          <email>felice.dellorletta@ilc.cnr.it</email>
        </contrib>
      </contrib-group>
      <abstract>
        <p>This paper presents a pilot study focused on Italian native speakers' perception of writing quality. A group of native speakers expressed their preferences on 100 pairs of essays extracted from an Italian corpus of compositions written by L1 students of lower secondary school. Analysing their answers, it was possible to identify a set of linguistic features characterizing essays perceived as well written and to assess the impact of students errors on the perception of text quality. The paper describes the crowdsourcing technique to collect data as well as the linguistic analysis and results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The institution of distance learning paradigms,
which has become crucial during the Covid-19
pandemic, showed the need to provide schools and
universities with Natural Language Processing
(NLP)based tools to assist students, teachers and
professors. Nowadays, language technologies are more
and more exploited to develop educational
applications, such as Intelligent Computer-Assisted
Language Learning (ICALL) systems
        <xref ref-type="bibr" rid="ref5">(Granger, 2003)</xref>
        and tools for automated essay scoring
        <xref ref-type="bibr" rid="ref1">(Attali and
Burstein, 2006)</xref>
        or automatic error detection and
correction
        <xref ref-type="bibr" rid="ref10">(Ng et al., 2013)</xref>
        . A fundamental
requirement for developing this kind of applications
is the availability of electronically accessible
corpora of learners’ productions. Corpora created so
far differ in many respects. For instance,
considering the types of examined learners, they can gather
productions written by L2 students or by native
speakers: the former have been built for many
languages (e.g. English, Arabic, German, Hungarian,
Basque, Czech, Italian), while the latter are mainly
available for English. In both cases, a peculiarity
of existing corpora is that they are cross-sectional
rather than longitudinal. A notable exception in
the context of Italian as L1 – which is the focus
of our contribution – is represented by CItA
(Corpus Italiano di Apprendenti L1), which was jointly
developed by the Institute for Computational
Linguistics of the Italian National Research Council
(CNR) of Pisa and the Department of Social and
Developmental Psychology at Sapienza University
of Rome
        <xref ref-type="bibr" rid="ref3">(Barbagli et al., 2016)</xref>
        : it is the first
digitalized collection of essays written by the same
group of Italian L1 learners in the first two years of
the lower secondary school1.
      </p>
      <p>The diachronic and longitudinal nature of CItA
makes it particularly suitable to study the evolution
of L1 writing competence over the two years,
assuming that many remarkable changes in writing
skills occur in this period. For instance, in their
recent work, Miaschi et al. (2021) showed that it is
possible to automatically learn the writing
development curve of students: they extracted a wide set of
linguistic features from the essays and used them
to train a binary classification algorithm able to
predict the chronological order of two productions
written by the same pupil at different times.</p>
      <p>The present study ranks among research based
on CItA, but chooses a different approach from the
one just mentioned: instead of tracking the
development of students’ writing competence, we focused
on the perception of writing quality by Italian L1
speakers with the aim of understanding whether it
is possible to find the linguistic features that are
crucially involved in the distinction between ‘better’
and ‘worse’ essays according to our target reader.
Contributions To the best of our knowledge, this
is the first paper that (i) introduces a dataset of
1The corpus is freely available for research goals at
http://www.italianlp.it/resources/cita-corpus-italiano-diapprendenti-l1/
evaluated essays in terms of perceived writing
quality by means of a crowdsourcing task, (ii) deals
with the correlation between linguistic features and
perceived quality of writing and (iii) assesses the
impact of students errors on quality perception.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Corpus Collection</title>
      <p>As previously mentioned, the starting point of our
study was the CItA corpus. It comprises 1,352
essays, written by 156 pupils of seven lower
secondary schools in Rome (three in the historical
center and four in the suburbs) during the school
years 2012-2013 and 2013-2014. The productions
respond to 124 writing prompts that pertain to vfie
textual typologies: reflexive, narrative, descriptive,
expository and argumentative. An additional
‘common prompt’ was presented at the end of each
school year, in which students were asked to write
a letter to advise a younger friend how to compose
better essays. The common prompts were aimed at
understanding how learners internalize the different
writing instructions given by teachers.</p>
      <p>Each essay contained in CItA is also provided by
a set of metadata tracking students’ biographical,
sociocultural and sociolinguistic information.
Beyond the longitudinal nature, the most significant
novelty introduced by CItA regards error
annotation, which was manually performed by a
middle school teacher according to a new three-level
schema including: the macro-class of error (i.e.
grammatical, orthographic and lexical); the class of
error (i.e. verbs, prepositions, monosyllabes); and
the corresponding type of modification required to
correct it. More details about the CItA collection
are reported in Barbagli et al. (2016).
2.1</p>
      <sec id="sec-2-1">
        <title>Essay Selection</title>
        <p>For the purpose of our investigation, we selected
200 essays from CItA to be submitted to human
evaluation. The essays ranged from a minimum
of 141 tokens to a maximum of 1153 tokens and
their average length was 359.4 tokens. Then, to
gather judgments on writing quality, we created ten
questionnaires, each one consisting of ten pairs of
essays of the same grade, and distribute them to
native speakers of all ages and cultural background.</p>
        <p>Table 1 reports the criteria we adopted to select
the pairs of essays. As it can be seen, Survey 1
allows the comparison between essays responding to
the common prompts written by students attending
the first or the second grades. In surveys 2-8, we
Survey</p>
        <p>Selection criteria
chosen essays pertaining to the same textual
typology – assuming that their similarity with regard to
the content could let the annotator focus on stylistic
issue to orient their judgment – and paired them
according to the school year in which they were
written. Instead, essays in questionnaires 9 and 10
were paired according to their number of errors:
for each year, we divided the range between the
minimum amount of errors (0) and the maximum
one (49 for the first year, 43 for the second one)
into ten error bins and designed the two surveys
choosing a couple of productions for each bin.
Surveys comparing essays with a similar amount of
errors were meant to understand which categories
of errors have a greater impact on human judgment.
After designing the surveys, we moved on to their
implementation using the QuestBase platform2.
We defined a three-section structure including the
iflling-in instructions, the personal data entry form
and the essays evaluation pages.</p>
        <p>Filling-in instructions. The first section reported
the following submission guidelines:</p>
        <p>Ciao!
Il presente sondaggio e` rivolto a partecipanti di
madrelingua italiana. La sua compilazione richiede
circa 20 minuti. Pima di proseguire, dando il consenso
alla partecipazione, ti spieghiamo in cosa consiste.
Nelle pagine che seguono leggerai dieci coppie di temi
scritti da studenti del primo e del secondo anno di scuola
media. I testi possono contenere un certo numero di
errori. Per ciascuna coppia ti chiediamo di indicare quale
dei due temi ritieni sia scritto meglio.</p>
        <p>Non esistono risposte giuste o sbagliate: conta
semplicemente quello che pensi! Tieni presente che i temi di
una stessa coppia possono trattare argomenti diversi, ma
questo non deve influire sul tuo giudizio.</p>
        <p>La tua partecipazione al sondaggio e` completamente
libera. Se in qualsiasi momento dovessi cambiare idea</p>
        <sec id="sec-2-1-1">
          <title>2https://story.questbase.com/</title>
          <p>e volessi interrompere il test, potrai farlo liberamente.
Un’ultima cosa: prima di iniziare il sondaggio, ti
chiediamo di darci alcune tue informazioni anagrafiche, che
serviranno solo a fini statistici. I dati rimarranno
completamente anonimi e in nessun modo le risposte verranno
associate alla tua persona.</p>
          <p>Se hai dubbi, curiosita` o proposte di miglioramento,
scrivimi all’indirizzo: a.cerulli1@studenti.unipi.it.
Buona lettura!</p>
          <p>For the sake of completeness, we also report an</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>English translation of the same guidelines:</title>
          <p>Hello!
This survey is addressed to Italian native speakers. Its
submission requires about 20 minutes. By completing
it, you give your consent to participation. Before going
on, we explain to you what it consists of.</p>
          <p>In the following pages you will read ten pairs of essays
written by Italian L1 learners during the first two years
of lower secondary school. The essays may contain
linguistic errors. For each pair, you are asked to choose
the best written of the two essays.</p>
          <p>No answers are right or wrong: you only have to
express your opinion! Bear in mind that the essays of
a pair can concern different topics, but this must not
affect your judgment.</p>
          <p>Your participation to the survey is completely free. You
may withdraw from it at any time.</p>
          <p>Before starting the survey, we ask you to provide some
personal information that will be used for statistical
purposes. Data will remain completely anonymous and
will not be connected to you in any way.</p>
          <p>If you have doubts, curiosities or
improvement proposals, please write me to the address:
a.cerulli1@studenti.unipi.it.</p>
          <p>Have a good read!</p>
          <p>Personal data entry form. The surveys were
obviously anonymous. However, as we mentioned
before, we asked the annotators to entry some
personal information (age, sex, education) for
statistical purposes.</p>
          <p>Essays evaluation. The third section comprised
ten pages, each occupied by two side by side essays
and a field to give the answer (Figure 1). The user
had to choose the label ‘1’ if they had preferred the
ifrst essay, ‘2’ otherwise.</p>
          <p>After carrying out a pilot study to test the
adequacy of the structure as well as the completeness
and clearness of the instructions, we started
collecting evaluations. Using Linktree3 we added the
ten questionnaires links to a single web page and
shared its link through WhatsApp, Facebook and
Instagram: clicking on it, users were redirected to
the page and could access every survey.
3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Analysis of Human Judgments</title>
      <p>We collected 223 annotations distributed quite
homogeneously among the ten surveys, except for the
ifrst one, submitted 28 times. It is worth to focus
on the heterogeneous composition of the readers
sample. Concerning sex, the large majority of
answers (183 units, equal to 82.1%) were given by
women, against the 38 (17%) by men; just two
people preferred not to specify their gender.</p>
      <p>Regarding age, we divided the group into six
bins (Figure 2). The most frequent class (97 units)
was ‘20-24 years’, followed by ‘25-29 years’ (64
units).</p>
      <sec id="sec-3-1">
        <title>This means that most readers (72.5%)</title>
        <p>ranged from 20 to 29 years of age. 35 evaluations
(15.8%) were made by natives between 30 and 39
years of age. People belonging to the remaining
bins contributed to the task for an overall 11.7%.</p>
        <p>Finally, Figure 3 shows the distribution of
submissions with respect to readers’ education: 91.9%
of annotations were given by people holding an
academic degree (118 units, equal to 53.2%) or
a high school diploma (86 units, equal to 38.7%).
12 annotators (5.4%) had a middle school
certificate; 4 (1,8%) held a doctoral degree; the last two
indicated a non-specific ‘Other’.
3.1</p>
        <sec id="sec-3-1-1">
          <title>Inter-Annotator Agreement</title>
          <p>
            At this point, we defined a selection function to
discard inaccurate annotations and obtain the same
number of coherent annotations for each survey.
Thus, we firstly built the average vector of every
survey as the set of ten values ‘1’ or ‘2’ chosen
according to the most assigned label to each pair
of essays; then, we calculated the distance between
each survey average vector and all its annotations.
We implemented the euclidean metric generalized
to the n-dimensional space that computes the
distance between two vectors as the square root of the
sum of their sizes squared difference:
(1)
(2)
der and calculated the Inter-annotator agreement
(IAA) of the first 15 and 20 annotations. We
implemented Krippendorff’s alpha (α ), a coefficient
that expresses IAA in terms of observed (Do) and
casual (De) disagreement
            <xref ref-type="bibr" rid="ref7">(Krippendorff, 2011)</xref>
            :
α = 1 − DDoe (3)
          </p>
          <p>We noticed that IAA values of the first 15
submissions ordered by their increasing weighted
distance were the highest. Thus, we took them into
account (150 total annotations) for the analysis and
discarded the remaining 734. It is noteworthy that
the selection led us to an average IAA of 0.26, that
is a much higher value than the initial 0.12.
Relying on the selected annotations, we established the
‘winning’ and ‘loser’ essay of each pair.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Data Analysis</title>
      <p>We carried out two evaluations: a first one was
meant to identify which linguistic features impact
more on the human assessment of the writing
quality; a second one focused on the impact of students
errors on annotators’ judgments. In what follows
we describe the approach underlying the two
perspectives and discuss our most interesting findings.
4.1</p>
      <sec id="sec-4-1">
        <title>Linguistic Profiling and Stylistic Analysis</title>
        <p>The first analysis relies on linguistic profiling , a
NLP-based methodology in which a large set of
linguistically-motivated features automatically
extracted from annotated texts are used to obtain a
vector-based representation of it. Such
representations can be then compared across texts
representative of different textual genres and varieties to
identify the peculiarities of each (Montemagni, 2013;
4The corpus of evaluated essays is available at
http://www.italianlp.it/EvaluatedEssays.zip
vu n
uX(pk − qk)2
t</p>
        <p>k=1</p>
        <p>To give relevance to the deviating degree of
answers differing from the average, we assigned every
pair a weight (wk) equal to the number of times in
which the ‘winning’ essay was chosen; then, we
computed the weighted distance between
annotations and average vectors.</p>
        <p>vu n
uX wk(pk − qk)2
t</p>
        <p>k=1</p>
        <p>Finally, we ranked weighted and unweighted
distance values of each survey in ascending
orvan Halteren, 2004). To perform the analysis, we
relied on Profiling-UD 5, a recently introduced tool
that allows the extraction of a wide set of lexical,
morpho-syntactic and syntactic features from texts
linguistically annotated according to the Universal
Dependencies (UD)6 formalism. These features,
described in details in Brunato et al. (2020), have
been shown to be involved in many tasks, all
related to modeling the form rather than the content
of a text, such as the assessment of text readability
and linguistic complexity and the identification of
stylistic traits of an author or groups of authors.</p>
        <p>We thus split our annotated corpus into two
sections: one comprised all ‘winning’ essays and the
other all ‘loser’ ones. Using Profiling-UD, we
extracted for each text of the two subsets a
featurebased vector representation. For each considered
feature we calculated the average value, the
standard deviation and the coefficient of variation ( ASvDg )
in the two subsets and we assessed whether the
variation between mean values was significant using
the Wilcoxon rank sum test.</p>
        <p>Table 2 shows the seven linguistic features
whose variation turned out to be statistically
significant ( p − value &lt; 0.05), ordered by
increasing p-values. It emerges that ‘winning’ essays
are on average longer (32.2 tokens more) than
the ‘losers’ (n tokens), a finding that may suggest
that longer compositions are evaluated as more
reasoned, structured and content-rich.
Interestingly, this also reflects the students’ perception
of school writing: Barbagli et al. (2015) showed
that two of the most frequent suggestions contained
in essays that respond to ‘common prompts’ are
Leggi/scrivi molto (“Read/write a lot”) and Lavora
sodo, fai vedere che ti impegni (“Work hard, show
your dedication”). Thus, pupils possibly write
more so as to show their dedication and get higher</p>
        <sec id="sec-4-1-1">
          <title>5http://linguistic-profiling.italianlp.it/ 6https://universaldependencies.org/</title>
          <p>
            Feature
verbs tense dist Fut
dep dist cop
dep dist flat:foreign
dep dist flat:name
dep dist det:predet
dep dist parataxis
obj pre
verb edges dist 0
verb edges dist 1
upos dist CCONJ
grades. Secondly, we noticed that a richer
vocabulary (ttr form chunks 100) plays a crucial role in
native’s judgment. This is in line with another
advice of the just mentioned ranking, Usa un
vocabolario ricco ed espressivo (“Use a rich and
expressive vocabulary”), that reflects teachers’
encouragement to use synonyms in order to write clearer
and more readable compositions. Values related
to the third feature (upos dist NOUN) reveal that
‘loser’ essays present a slightly higher distribution
of nouns. A predominant use of nouns is typical
of highly informative texts (e.g. newspaper
articles, laws), while genres closer to speech contain
more verbs
            <xref ref-type="bibr" rid="ref9">(Montemagni, 2013)</xref>
            . Belonging to the
second category, a school essay with fewer nouns
is probably perceived as more coherent with its
genre. Concerning verbal inflection, ‘better’
productions include, on average, 0.28% more future
verbs (verbs tense dist Fut), 0.81% more gerund
verbs (verbs form dist Ger) and 1.93% more
subjunctive auxiliary verbs (aux mood dist Sub).
Verbal tenses differing from present and moods
differing from indicative require elevated
linguistic skills, which positively influence annotators’
choices. The last feature significantly varying
between the two groups is the number of prepositional
chains (n prepositional chains): ‘winning’
compositions have, on average, 1.2 more of them.
          </p>
          <p>A further study was focused on the variability
degree of linguistic features in the two essay groups.
For each subset, we ordered the features by their
increasingly coefficients of variation; then, we
calculated the difference between the two rankings in
order to identify the features that were maximally
uniformly distributed in ‘better’ essays as compared
to the ‘worse’ ones (Table 3). It can be noticed
that future verbs (verbs tense dist Fut) are very
uniformly distributed among ‘better’ essays. We
have previously commented that their frequency
is higher in the ‘winners’; it proves again that
natives interpret the use of complex verbal forms
as an indicator of higher skills. Also parataxis
distribution (dep dist parataxis) is quite uniform
in ‘winning’ essays; however, its average value is
higher in the ‘loser’ ones. It can be deduced that
annotators prefer hypotaxis but this is not
surprising: hypotactic periods are more structured and
elegant and require refined abilities to be built. The
same evidence is given on the morphosyntactic
level (upos dist CCONJ), since ‘better’
compositions include 0.34% less coordinating conjunctions.
Curiously, ‘better’ essays have, on average, 0.1%
more foreign terms (dep dist flat:foreign ); this may
suggest that annotators appreciate these
expressions. Finally, it is worth highlighting a higher and
more uniform percentage of verbs with few
modifiers in the ‘winning’ essays ( verb edges dist 0,
verb edges dist 1).
4.2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Students Errors Impact</title>
        <p>The last analysis was aimed at assessing whether
and in what measure students errors impact on
human judgments. We counted the pairs of essays
whose ‘winning’ composition had a lower number
of errors, those in which the ‘loser’ one had more
mistakes and those with an equal number of errors.
We noticed that essays with fewer errors had won
in 56% cases, reaching the 79% if including pairs
with the same number of errors. This procedure
gave a first empirical answer to our starting
question: errors substantially affect human assessment.</p>
        <p>At this point, we focused on error categories to
identify which ones affect more the perception of
writing quality. For each category, we calculated
the average number of errors and their standard
deviation in both subsets; then, relying on Wilcoxon
rank sum test, we found out that grammatical and
orthographic mistakes vary significantly between
the two groups (Table 4). As expected, ‘loser’
essays have, on average, 1.29 more grammatical
errors and 0.85 more orthographic errors. It is
worth to add that orthographic mistakes variation
(p − value = 0.007) is more significant than the
other (p − value = 0.029). This could mean that
natives judge deviations in orthography worse than
those in grammar. Once again, our findings are in
line with Barbagli et al. (2015): Usa una corretta
ortografia (“Use correct orthography”) is the 2nd of
the most frequent suggestions given in the second
Category
Grammar
Orthography
year; moreover, Errori di ortografia (“Orthography
errors”) occupies the 6th and the 1st position among
the most salient terms respectively of the first and
the second year. The non-significant variations of
lexical (p − value = 0.581) and punctuation
errors (p − value = 0.617) are probably due to their
scarce amount in the analysed essays.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>We presented a pilot study towards the
identification of the linguistic features that are own of well
written perceived essays. We collected Italian
natives’ preferences on 100 pairs of essays written by
L1 students, that we analysed in terms of linguistic
profiling and errors distribution. Our results reveal
an interesting correspondence between annotators’
judging criteria and writing instructions that L1
learners receive by teachers. Our findings could be
interpreted as an indicator of the reliability of our
data and, more in general, could suggest the
effectiveness of crowdsourcing methods to quickly build
large and reliable datasets. Considering the lack
of Italian corpora of graded essays, such datasets
could be valuable resources for the development of
Computer-Assisted Learning Systems.</p>
      <p>The limited size of our dataset certainly reduced
the amount of results. Thus, we have to expand it (i)
by collecting more annotations for the already
existing surveys and (ii) by creating and distributing new
surveys in order to gather judgments on new pairs
of essays. Analysis on the enlarged dataset could
provide more features that are own of good essays.
Following the model of Miaschi et al. (2021), we
could use the results to train a classifier that, given
a pair of essays, recognizes the best written one.</p>
      <p>The tool would not presume to replace teachers,
but it could be a valuable teaching aid. Students
could use it to get an immediate and preliminary
self-assessment on their written productions so as
to better understand their mistakes and hopefully
avoid repeating them. Such tools can be very useful
if integrated into educational processes based on
distance learning paradigms, which need adequate
technological infrastructures to be really efficient.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Yigal</given-names>
            <surname>Attali</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jill</given-names>
            <surname>Burstein</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Automated Essay Scoring With e-rater® V. 2</article-title>
          .
          <source>The Journal of Technology, Learning, and Assessment</source>
          ,
          <volume>4</volume>
          (
          <issue>3</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Alessia</given-names>
            <surname>Barbagli</surname>
          </string-name>
          , Pietro Lucisano, Felice Dell'Orletta,
          <string-name>
            <given-names>and Giulia</given-names>
            <surname>Venturi</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Il ruolo delle tecnologie del linguaggio nel monitoraggio dell'evoluzione delle abilita` di scrittura: primi risultati</article-title>
          .
          <source>Italian Journal of Computational Linguistics (IJCoL)</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>99</fpage>
          -
          <lpage>117</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Alessia</given-names>
            <surname>Barbagli</surname>
          </string-name>
          , Lucisano Pietro, Felice Dell'Orletta,
          <string-name>
            <given-names>Simonetta</given-names>
            <surname>Montemagni</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Giulia</given-names>
            <surname>Venturi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Cita: an L1 Italian Learners Corpus to Study the Development of Writing Competence</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)</source>
          , pages
          <fpage>88</fpage>
          -
          <lpage>95</lpage>
          , Portorozˇ, Slovenia.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Brunato</surname>
          </string-name>
          , Andrea Cimino, Felice Dell'Orletta,
          <string-name>
            <given-names>Simonetta</given-names>
            <surname>Montemagni</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Giulia</given-names>
            <surname>Venturi</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Profiling-UD: a Tool for Linguistic Profiling of Texts</article-title>
          .
          <source>In Proceedings of the 12th Conference of Language Resources and Evaluation (LREC</source>
          <year>2020</year>
          ), pages
          <fpage>7145</fpage>
          -
          <lpage>7151</lpage>
          , Marseille, France.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Sylviane</given-names>
            <surname>Granger</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Error-tagged learner corpora and CALL: A promising synergy</article-title>
          .
          <source>CALICO Journal</source>
          ,
          <volume>20</volume>
          (
          <issue>3</issue>
          ):
          <fpage>465</fpage>
          -
          <lpage>480</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Hans van Halteren</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Linguistic profiling for author recognition and verification</article-title>
          .
          <source>In Proceedings of the Association for Computational Linguistics</source>
          , pages
          <fpage>200</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Klaus</given-names>
            <surname>Krippendorff</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Computing Krippendorff's Alpha-Reliability</article-title>
          .
          <source>Technical report</source>
          , University of Pennsylvania.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Alessio</given-names>
            <surname>Miaschi</surname>
          </string-name>
          ,
          <source>Dominique Brunato, and Felice Dell'Orletta</source>
          .
          <year>2021</year>
          .
          <article-title>A NLP-based stylometric approach for tracking the evolution of l1 written language competence</article-title>
          .
          <source>Journal of Writing Research (JoWR)</source>
          ,
          <volume>13</volume>
          (
          <issue>1</issue>
          ):
          <fpage>71</fpage>
          -
          <lpage>105</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Simonetta</given-names>
            <surname>Montemagni</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Tecnologie linguisticocomputazionali e monitoraggio della lingua italiana</article-title>
          .
          <source>Studi Italiani di Linguistica Teorica e Applicata (SILTA)</source>
          , pages
          <fpage>145</fpage>
          -
          <lpage>172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Hwee</given-names>
            <surname>Tou</surname>
          </string-name>
          <string-name>
            <given-names>Ng</given-names>
            , Siew Mei Wu,
            <surname>Yuanbin Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Hadiwinoto</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Joel</given-names>
            <surname>Tetreault</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The CoNLL2013 Shared Task on Grammatical Error Correction</article-title>
          .
          <source>In Proceedings of the Seventeenth Conference on Computational Natural Language Learning: Shared Task</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          , Sofia, Bulgaria. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>