<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AcCompl-it @ EVALITA2020: Overview of the Acceptability &amp; Complexity Evaluation Task for Italian</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dominique Brunato</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cristiano Chesi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felice Dell'Orletta</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simonetta Montemagni</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giulia Venturi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Zamparelli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ILC-CNR</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Via G. Moruzzi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>NETS-IUSS</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>P.zza Vittoria</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pavia</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>CIMeC-UNITRENTO</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Corso Bettini</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rovereto Italy [name.surname]@ilc.cnr.it</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>cristiano.chesi@iusspavia.it</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>roberto.zamparelli@unitn.it</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>The Acceptability and Complexity evaluation task for Italian (AcCompl-it) was aimed at developing and evaluating methods to classify Italian sentences according to Acceptability and Complexity. It consists of two independent tasks asking participants to predict either the acceptability or the complexity rate (or both) of a given set of sentences previously scored by native speakers on a 1-to-7 points Likert scale. In this paper, we introduce the datasets distributed to the participants, we describe the different approaches of the participating systems and provide a first analysis of the obtained results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The availability of annotated resources and
systems aimed at predicting the level of
grammatical acceptability or linguistic complexity of a
sentence (see, among others,
        <xref ref-type="bibr" rid="ref2 ref23">(Warstadt et al., 2018;
Brunato et al., 2018)</xref>
        ) is becoming increasingly
relevant for different research communities that
focus on the study of language. From the Natural
Language Processing (NLP) perspective, the
interest has been recently prompted by automatic
generation systems (e.g. Machine Translation, Text
Simplification, Summarization) mostly based on
Deep Neural Networks algorithms
        <xref ref-type="bibr" rid="ref10 ref15 ref23 ref24 ref4">(Gatt and
Krahmer, 2018)</xref>
        . In this scenario, resources and
methods able to assess the quality of automatically
generated sentences or devoted to investigate the
ability of artificial neural networks to score linguistic
phenomena on the acceptability and complexity
scales are of pivotal importance. From the
theoretical linguistics perspectives, controlled datasets
containing acceptability judgments and analyzed
with machine learning techniques can be useful
to test the extent to which syntactic and semantic
deviance can be induced from corpus data alone,
especially for low frequency phenomena
        <xref ref-type="bibr" rid="ref10 ref12 ref15 ref23 ref24 ref24 ref4">(Chowdhury and Zamparelli, 2018; Gulordava et al., 2018;
Wilcox et al., 2018)</xref>
        , while the same data, seen
from a psycholinguistic angle, can shed light on
the relation between complexity and
acceptability
        <xref ref-type="bibr" rid="ref3 ref5">(Chesi and Canal, 2019)</xref>
        , and on the extent to
which measures of on-line perplexity in artificial
language models can track human parsing
preferences
        <xref ref-type="bibr" rid="ref13 ref9">(Demberg and Keller, 2008; Hale, 2001)</xref>
        .
      </p>
      <p>
        The Acceptability &amp; Complexity evaluation
task for Italian (AcCompl-it) at EVALITA 2020
        <xref ref-type="bibr" rid="ref1">(Basile et al., 2020)</xref>
        is in line with this
emerging scenario. Specifically, it is aimed at
developing and evaluating methods to classify Italian
sentences according to Acceptability and
Complexity, which can be viewed as two simple
numeric measures associated with linguistic
productions. Among the outcomes of the task, we also
include the creation of a set of sentences annotated
with acceptability and complexity human
judgments that we are going to share with the
linguistic community. While datasets annotated for
acceptability exist for English, see in particular
the COLA dataset
        <xref ref-type="bibr" rid="ref23">(Warstadt et al., 2018)</xref>
        , to our
knowledge the present dataset is a first for Italian,
and is also the first one to combine judgments of
acceptability and complexity.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Definition of the task</title>
      <p>
        We conceived AcCompl-it as a prediction task
where participants were asked to estimate the
average acceptability and complexity score of a set of
sentences previously rated by native speakers on
a 1-7 Likert scale and, if possible, to predict the
actual standard error (SE) among the annotations.
SE gives an estimation of the actual agreement
between human annotators: the highest the SE, the
lowest the agreement. The task is articulated in
three subtasks, as follows:
• the Acceptability prediction task (ACCEPT),
where participants have to estimate the
acceptability score of sentences (along with
their standard error); in this case, 1
corresponds to the lowest degree of acceptability,
while 7 corresponds to the highest level. The
assignment of a score on a gradual scale is in
line with the definition of perceived
acceptability that we intend to empirically inspect.
According to the literature, in fact,
acceptability is a concept closely related to
grammaticality but with some major differences
(see, among others,
        <xref ref-type="bibr" rid="ref20 ref21">(Sprouse, 2007; Sorace
and Keller, 2005)</xref>
        ). While the latter is a
theoretical construction corresponding to
syntactic wellformedness and it is typically
interpreted as a binary property (i.e., a sentence
is either grammatical or ungrammatical),
acceptability can depend on many factors, such
as syntactic, semantic, pragmatic, and non–
linguistic factors;
• the Complexity prediction task (COMPL),
where participants have to estimate the
complexity score of sentences (along with their
standard error); in this case, 1 corresponds to
the lowest possible level of complexity, while
7 indicates the highest degree. Similarly to
the Acceptability prediction task, the use of
the Likert scale as a tool to collect perceived
values is motivated by the assumption that
sentence complexity is a gradient, rather than
binary, concept;
• the Open task, where participants are
requested to model linguistic phenomena
correlated with the human ratings of sentence
acceptability and/or complexity in the datasets
provided.
      </p>
      <p>The three subtasks were independent and
participants could decide to participate in any one of
them, though we encouraged participation in
multiple subtasks, since the complexity metrics might
be influenced by the grammatical status of an
expression and vice versa. In line with this intuition,
we distributed a subset of sentences annotated
with both acceptability and complexity scores in
order to investigate whether and to what extent
there is a correlation between the two phenomena.</p>
      <p>In all subtasks, participants were free to use
external resources, and they were evaluated against a
blind test set.
3
3.1</p>
    </sec>
    <sec id="sec-3">
      <title>Dataset</title>
      <sec id="sec-3-1">
        <title>Composition</title>
        <p>Acceptability dataset: it contains 1,683 Italian
sentences annotated with human judgments of
acceptability on a 7–point Likert scale. The
number of annotations per sentence ranges from 10
to 85, with an average of 16.38. The dataset
is constructed by merging the data of four
psycholinguistic studies on minimal variations of
controlled linguistic oppositions with different levels
of grammaticality with a subset of 672 sentences
generated from templates.</p>
        <p>
          The first subset (128 sentences), taken from
          <xref ref-type="bibr" rid="ref3 ref5">(Chesi and Canal, 2019)</xref>
          , focuses on person
features oppositions in object clefts dependencies
where Determiner Phrases (DPs) are either
introduced by determiners or by pronouns used as
determiner as in (1).
(1) {Sono j siete} {gli j voi} architetti che {gli j
{are3Ppl j are2Ppl} {the j you} architects that {the j
voi} ingegneri {hanno j avete} consultato.
you} engineers {have3Ppl j have2Ppl} consulted
‘it is {thejyou} architects that {thejyou}
engineers have consulted’
The second subset (515 sentences) is taken from
the studies presented in
          <xref ref-type="bibr" rid="ref11">(Greco et al., 2020)</xref>
          involving copular constructions (e.g. canonical (2a)
vs. inverse (2b)
          <xref ref-type="bibr" rid="ref16">(Moro, 1997)</xref>
          .
(2) a. Le foto del muro sono la causa della
the pictures of_the wall are the cause of_the
rivolta.
        </p>
        <p>riot
b. La causa della rivolta sono le foto
the cause of_the riot are the pictures
del muro.</p>
        <p>of_the wall
This subset also contains declarative and
interrogative (yes/no) sentences with a minimal
verbal structure (contrasting preverbal vs postverbal
subject position in unergatives (3a), unaccusatives
(3b) and transitive predicates (3c))
(3) a. I cani hanno abbaiato j Hanno abbaiato i
the dogs have barked j have barked the
cani.
dogs
b. Gli autobus sono partiti j Sono partiti gli
The buses have left j have left the
autobus.</p>
        <p>
          buses
c. Le bambine hanno mangiato il dolce j
the girls have eaten the dessert j
Hanno mangiato le bambine il dolce
have eaten the girls the dessert
The third set (320 sentences) is based on a study in
which number and person subject-verb agreement
and unagreement cases are tested
          <xref ref-type="bibr" rid="ref15">(Mancini et al.,
2018)</xref>
          :
(4) Qualcuno ha detto che io1Psg {scrivo1Psg j*
Somebody has said that I1Psg {write1Psg j
scriviamo1Ppl} una lettera.
        </p>
        <p>
          *write1Psg} a letter
The fourth one (48 sentences) contains
experimental items from
          <xref ref-type="bibr" rid="ref22">(Villata et al., 2015)</xref>
          involving
different types of wh-islands violations.
(5) {Cosa j Quale edificio}i ti chiedi {chi j
{What j Which building}i do you wonder {who j
quale ingegnere} abbia costruito i?
which engineer} has built i?
The last set of 672 sentences was generated by
creating all the possible content word combinations
from various structural templates designed to test
acceptability patterns due to: (i) extra or missing
gaps in Wh-extractions (6a) vs. topic
constructions (6b).
(6) a. {Cosa j Quale problema}i lo studente
{what j which problem}i the student
dovrebbe descriver(e) { ij -loi j questo
should describe { ij it j this
problema}?
problem}
b. Questo problemai, lo studente dovrebbe
this problem, the student should
descriver(e) { ij -loi j questo problema}
describe { ij iti j this problem}
(ii) Wh- and relative clauses with gaps inside VP
conjunctions (in all conjuncts, i.e. "Across the
Board", in only one conjunct, or not at all, see e.g.
(7)).
(7) Chii ... Maria vuole chiamar(e) { ij -lo} e
whoi ... Mary wants callinf { ij him} and
il dottore medicar(e) { ij -lo}?
the doctor cure { ij him}?
(iii) embedded Wh-clauses and the possibility of
subextractions from them (similar to (5)).
(8) Quale provvedimento Maria ha saputo {che j
Which measure M. has heard {that j
dove j perché j quando} il ministro prenderà?
where j why j when} il ministro prenderà?
(iv) extractions from VPs in subject vs. object
positions (9) (cf. (2)).
(9) Carlo conosceva bene il compagnoi di classe
Carlo knew well the classmatei
che {incontrare i divertiva sempre Anna j Anna
that {meetinf i amused always Anna j Anna
voleva sempre incontrare i}
wanted always meetinf i}
(v) NEGPOLs (nessuno, alcunché, mai ‘any,
anything, ever’) that are licensed by a higher
negation, by a question, or not licensed, in simple or
(deeply) embedded sentences (e.g. (10)).
(10) {Maria j Nessuno} si aspetta che qualcuno
{M. j No-one} self expects that someone
possa aver {già j mai} finito questo
could have {already j never} completed this
esercizio (?)
exercise (?)
The use of expanded templates was designed to
minimize the potential effect of collocations or
specific lexical choices.
        </p>
        <p>Whenever possible each sentence was also
manually annotated according to the
linguistictheoretic expectations for “grammaticality”, on a
4-points scale: * (ungrammatical, coded as 0), ??
(very marginal, coded as 0.66), ? (marginal, coded
as 0.33) and OK (grammatical, coded as 1).</p>
        <p>
          Complexity dataset: it comprises 2,530
Italian sentences annotated with human judgments of
perceived complexity on a 7–point Likert scale
as for the acceptability dataset. The number of
annotations per sentence ranged from 11 to 20,
with an average of 16.753. The corpus was
internally subdivided into two subsets
representative of two different typologies of data, i.e. 1,858
naturalistic sentences extracted from corpora and
672 artificially-generated sentences drawn from
the Acceptability dataset, and chosen to cover the
range of linguistic phenomena represented in its
templates. The first subset contains sentences
taken from the Universal Dependency (UD)
treebanks
          <xref ref-type="bibr" rid="ref17">(Nivre et al., 2016)</xref>
          available for Italian,
representative of different text genres and
domains. In this regard, the largest portion
contains 1,128 sentences taken from the newswire
section of the Italian Stanford Dependency
Treebank (ISDT) (Bosco et al., ), annotated with
complexity judgments by Brunato et al. (2018).
Beside these, we chose to include in this corpus
smaller subsets of sentences representative of a
non-standard language variety and of specific
constructions, i.e. Wh-questions and direct speech.
Non-standard sentences (for a total of 323) are
in the form of generic tweets and tweets labelled
for irony taken from two representative treebanks,
i.e. PoSTWITA and TWITTIRÒ
          <xref ref-type="bibr" rid="ref18 ref5">(Sanguinetti et
al., 2018; Cignarella et al., 2019)</xref>
          . Wh-questions
(164 sentences) were extracted from a dedicated
section (prefixed by the string ‘quest’) included in
ISDT. Direct speech sentences (243) mainly
include transcripts of European parliamentary
debates (taken from the ‘europarl’ section of ISDT)
and extracts from literary texts (mostly contained
in the UD Italian VIT
          <xref ref-type="bibr" rid="ref7">(Delmonte et al., 2007)</xref>
          ).
The choice of annotating a shared portion of data
with both acceptability and complexity scores was
explicitly motivated by the attempt to empirically
investigate whether there is a correlation between
the two sentence properties, and whether
complexity is judged differently in the case of ill-formed
constructions.
        </p>
        <p>For the purpose of the task, both datasets were
split into training and validation samples with a
proportion of 80% to 20%, respectively.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Annotation with Human Judgments</title>
        <p>
          For the collection of judgments of sentence
acceptability and complexity by Italian native
speakers we relied on crowdsourcing techniques using
different platforms. More specifically, for
Acceptability, the set of sentences drawn from the
psycholinguistic studies described in Section 3.1
was annotated using an on-line platform based on
jsPsych scripts
          <xref ref-type="bibr" rid="ref6">(De Leeuw, 2015)</xref>
          . For the
Complexity dataset, the annotation of the subcorpus of
sentences taken from
          <xref ref-type="bibr" rid="ref2">(Brunato et al., 2018)</xref>
          was
performed through the CrowdFlower platform1
(more details are reported in the reference paper),
while the remaining sentences in this dataset were
annotated using Prolific2. To make the
annotation process comparable to the one followed by
          <xref ref-type="bibr" rid="ref2">(Brunato et al., 2018)</xref>
          , the whole process was split
into different tasks, each one consisting in the
annotation of about 200 sentences randomly mixed
for the various typologies. For all tasks, workers
1Now known as Figure Eight, https://appen.com/
2www.prolific.co
were asked to read each sentence and answer the
following question:
“Quanto è complessa questa frase da 1
(semplicissima) a 7 (molto difficile)?”
‘How difficult is this sentence from 1
(very easy) to 7 (very difficult)?’
        </p>
        <p>Beyond complexity, the 672
artificiallygenerated sentences were also labelled for
perceived acceptability according to the following
question:
“Quanto è accettabile questa frase da
1 (completamente agrammaticale) a 7
(perfettamente grammaticale)?” ‘How
acceptable is this sentence from 1
(completely ungrammatical) to 7 (completely
grammatical)?’</p>
        <p>After collecting all annotations, we excluded
workers who performed the assigned task in less
than 10 minutes, which we set as the minimum
threshold to accurately complete the survey.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Analysis of Judgments across Corpora</title>
        <p>Table 1 shows the average value, standard
deviation and minimum and maximum score of
complexity and acceptability labels for the whole
dataset. As it can be noticed, complexity values
are on average lower and less scattered than the
acceptability ones. For this corpus, the lowest value
on the Likert scale (1) – which should have been
used to label sentences perceived as very easy, in
line to the task question – was given only twice,
specifically to the following sentences:
(11) Dimmi il nome di una città finlandese.
tell me the name of a town Finnish
‘Tell me the name of Finnish town’
(12) Quali uve si usano per produrre vino?</p>
        <p>Which grapes PRT they_use to make wine?
Conversely, for the acceptability corpus, the
highest value on the Likert scale (i.e. 7, meaning in
this case completely acceptable) was attributed to
26 sentences. For space reasons, we report here
only two examples:
(13) Le sorelle sono sopravvissute.</p>
        <p>The sisters are survived.</p>
        <p>‘the sisters have survived’
(14) I lupi hanno ululato.</p>
        <p>The wolves have howled.</p>
        <p>With respect to the ‘worst’ values, two sample
sentences judged respectively as the most complex
(i.e. 6.46 on the Likert scale) and (among) the least
acceptable (1.55) in each dataset are the following
ones, respectively:
(15) Chi è che lui ha affermato che il professore
who is that he has claimed that the professor
aveva detto che lo studente avrebbe dovuto
had said that the student hadsubj must
considerare questo candidato?
consider this candidate?
(16) Il falegname è arrivato mentre noi
The carpenter has arrived while we
montavo la mensola.</p>
        <p>were_assembling1Psg the shelf.
min
max</p>
      </sec>
      <sec id="sec-3-4">
        <title>COMPL</title>
        <p>SCORE
3.12
1.04
1
6.46</p>
        <p>SE
0.332
0.08</p>
        <p>0
0.63</p>
      </sec>
      <sec id="sec-3-5">
        <title>ACCEPT</title>
        <p>SCORE
4.45
1.7
1.13
7</p>
        <p>SE
0.36
0.14</p>
        <p>0
0.74</p>
        <p>If we consider the internal composition of the two
datasets we can see a more articulated picture
depending on its various subparts (see Table 2 and
3). For complexity, average scores are higher for
sentences created to display specific acceptability
patterns, thus proving that acceptability does
affect the perception of complexity. Note that the
most complex sentence (reported in (15)) is
contained in this set, and is ungrammatical (no gap).</p>
        <p>Among the treebank sentences, those extracted
from journalistic texts (ISDT_news) were judged
on average as the most complex, questions as the
easiest ones. Twitter and direct speech sentences
obtained scores in between the highest and the
lowest value and very close to each other. This is
in line with stylistic and linguistic analysis
showing that the language of social media inherits many
features from spoken language.</p>
        <p>For the whole acceptability dataset, the
Spearman’s rank correlation coefficient between
theoretically-driven grammaticality and mean
acceptability labels is very strong (r(656)=.83,
p&lt;.001). While this could be somehow expected,
when we focus only on the 672 sentences
annotated for both complexity and acceptability, we
still observe a significant but lower correlation
between expected grammaticality and mean
complexity (r(656)=.34, p&lt;.001). Still considering
this subset, an additional outcome is the moderate
(and negative) correlation between the two metrics
(r(672)=.49, p&lt;.001), further suggesting that the
more a sentence is perceived as complex, the less
acceptable it is.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation measures</title>
      <p>For both the ACCEPT and COMPL Task, the
evaluation metric was based on Spearman’s rank
correlation coefficient between the participants’
scores and the test set scores. For each task, two
different ranks were produced according to the
prediction of the relative scores and to standard
errors. In each task a different baseline was defined:
• in the ACCEPT task, it corresponds to the
score assigned by a SVM linear regression
using unigram and bigram of words as
features;
• in the COMPL task, it corresponds to the
score assigned by a SVM linear regression
using sentence length as its sole feature.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Participation and results</title>
      <p>The AcCompl-it task received three submissions
for each subtask from two different participants,
for a total of 6 runs. Unfortunately, neither
participant took part in the Open Task. Results for the
other ones are reported in Tables 4 and 5.</p>
      <p>
        The systems from the two participants in the
task follow very different approaches: one is based
on deep learning and trained on raw texts
        <xref ref-type="bibr" rid="ref19">(Sarti,
2020)</xref>
        , the other relies on (heuristic) rules applied
to semantic and syntactic features automatically
extracted from sentences
        <xref ref-type="bibr" rid="ref8">(Delmonte, 2020)</xref>
        . In
spite of their very different nature, the two
approaches also present some commonalities, such
as the reliance on external resources. In
particular, both make use of additional sentences taken
from existing Italian treebanks, either to enrich the
original training sets with additional annotated
examples (Sarti’s case) or to check the frequency of
a given construction and use this info among the
features of the proposed system (ItVenses).
      </p>
      <p>Sarti’s systems obtained the best performance
on both tasks using a similar multi-task learning
(MTL) approach, which consists in leveraging the
predictions of a state-of-the-art neural language
model for Italian (i.e. UmBERTO3) fine-tuned on
the two downstream tasks to augment the original
development sets with a large set of unlabeled
ex3https://github.com/musixmatchresearch/umberto</p>
      <sec id="sec-5-1">
        <title>PARTICIPANT</title>
        <p>UmBERTO-MTSA (Sarti)
ItVenses-run1 (Delmonte)
ItVenses-run2 (Delmonte)
Baseline
amples extracted from available Italian treebanks.
The bigger dataset was then split into different
portions to train an ensemble of classifiers. The
resulting MTL model was finally used to predict
the complexity/acceptability labels on the original
test sets.</p>
        <p>Delmonte’s ItVenses system parses the
sentences to obtain a sequence of constituents and a
set of sentence-level semantic features (presence
of agreement, negation markers, speech act and
factivity). These features, along with constituent
triples and their frequency in the training set and
in the Venice Italian Treebank are weighed with
various heuristics and used to derive a
prediction. Agreement mismatches were checked using
morphological analysis of verb and subject, while
the argumental structure is inferred using a deep
parser. The two versions of the system (run1 and
run2) differ only in their use of features (run2
dispenses with proposional negation and certain verb
agreement features).</p>
        <p>As it can be seen, ItVenses’s performance were
considerably lower than Sarti’s system (lower, in
fact, than the baseline based on sentence length, in
the COMPL prediction task). However, as better
explained in the following section, in the artificial
data subset, which has complex but far less diverse
structures, the gap with the winning system is
reduced in the COMPL task (cfr. Table 7) and, even
more robustly, in the ACCEPT task (Table 6).
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>
        The extremely good performance of the winning
system in both tasks is not wholly unexpected in
light of the impressive results obtained by
current neural networks models across a variety of
NLP tasks. In this regard, it is worth noticing
that, in his report, the author compared the
performance of the best system based on multi-task
learning to the one obtained by a simpler version
of the UmBERTO-based model with standard
finetuning on the two downstream tasks, achieving
already very good results (.90 and .84 for
acceptability and complexity predictions on the training
corpus, respectively). Similarly, and especially
for the automatic assessment of sentence
acceptability, the scores obtained by the winning system
(.88) are in line with those reported in
        <xref ref-type="bibr" rid="ref14">(Linzen et
al., 2016)</xref>
        , who train a classifier to detect
subjectverb agreement mismatches from the hidden states
of an LSTM, achieving a .83 score. Most other
systems at work on the ability of neural models to
detect acceptability or grammaticality in a broader
range of cases report much lower scores, but they
try to read (minimal pair) judgments from metrics
associated to the performance of systems that have
not been expressly trained on giving judgments,
reasoning that ‘judgment giving’ is not a task
humans have a life-long training for, but which is
nonetheless feasible.
      </p>
      <p>To have a better understanding of the potential
impact of different types of data on the
predictive capabilities of the two systems, we further
inspected the final results by testing each system on
sentences representative of diverse linguistic
phenomena and textual genres. To this end, we split
the whole test set into the distinct subsets defined
in the corpus collection process (cfr. Section 3.1)
and we assessed the correlation score between
predicted and real labels for each type: note that, for
the ACCEPT predictions, this analysis was
performed considering only two ‘macro–classes’, i.e.
artificial vs psycholinguistics-related data, in
order to have a significant number of examples in
the test set. Similarly, for COMPL, we
distinguished the artificially–generated sentences from
sentences drawn from all treebanks. Results of this
fine-grained analysis are shown in Tables 6 and 7.</p>
      <p>Interestingly, although the gap between the two
systems is still evident, we observed that artificial
data have an opposite effect on their performance.
In particular, as anticipated in the previous section,
ItVenses is more accurate in predicting both the
complexity and, especially, the acceptability level
of this group of sentences. The opposite holds for
Sarti’s system, which although still very good in
both tasks, achieves lower correlation scores when
tested against artificial data.</p>
      <p>Running an exploratory analysis based on
expected grammaticality, we observed that Sarti’s
system performs much better in predicting the
acceptability score on expected grammatical
sentences (r=.80, p&lt;.001) than on expected
ungram</p>
      <sec id="sec-6-1">
        <title>PARTICIPANT SCORE</title>
      </sec>
      <sec id="sec-6-2">
        <title>Psycholinguistics related</title>
        <p>UmBERTO-MTSA (Sarti) 0.90**
ItVenses-run1 (Delmonte) 0.42**
ItVenses-run2 (Delmonte) 0.50**</p>
      </sec>
      <sec id="sec-6-3">
        <title>Artificial data</title>
        <p>UmBERTO-MTSA (Sarti) 0.74**
ItVenses-run1 (Delmonte) 0.50**
ItVenses-run2 (Delmonte) 0.46**
0.55**
0.24**
0.48**
0.33**
0.20*
0.25*
matical ones (r=.76, p&lt;.001). Similarly, but less
robustly, the same numerical asymmetry is
observed in both Delmonte’s runs: for grammatical
predictions, RUN1 r=.33, RUN2 r=.35; for
ungrammatical ones RUN1 r=.32, RUN2 r=.34, all
correlations being equally significant (p&lt;.001).</p>
      </sec>
      <sec id="sec-6-4">
        <title>PARTICIPANT SCORE</title>
      </sec>
      <sec id="sec-6-5">
        <title>Treebank sentences</title>
        <p>UmBERTO-MTSA (Sarti) 0.86**
ItVenses-run1 (Delmonte) 0.25**
ItVenses-run2 (Delmonte) 0.24**</p>
      </sec>
      <sec id="sec-6-6">
        <title>Artificial data</title>
        <p>UmBERTO-MTSA (Sarti) 0.70**
ItVenses-run1 (Delmonte) 0.44**
ItVenses-run2 (Delmonte) 0.51**
SE
0.61**
0.13*
0.10*
0.06
-0.07
-0.11</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Evalita 2020: Overview of the 7th evaluation campaign of natural language processing and speech tools for italian</article-title>
          .
          <source>In Valerio Basile</source>
          , Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Brunato</surname>
          </string-name>
          , Lorenzo De Mattei, Felice Dell'Orletta,
          <string-name>
            <given-names>Benedetta</given-names>
            <surname>Iavarone</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Giulia</given-names>
            <surname>Venturi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Is this sentence difficult? do you agree</article-title>
          ?
          <source>In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>2690</fpage>
          -
          <lpage>2699</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Cristiano</given-names>
            <surname>Chesi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Canal</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Person features and lexical restrictions in Italian clefts</article-title>
          . Frontiers in Psychology,
          <volume>10</volume>
          :
          <fpage>2105</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Shammur</given-names>
            <surname>Absar</surname>
          </string-name>
          Chowdhury and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Zamparelli</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Rnn simulations of grammaticality judgments on long-distance dependencies</article-title>
          .
          <source>In Proceedings of the 27th international conference on computational linguistics</source>
          , pages
          <fpage>133</fpage>
          -
          <lpage>144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Alessandra</given-names>
            <surname>Teresa</surname>
          </string-name>
          <string-name>
            <surname>Cignarella</surname>
          </string-name>
          , Cristina Bosco, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Presenting TWITTIRÒ-UD: An italian twitter treebank in universal dependencies</article-title>
          .
          <source>In Proceedings of the Fifth International Conference on Dependency Linguistics (Depling</source>
          , SyntaxFest
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Joshua R De Leeuw</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>jspsych: A javascript library for creating behavioral experiments in a web browser</article-title>
          .
          <source>Behavior research methods</source>
          ,
          <volume>47</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Rodolfo</given-names>
            <surname>Delmonte</surname>
          </string-name>
          , Antonella Bristot, and
          <string-name>
            <given-names>Sara</given-names>
            <surname>Tonelli</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>VIT - Venice Italian Treebank: Syntactic and quantitative features</article-title>
          .
          <source>In Proceedings of the Sixth International Workshop on Treebanks and Linguistic Theories.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Rodolfo</given-names>
            <surname>Delmonte</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Venses@AcCompl-it: Computing complexity vs acceptability with a constituent trigram model and semantics</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Vera</given-names>
            <surname>Demberg</surname>
          </string-name>
          and Frank Keller.
          <year>2008</year>
          .
          <article-title>Data from eyetracking corpora as evidence for theories of syntactic processing complexity</article-title>
          .
          <source>Cognition</source>
          ,
          <volume>109</volume>
          (
          <issue>2</issue>
          ):
          <fpage>193</fpage>
          -
          <lpage>210</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Albert</given-names>
            <surname>Gatt</surname>
          </string-name>
          and
          <string-name>
            <given-names>Emiel</given-names>
            <surname>Krahmer</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Survey of the state of the art in natural language generation: Core tasks, applications and evaluation</article-title>
          .
          <source>Journal of Artificial Intelligence Research</source>
          ,
          <volume>61</volume>
          :
          <fpage>65</fpage>
          -
          <lpage>170</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Greco</surname>
          </string-name>
          , Paolo Lorusso, Cristiano Chesi, and
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Moro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Asymmetries in nominal copular sentences: Psycholinguistic evidence in favor of the raising analysis</article-title>
          .
          <source>Lingua</source>
          ,
          <volume>245</volume>
          :
          <fpage>102926</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Kristina</given-names>
            <surname>Gulordava</surname>
          </string-name>
          , Piotr Bojanowski, Edouard Grave, Tal Linzen, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Colorless green recurrent networks dream hierarchically</article-title>
          . arXiv preprint arXiv:
          <year>1803</year>
          .11138.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>John</given-names>
            <surname>Hale</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>A probabilistic earley parser as a psycholinguistic model. In Second Meeting of the North American Chapter of the Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Linzen</surname>
          </string-name>
          , E. Dupoux, and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Goldberg</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Assessing the ability of LSTMs to learn syntax-sensitive dependencies</article-title>
          .
          <source>In Transactions of the Association for Computational Linguistics</source>
          , volume
          <volume>4</volume>
          , pages
          <fpage>521</fpage>
          -
          <lpage>535</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Simona</given-names>
            <surname>Mancini</surname>
          </string-name>
          , Paolo Canal, and
          <string-name>
            <given-names>Cristiano</given-names>
            <surname>Chesi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>The acceptability of person and number agreement/disagreement in Italian: an experimental study</article-title>
          . Lingbuzz preprint: https://ling.auf.net/lingbuzz/005514.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Moro</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>The raising of predicates: Predicative noun phrases and the theory of clause structure</article-title>
          , volume
          <volume>80</volume>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Joakim</given-names>
            <surname>Nivre</surname>
          </string-name>
          ,
          <string-name>
            <surname>Marie-Catherine De</surname>
            <given-names>Marneffe</given-names>
          </string-name>
          , Filip Ginter, Yoav Goldberg, Jan Hajic,
          <string-name>
            <surname>Christopher D Manning</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ryan</surname>
            <given-names>McDonald</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Slav</given-names>
            <surname>Petrov</surname>
          </string-name>
          , Sampo Pyysalo,
          <string-name>
            <given-names>Natalia</given-names>
            <surname>Silveira</surname>
          </string-name>
          , et al.
          <year>2016</year>
          .
          <article-title>Universal dependencies v1: A multilingual treebank collection</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)</source>
          , pages
          <fpage>1659</fpage>
          -
          <lpage>1666</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Cristina Bosco, Alberto Lavelli, Alessandro Mazzei, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Tamburini</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>PoSTWITA-UD: an Italian Twitter Treebank in universal dependencies</article-title>
          .
          <source>In Proceedings of the Eleventh Language Resources and Evaluation Conference (LREC</source>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Gabriele</given-names>
            <surname>Sarti</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>UmBERTo-MTSA @ AcComplit: Improving complexity and acceptability prediction with multi-task learning on self-supervised annotations</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Antonella</given-names>
            <surname>Sorace</surname>
          </string-name>
          and Frank Keller.
          <year>2005</year>
          .
          <article-title>Gradience in linguistic data</article-title>
          .
          <source>Lingua</source>
          ,
          <volume>115</volume>
          :
          <fpage>1497</fpage>
          -
          <lpage>1524</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Jon</given-names>
            <surname>Sprouse</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Continuous acceptability, categorical grammaticality, and experimental synta</article-title>
          .
          <source>Biolinguistics</source>
          , pages
          <fpage>1123</fpage>
          -
          <lpage>134</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Sandra</given-names>
            <surname>Villata</surname>
          </string-name>
          , Paolo Canal, Julie Franck, Andrea Carlo Moro, and
          <string-name>
            <given-names>Cristiano</given-names>
            <surname>Chesi</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Intervention effects in wh-islands: An eye-tracking study</article-title>
          .
          <source>In Architectures and Mechanisms for Language Processing (AMLaP</source>
          <year>2015</year>
          ), pages
          <fpage>195</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Alex</given-names>
            <surname>Warstadt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Amanpreet</given-names>
            <surname>Singh</surname>
          </string-name>
          , and
          <string-name>
            <surname>Samuel R Bowman</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Neural network acceptability judgments</article-title>
          . arXiv preprint arXiv:
          <year>1805</year>
          .12471.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Ethan</given-names>
            <surname>Wilcox</surname>
          </string-name>
          ,
          <string-name>
            <surname>Roger Levy</surname>
          </string-name>
          , Takashi Morita, and Richard Futrell.
          <year>2018</year>
          .
          <article-title>What do rnn language models learn about filler-gap dependencies? arXiv preprint</article-title>
          arXiv:
          <year>1809</year>
          .00042.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>