<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Fixing Comma Splices in Italian with BERT</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rene´e E. D'Aoust</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Daniele Puccinelli DTI-ISIN University of Applied Sciences of Southern Switzerland Manno</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>North Idaho College Sandpoint, Idaho, USA Casper College Casper</institution>
          ,
          <addr-line>Wyoming</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Silvia Demartini DFA-DILS University of Applied Sciences of Southern Switzerland Locarno</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We propose a fully unsupervised strategy to fix comma splices. Leveraging the pre-training of Bidirectional Encoder Representations from Transformers (BERT), our strategy is to mask out commas and let BERT guess what to replace them with. Our strategy achieves promising results on a challenging targeted corpus of awkwardly worded sentences from Italian-language college student essays.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Comma splices can be defined as independent
clauses joined by a comma without a
coordinating conjunction
        <xref ref-type="bibr" rid="ref7">(Hacker, 2009)</xref>
        . Comma
splices are frequent in both English and
Italian and typically suggest a lack of basic
understanding of sentence structure. As we will
show, they come in various flavors, and there
exist subtle differences between how they
occur in English and Italian.
      </p>
      <p>Comma splices are generally detected by
commercial grammar and style checkers, but
their automated correction has only been
addressed by a few studies specific to English.
Because the common denominator shared by
such studies is the use of supervised machine
learning techniques, the key research question
that motivated the present study is whether we
can use transfer learning to correct comma
splices automatically in a completely
unsupervised fashion and in languages other than
English.</p>
      <p>
        Thanks to contextualized word
embeddings, and, in particular, thanks to BERT
        <xref ref-type="bibr" rid="ref3">(Devlin et al., 2019)</xref>
        , we show that it is possible
to correct common cases of comma splices in
Italian. We also discuss the limitations of our
unsupervised approach.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Comma splices in Italian</title>
      <p>
        Comma splices are widespread in
contemporary written Italian language usage due to a
tendency to over-extend the use of commas
        <xref ref-type="bibr" rid="ref2 ref2 ref5 ref6 ref6">(Ferrari, 2017, 2018; Demartini and Ferrari,
2018)</xref>
        . Several authors have studied this
tendency in recent years. Some preserve the
English language designation; this is the case
in
        <xref ref-type="bibr" rid="ref1">(Corno, 2019)</xref>
        , where the expression frasi
fuse (fused sentences) is also employed.
Others employ alternate designations, such as
virgola passe-partout (passe-partout comma) in
        <xref ref-type="bibr" rid="ref14">(Tonani, 2010)</xref>
        and virgola tuttofare (factotum
comma) in
        <xref ref-type="bibr" rid="ref13">(Serianni and Benedetti, 2009)</xref>
        .
      </p>
      <p>Comma splices are one of the most
frequent comma usage errors in Italian,
especially among inexperienced L1 and L2
writers. Comma splices are also one of the
principal and most common problems in the writing
of university students, especially in science
and engineering. Usually, these writers have
failed to develop any linguistic awareness for
text segmentation and organization, and they
mistakenly assume that a comma can convey
multiple functions, working both as a linker
or as a strong stop.</p>
      <p>There are some similarities and some
differences compared to English usage, due to
the fact that Italian punctuation is more
communicative and less morphosyntactic. In
general, there are two main kinds of comma
splices in Italian that are caused by the use of
a comma where we would expect:
1. a logical connector to join two sentences
that have a particular relationship;
2. a stronger punctuation mark to mark a
logical-syntactic connection (colon) or
break (semicolon or period).</p>
      <p>
        According to
        <xref ref-type="bibr" rid="ref4">(Ferrari, 2014)</xref>
        , comma
splices reflect a deep inability to handle both
basic syntactic structures and text
construction: if a text is characterized by coherence,
cohesion, and topical organization, comma
splices deconstruct these properties from the
inside. For this reason, analyzing comma
splices is extremely important in the context
of improving language teaching.
      </p>
      <p>Comma splices can be fixed in various
ways, depending on the context and on the
kinds of clauses involved. In the most
straightforward cases, the comma can be
replaced by a period or a semi-colon that
explicitly separates the clauses on either side of
the comma. In other cases, the comma can be
replaced by an element that links the clauses,
such as a colon, a conjunction, or a
conjunctive adverb. Care must be exercised if
sentences are more complex (i.e. with
parenthetical elements) or syntactically inaccurate.</p>
      <p>Due to the lack of an Italian-language
corpus dedicated to comma splices, the authors
have assembled a small corpus of 100
sentences containing a wide array of comma
splices collected from college student
writings (mostly in the field of engineering) at the
Universita` del Piemonte Orientale (UPO) and
the University of Applied Sciences of
Southern Switzerland (SUPSI) in the mid-to-late
2010s. In the remainder of the paper, we
will employ this UPO-SUPSI-SPLICE corpus
(henceforth USS corpus) to evaluate the
potential of our proposed method. Aside from
containing comma splices, many USS
sentences are poorly worded, syntactically
inaccurate, and often unclear.</p>
    </sec>
    <sec id="sec-3">
      <title>Related work</title>
      <p>
        In the active research thread on automated
grammar and style correction, the studies that
are most closely related to ours are
        <xref ref-type="bibr" rid="ref9">(Lee
et al., 2014)</xref>
        on the automated detection of
comma splices and, most recently, (Zheng
et al., 2018) on the automated correction of
run-on sentences. The techniques proposed in
these studies, which are specific to English,
rely on supervised learning techniques that
require relatively extensive training sets. To the
best of our knowledge, ours is the first
investigation of the automated correction of
Italianlanguage comma splices using unsupervised
learning.
      </p>
      <p>
        Our proposed unsupervised strategy
leverages the rich research thread on word
embeddings. Dense word embeddings went
mainstream with Word2Vec
        <xref ref-type="bibr" rid="ref10">(Mikolov et al., 2013)</xref>
        and gained traction in the mid-to-late 2010s
in spite of their key limitation that a word
type has the same word embedding regardless
of context. Because words also have
different aspects depending on semantics,
syntactic behavior, and register/connotations,
contextualized word embeddings have emerged
as an elegant solution to capture word
semantics across different contexts. TagLM
        <xref ref-type="bibr" rid="ref11">(Peters et al., 2017)</xref>
        uses the hidden state
of the bidirectional long-short term memory
(LSTM)
        <xref ref-type="bibr" rid="ref8">(Hochreiter and Schmidhuber, 1997)</xref>
        as a contextual word embedding. Instead of
just using the output of the LSTM, ELMo
        <xref ref-type="bibr" rid="ref12">(Peters et al., 2018)</xref>
        uses all the available hidden
layers and combines them in a task-specific
way with task-dependent trainable weights
that can be learned for each task. ELMo
embeddings have been shown to improve the
state-of-the-art on a wide variety of
challenging NLP tasks, but even more significant
improvements have been shown with BERT
        <xref ref-type="bibr" rid="ref3">(Devlin et al., 2019)</xref>
        . Based on Transformer
encoders
        <xref ref-type="bibr" rid="ref15">(Vaswani et al., 2017)</xref>
        , which are
essentially a multi-headed attention stack where
depth serves to compensate for the lack of
recurrence, BERT pre-trains bidirectional
representations by jointly conditioning on both the
left and right context of individual tokens, and
allows for low-cost task-specific fine-tuning.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Fixing Comma Splices with BERT</title>
      <p>While bidirectionality comes naturally to
LSTM-based models, it is challenging to
achieve it with Transformer-based models,
because bidirectional conditioning with
multiple layers inherently allows each word to
see itself. BERT’s solution is to mask a
relatively small portion of the tokens in the
pretraining data and to train a bidirectional
language model to guess them. If too few words
are masked, training is too expensive, while
if too many words are masked, BERT fails to
learn about language; it was determined
empirically that masking 15% of all tokens
represents a reasonable compromise.</p>
      <p>This specific aspect of BERT’s pre-training
means that a pre-trained BERT model has the
ability to predict missing tokens out of the
box, i.e., with no task-specific fine-tuning and,
therefore, no need for task-specific training
data. For our purposes, this translates into
a straightforward strategy to correct comma
splices: mask all commas and use BERT
to guess what they should be. In principle,
if a masked comma is legitimate, we expect
BERT to guess it is indeed a comma, while
if it is not, as in a comma splice, we expect
BERT to replace it with a more appropriate
token.</p>
      <p>BERT naturally lends itself to this task
because it outputs an empirical probability
distribution over a set of potential replacement
tokens. Such tokens can be drawn out of
the entire dictionary (including word pieces)
or over a controlled subset. Jointly with the
probabilistic nature of its output, BERT’s
inherent bidirectionality may be directly
harnessed by making predictions based on both
the left and the right context of a masked
comma and choosing the set of predictions
associated with the highest probability.</p>
      <p>If the array of potential replacement token
is unrestricted, in complex sentences BERT
may elect to replace commas with tokens
belonging to inappropriate word classes, such as
nouns or verbs. This can be avoided by
reStrategy Accuracy
Baseline 0.41
BERT - left context only 0.77
BERT - left &amp; right context 0.81</p>
      <p>BERT - PoS + left &amp; right 0.87
stricting the eligible potential replacement
tokens to reasonable word classes.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>As a proof of concept, we perform an
empirical evaluation of our BERT-based
strategy on the USS corpus, which contains
sentences with at least one comma splice and
a total number of commas ranging from one
to seven. To the best of our knowledge, no
directly comparable technique to fix
Italianlanguage comma splices programmatically is
freely available at the time of writing. To get a
rough idea of the potential of our strategy, we
use a simple baseline that replaces all commas
with periods. While this baseline fails each
time a sentence contains multiple commas, it
fixes over 90% of the USS sentences that
contain exactly one comma (41 out of 45). Aside
from setting a performance floor, this baseline
also offers a quick idea of the complexity of
the sentences in the corpus.</p>
      <p>As for our BERT-based strategy to fix
comma splices, we make the following
choices for the sake of simplicity:
we employ
bert-multilingualuncased (and normalize all tokens to
lower case);
we draw potential replacement tokens
out of the entire dictionary (aside from
the PoS-based restrictions described
below), but only consider potential
replacement tokens with an estimated
probability greater than 0.01 (arbitrary
threshold);
we make predictions based on both the
left and the right context of the masked
tokens and choose the prediction
associated with the highest probability,
computed as the product of the probabilities
of the most probable token replacement
for each comma occurrence (we always
mask out one comma at a time);
we use PoS tags to exclude potential
replacement tokens from word classes
other than conjunctions and punctuation
marks.</p>
      <p>We employ TreeTagger to determine the PoS
tags and use pre-trained BERT by way of
pytorch pretrained bert.</p>
      <p>We use sentence-level accuracy as our
figure of merit and compute it as the fraction
of error-free corrected sentences. A sentence
is considered to be error-free by our strategy
and/or by the baseline if the corrected
version is acceptable according to two L1 human
annotators. The corrected versions of
sentences with multiple commas are only
considered error-free if they contain no anomalies;
while this is overly penalizing for our
strategy in multi-comma sentences where a single
mistake is made, it offers a conservative
estimate of the performance of our BERT-based
strategy.</p>
      <p>As shown in Table 1, our BERT-based
strategy is able to correct a total of 87 of the 100
sentences in the USS corpus to the
satisfaction of the two L1 human annotators. An
additional sentence is also corrected, but only if
our strategy operates unidirectionally.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
    </sec>
    <sec id="sec-7">
      <title>Commas per sentence. The mean number</title>
      <p>of commas is 2.1 in the sentences where our
strategy succeeds, while it is as high as 3.5
in the 12 sentences where our strategy fails.
While multi-comma sentences are inherently
more challenging, there doesn’t seem to be a
hard limit to the number of commas per
sentence that our strategy can handle. Notably,
our corpus contains a 7-comma excerpt:
Di solito, chi scrive senza conoscere
le fasi della scrittura, scrive di
getto, seguendo i propri
ragionamenti senza un ordine, cos`ı facendo,
rischia di non scrivere un testo
idoneo e fluente, dobbiamo essere
attenti alle punteggiature, non
scrivere le frasi molto lunghe e dividere
in modo adeguato i capoversi.
which is fixed as</p>
      <sec id="sec-7-1">
        <title>Di solito, chi scrive . . . scrittura,</title>
        <p>scrive di getto, seguendo . . . ordine.</p>
      </sec>
      <sec id="sec-7-2">
        <title>Cos`ı facendo, rischia . . . fluente.</title>
      </sec>
      <sec id="sec-7-3">
        <title>Dobbiamo . . . punteggiature, non</title>
        <p>. . . capoversi.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Failures in single-comma sentences.</title>
      <p>There are two single-comma sentences where
our strategy fails: one contains a run-on
sentence and also causes the baseline strategy
to fail, while the other one has a mild form of
comma splice:</p>
      <sec id="sec-8-1">
        <title>Successivamente avviene la documentazione, si raccolgono e si scelgono le informazioni da fonti attendibili e si pianifica come esporle.</title>
        <p>This sentence is the only instance in USS
where our strategy fails and the baseline
strategy succeeds. BERT chooses not to replace
the comma, keeping the (borderline
acceptable) comma splice unaltered. This happens
due to the relative values of the
probabilities assigned by BERT to a comma and a
colon. Curiously, replacing Successivamente
with the equivalent expression Al passo
successivo is enough to nudge BERT in the right
direction and assign a higher probability to
a colon. This suggests that modifying
individual tokens in a small corpus such as USS
would be a meaningful dataset augmentation
technique.</p>
        <p>Left and right context. For 77 out of 100
sentences, a unidirectional pass based on the
left context of the missing tokens is sufficient
for our strategy to succeed. Only one of these
77 sentences can only be corrected
unidirectionally; five other sentences can also be
corrected by looking at the right context of the
missing tokens in a backward pass, which
helps avoid blatantly erroneous replacements.
Therefore, our strategy should be used with
both left and right context. As an example,
consider the sentence:</p>
      </sec>
      <sec id="sec-8-2">
        <title>Essa consiste nel fatto che non</title>
        <p>c’e` alcun legame naturalmente
motivato, il significante cane non ha
di per se´ nulla che rimandi al suo
nome, che faccia s`ı che quella cosa
si possa chiamare cos`ı.</p>
        <p>If BERT only relies on the left context of
missing commas, the sentence is awkwardly
split into three parts, with a striking error at
the end:</p>
      </sec>
      <sec id="sec-8-3">
        <title>Essa . . . motivato. Il significante</title>
        <p>. . . al suo nome. Che faccia s`ı che
quella cosa si possa chiamare cos`ı.</p>
        <p>With both left and right context, instead,
our strategy offers an acceptable correction:</p>
      </sec>
      <sec id="sec-8-4">
        <title>Essa . . . motivato. Il significante . . . al suo nome e che faccia . . . cos`ı.</title>
        <p>PoS filtering. A further six USS sentences
can be fixed by combining left &amp; right
context and PoS filtering, which serves to avoid
replacement tokens from implausible word
classes and prevent awkward errors, such as
the replacement of a comma with a
preposition, a che, or a negation. Comma
replacements with negations are particularly critical
because they modify the meaning of the
corrected sentence. PoS filtering is also
helpful to prevent BERT from replacing commas
with word pieces, which may occur with
awkwardly worded sentences.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Unacceptable replacements. We have ob</title>
      <p>served a limited number of unacceptable
replacements of commas with colons, all of
which occur in long-winded multi-comma
sentences. Consider the six comma sentence:</p>
      <sec id="sec-9-1">
        <title>Per la creazione della piattaforma web, il committente ha desiderato utilizzare una web application in Java, avendo la possibilit a`</title>
        <p>which becomes:
di scegliere tra due framework,</p>
      </sec>
      <sec id="sec-9-2">
        <title>Spring e Struts, si e` optato per</title>
        <p>l’utilizzo di Spring, siccome e` uno
strumento gi a` utilizzato
precedentemente, possiede un’ottima
documentazione.</p>
        <p>Per la creazione della piattaforma
web. Il committente . . . in Java,
avendo la possibilit a` di scegliere tra
due framework: Spring e Struts. Si
e` optato per l’utilizzo di Spring,
siccome e` uno strumento gi a` utilizzato
precedentemente e possiede . . .</p>
        <p>The first comma is erroneously replaced with
a period, and the third one is questionably
replaced with a colon. Replacing the fourth
comma with a period is acceptable, as is
preserving the second and fifth commas and
turning the sixth comma into an e. Though our
strategy makes four correct decisions out of
six, this example is considered incorrect for
our sentence-level quantitative analysis.</p>
        <p>Interestingly, we have not observed any
replacements with semicolons. We conjecture
that BERT’s strong preference for colons may
be due to the relative frequency of colons
versus semicolons in the pre-training text.
7</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Conclusion</title>
      <p>While the main limitation of the present study
is the limited size of the USS corpus, we
believe that the challenging nature of the writing
excerpts in the USS corpus has enabled us to
stress-test our strategy and to deliver a solid
proof of concept that leverages the power of
BERT-style contextualized word embeddings
for automated style correction. Our future
plans include using our BERT-based strategy
to correct comma splices in English-language
L1 and L2 student writing and to correct
runon sentences.
J.</p>
      <p>Zheng, C. Napoles, J. Tetreault, and
K. Omelianchuk. 2018. How do you correct
run-on sentences it’s not as easy as it seems. In
The 4th Workshop on Noisy User-generated Text
W-NUT. Brussels, Belgium.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Corno</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Scrivere e comunicare. La scrittura italiana in teoria e in pratica</article-title>
          . Pearson.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Demartini</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Ferrari</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>La virgola splice nei testi di studenti universitari: un problema solo in apparenza superficiale</article-title>
          .
          <source>In La Punteggiatura Italiana Contemporanea nella Varieta` dei Testi Comunicativi.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In NAACL '19.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferrari</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Linguistica del testo</article-title>
          . Principi, fenomeni, strutture..
          <source>Carocci.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferrari</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Usi ”estesi” del punto e della virgola nella scrittura italiana contemporanea</article-title>
          .
          <source>La lingua italiana. Storia</source>
          , strutture, testi XIII:
          <fpage>137</fpage>
          -
          <lpage>153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferrari</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>La punteggiatura italiana contemporanea</article-title>
          .
          <source>Unanalisi comunicativo-testuale. Carocci.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Hacker</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>The Bedford Handbook for Writers</article-title>
          . Bedford Books.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Long shortterm memory</article-title>
          .
          <source>Neural Comput</source>
          .
          <volume>9</volume>
          (
          <issue>8</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. Y.</given-names>
            <surname>Yeung</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Chodorow</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Automatic detection of comma splices</article-title>
          .
          <source>In 28th Pacific Asia Conference on Language, Information and Computation pages (PACLIC '14).</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In NIPS '13.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ammar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bhagavatula</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Power</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Semi-supervised sequence tagging with bidirectional language models</article-title>
          .
          <source>In ACL'17</source>
          . Vancouver, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Neumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Iyyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gardner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>In NAACL'18</source>
          . New Orleans, Louisiana.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>L.</given-names>
            <surname>Serianni</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Benedetti</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Scritti sui banchi. L'italiano a scuola fra alunni e insegnanti</article-title>
          . Carocci.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>E.</given-names>
            <surname>Tonani</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Il romanzo in bianco e nero. Ricerche sull'uso degli spazi bianchi e dell'interpunzione nella narrativa italiana dall'Ottocento a oggi</article-title>
          .
          <source>Cesati.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I.</given-names>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Attention is all you need</article-title>
          .
          <source>In NIPS '17.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>