<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Italian and English Sentence Simplification: How Many Differences?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Martina Fieromonte</string-name>
          <email>eromonte@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dominique Brunato</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felice Dell'Orletta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giulia Venturi</string-name>
          <email>giulia.venturig@ilc.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Pavia Istituto di Linguistica Computazionale “Antonio Zampolli” (ILC-CNR) ItaliaNLP Lab -</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The paper proposes a cross-linguistic analysis of two parallel monolingual corpora conceived for automatic text simplification in two languages, Italian and English. The aim is to find similarities and differences in the process of simplification in two typologically different languages. To carry out the comparison, 1,000 sentences were extracted from the two corpora and annotated with a scheme previously used to annotate simplification phenomena.1</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        In recent years, the availability of parallel
monolingual corpora has boosted the adoption of
datadriven techniques for the task of automatic text
simplification (ATS). These corpora are in general
aligned at sentence level and consist of complex
sentences paired with their simple version.
However, except for English which can rely on two
large parallel corpora, i.e. the Parallel Wikipedia
Corpus2
        <xref ref-type="bibr" rid="ref4">(Coster and Kauchak, 2011)</xref>
        (ParWik) and
the Newsela corpus3
        <xref ref-type="bibr" rid="ref7">(Xu et al., 2015)</xref>
        , these
corpora are scarce or rather small in other languages.
To reduce time and effort required for the
construction of parallel corpora, some works tried new
approaches to automatically or semi-automatically
collect such resources, e.g. Coster and Kauchak
(2011),Yatskar et al. (2010), Brunato et al. (2016),
Tonelli et al. (2016). Moreover to take
advantage of empirical data, most of these resources
were annotated with rules aimed at identifying
the typologies of modifications an original
sentence goes through during the process of
simplification. The inspection can be considered
use1Copyright c 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
      </p>
      <p>
        2http://www.cs.pomona.edu/ dkauchak/simplification/
3https://newsela.com/data/
ful for several reasons: it permits i) to detect and
classify a set of necessary transformations in TS,
ii) to assess if a given corpus complies with user
requirements and simplification tasks and iii) to
evaluate the impact of simplification operations
on target populations. If the corpus investigation
also encompasses a cross-linguistic comparison,
it might also shed light on peculiarities and
similarities underlying the process of simplification
across languages. However, so far this last
issue has been rather ignored with the except
        <xref ref-type="bibr" rid="ref5">ion of
Gonzalez-Dios et al. (2018</xref>
        ), who compared how
macro-simplification operations derived from
different annotation schemes are distributed in
Italian, Basque and Spanish parallel corpora. This
paper intends to explore this under-investigated
perspective and proposes a cross-linguistic analysis of
two parallel monolingual corpora, i.e. the Italian
corpus PaCCSS-IT (Parallel Corpus of Complex–
Simple Aligned Sentences for ITalian)
        <xref ref-type="bibr" rid="ref3">(Brunato et
al., 2016)</xref>
        and the English Parallel Wikipedia
Corpus
        <xref ref-type="bibr" rid="ref4">(Coster and Kauchak, 2011)</xref>
        . Through this
comparison, the paper tries to answer the
following three questions:
      </p>
      <p>1. To what extent can an annotation scheme
conceived for the annotation of simplification in
one language be used to annotate simplifications
in other language?</p>
      <p>2. Are there any differences or similarities in the
distribution and nature of simplification operations
in the two languages?</p>
      <p>3. If we find differences, to what extent do they
depend on language only, or on the type of
corpora?</p>
      <p>To answer these questions, 1,000 paired
sentences were extracted from the two corpora and
annotated with the scheme described in Brunato
et al. (2016). This allows us to carry out a
quantitative and qualitative analysis focused on
understanding the nature of the modifications occurring
in the datasets.</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>Given the relevance of parallel monolingual
corpora in ATS, many projects have driven their
attention on the development of these resources. The
main approaches in the literature vary from the
manual simplification of original texts carried out
by experts (see e.g. Xu et al. (2015) in English,
Bott and Saggion (2014) in Spanish, Brunato et al.
(2015) in Italian), to the alignment of already
existing text collections, containing same-topic
documents written in two different styles, a complex
and a simple one. It is the case of e.g. Coster and
Kauchak (2011) and Tonelli et al. (2016), both
relying on the Wikipedia corpus but in a different
way. The first is based on the alignment between
articles extracted from the standard and the
Simple English Wikipedia, a project started in 2001
containing English Wikipedia pages written in
basic English; the latter relies on the edits that users
had made on the Italian Wikipedia and explicitly
marked as instances of simplification. A further
strategy was envisaged by Brunato et al. (2016),
who first collected a corpus of sentences sharing
the same meaning from a large web corpus, and
then ranked the most similar pairs according to
their linguistic complexity assigned by an
automatic readability assessment system.</p>
      <p>In many cases, existing ATS corpora were also
annotated with rules to make explicit the most
frequent operations occurring in the process of
sentence simplification and distinguishing
different typologies of linguistic phenomena involved
in sentence transformation. The classification
of simplification operations is typically two-level
based, i.e. it contains a few macro-level
operations and for some of them a more specific
subclass which can depend on the size of the unit
affected (e.g. sentence, phrase or word) or the
linguistic level at which the operation applies (i.e.
lexical, syntactic, discourse). Comparing ParWik
with the manually simplified corpus Newsela, Xu
et al. (2015) also noticed that the approach adopted
to construct ATS resources has an impact on the
type of simplification phenomena. For instance,
there are more differences between paired
sentences before and after simplification in Newsela,
suggesting that complex linguistic structures are
often retained in ParWik. Simple sentences in
ParWik contains also longer words, together with a
greater number of function words and punctuation.
Similar differences related to the approach
underlying the construction of parallel corpora were also
observed in Italian. For example, the comparison
reported in Tonelli et al. (2016) between a
corpus of Wikipedia edit stories and two corpora of
heterogeneous texts for young readers manually
simplified according to different strategies (i.e. a
structural and an intuitive one) proved the
existence of differences in terms of the linguistic level
affected by simplification. They concern for
instance the distribution of some simplification
operations and the average of operations per sentence.
As regards the first aspect, in manually
simplified corpora, editors opted for a word-level lexical
substitution, while Wikipedia editors for a
phraselevel substitution. As regards the second aspect,
the Wikipedia edit story corpus contains an
average lower distribution of simplification per
sentence. Though related to these works, our
contribution differs in that it adds a cross-linguistic
level of comparison and also tries to provide an
overview of possible factors affecting the
distribution and the nature of simplification operations in
ATS corpora.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Corpora and annotation scheme</title>
      <p>Corpora. The corpora used in the analysis are the
Italian corpus PaCCSS-IT and the English
Parallel Wikipedia Corpus (ParWik). PaCCSS-IT is a
parallel corpus composed of about 63,000 paired
sentences, obtained crawling the web. The
corpus is the result of a three-step approach strongly
shaped by the level of simplification under
investigation, i.e. syntactic simplification, consisting in:
i) an unsupervised step in which a great amount
of sentences with overlapping lexicon and
different syntactic structure was clustered according to
a similarity metric and automatically aligned 4; ii)
a supervised step aimed to train a classifier to
predict the sentence alignment and iii) a readability
assessment step aimed at assigning a readability
score to the sentences in each pair. ParWik
instead was obtained aligning two already existing
text collections: the English Wikipedia and the
Simple English Wikipedia. The authors aligned
paragraphs whose TF*IDF cosine similarity was
over a threshold of 0.5. The final corpus consists
of 167,000 aligned sentence pairs.</p>
      <p>To summarize, the two corpora differ in the
fol4To be part of a cluster a sentence had to share all lemmas
with PoS ‘noun’, ‘verb’, ‘numeral’, ‘personal pronoun’ and
‘negative adverb’.
lowing aspects: i) language; ii) corpus collection
approach iii) domain of texts; iv) level of
simplification under investigation.</p>
      <p>i)
ii)
iii)
iv)</p>
      <p>PaCCSS-IT
Italian
Web crawling
Web corpus
Mainly syntax</p>
      <p>ParWik
English
Wiki-based alignment
Encyclopedic</p>
      <p>Lexicon+Syntax</p>
      <sec id="sec-3-1">
        <title>Annotation of simplification operations. The</title>
        <p>comparison was conducted on 1,000 sentence
pairs randomly extracted from the two corpora. To
make possible the comparison, the sentences were
annotated with the scheme in Table 2, previously
conceived to annotate PaCCSS-IT.</p>
        <p>Simplification operations
Deletion
Insertion
Verbal Features
Lexical Substitution
Reordering
Sentence Type</p>
        <p>Residual</p>
        <p>
          The manual annotation was carried out by one
of the authors using the web-based annotation tool
Brat5. As reported in the next section, the results
of the manual annotation process provide an
answer to the first question. The adopted schema
originally designed to identify simplification
operations within different typologies of parallel
corpora in another language is able to cover almost
all transformations in ParWik. The main limit is
that the scheme does not take into account one of
the more typical simplification operations, that is
splitting long and complex sentences into one or
more shorter ones
          <xref ref-type="bibr" rid="ref6">(Narayan et al., 2017)</xref>
          . This
is because it was conceived to make explicit the
transformations occurring in the PaCCSS-IT
corpus, which only includes 1:1 pairs, i.e. for each
‘complex’ sentence only one ‘simple’ version
exists. To annotate this operation in ParWik, we used
the tag residual.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Corpora analysis 4</title>
      <p>4.1</p>
      <sec id="sec-4-1">
        <title>Distribution of simplification operations</title>
        <p>Figure 1 reports the average distribution of
simplification operations in the two corpora. As we
5https://brat.nlplab.org/
can see, the first three most frequent operations in
PaCCSS-IT are: ‘deletion’, ‘verbal features’ and
‘insertion’ and in ParWik ‘deletion’, ‘lexical
substitution’ and ‘insertion’. Excluding deletion, the
differences resulted to be statistically significant
for all operations, according to the Chi-squared
test (p value &lt;0.05).</p>
        <p>At first glance, these results seem to suggest that
language-specific factors affect the process of
simplification. However, it is interesting to note that a
qualitative analysis of these findings partially rules
out this hypothesis, suggesting instead to interpret
the differences also in view of the other criteria
reported in Table 1. Specifically, the impact of
language is limited to the different distribution of
the ‘verbal feature’ operation. In PaCCSS-IT, it
represents 29% of the total number of annotated
operations while it is much less frequent in
ParWik (&lt;5 %). In particular, the distribution of this
operation in the Italian corpus is mainly due the
higher number of verbs at the conditional mood,
which are transformed into indicative in the
simplified sentence. As expected, verbs in ParWik
are mostly at the indicative in both versions of the
sentence. However, this different distribution has
to be read also in view of another factor, i.e. the
domain of texts in the corpora. Since it has been
crawled from the web, PaCCSS-IT contains
heterogeneous domains and many complex sentences
belong to a ‘written to be spoken’ style, which
implies the use of polite forms, expressed in Italian
with the conditional mood. As a consequence of
the different domain of texts contained in the two
corpora, we can also observe a gap concerning the
frequency of ‘insertion’. Specifically, the
encyclopedic nature of texts in ParWik may require the
insertion of glosses and explanations to improve
the understanding of complex terms. The lower
frequency of lexical substitution operations in the
Italian corpus (8.9% vs 23.9%) is easily explained
if one considers the main purpose for which the
corpus was designed, i.e. the investigation of
syntactic simplification. On the contrary, editors of
Simple Wikipedia are explicitly recommended “to
write using Basic English words”6.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Linguistic analysis</title>
        <p>The diversity between the two corpora affects also
the nature of the linguistic phenomena subjected
6https://simple.wikipedia.org/wiki/Wikipedia:How to
write Simple English pages
contribute to simplicity of this type of insertion is
quite clear in:</p>
        <p>C: According to the Armenian tradition, Saint Jude
suffered martyrdom about 65 AD in Beirut, in the Roman
province of Syria, together with the apostle Simon the
Zealot, with whom he is usually connected.</p>
        <p>S: St. Jude was martyred, killed for his beliefs, with
another apostle, Simon the Zealot in Beirut, Lebanon,
around AD 65.
to simplification. This means that the type of
linguistic elements which are, for example, deleted,
inserted or substituted might be different. Again
this variance is poorly attributable to the
proprieties of the languages at play. In the following, we
will try to outline a categorization of the
linguistic elements subjected to modifications in the two
corpora, providing an example for each case.</p>
        <p>Deletion. This operation involves the deletion
of single words or clauses. In particular, we
observe a similar trend in the two corpora with the
deletion of functional words, modal adverbs and
adjectives alone or entire clauses containing these
parts of speech.</p>
        <p>C: The main bar at King’s is far older and is the site of
more informal meeting between students. [ParWik]
S: The main bar at King’s is far older. [ParWik]
C: Probabilmente sospetto che non sarebbe
comunque una buona idea. (Probably, I suspect that it
would not be however a good idea.) [PaCCSS-IT]
S: Non fu una buona idea. (It was not a good idea.)
[PaCCSS-IT]</p>
        <p>Insertion. In both corpora auxiliaries and full
verbs are inserted. Moreover in ParWik also nouns
and pronouns are inserted as subjects of the new
sentence, typically as a consequence of a split.
As said, this does not occur in PaCCSS-IT, where
however, implicit-explicit clause transformation
implies the insertion of explicit elements, such as
articles and verbs.</p>
        <p>C: Spese del presente grado di giudizio compensate tra
le parti costituite. (Expense of the present level of
justice compensated among the parts) [PaCCSS-IT]
S: Le spese del presente grado di giudizio possono
essere compensate tra le parti. (The expense of the
present level of justice can be compensated among the
parts). [PaCCSS-IT]</p>
        <p>As said before, ParWik editors tend to insert
explanations of complex terms and concepts. The
C: Velvet Revolver is an American hard rock
supergroup consisting of former Guns N’ Roses members
Slash, Duff McKagan, and Matt Sorum, alongside
Dave Kushner formerly of punk band Wasted Youth.
S: Velvet Revolver, VR, is a Grammy Award-winning
rock supergroup. The members of the band are Slash
guitarist, Duff McKagan bassist, backing vocals,
Matt Sorum drums of Guns N’ Roses, Scott Weiland
lead vocals of Stone Temple Pilots and Dave Kushner
guitarist of Wasted Youth.</p>
        <p>Lexical substitution the more striking
difference between the two corpora concerns this
operation, not only in terms of frequency but also in
respect of the type of substitution. In
PaCCSSIT, the operation affects only the substitution of
words whose PoS was not considered in the
clustering step, e.g. adjectives, adverbs and articles,
etc.</p>
        <p>Moreover the substitution does not always
contribute to the simplification of the sentence:
this means that in some cases the complex term
may be not replaced with a simpler synonym. In
ParWik instead the operation affects phrase and
sentence level, yielding to real paraphrases.</p>
        <p>C: Il concorrente e` preventivamente stato avvertito per
assistere all’operazione (The concurrent had been
informed in advance to assist to the operation
[PaCCSSIT]
S: Il concorrente e` stato avvertito preventivamente,
affinche´ possa assistere all’operazione. (The
concurrent had been informed in advance in order to assist to
the operation) [PaCCSS-IT]
C: Sporting venues in the city include the
Millennium Stadium the national stadium for the Wales
national rugby union team and the Wales national
football team, SWALEC Stadium the home of Glamorgan
County Cricket Club, Cardiff City Stadium the home
of Cardiff City football team and Cardiff Blues rugby
union team, Cardiff International Sports Stadium the
home of Cardiff Amateur Athletic Club and Cardiff
Arms Park the home of Cardiff Rugby Club. [ParWik]
S: Cardiff has one of the largest stadiums in the United
Kingdom, the Millennium Stadium, where important
world sports matches and concerts happen. Other
big stadiums in the city are the Cardiff City Stadium,
where the main football and rugby teams play, and the
SWALEC Stadium where cricket is played. [ParWik]
Verbal features As said before, the Italian
‘conditional!indicative’ transformation does not
occur in the English corpus, where instead the tag
‘verbal features’ was assigned to mark voice
modification and ‘indefinite!finite’ mood
transformations.</p>
        <p>C: Salve, avrei bisogno di una informazione piuttosto
urgente. (Good morning, I would need a rather urgent
information.) [PaCCSS-IT]
S: Ho bisogno di una informazione urgente. (I need a
urgent information.) [PaCCSS-IT]
C: It is most often black but can come in a variety of
colors including clear, allowing the top of the deck to
be decorated. [ParWik]
S: However, it can come in many different colors like
clear. Clear allows the top of the deck to be decorated.
[ParWik]</p>
        <p>Reordering In general, in PaCCSS-IT,
reordering implies the resetting of the canonical word
order, while in ParWik there is a tendency to
transform noun pre-modifiers in appositive phrases. As
regards the position of subordinate clauses, neither
of the two corpora assign to them a fixed position,
i.e. before or after the main clause, although in
ParWik embeddings are often extracted to form a
new sentence.</p>
        <p>C: Un’unica cosa vorrei aggiungere. (Only a thing I
would like to add.) [PaCCSS-IT]
S: Volevo aggiungere solo una cosa. (I wanted to add
only a thing.) [PaCCSS-IT]
C: The United States presidential election of 1992 had
three major candidates: Incumbent Republican
President George H. W. Bush; Democratic Arkansas
Governor Bill Clinton, and independent Texas
businessman Ross Perot. [ParWik]
S: The United States presidential election of 1992 was
on November 3, 1992 in the United States. The three
main people running were: George H. W. Bush, a
Republican from Texas and the President; Bill Clinton,
who was a Democrat and Governor of Arkansas; and
Ross Perot an Independent candidate. [ParWik]</p>
        <p>Sentence type. Three main phenomena fall
under this tag: i) passive-active modification,
ii) implicit-explicit clause modification and iii)
verbalization-nominalization modification. While
the first two modifications occur in both corpora,
the third was found only in ParWik. Again,
this difference is partly affected by
languagedependent factors but it also depends on specific
corpus-dependent constraints.</p>
        <p>C: Il presidente, ricordato che nella seduta di ieri si e`
svolta la relazione, dichiara aperta la discussione
generale. (The president, reminded that the reporting was
held in the yesterday part-session, declares open the
general discussion.) [PaCCSS-IT]
S: Il presidente ricorda che nella seduta di ieri e` stata
svolta la relazione introduttiva e dichiara quindi aperta
la discussione generale. (The president reminds that
in the yesterday part-session was held the introductory
reporting and declares open the general discussion.)
[PaCCSS-IT]
C: Findings of coins indicate that the Romans were in
Buxton throughout their occupation. [ParWik]</p>
        <p>S: Roman coins have been found in Buxton. [ParWik]
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and future works</title>
      <p>The paper proposed a cross-linguistic
comparison between two monolingual parallel corpora
for ATS. The comparison tried to answer three
main questions. As regards question 1, the
annotation stage proved the possibility to use,
except few modifications, a language-specific
annotation scheme for another language. More
than language-specific factors, an in-depth
analysis of the annotated pairs of sentences highlighted
that the observed differences are due to linguistic
phenomena characterizing different textual
genres. This is the case for example of
modifications due to the insertion of glosses, which is
driven by the encyclopedic nature of Wikipedia
pages rather than to the specific language.
Similarly, textual genre has an impact on the
linguistic level involved in the lexical substitution.
The higher occurrence of substitutions at phrase
level, rather than at word-level, reflects the attempt
of Wikipedia editors to make scientific contents
clearer and simpler for a wide target population.
Corpus-design differences, especially those
occurring between manually and automatically derived
corpora, may affect the distribution of the
simplification operations also within the same genre. This
is one of the possible directions that we want to
explore in the near future.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was partially supported by the
2year project ADA, Automatic Data and
documents Analysis to enhance human-based
processes, funded by Regione Toscana (BANDO
POR FESR 2014-2020).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Bott</surname>
          </string-name>
          and
          <string-name>
            <given-names>Horacio</given-names>
            <surname>Saggion</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Text Simplication Resources for Spanish. Language Resources and Evaluation</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>48</volume>
          (
          <issue>1</issue>
          ):
          <fpage>93</fpage>
          -
          <lpage>120</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Brunato</surname>
          </string-name>
          , Felice Dell'Orletta,
          <string-name>
            <given-names>Giulia</given-names>
            <surname>Venturi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Simonetta</given-names>
            <surname>Montemagni</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Design and Annotation of the Frst Italian Corpus for Text Simplication</article-title>
          .
          <source>Proceedings of the 9th Linguistic Annotation Workshop (LAW15)</source>
          , Denver, Colorado, USA.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Brunato</surname>
          </string-name>
          , Andrea Cimino, Felice Dell'Orletta and
          <string-name>
            <given-names>Giulia</given-names>
            <surname>Venturi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>PaCCSS-IT: A Parallel Corpus of ComplexSimple Sentences for Automatic Text Simplication</article-title>
          .
          <source>Methods in Natural Language Processing (EMNLP</source>
          <year>2016</year>
          ), pages
          <fpage>1018</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>William</given-names>
            <surname>Coster</surname>
          </string-name>
          and
          <string-name>
            <given-names>David</given-names>
            <surname>Kauchak</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Simple English Wikipedia: a new text simplification task</article-title>
          .
          <source>Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>I.</given-names>
            <surname>Gonzalez-Dios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Aranzabe</surname>
          </string-name>
          , and A. D´ıaz de Ilarraza.
          <year>2018</year>
          .
          <article-title>The corpus of Basque simplified texts (CBST)</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>52</volume>
          (
          <issue>1</issue>
          )
          <fpage>217</fpage>
          -
          <lpage>47</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Shashi</given-names>
            <surname>Narayan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Claire</given-names>
            <surname>Gardent</surname>
          </string-name>
          and
          <string-name>
            <given-names>Shay B.</given-names>
            <surname>Cohen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Anastasia</given-names>
            <surname>Shimorina</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Split and Rephrase</article-title>
          .
          <source>Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Wei</surname>
            <given-names>Xu</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chris</surname>
            Callison-Burch, and
            <given-names>Courtney</given-names>
          </string-name>
          <string-name>
            <surname>Napoles</surname>
          </string-name>
          .
          <year>2015</year>
          . Problems in Current Text Simplication Research:
          <article-title>New Data can Help</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          ,
          <volume>3</volume>
          :
          <fpage>283</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Sara</given-names>
            <surname>Tonelli</surname>
          </string-name>
          , Alessio Palmero Aprosio,
          <string-name>
            <given-names>Francesca</given-names>
            <surname>Saltori</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>SIMPITIKI: a Simplification corpus for Italian</article-title>
          .
          <source>Proceedings of the Third Italian Conference on Computational Linguistics</source>
          , Naples, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Mark</given-names>
            <surname>Yatskar</surname>
          </string-name>
          , Bo Pang, Cristian DanescuNiculescuMizil, and
          <string-name>
            <given-names>Lillian</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>For the Sake of Simplicity: Unsupervised Extraction of Lexical Simplications from Wikipedia</article-title>
          . In Human Language Technologies:
          <article-title>The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics</article-title>
          ,
          <source>HLT '10</source>
          , pages
          <fpage>365</fpage>
          -
          <lpage>368</lpage>
          , Stroudsburg, PA, USA. Association for Computational Linguistic.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>