<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Personalised Simplification based on L2 Learners' Native Language</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessio Palmero Aprosioy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Meniniy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sara Tonelliy Luca Ducceschiz</string-name>
          <email>luca.ducceschi@unitn.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leonardo Herzogz</string-name>
          <email>leonardo.herzog@studenti.unitn.it</email>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>25</fpage>
      <lpage>28</lpage>
      <abstract>
        <p>English. We present an approach to improve the selection of complex words for automatic text simplification, addressing the need of L2 learners to take into account their native language during simplification. In particular, we develop a methodology that automatically identifies 'difficult' terms (i.e. false friends) for L2 learners in order to simplify them. We evaluate not only the quality of the detected false friends but also the impact of this methodology on text simplification compared with a standard frequency-based approach.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. In questo contributo
presentiamo un approccio per selezionare le
parole complesse da semplificare in modo
automatico, tenendo conto della lingua
madre dell’utente. Nello specifico, la nostra
metodologia identifica i termini ‘difficili’
(falsi amici) per l’utente per proporne la
semplificazione. In questo contesto, viene
valutata non soltanto la qualita` dei falsi
amici individuati, ma anche l’impatto che
questa semplificazione personalizzata ha
rispetto ad approcci standard basati sulla
frequenza delle parole.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        The task of automated text simplification has been
investigated within the NLP community for
several years with a number of different approaches,
from rule-based ones
        <xref ref-type="bibr" rid="ref11 ref18 ref3">(Siddharthan, 2010;
Barlacchi and Tonelli, 2013; Scarton et al., 2017)</xref>
        to supervised
        <xref ref-type="bibr" rid="ref15 ref4 ref6">(Bingel and Søgaard, 2016;
AlvaManchego et al., 2017)</xref>
        and unsupervised ones
        <xref ref-type="bibr" rid="ref15 ref4 ref6">(Paetzold and Specia, 2016)</xref>
        , including recent
studies using deep learning
        <xref ref-type="bibr" rid="ref13 ref19 ref22">(Zhang and Lapata,
2017; Nisioi et al., 2017)</xref>
        . Nevertheless, only
recently researchers have started to build
simplification systems that can adapt to users, based on
the observation that the preceived simplicity of a
document depends a lot on the user profile,
including not only specific disabilities but also
language proficiency, age, profession, etc. Therefore
in the last few months the first approaches to
personalised text simplification have been proposed
at major conferences, with the goal of simplifying
a document for different language proficiency
levels
        <xref ref-type="bibr" rid="ref10 ref10 ref17 ref17 ref2 ref2 ref5 ref5 ref5">(Scarton and Specia, 2018; Bingel et al., 2018;
Lee and Yeung, 2018)</xref>
        .
      </p>
      <p>
        Along this research line, we present in this
paper an approach to perform automated lexical
simplification for L2 learners, able to adapt to the user
mother tongue. To our knowledge, this is the first
work taking into account this aspect and
presenting a solution that, given an Italian document and
the user’s mother tongue as input, selects only the
words that the user may find difficult given his/her
knowledge of another language. Specifically, we
detect and simplify automatically the terms that
may be misleading for the user because they are
false friends, while we do not simplify those that
have an orthographically and semantically similar
translation in the user native language (so-called
cognates). In multilingual settings, for instance
while teaching, learning or translating a foreign
language, these two phenomena have proven to be
very relevant
        <xref ref-type="bibr" rid="ref16">(Ringbom, 1986)</xref>
        , because the
lexical similarities between the two languages in
contact have proven to create interferences, favouring
or hindering the course of learning.
      </p>
      <p>We compare our approach to the selection of
words to be simplified with a standard
frequencybased one, in which only the terms that are not
listed in De Mauro’s Dictionary of Basic
Italian1 are simplified, regardless of the user native
1https://dizionario.internazionale.it/
language. Our experiments are evaluated on the
Italian-French pair, but the approach is generic.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Approach description</title>
      <p>Given a document Di to be simplified, and a
native language L1 spoken by the user, our approach
consists of the following steps:
1. Candidate selection: for each content word2
wi in Di, we automatically generate a list
of words W1 L1 which are
orthographically similar to wi. In this phase, several
orthographical similarity metrics are evaluated.</p>
      <p>We keep the 5 most-similar terms to wi.
2. False friend and cognate detection: for
each of the 5 most similar words in W1, we
classify whether it is a false friend of wi or
not.
3. Simplification choice: Based on the output
of the previous steps, the system marks wi
as difficult to understand for the user if there
are corresponding false friends in L1.
Otherwise, wi is left in its original form. When a
word is marked as difficult, a subsequent
simplification module (not included in this work)
should try to find an alternative form (such as
a synonym, or a description) to make the term
more understandable to the user.
2.1</p>
      <p>
        Candidate Selection
A number of similarity metrics have been
presented in the past to identify candidate cognates
and false friends, see for example the evaluation
in Inkpen and Frunza (2005). We choose three of
them, motivated by the fact that we want to have at
least one ngram-based metric (XXDICE) and one
non ngram-based (Jaro/Winkler). To that, we add
a more standard metric, Normalized Edit Distance
(NED). The three metrics are explained below:
XXDICE
        <xref ref-type="bibr" rid="ref7">(Brew et al., 1996)</xref>
        . It takes in
consideration the shared number of extended
bigrams3 and their position relative to two
nuovovocabolariodibase
      </p>
      <p>
        2Content words are words that have a meaning such as
names, adjectives, verbs and adverbs. To extract this
information, we use the POS tagger included in the Tint pipeline
        <xref ref-type="bibr" rid="ref10 ref17 ref2 ref5">(Aprosio and Moretti, 2018)</xref>
        .
      </p>
      <p>3An extended bigram is an ordered letter pair formed by
deleting the middle letter from any three letter substring of
the word.
strings S1 and S2. The formula is:</p>
      <p>XX(S1; S2) =</p>
      <p>P</p>
      <p>
        2
B 1+(pos(x) pos(y))2
xb(S1) + xb(S2)
where B is the set of pairs of shared extended
bigrams (x; y), x in S1 and y in S2. The
functions pos(x) and xb(S) return the
position of extended bigram x and the number of
extended bigrams in string S respectively.
NED, Normalized Edit Distance
        <xref ref-type="bibr" rid="ref20">(Wagner and
Fischer, 1974)</xref>
        . A regular Edit Distance
calculates the orthographic difference between
two strings assigning a cost to any minimum
number of edit operations (deletion,
substitution and insertion, all with cost of 1) needed
to make them equal. NED is obtained by
dividing the edit cost by the length of the
longest string.
      </p>
      <p>
        Jaro/Winkler
        <xref ref-type="bibr" rid="ref21">(Winkler, 1990)</xref>
        . The Jaro
similarity metric for two strings S1 and S2 is
computed as follows:
J (S1; S2) =
where m is the number of characters in
common, provided that they occur in the same
(not interrupted) sequence, and T is the
number of transpositions of character in S1 to
obtain S2. The Winkler variation of the metric
adds a bias if the two strings share a prefix.
J W (S1; S2) = J (S1; S2)+(1 J (S1; S2))lp
where l is the number of characters of the
common prefix of the two strings, up to four,
and p is a scaling factor, usually set to 0:1.
      </p>
      <p>
        Each of these three measures has some
disadvantages. For example, we found that
Jaro/Winkler metric boosts the similarity of words
with the same root. On the other hand, applying
NED leads to several pairs of words having the
same similarity score. As a result, two words that
are close according to a metric can be far using
another metric. To overcome this limitation, we
balance the three metrics by computing a weighted
average of the three scores tuned on a training set.
For details, see Section 3.
As for false friend and cognate detection, we rely
on a SVM-based classifier and train it on a single
feature obtained from a multilingual embedding
space
        <xref ref-type="bibr" rid="ref11">(Mikolov et al., 2013)</xref>
        , where the user
language L1 and the language of the document to be
simplified L2 are aligned. In particular, the feature
is the cosine distance between the embeddings of a
given content word wi in the language L2 and the
embedding of its candidate false friends or
cognates in L1. The intuition behind this approach
is that two cognates have a shared semantics and
therefore a high cosine similarity, as opposed to
false friends, whose meanings are generally
unrelated. While past approaches to false friend and
cognate detection have already exploited
monolingual word embeddings
        <xref ref-type="bibr" rid="ref19">(St Arnaud et al., 2017)</xref>
        ,
we employ for our experiments a multilingual
setting, so that the semantic distance between the
candidate pairs can be measured in their original
language without a preliminary translation.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Experimental Setup</title>
      <p>In our experiments, we consider a setting in which
French speakers would like to make Italian
documents easier for them to read. Nevertheless,
the approach can be applied to any language pair,
given that it requires minimal adaptation.</p>
      <p>In order to tune the best similarity metrics
combination and to train the SVM classifier, a
linguist has manually created an Italian-French gold
standard, containing pairs of words marked as
either cognates or false friends. These terms were
collected from several lists available on the web.
Overall, the Ita-Fr dataset contains a training set
of 1,531 pairs (940 cognates and 591 false friends)
and a test set of 108 pairs (51 cognates and 57 false
friends).</p>
      <p>
        For the candidate selection step, the goal is to
obtain for each term wi in Italian, the 5 French
terms with the highest orthographic similarity.
Therefore, given wi, we compute its similarity
with each term in a French online dictionary4
        <xref ref-type="bibr" rid="ref12">(New, 2006)</xref>
        using the three scores described in the
previous section. The lemmas were normalized
for accents and diacritics, in order to avoid poor
results of the metrics in cases like ge´ne´ral and
generale, where the accented e´ character would be
considered different with respect to e.5
4http://www.lexique.org/
5For example, NED between ge´ne´ral and generale returns
In order to identify the best way to combine the
three similarity metrics detailed in Section 2.1., we
compute all the possibile combinations of weights
on 10 groups of 200 word pairs randomly
extracted from the 1,531 pairs in the training set, and
then keep the combination that scores the highest
average similarity.
      </p>
      <p>In Table 1 we report the percentage of times in
which the cognate or false friend of wi in the
training set would appear among the 5 most-similar
terms extracted from the French online dictionary
according to the three different scores in isolation:
XX for XXDICE, JW for Jaro/Winkler and NED
for Normalized Edit Distance. We also report the
best configuration of the three metrics with the
corresponding weight to maximise the presence of
a cognate or false friend among the 5 most
similar terms. We observe that, while the three metrics
in isolation yield a similar result, combining them
effectively increases the presence of cognates and
false friends among the top candidates. This
confirms that the metrics capture three different types
of similarity, and that it is recommended to take
them all into account when performing candidate
selection: an approach where evey metric
contributes to detecting false friend / cognate
candidates outperforms the single metrics.</p>
      <p>XX
1.0
0.2</p>
      <p>JW</p>
      <p>1.0</p>
      <p>0.4</p>
      <p>NED
1.0
0.4</p>
      <p>For false friends and cognates detection, we
proceed as follows. Given a word wi in Italian, we
identify the 5 most similar words in French using
the 0.2-0.4-0.4 score introduced before. In case
of ties in the 5th positon, we extend the selection
to all the candidates sharing the same similarity
value.</p>
      <p>
        Each word pair including wi and one of the
5 most similar words is then classified as false
friend or cognate with a SVM using a radial kernel
trained on the 1,531 word pairs in the training set.
For the multilingual embeddings used to compute
0.375 when the two strings are not normalized and 0.125
when they are.
the semantic similarity between the Italian words
and their candidates, we use the vectors from
Bojanowski et al. (2016)6 trained on Wikipedia data
with fastText
        <xref ref-type="bibr" rid="ref6 ref9">(Joulin et al., 2016)</xref>
        . We chose these
resources since they are available both for Italian
and French (and several other languages). For the
alignement of the semantic spaces of the two
languages we use 22,767 Italian-French word pairs
collected from an online dictionary.7
4
      </p>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>We perform two types of evaluation. In the first
one, the goal is to assess whether the system can
correctly identify false friends and cognates in a
text. In the second one, we want to check what
is the difference between the terms simplified by
a system with our approach compared with a
standard frequency-based simplification system.</p>
      <p>For the first evaluation, we manually create a
set of 108 Italian sentences containing one false
friend or cognate for French speakers taken from
the test set. On each term, we run our algorithm
and we consider a term a false friend according
to two strategies: a) if all 5 most similar words
in French are classified as false friends, or b) if
the majority of them are classified as false friends.
Results are reported in Table 2.</p>
      <p>false friends (a)
false friends (b)</p>
      <p>The evaluation shows that the two settings lead
to two different outcomes. In general terms, the
first strategy is more conservative and favours
Precision, while the second boosts Recall and F1.</p>
      <p>As for the second evaluation, on the same set of
sentences, we run our algorithm again, this time
trying to classify any content word as being a false
friend for French speakers or not. We evaluate this
component as being part of a simplification
system that simplifies only false friends, and we
compare this choice with a more standard approach,
in which only ‘unusual’ or ‘unfrequent’ terms are
simplified. This second choice is taken by
com6https://github.com/facebookresearch/
fastText/blob/master/pretrained-vectors.
md
7http://dizionari.corriere.it/
paring each content word with De Mauro’s
Dictionary of Basic Italian and simplifying only those
that are not listed among the 7,000 entries of the
basic vocabulary.</p>
      <p>This evaluation shows that out of 1,035
content words in the test sentences, our simplification
approach based on a) would simplify 367 words,
and 823 if we adopt the strategy b). Based on
De Mauro’s dictionary, instead, 240 terms would
be simplified. Furthermore, there would be only
76 terms simplified using both strategy a) and De
Mauro’s list, and 154 overlaps for strategy b). This
shows that the two approaches are rather
complementary and based on different principles. This
is evident also looking at the evaluated sentences:
while considering frequency lists like De Mauro’s,
terms such as accademico and speleologo should
be simplified because they are not frequently used
in Italian, our approach would not simplify them
because they have very similar French translations
(acade´mique and spe´le´ologue respectively), and
are not classified as false friends by the system.
On the other hand, vedere would not be
simplified in a standard frequency-based system because
it is listed among the 2,000 fundamental words in
Italian. However, our approach would identify it
as a false friend to be simplified because vider in
French (transl. svuotare) is orthographically very
similar to vedere but has a completely different
meaning.
5</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this work, we have presented an approach
supporting personalized simplification in that it
enables to adapt the selection of difficult words for
lexical simplification to the native language of L2
learners. To our knowledge, this is the first
attempt to deal with this kind of adaptation. The
approach is relatively easy to apply to new languages
provided that they have a similar alphabet, since
multilingual embeddings are already available and
lists of cognates and false friends, although of
limited size, can be easily retrieved online.8</p>
      <p>
        The work will be extended along different
research directions: first, we will evaluate the
approach on other language pairs. Then, we will add
a lexical simplification module selecting only the
words identified as complex by our approach. For
8See for example the Wiktionary entries at
https://en.wiktionary.org/wiki/Category:
False_cognates_and_false_friends
this, we can rely on existing simplification tools
        <xref ref-type="bibr" rid="ref14">(Paetzold and Specia, 2015)</xref>
        , which could be tuned
to adapt also the simplification choices to the user
native language, for example by changing the
candidate ranking algorithm. Finally, it would be
interesting to involve L2 learners in the evaluation,
with the goal to measure the effectiveness of
different simplification strategies in a real setting.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work has been supported by the European
Commission project SIMPATICO
(H2020-EURO6-2015, grant number 692819). We would like to
thank Francesca Fedrizzi for her help in creating
the gold standard.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Fernando</given-names>
            <surname>Alva-Manchego</surname>
          </string-name>
          , Joachim Bingel, Gustavo Paetzold, Carolina Scarton, and
          <string-name>
            <given-names>Lucia</given-names>
            <surname>Specia</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Learning how to simplify from explicit labeling of complex-simplified text pairs</article-title>
          .
          <source>In Greg Kondrak and Taro Watanabe</source>
          , editors,
          <source>Proceedings of the Eighth International Joint Conference on Natural Language Processing, IJCNLP</source>
          <year>2017</year>
          , Taipei, Taiwan,
          <source>November 27 - December 1</source>
          ,
          <fpage>2017</fpage>
          - Volume
          <volume>1</volume>
          :
          <string-name>
            <given-names>Long</given-names>
            <surname>Papers</surname>
          </string-name>
          , pages
          <fpage>295</fpage>
          -
          <lpage>305</lpage>
          .
          <article-title>Asian Federation of Natural Language Processing</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Alessio</given-names>
            <surname>Palmero</surname>
          </string-name>
          Aprosio and
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Moretti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Tint 2.0: an all-inclusive suite for nlp in italian</article-title>
          .
          <source>In Proceedings of the Sixth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2018</year>
          ), Torino, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Gianni</given-names>
            <surname>Barlacchi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sara</given-names>
            <surname>Tonelli</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>ERNESTA: A Sentence Simplification Tool for Children's Stories in Italian</article-title>
          . In Alexander Gelbukh, editor,
          <source>Computational Linguistics and Intelligent Text Processing: 14th International Conference, CICLing</source>
          <year>2013</year>
          , Samos, Greece, March
          <volume>24</volume>
          -30,
          <year>2013</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>II</given-names>
          </string-name>
          , pages
          <fpage>476</fpage>
          -
          <lpage>487</lpage>
          , Berlin, Heidelberg. Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Joachim</given-names>
            <surname>Bingel</surname>
          </string-name>
          and
          <string-name>
            <given-names>Anders</given-names>
            <surname>Søgaard</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Text simplification as tree labeling</article-title>
          .
          <source>In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          , pages
          <fpage>337</fpage>
          -
          <lpage>343</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Joachim</given-names>
            <surname>Bingel</surname>
          </string-name>
          , Gustavo Paetzold, and
          <string-name>
            <given-names>Anders</given-names>
            <surname>Søgaard</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Lexi: A tool for adaptive, personalized text simplification</article-title>
          .
          <source>In Proceedings of COLING. Association for Computational Linguistics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>CoRR</source>
          , abs/1607.04606.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Chris</given-names>
            <surname>Brew</surname>
          </string-name>
          ,
          <string-name>
            <surname>David McKelvie</surname>
          </string-name>
          , et al.
          <year>1996</year>
          .
          <article-title>Word-pair extraction for lexicography</article-title>
          .
          <source>In Proceedings of the 2nd International Conference on New Methods in Language Processing</source>
          , pages
          <fpage>45</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Diana</given-names>
            <surname>Inkpen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Oana</given-names>
            <surname>Frunza</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Automatic identification of cognates and false friends in french and english</article-title>
          .
          <source>In Proceedings of RANLP</source>
          , pages
          <fpage>251</fpage>
          -
          <lpage>257</lpage>
          ,
          <fpage>01</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Armand</given-names>
            <surname>Joulin</surname>
          </string-name>
          , Edouard Grave, Piotr Bojanowski, Matthijs Douze,
          <article-title>He´rve Je´gou, and</article-title>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Fasttext.zip: Compressing text classification models</article-title>
          .
          <source>arXiv preprint arXiv:1612</source>
          .
          <fpage>03651</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>John</given-names>
            <surname>Lee</surname>
          </string-name>
          and Chak Yan Yeung.
          <year>2018</year>
          .
          <article-title>Personalizing lexical simplification</article-title>
          .
          <source>In Proceedings of the 27th International Conference on Computational Linguistics</source>
          , pages
          <fpage>224</fpage>
          -
          <lpage>232</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Quoc V Le,
          <string-name>
            <given-names>and Ilya</given-names>
            <surname>Sutskever</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Exploiting similarities among languages for machine translation</article-title>
          .
          <source>arXiv preprint arXiv:1309</source>
          .
          <fpage>4168</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Boris</given-names>
            <surname>New</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Lexique 3: Une nouvelle base de donne´es lexicales</article-title>
          . In Actes de la Confe´
          <article-title>rence Traitement Automatique des Langues Naturelles (TALN</article-title>
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Sergiu</given-names>
            <surname>Nisioi</surname>
          </string-name>
          , Sanja Stajner, Simone Paolo Ponzetto, and
          <string-name>
            <given-names>Liviu P.</given-names>
            <surname>Dinu</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Exploring neural text simplification models</article-title>
          . In Regina Barzilay and
          <string-name>
            <surname>Min-Yen</surname>
            <given-names>Kan</given-names>
          </string-name>
          , editors,
          <source>Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          <year>2017</year>
          , Vancouver, Canada,
          <source>July 30 - August 4</source>
          , Volume
          <volume>2</volume>
          :
          <string-name>
            <given-names>Short</given-names>
            <surname>Papers</surname>
          </string-name>
          , pages
          <fpage>85</fpage>
          -
          <lpage>91</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Gustavo</given-names>
            <surname>Paetzold</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lucia</given-names>
            <surname>Specia</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Lexenstein: A framework for lexical simplification</article-title>
          .
          <source>In ACLIJCNLP 2015 System Demonstrations, ACL</source>
          , pages
          <fpage>85</fpage>
          -
          <lpage>90</lpage>
          , Beijing, China.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Gustavo H.</given-names>
            <surname>Paetzold</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lucia</given-names>
            <surname>Specia</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Unsupervised lexical simplification for non-native speakers</article-title>
          . In Dale Schuurmans and Michael P. Wellman, editors,
          <source>Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, February 12-17</source>
          ,
          <year>2016</year>
          , Phoenix, Arizona, USA., pages
          <fpage>3761</fpage>
          -
          <lpage>3767</lpage>
          . AAAI Press.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Ringbom</surname>
          </string-name>
          .
          <year>1986</year>
          .
          <article-title>Crosslinguistic influence and the foreign language learning process</article-title>
          . In E. Kellerman and Smith Sharwood M., editors,
          <source>Crosslinguistic Influence in Second Language Acquisition</source>
          . Pergamon Press, New York.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Carolina</given-names>
            <surname>Scarton</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lucia</given-names>
            <surname>Specia</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Learning simplifications for specific target audiences</article-title>
          .
          <source>In ACL (2)</source>
          , pages
          <fpage>712</fpage>
          -
          <lpage>718</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Advaith</given-names>
            <surname>Siddharthan</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Complex lexico-syntactic reformulation of sentences using typed dependency representations</article-title>
          .
          <source>In Proceedings of the 6th International Natural Language Generation Conference (INLG</source>
          <year>2010</year>
          ), Dublin, Ireland.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Adam</given-names>
            <surname>St Arnaud</surname>
          </string-name>
          , David Beck,
          <string-name>
            <given-names>and Grzegorz</given-names>
            <surname>Kondrak</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Identifying cognate sets across dictionaries of related languages</article-title>
          .
          <source>In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>2519</fpage>
          -
          <lpage>2528</lpage>
          , Copenhagen, Denmark, September. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Robert A Wagner and Michael J Fischer</surname>
          </string-name>
          .
          <year>1974</year>
          .
          <article-title>The string-to-string correction problem</article-title>
          .
          <source>Journal of the ACM (JACM)</source>
          ,
          <volume>21</volume>
          (
          <issue>1</issue>
          ):
          <fpage>168</fpage>
          -
          <lpage>173</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>William E</given-names>
            <surname>Winkler</surname>
          </string-name>
          .
          <year>1990</year>
          .
          <article-title>String comparator metrics and enhanced decision rules in the fellegi-sunter model of record linkage</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Xingxing</given-names>
            <surname>Zhang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mirella</given-names>
            <surname>Lapata</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Sentence simplification with deep reinforcement learning</article-title>
          .
          <source>In Martha Palmer</source>
          , Rebecca Hwa, and Sebastian Riedel, editors,
          <source>Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP</source>
          <year>2017</year>
          , Copenhagen, Denmark, September 9-
          <issue>11</issue>
          ,
          <year>2017</year>
          , pages
          <fpage>584</fpage>
          -
          <lpage>594</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>