<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The impact of phrases on Italian lexical simplification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sara Tonelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>satonelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>aprosiog@fbk.eu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Marco Mazzon Dept. of Psychology and Cognitive Science University of Trento</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. Automated lexical simplification has been performed so far focusing only on the replacement of single tokens with single tokens, and this choice has affected both the development of systems and the creation of benchmarks. In this paper, we argue that lexical simplification in real settings should deal both with single and multi-token terms, and present a benchmark created for the task. Besides, we describe how a freely available system can be tuned to cover also the simplification of phrases, and perform an evaluation comparing different experimental settings.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Italiano. La semplificazione lessicale
automatica e` stata affrontata fino ad ora
dalla comunita` di ricerca TAL
concentrandosi sulla sostituzione di parole singole
con altre parole singole. Questa modalita`
ha condizionato sia lo sviluppo di
sistemi di semplificazione che la creazione
di benchmark per la valutazione. In
questo articolo, sosteniamo che la
semplificazione lessicale in contesti reali debba
includere sia parole singole che
espressioni composte da piu` parole, e
presentiamo un benchmark creato a questo fine.
Inoltre, descriviamo come adattare un
sistema disponibile per la semplificazione
lessicale in modo che supporti anche la
semplificazione di sintagmi, e presentiamo
una valutazione confrontando diversi
setting sperimentali.
clarity and readability. Thanks to the
development of benchmarks
        <xref ref-type="bibr" rid="ref1 ref13 ref5 ref8 ref9">(Paetzold and Specia, 2016a)</xref>
        and freely available tools for lexical simplification
        <xref ref-type="bibr" rid="ref4 ref7">(Paetzold and Specia, 2015)</xref>
        , a number of works
have focused on this challenge, see for
example the systems participating in the simplification
shared task at SemEval-2012
        <xref ref-type="bibr" rid="ref11">(Specia et al., 2012)</xref>
        .
However, the task has been designed as an
exercise to replace complex single tokens with simpler
single tokens, and most widely used benchmarks
and systems all follow this paradigm. We believe,
however, that this setting covers only a limited
number of lexical simplifications as they would be
performed in a real scenario. In particular, we
advocate the need to shift the lexical simplification
paradigm from single tokens to phrases, and to
develop datasets and tools that deal also with these
cases. This is mainly the contribution of this work,
which covers four main points:
      </p>
      <p>We analyse existing corpora of simplified
texts, not specifically developed for a shared
task or for system evaluation, and we
measure the impact of phrases in lexical
simplifications
We modify a state-of-the-art tool for lexical
simplification in order to support phrases
We compare different strategies for phrase
extraction and evaluate them over a
benchmark
We perform all the above on Italian, for
which there was no lexical simplification
system available.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>Lexical simplification is a well-studied topic
within the NLP community, dealing with the
automatic replacement of complex terms with
simpler ones in a sentence, in order to improve its
Besides, we make freely available the first
benchmark for the evaluation of Italian lexical
simplification, with the goal to support research
on this task and to foster the development of
Italian simplification systems.</p>
    </sec>
    <sec id="sec-3">
      <title>Corpus analysis and Benchmark creation</title>
      <p>
        We first analyse existing simplification corpora in
Italian to study the impact of phrases on lexical
simplification. There are only two such manually
created corpora, which contain different types
of data but have been annotated following the
same scheme: the Simpitiki corpus
        <xref ref-type="bibr" rid="ref13">(Tonelli et al.,
2016)</xref>
        and the one developed by the ItaNLP Lab
in Pisa
        <xref ref-type="bibr" rid="ref3">(Brunato et al., 2015)</xref>
        . The former contains
1,163 sentence pairs1, where one is the original
sentence and the other is the simplified one. The
pairs were created starting from Wikipedia edits
and from documents in the public administration
domain. The ItaNLP corpus, instead, contains
1,393 pairs extracted from children’s stories and
from educational material. Both corpora were
annotated following the scheme proposed in
        <xref ref-type="bibr" rid="ref3">(Brunato et al., 2015)</xref>
        , in which simplifications
were classified as Split, Merge, Reordering, Insert,
Delete and Transformation (plus a set of
subclasses for the Insert, Delete and Transformation
cases). Since our goal was to isolate a benchmark
of pairs containing only the lexical cases, we
discarded the classes not compatible with lexical
simplifications (e.g. Delete, Reordering) and
then manually checked the others to identify the
cases of interest. When, as in the majority of
cases, a lexical simplification was present together
with other simplification types, we re-wrote the
target sentence in order to retain only lexical
cases. For example, in the examples below, a)
is the original sentence and b) is the simplified
one in the Simpitiki corpus, which contains a
lexical simplification of ‘include’ and a shift of
position of ‘per convenzione’. We created version
c), so that only the lexical simplification is present:
a) Eurasia e` il termine con cui per convenzione si
definisce la zona geografica che include l’Europa
e l’Asia.
b) Eurasia e`, per convenzione, il termine con cui
si definisce la zona geografica che comprende
l’Europa e l’Asia.
c) Eurasia e` il termine con cui per convenzione
si definisce la zona geografica che comprende
l’Europa e l’Asia.
      </p>
      <p>1The number is slightly different from what was reported
in the original paper because the corpus was revised after the
first release.</p>
      <p>This revision process led to the creation of a
benchmark with pairs extracted from the two
original corpora, where only cases of lexical
simplification are present2. Some statistics related to the
benchmark are reported in Table 1. We identify
four possible lexical simplification types: a
single token is replaced by a single token (ST!ST),
a single token is simplified through a phrase
(ST!P), a phrase is simplified through a single
token (P!ST), and a phrase is replaced by another
phrase (P!P).</p>
      <sec id="sec-3-1">
        <title>ItaNLP</title>
      </sec>
      <sec id="sec-3-2">
        <title>Simpitiki</title>
      </sec>
      <sec id="sec-3-3">
        <title>Total</title>
        <p>We observe that the most frequent lexical
simplification type is ST!ST, on which most systems
and shared tasks are based. However, this
simplification type covers only half of the cases included
in our benchmark. This confirms the need to
include cases of phrase-based simplification in the
creation of benchmarks. It corroborates also the
importance of developing systems for lexical
simplification that support phrase replacement, so as
to make them work in real settings and not only
on ad-hoc test sets. Another interesting remark is
that single tokens are not necessarily simpler than
phrases, or vice versa: in our data, there are 136
ST!P and 169 P!ST, showing that no general
rule can be applied to favour (or demote) Ps over
STs.</p>
        <p>We use the final benchmark3, containing 901
sentence pairs, to evaluate a system for lexical
simplification taking into account phrases, as
described in the following Section.
3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Automated lexical simplification</title>
      <p>In this Section we describe the experiments we
carried out to perform automated lexical
simplification using the benchmark presented in Section
2. We describe the tool used and how it was
mod2In Simpitiki we focused only on the pairs in the public
administration domain due to project constraints. We plan
to include the pairs from Wikipedia in the next benchmark
version.</p>
      <p>3Available at https://drive.google.com/
file/d/0B4QAWZllD-egYS0yNWZ5dTdYQVE/
view?usp=sharing
ified to deal with phrases. We also detail the
resources (language model and word embeddings)
created for the task.
3.1</p>
      <sec id="sec-4-1">
        <title>The Lexenstein system</title>
        <p>
          We use Lexenstein
          <xref ref-type="bibr" rid="ref4 ref7">(Paetzold and Specia, 2015)</xref>
          ,
an open source tool for lexical simplification, to
collect a list of candidates that should replace a
given word in the text. In particular, the Paetzold
generator
          <xref ref-type="bibr" rid="ref1 ref13 ref5 ref8 ref9">(Paetzold and Specia, 2016b)</xref>
          is based
on an unsupervised approach to produce
simplification candidates using a context-aware word
embeddings model: features used for the
selection include word2vec vectors
          <xref ref-type="bibr" rid="ref6">(Mikolov et al.,
2013)</xref>
          , language model created by SRILM
          <xref ref-type="bibr" rid="ref12">(Stolcke, 2002)</xref>
          , and conditional probability of a
candidate given the PoS tag of the target word. So far,
no evaluation on Lexestein for Italian is available.
        </p>
        <p>
          For each complex word, five candidate
replacements are first retrieved, ranked according to
several features, such as n-gram frequencies and word
vector similarity with the target word, and then
reranked according to their average rankings
          <xref ref-type="bibr" rid="ref4 ref7">(Glavasˇ
and Sˇtajner, 2015)</xref>
          .
        </p>
        <p>
          Since we wanted to test different strategies to
create the embeddings (i.e. with and without
phrases), we created the word/phrase vectors and
the language model starting from freely available
corpora (1.3 billion words in total): the Italian
Wikipedia,4 OpenSubtitles2016
          <xref ref-type="bibr" rid="ref1 ref13 ref5 ref8 ref9">(Lison and
Tiedemann, 2016)</xref>
          ,5 PAISA` ,6 and the Gazzetta
Ufficiale,7 a collection of Italian laws. Due to the
size of the data, both the corpus and the model are
available upon request to the authors.
3.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Experimental Setup</title>
        <p>We conduct several experiments to evaluate the
quality of lexical simplification when taking into
account phrases (or not), and compare different
strategies for phrase recognition. We compare
different variants to create the embeddings and the
language model (LM) that were then used by
Lexenstein.</p>
        <p>The first baseline model relies on the standard
Lexenstein setting: word embeddings are created
using the word2vec package, and the LM
considers each token separately.</p>
        <p>
          4https://it.wikipedia.org/wiki/Pagina_
principale
5http://www.opensubtitles.org/
6http://www.corpusitaliano.it/
7http://www.gazzettaufficiale.it/
The first system variant (word2phrase) includes
phrase recognition, i.e. before extracting the
embeddings and creating the LM, the documents
are analysed by the word2phrase module in the
word2vec package. This is an implementation of
the algorithm presented in
          <xref ref-type="bibr" rid="ref6">(Mikolov et al., 2013)</xref>
          ,
which basically identifies words that appear
frequently together, and infrequently in other
contexts, and treats them as single tokens (connected
by an underscore).
        </p>
        <p>
          The second system variant
(word2phrase+LemmaPos) adds another
information layer, in that each document is first
lemmatized and PoS tagged using the Tint NLP
Suite
          <xref ref-type="bibr" rid="ref1 ref13 ref5 ref8 ref9">(Aprosio and Moretti, 2016)</xref>
          , that works at
token level; then word2phrase is run, and then the
embeddings and the LM are created. In this way,
we obtain so-called ‘context-aware’ embeddings,
which is the recommended setting in
          <xref ref-type="bibr" rid="ref1 ref13 ref5 ref8 ref9">(Paetzold
and Specia, 2016b)</xref>
          .
4
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>The evaluation of automated simplification is an
open issue since, similar to machine translation,
there may be different acceptable simplifications
for a term, while a benchmark usually presents
only one solution. Therefore, we perform two
evaluations: the first is based on an automated
comparison between Lexenstein output and the
gold simplifications in the benchmark. The
second is a manual evaluation aimed at scoring
fluency, adequacy and simplicity of the output.</p>
      <p>For the first evaluation, we compute the Mean
Reciprocal Rank (MRR), which is usually adopted
to evaluate a list of possible responses ordered by
probability of correctness against a gold answer.
We use this metrics because Lexenstein returns 5
possible simplifications, ranked by relevance, and
with MRR it is possible to weight the response
matching with the gold simplification according to
its rank. In particular, MRR is computed as:
MRR =
1 jQj</p>
      <p>X</p>
      <p>1
jQj i=1 ranki
where Q is the number of simplifications to be
performed (901) and ranki is the position of the
correct simplification in the rank returned by
Lexenstein.</p>
      <p>We run the system in the three configurations
described in Section 3.2 on each source sentence
in the benchmark. The single or multi-token term
to be simplified is given. If it is found in the LM,
the system suggests 5 ranked simplification
candidates. Otherwise, no output is given.</p>
      <p>Results show that the baseline model, i.e. the
standard Lexenstein configuration replacing only
single tokens with single tokens, yields MRR =
0:036. The one using word2phrase achieves
MRR = 0:042, while the version including
also lemma and PoS information yields MRR =
0:050. A detailed evaluation is reported in Table
2: for each of the three experimental settings, we
report the number of cases in which the gold
simplification matches the first ranked replacement
returned by Lexenstein (1st), the second, the third,
and so on. In the last column, we report how many
times (out of 901) the rank returned by Lexenstein
does not contain the gold simplification present in
the benchmark.</p>
      <p>Baseline
word2phrase
+LemmaPos</p>
      <p>This evaluation shows that, although limited,
using word2phrase in combination with lemma
and PoS information yields an improvement over
the baseline. However, the informativeness of this
automated simplification is limited because the
cases labeled as ‘none’ include both wrong
simplifications and correct simplifications that are not
present in the benchmark. Besides, they include
also cases in which the word to be simplified was
not found in the LM.</p>
      <p>
        In order to better understand where the
approach fails, we also perform a manual evaluation.
Following the standard scheme for human
evaluation of automatic text simplification
        <xref ref-type="bibr" rid="ref10">(Saggion and
Hirst, 2017)</xref>
        , we judge Fluency (grammaticality),
Adequacy (meaning preservation) and Simplicity
of lexical simplifications using a five-point Likert
scale (the higher the score, the better the output).
For the setting using lemma and PoS, we do not
judge Fluency, since the output is lemmatized and
not converted in the original form of the source
term (we plan to add this in the near future).
Evaluation is performed using a set of 150 sentence
pairs randomly extracted from the benchmark.
We introduce also this kind of evaluation in order
to have a fine-grained analysis of system output.
For example, in the original sentence d) (see
below), ‘tempestivamente’ was simplified with
‘periodicamente’, which is grammatically correct
(high Fluency) but does not preserve the meaning
of the original sentence (low Adequacy).
d) Il richiedente dovra` comunicare
tempestivamente l’esattezza dei recapiti
forniti.
      </p>
      <p>When using word2phrase without
lemmatization, the average Fluency is 3:72, Adequacy is
2:60 and Simplicity is 2:95. This shows that, while
PoS and form of a simplified term are generally
correct also without any processing, the
preservation of the meaning is a critical issue.
Simplicity achieves better scores than Adequacy, but it
still needs improvements. Results obtained using
lemma and PoS in combination with word2phrase
are slightly better, with 2:64 Adequacy and 3:01
Simplicity. In general, the above evaluations show
that using word2phrase with lemma and PoS
information is a promising approach to improve the
performance of lexical simplification in real
settings. The performance of Lexenstein could be
further improved by adding other corpora to the
LM and post-process the output of the system, so
as to discard inconsistent simplifications, for
example when a verb is simplified through an
adverb. However, some linguistic phenomena like
non-local dependencies cannot be addressed using
this approach, and a separate strategy to simplify
them should be taken into account.
5</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>
        In this work, we presented a first analysis of the
role of phrases in Italian lexical simplification.
We also introduced the adaptation of Lexenstein,
an existing lexical simplification system, so as
to take phrases into account. In the future, we
plan to test other approaches for the extraction
of phrases, for example by applying algorithms
for recognising multiword expressions. We also
plan to integrate our best model for phrase
simplification in ERNESTA
        <xref ref-type="bibr" rid="ref2">(Barlacchi and Tonelli,
2013)</xref>
        , a system for syntactic simplification of
Italian documents. Furthermore, within the H2020
SIMPATICO project, we will integrate our phrase
simplification approach in the existing services
of Trento Municipality and perform a pilot study
with real users.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>The research leading to this paper was supported
by the EU Horizon 2020 Programme via the
SIMPATICO Project (H2020-EURO-6-2015, n.
692819).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Alessio</given-names>
            <surname>Palmero</surname>
          </string-name>
          Aprosio and
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Moretti</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Italy goes to Stanford: A collection of CoreNLP modules for Italian</article-title>
          . CoRR, abs/1609.06204.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Gianni</given-names>
            <surname>Barlacchi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sara</given-names>
            <surname>Tonelli</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>ERNESTA: A Sentence Simplification Tool for Children's Stories in Italian</article-title>
          . In Alexander Gelbukh, editor,
          <source>Computational Linguistics and Intelligent Text Processing: 14th International Conference, CICLing</source>
          <year>2013</year>
          , Samos, Greece, March
          <volume>24</volume>
          -30,
          <year>2013</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>II</given-names>
          </string-name>
          , pages
          <fpage>476</fpage>
          -
          <lpage>487</lpage>
          , Berlin, Heidelberg. Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Brunato</surname>
          </string-name>
          , Felice Dell'Orletta,
          <string-name>
            <given-names>Giulia</given-names>
            <surname>Venturi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Simonetta</given-names>
            <surname>Montemagni</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Design and Annotation of the First Italian Corpus for Text Simplification</article-title>
          .
          <source>In Proceedings of The 9th Linguistic Annotation Workshop</source>
          , pages
          <fpage>31</fpage>
          -
          <lpage>41</lpage>
          , Denver, Colorado, USA, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>Goran Glavasˇ and Sanja Sˇ tajner</article-title>
          .
          <year>2015</year>
          .
          <article-title>Simplifying lexical simplification: Do we need simplified corpora? In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th</article-title>
          <source>International Joint Conference on Natural Language Processing (Volume 2: Short Papers)</source>
          , pages
          <fpage>63</fpage>
          -
          <lpage>68</lpage>
          , Beijing, China, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Pierre</given-names>
            <surname>Lison</surname>
          </string-name>
          and Jo¨rg Tiedemann.
          <year>2016</year>
          .
          <article-title>Opensubtitles2016: Extracting large parallel corpora from movie and TV subtitles</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation LREC</source>
          <year>2016</year>
          ,
          <article-title>Portorozˇ</article-title>
          , Slovenia, May
          <volume>23</volume>
          -28,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Ilya Sutskever, Kai Chen, Gregory S. Corrado, and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8</source>
          ,
          <year>2013</year>
          ,
          <string-name>
            <given-names>Lake</given-names>
            <surname>Tahoe</surname>
          </string-name>
          , Nevada, United States., pages
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Gustavo</given-names>
            <surname>Paetzold</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lucia</given-names>
            <surname>Specia</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Lexenstein: A framework for lexical simplification</article-title>
          .
          <source>In ACLIJCNLP 2015 System Demonstrations, ACL</source>
          , pages
          <fpage>85</fpage>
          -
          <lpage>90</lpage>
          , Beijing, China.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Gustavo</given-names>
            <surname>Paetzold</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lucia</given-names>
            <surname>Specia</surname>
          </string-name>
          . 2016a.
          <article-title>Benchmarking lexical simplification systems</article-title>
          .
          <source>In Nicoletta Calzolari (Conference Chair)</source>
          , Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, and Stelios Piperidis, editors,
          <source>Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ), Paris, France, may.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Gustavo H.</given-names>
            <surname>Paetzold</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lucia</given-names>
            <surname>Specia</surname>
          </string-name>
          . 2016b.
          <article-title>Unsupervised lexical simplification for non-native speakers</article-title>
          . In Dale Schuurmans and Michael P. Wellman, editors,
          <source>Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, February 12-17</source>
          ,
          <year>2016</year>
          , Phoenix, Arizona, USA., pages
          <fpage>3761</fpage>
          -
          <lpage>3767</lpage>
          . AAAI Press.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Hirst</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Automatic Text Simplification</article-title>
          .
          <source>Synthesis Lectures on Human Language Technologies</source>
          . Morgan &amp; Claypool Publishers.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Lucia</given-names>
            <surname>Specia</surname>
          </string-name>
          , Sujay Kumar Jauhar, and
          <string-name>
            <given-names>Rada</given-names>
            <surname>Mihalcea</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Semeval-2012 task 1: English lexical simplification</article-title>
          .
          <source>In Proceedings of the First Joint Conference on Lexical and Computational Semantics - Volume 1: Proceedings of the Main Conference and the Shared Task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation, SemEval '12</source>
          , pages
          <fpage>347</fpage>
          -
          <lpage>355</lpage>
          , Stroudsburg, PA, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Stolcke</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Srilm - an extensible language modeling toolkit</article-title>
          . pages
          <fpage>901</fpage>
          -
          <lpage>904</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Sara</given-names>
            <surname>Tonelli</surname>
          </string-name>
          , Alessio Palmero Aprosio, and
          <string-name>
            <given-names>Francesca</given-names>
            <surname>Saltori</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>SIMPITIKI: a Simplification corpus for Italian</article-title>
          .
          <source>In Proceedings of the 3rd Italian Conference on Computational Linguistics (CLiC-it)</source>
          , volume
          <volume>1749</volume>
          <source>of CEUR Workshop Proceedings.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>