<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploration of Spanish Word Embeddings for Lexical Simplification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rodrigo Alarcón</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lourdes Moreno</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paloma Martínez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science and Engineering Department, Universidad Carlos III de Madrid</institution>
          ,
          <addr-line>28911 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <fpage>29</fpage>
      <lpage>41</lpage>
      <abstract>
        <p>Lexical simplification systems rely heavily on handcrafted databases or parallel corpora which represents a high cost of production. In this paper we present alternatives for every step in the lexical simpliifcation process individually for the Spanish language by exploring the potential that word embeddings can ofer. This study covers the entire pipeline in lexical simplification, from the task of complex word identification (CWI) to substitute generation, selection and ranking (SG/SS/SR). Taking advantage of the diferent applications of BERT models, we fine-tune two pre-trained models to detect unusual words with the help of available Spanish datasets. Next, we compare features that diferent types of embedding can give to find the best candidate for replacement for a target word. The resulting models in the CWI step show a fair result compared to other systems that used the same datasets. Also, we found better results than previous works by analyzing the similarity of words in context when evaluating embedding models.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;lexical simplification</kwd>
        <kwd>word embedding</kwd>
        <kwd>BERT</kwd>
        <kwd>Spanish</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Lexical simplification (LS) is an important subfield of text simplification that gives attention to
the complexity of words, and particularly how to measure readability and reduce the complexity
using alternative replacements. Most current approaches to LS heavily rely on corpus statistics
and surface-level features, such as word length and corpus-based word frequencies [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
most popular LS systems still predominantly use a set of rules for substituting complex words
with their frequent synonyms from carefully handcrafted databases or automatically induced
from parallel corpora. However, language resources are scarce or expensive to produce, such as
WordNet and Simple Wikipedia. Also, the scarcity of these resources may be greater according
to the language, as it is the case of Spanish versus English.
      </p>
      <p>
        In order to find alternative solutions to these costly procedures, recent works use word
embeddings to extract important information from a text with less efort [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This information
can help in various aspects of the simplification process, such as extracting word vectors to
detect unusual words [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or determining the similarity of words in a given context.
      </p>
      <p>In this paper, we explore applications that word embedding can ofer in each of the stages of the
LS process individually. We present a simple way of fine-tuning Spanish and multilingual BERT
models for the detection of unusual or complex words. As a next step, to provide replacements
that match the context of the original word, we use the similarity information between the
word vectors of diferent types of embeddings. Finally, we combine this information to rank
words in terms of simplicity.</p>
      <p>The remainder of this paper is organized as follows. In Section 2, we review the work related
to the text simplification process. In Section 3, we describe the datasets to train and evaluate
the procedures proposed in each step of LS. Sections 4, 5, 6 and 7 provide the procedures and
evaluation to the complex word identification (CWI), substitute generation/selection/ranking
(SG/SS/SR) modules. Finally, Section 8 ofers conclusions and future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Natural language processing is a discipline dedicated to developing technology capable of
understanding natural language in a way similar to human beings. One area in which this
could be applied is the development of technology that improves accessibility for individuals
with disabilities. LS aims at replacing complex words with simpler alternatives which can help
various groups of people, including people with autism [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], aphasia [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], low vision
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], dyslexia [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] or people with intellectual disabilities [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. According to
studies [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], LS is an essential task because a person needs to know 95% of the vocabulary to
understand a text at a basic level. Therefore, this suggests that replacing words that are unusual
for a person can improve the accessibility of a given text.
      </p>
      <p>
        There are three approaches to LS, from supervised machine learning algorithms to
unsupervised algorithms or even hybrid approaches which combine the advantages of both. Paetzold
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] proposed four stages to achieve LS, which are: CWI, SG, SS and SR. This paper follows these
stages and explores whether embeddings can help at each stage.
      </p>
      <p>
        CWI aims to select the complex words candidates to be simplified in a given text. Several ways
to accomplish this task have been proposed, however, approaches based on machine learning
have proven to be the most suitable. Shardlow [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] compared binary support vector machine
(SVM), threshold-based, and łSimplify Everythingž approaches, where in the latter, it is assumed
that all words in a sentence can be simplified. The results demonstrated that the SVM approach
outperforms the others in terms of precision. This is confirmed in competitions focused on this
task, such as the BEA workshop (2018) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], where machine learning-based approaches proved
to have the best F1 scores [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In this research work, we exploit the versatility of BERT and
perform fine-tuning of a model with the help of information from two Spanish datasets (see
Section 3) with the aim of performing Named Entity Recognition (NER) to detect complex and
simple words in a given text described at Section 4.
      </p>
      <p>
        Moving on to the next step, SG involves producing substitute candidates for the complex
words detected. Two approaches have been propose, which are linguistic database querying and
automatic generation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The first one obtains candidates from databases manually constructed
by professionals, thus providing reliable data on words associated with their candidates [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
Nevertheless, it has the disadvantages of being a time-consuming task and not having a wide
coverage, especially in Spanish. Automatic generation focuses on overcoming this disadvantage
and seeks to gather extracted candidates from less expensive resources. For example, the simple
multilingual Paraphrase Database (PPDB) [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] where they initially annotated a paraphrase
dataset to train a model to classify simplified paraphrases resulting in a database of more than a
billion of paraphrases for diferent languages. Due to the limited amount of resources available
for the Spanish language, in this work we explore another source that has the ability to provide
replacements for a target word, such as the case of word embeddings. In Section 5, we compare
diferent models to evaluate which type of embedding performs better at this step.
      </p>
      <p>
        In the third step (SS), in which a substitute is selected from the set of synonyms extracted
from the previous step, the most suitable synonym is chosen according to its simplicity and
context. In recent years, several strategies have been proposed, for example, works that took
this step as a task of Word Sense Disambiguation (WSD) [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Moreover, in languages where
WSD resources are sparse or unavailable, Part-of-speech (POS) strategies were proposed, as
in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], where the words are filtered using a set of rules, including among others, the POS tag
of the candidate. Unfortunately, this approach showed poor results when dealing with highly
ambiguous words. Therefore, to address these problems, recent works incorporate similarity
metrics in the selectors where authors picked out the final synonym using the cosine distance in
a word embedding model. Given a word to be simplified, the word with the closest vector based
on cosine similarity was chosen [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. In Section 6 we perform a similar procedure and explore
which type of embedding has better results when selecting which words are more appropriate
for a given context.
      </p>
      <p>
        Finally, the SR step consists in deciding which of the candidate substitutions that fit the
context of a complex word is the simplest. One of the most commonly used and simplest
strategies for dealing with this task is frequency-based procedures. These suggest that the
more a word is used, the more familiar it is to a user; these word frequencies can be extracted
from very large corpora [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] and can be quite efective in front of other approaches. As in the
other steps, machine learning assisted approaches have been adopted lately. For example, a
support vector machine accompanied with additional metrics to sort words according to their
simplicity [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Later, more sophisticated works such as neural approaches were presented [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
such is the case of [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], where a supervised neural ranking model is presented. This ranker
receives a set of features (n-gram probabilities) for a pair of candidates as input, and produces as
output the simplicity diference between them. Recently, some works have combined resources
obtained from the strategies described above. Such is the case of [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], which uses a weighting
system where it takes features extracted from word embeddings, language models and word
frequencies. In this work, we follow this idea by incorporating our own features adapted to
Spanish, extracted from word embeddings and a frequency dictionary (described at Section 7).
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Datasets</title>
      <p>
        For experimentation, diferent datasets for training and testing are used. These datasets are
described below.
The dataset is composed of annotated Spanish Wikipedia pages proposed in the BEA Workshop
2018 for the CWI 1 task. As shown in Table 3, a total of 17603 instances were annotated by 54
Spanish speakers, most of whom were native. Each instance contains a uniword/multiword
target which is selected by annotators. A target is marked as complex if at least one annotator
designates it as complex.
3.2. EASIER DATASETS
These datasets [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] are part of the EASIER corpus 2 3, which were developed by the authors
of this work to ofer evaluation support for CWI tasks and fitting SG/SS tasks contextually. A
linguist expert in easy-to-read and plain language guidelines annotated 260 news documents.
Later on, two additional experts and a target audience analysed the resulting corpus to assure
the quality of the data provided.
      </p>
      <p>Table 1 and 2, show examples of instances found in the CWI and substitute datasets,
respectively. In the dataset for CWI, information such as a sentence, a target word, the word ofset
and the gold-standard label for the binary task can be found. While, for the substitutes dataset,
a sentence, a target word and proposed substitutes can be found.</p>
      <p>Sentence
La importancia de leer bien el etiquetado antes de comprar un alimento.
(The importance of carefully reading the labelling before purchasing foodstufs.)
La importancia de leer bien el etiquetado antes de comprar un alimento.
(The importance of carefully reading the labelling before purchasing foodstufs.)
La importancia de leer bien el etiquetado antes de comprar un alimento.
(The importance of carefully reading the labelling before purchasing foodstufs.)
La importancia de leer bien el etiquetado antes de comprar un alimento.
(The importance of carefully reading the labelling before purchasing foodstufs.)
La importancia de leer bien el etiquetado antes de comprar un alimento.
(The importance of carefully reading the labelling before purchasing foodstufs.)
Start ofset End ofset Word
3 14 I(mimppoorrttaanncciae) 0
18 22 (Lreeearding) 0
31 41 (Eltaibqeulelitnagd)o 1
51 58 (Cpoumrcphraasring) 0
62 70 (Afoliomdesntutofs) 0</p>
      <p>Label</p>
      <p>Also to complement the above information, Table 3 shows additional information on the size
of the resources. The BEA dataset contains more than 17,000 instances where more than 7,000
complex words are found. Whereas, the CWI dataset of the EASIER corpus contains more than</p>
      <sec id="sec-3-1">
        <title>1google.com/view/cwisharedtask2018 2https://data.mendeley.com/datasets/ywhmbnzvmx/2 3github.com/LURMORENO/EASIER_CORPUS</title>
        <p>44,000 instances where more than 8,000 complex words are found. Similarly, the EASIER corpus
substitutes dataset contains 5,130 instances resulting in more than 7,000 proposed substitutes.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Complex Word Identification</title>
      <p>In this stage, we need to distinguish which words are complex and which are not for a certain
audience.</p>
      <p>
        We propose BERT [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] for NER because is a powerful NLP model but using it for NER
without fine-tuning it on NER dataset won’t give good results. In this work, we fine-tune two
models with the help of the datasets described at Section 3 to perform the CWI task: a Google´s
multilingual BERT pre-trained model (mBERT) 4 [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] and an Spanish BERT pre-trained model
(BETO)5 [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. We took the original idea from an implementation 6 for CoNLL-2003 [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] and
then modified it so that the model can predict the entities "COMPLEX" and "SIMPLE" in a given
text. The default parameters for the fine-tuning are the following:
• train_batch_size: 32
• max_seq_length: 128
• learning_rate: 2e-5
• num_train_epochs: 4.0
• do_lower_case: False
• Crf: True
      </p>
      <sec id="sec-4-1">
        <title>4https://github.com/google-research/bert/blob/master/multilingual.md 5https://github.com/dccuchile/beto 6github.com/kyzhouhzau/BERT-NER</title>
        <p>Also, by analyzing errors, we noticed that many of the false positives were instances where
a multiword was the target, however, the data for training/testing used in fine-tuning, only
correspond to instances that have uniwords, while SVM classifies instances where the target is
uniwords and multiwords. We believe that by incorporating these instances in the classification
of the BERT model, the score can be improved and compared to the BEA score.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Substitute Generation</title>
      <p>
        The SG stage generates substitution candidates for complex words, considering all the contexts
in which they may appear. We tested the performance of diferent embedding models by
extracting and evaluating the nearest neighbors of each target word (top-50 neighbors). The
tested models are the following:
• Word2vec model: pre-trained on The Spanish Billion Words Corpus 7.
• Sense2Vec model: Since there are no Spanish Sense2Vec models [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]. We created a model
trained on The Spanish Billion Words Corpus [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. A sense is a word combined with a
label that represents the context of a word (in this case we use the POS tag as a label). The
main diference of Sense2Vec and Word2Vec vectors is that the latter fail to encode the
context by assigning a single key regardless of the context in which it appears. This does
not happen in a Sense2vec model, because it generates vectors of words with contextual
keys (i.e., one vector for each sense of the word).
• FastText model: pre-trained on Wikipedia with the FastText tool with character n-grams
of length 5 8 [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ].
      </p>
      <p>• BERT model: Pytorch BETO model described at Section 4.</p>
      <p>The gold set is part of the EASIER corpus, which is represented by 575 instances in which
for each instance a target word has three proposed substitutes. For this evaluation the first
500 instances are taken for the test set. In addition, we compare the results of the models</p>
      <sec id="sec-5-1">
        <title>7https://crscardellino.ar/SBWCE/ 8https://fasttext.cc/docs/en/crawl-vectors.html</title>
        <p>
          with a previous approach which performs a linguistic database strategy proposed in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] where
we developed the same task by extracting replacements for a target word from Babelnet [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ],
Thesaurus 9 and PPDB [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
        </p>
        <p>
          The evaluation metrics used are those found in the work of Paetzold [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], which are as follows:
• Potential: The proportion of instances for which at least one of the candidates generated
is contained within the gold standard.
• Precision: The proportion of generated substitutions that are contained within the gold
standard.
• Recall: The proportion of gold-standard substitutions that are among the generated
substitutions.
        </p>
        <p>• F-1: The harmonic average between precision and recall.</p>
        <p>
          Table 5 includes the results obtained for this step. At this stage, potential and recall are
important measures because, according to its definition, it is required to obtain the widest
coverage in the contexts in which a word may appear. The approach developed in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] showed a
higher performance than the result of the embedding models, obtaining a potential and recall of
0.898 and 0.597 respectively, versus the second best being the Sense2Vec model with a potential
of 0.506 and a recall of 0.298. When analyzing the negative results we found cases in which
the output was a repeated candidate but in diferent grammatical forms. In turn, because these
models provide semantic similarity of words, in many cases, apart from synonyms, antonyms
were found in the lists.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Substitute Selection</title>
      <p>The SS stage takes the list of synonyms extracted from the previous step and selects the most
suitable synonym according to its simplicity and context. As the core resource in this step,
we use diferent types of word embedding models, from static to contextualized. We use the
same embedding models as in Section 5. These models allow us to calculate the cosine distance
between word vectors to perform the following procedures:
• No selections : selects all candidates.
• Lexical window : obtains three similarity values (candidate and target word, candidate
and target word’s context words in the sentence (previous and subsequent words)). Next,
these values are added and stored. Finally, this process is repeated for every candidate,
and the selector picks the three candidates with the highest values.</p>
      <p>For the evaluation of this stage, we use the same data set and metrics as in Section 5. On
the selector to evaluate, each selector needed candidates to rank, so we use the generator with
the best potential ranking from the previous step described in Section 5 and then randomly
insert the correct substitutes for each of the instances from the gold set. Furthermore, in this
evaluation each selector had to propose the top 3 substitutes per instance. We believe that in
this way we can easily determine the efectiveness of the selector based on which selector yields
the highest number of potential correct answers.</p>
      <p>
        Table 6 illustrates the results. Unlike the previous step, in the substitute selection a higher
precision is pursued. As expected, performing no selection results in high potential and recall,
however, the precision is very low. With the closest score in potential is the FastText model,
which showed a precision of 0.364 being higher than the Word2Vec model used in previous
works [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] with a precision score of 0.315. We assume that this higher score was obtained
because the FastText model provides char and ngrams embeddings to face the problem of OOV
(Out-of-vocabulary) words.
      </p>
      <p>On the other hand, about the results of the BERT model, it is worth noting that word-level
similarity comparisons are not appropriate with BERT embeddings because these embeddings
are contextually dependent, meaning that the word vector changes depending on the sentence
it appears in. A better intuition for this stage with the BERT model would be to evaluate the
similarity between the sentences in which a candidate is found.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Substitute Ranking</title>
      <p>The SR stage takes the list of synonyms extracted from the previous step and chooses which
candidate that fits the context is the simplest, taking into account the target user. At this stage, a
combination of frequency-based and machine learning-assisted strategies has been implemented
by developing a weighting module that uses the diferent features to rank a word.</p>
      <p>Table 7 shows the results of diferent combinations of these features. To our knowledge there
are no datasets in Spanish to evaluate this procedure, therefore, the decision to adapt these
procedures to evaluate it with English language datasets was made, specifically, datasets from
the English Lexical Simplification task of SemEval 2012 [ 35]. The trial set is composed of 300
instances, and the test set, 1, 710 instances. Each instance contains a sentence, a target complex
word, and candidates ranked by their simplicity.</p>
      <p>The evaluation metric is the TRank measure, proposed in the shared task. This metric
calculates the proportion of instances for which the highest ranked candidate produced by a
ranker is the same as the one in the gold-standard. In addition, the Table 7 also shows results
of the best ranker presented in the shared task. The ranker must make the decision to choose
the simplest candidate based on the candidates that obtained the best results in each of the
following features:
• BERT prediction: Probability distribution of the candidate. This can be obtained from
the vocabulary corresponding to the mask word. The higher the probability, the more
relevant the candidate for the original sentence. The BETO model described above is used
for Spanish and a multilingual DistilBERT [36] model is used for English.
• Semantic similarity: Cosine distance between the original word vectors and the
candidate vectors in the list. The shorter the distance, the more similar the two words. To
extract these vectors we test diferent embedding models. For the Spanish language, the
classic embedding models described above are used and for English, pre-trained models
for that language are used 10 11 12.
• Frequency Feature: Because frequency-based approaches have shown good results at
this stage, the decision was made to incorporate it as a feature in the ranker. The more
frequent a word is, we assume that a word is simpler. For Spanish, we used a dictionary
of the Real Academia de la Lengua Española (RAE)13 to extract the frequency of each
candidate, which is made up of 10000 terms ordered by their frequency. As for the English
language, a portion of about 5000 instances of the Corpus of Contemporary American
English (COCA)14 has been used.</p>
      <p>As shown in Table 7, the frequency-based approach alone obtained good results with a
TRank of 0.513, outperforming a strong baseline with TRank of 0.454 and being close in TRank
to the best team (UOW-SHEF-SimpLex) presented in the task which developed a supervised
approach with contextual and psycholinguistic features. On the other hand, the proposed
embedding approaches did not show good results in detecting the simplicity of the words,
moreover, when combined with the other features, they showed a lower TRank score than the
individual frequency feature (0.37). When analyzing errors, problems were detected with the
classification of multiwords, because the classical embedding models receive uniwords as inputs,
they did not assign a weight to the multiwords, consequently, classifying it as the most complex
term in the list and therefore, obtaining wrong results in many cases. In the case of the results
for the BERT model, we believe that by performing a fine-tuning process as was done in the
CWI stage, it could improve the results in this task.</p>
      <p>10https://mccormickml.com/2016/04/12/googles-pretrained-word2vec-model-in-python/
11https://fasttext.cc/docs/en/crawl-vectors.html
12https://github.com/explosion/sense2vec
13http://corpus.rae.es/lfrecuencias.html
14https://www.english-corpora.org/coca/</p>
    </sec>
    <sec id="sec-8">
      <title>8. Conclusions and Future Work</title>
      <p>The main objective of this work is to explore the possible uses of recent Spanish word embeddings
for each of the stages of the LS task in the Spanish language, which has limited resources.</p>
      <p>To achieve this goal, as a first step, we fine-tuned BETO and mBERT models to perform
the NER task to discern between complex and simple words. The experiments showed a fair
result against other supervised approaches, however, there is room for improvement. On the
generator side, the performance of diferent types of embeddings was explored. By analyzing
the results, we can understand that embeddings are not recommended for this stage, due to the
presence of antonymy in the near neighbors of a target word. In contrast, when evaluating the
similarity between words, the embeddings models showed better results in the selectors, such is
the case of the FastText model that obtained a higher precision than in previous works. Finally,
in the SR stage, a weighting system using information extracted from frequency dictionaries
and embedding models was proposed to choose the simplest candidate. When adapted and
evaluated in English, the best results were obtained for the frequency features and showed
room for improvement with the embedding features.</p>
      <p>As future work, the incorporation of multiwords in the fine-tuning process should be
contemplated for the task of CWI and SR, because they were the main cause of the drop in the
respective scores for each task. Also, a round-trip evaluation is necessary to determine the
results of a complete system.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>This work has been supported by the Madrid Government (Comunidad de Madrid-Spain)
under the Multiannual Agreement with UC3M in the line of Excellence of University
Professors (EPUC3M17), and in the context of the V PRICIT (Regional Programme of Research and
Technological Innovation).
in: Proceedings of the 48th annual meeting of the association for computational linguistics,
2010, pp. 216ś225.
[35] L. Specia, S. K. Jauhar, R. Mihalcea, SemEval-2012 task 1: English lexical simplification,
in: *SEM 2012: The First Joint Conference on Lexical and Computational Semantics
ś Volume 1: Proceedings of the main conference and the shared task, and Volume 2:
Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012),
Association for Computational Linguistics, Montréal, Canada, 2012, pp. 347ś355. URL:
https://aclanthology.org/S12-1046.
[36] V. Sanh, L. Debut, J. Chaumond, T. Wolf, Distilbert, a distilled version of bert: smaller,
faster, cheaper and lighter, ArXiv abs/1910.01108 (2019).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Maddela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <article-title>A word-complexity lexicon and a neural readability ranking model for lexical simplification</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>05754</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G. H.</given-names>
            <surname>Paetzold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Specia</surname>
          </string-name>
          ,
          <article-title>A survey on lexical simplification</article-title>
          ,
          <source>Journal of Artificial Intelligence Research</source>
          <volume>60</volume>
          (
          <year>2017</year>
          )
          <article-title>549ś593</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Alarcon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Martínez</surname>
          </string-name>
          ,
          <article-title>Lexical simplification system to improve web accessibility</article-title>
          ,
          <source>IEEE Access 9</source>
          (
          <year>2021</year>
          )
          <article-title>58755ś58767</article-title>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2021</year>
          .
          <volume>3072697</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E.</given-names>
            <surname>Barbu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Martín-Valdivia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A. U.</given-names>
            <surname>Lopez</surname>
          </string-name>
          ,
          <article-title>Open book: a tool for helping asd users' semantic comprehension</article-title>
          ,
          <source>in: Proceedings of the Workshop on Natural Language Processing for Improving Textual Accessibility</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>11ś19</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Orasan</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Dornescu</surname>
          </string-name>
          ,
          <article-title>An evaluation of syntactic simplification rules for people with autism</article-title>
          ,
          <source>Association for Computational Linguistics</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Barbu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Martín-Valdivia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Martínez-Cámara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Urena-López</surname>
          </string-name>
          ,
          <article-title>Language technologies applied to document simplification for helping autistic people</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>42</volume>
          (
          <year>2015</year>
          )
          <article-title>5076ś5086</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Carroll</surname>
          </string-name>
          , G. Minnen,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Canning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tait</surname>
          </string-name>
          ,
          <article-title>Practical simplification of english newspaper text to assist aphasic readers</article-title>
          ,
          <source>in: Proceedings of the AAAI-98 Workshop on Integrating Artificial Intelligence and Assistive Technology</source>
          ,
          <year>1998</year>
          , pp.
          <fpage>7ś10</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , G. Unthank,
          <article-title>Helping aphasic people process online information</article-title>
          ,
          <source>in: Proceedings of the 8th International ACM SIGACCESS Conference on Computers and Accessibility</source>
          ,
          <year>2006</year>
          , pp.
          <fpage>225ś226</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L.</given-names>
            <surname>Sauvan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Stolowy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Aguilar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>François</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Matonti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Castet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Calabrese</surname>
          </string-name>
          ,
          <article-title>Text simplification to help individuals with low vision read more fluently</article-title>
          ,
          <source>in: LREC 2020-Language Resources and Evaluation Conference</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>11ś16</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Wilkens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Oberle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Todirascu</surname>
          </string-name>
          ,
          <article-title>Coreference-based text simplification</article-title>
          ,
          <source>in: Proceedings of the 1st Workshop</source>
          on Tools and
          <article-title>Resources to Empower People with REAding DIficulties (READI</article-title>
          ),
          <year>2020</year>
          , pp.
          <fpage>93ś100</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Rello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Baeza-Yates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          ,
          <article-title>Simplify or help? text simplification strategies for people with dyslexia</article-title>
          ,
          <source>in: Proceedings of the 10th International Cross-Disciplinary Conference on Web Accessibility</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>1ś10</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Gómez-Martínez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Etayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Anula</surname>
          </string-name>
          , L. Bourg,
          <article-title>Text simplification in simplext: Making texts more accessible, Procesamiento del lenguaje natural (</article-title>
          <year>2011</year>
          ) 341ś
          <fpage>342</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Drndarević</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          ,
          <article-title>Can spanish be simpler? lexsis: Lexical simplification for spanish</article-title>
          ,
          <source>in: Proceedings of COLING</source>
          <year>2012</year>
          ,
          <year>2012</year>
          , pp.
          <fpage>357ś374</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          ,
          <article-title>Automatic text simplification: Synthesis lectures on human language technologies</article-title>
          , vol.
          <volume>10</volume>
          (
          <issue>1</issue>
          ), California, Morgan &amp; Claypool Publishers (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Shardlow</surname>
          </string-name>
          ,
          <article-title>A comparison of techniques to automatically identify complex words</article-title>
          .,
          <source>in: 51st Annual Meeting of the Association for Computational Linguistics Proceedings of the Student Research Workshop</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>103ś109</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Yimam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Stajner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Riedl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Biemann</surname>
          </string-name>
          ,
          <article-title>Multilingual and cross-lingual complex word identification</article-title>
          .,
          <source>in: RANLP</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>813ś822</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Yimam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Biemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. H.</given-names>
            <surname>Paetzold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Specia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Štajner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <source>A report on the complex word identification shared task</source>
          <year>2018</year>
          , arXiv preprint arXiv:
          <year>1804</year>
          .
          <volume>09132</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Burstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sabatini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-W.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ventura</surname>
          </string-name>
          ,
          <article-title>The automated text adaptation tool, in: Proceedings of Human Language Technologies: The Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT</article-title>
          ),
          <year>2007</year>
          , pp.
          <fpage>3ś4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>E.</given-names>
            <surname>Pavlick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Callison-Burch</surname>
          </string-name>
          ,
          <article-title>Simple ppdb: A paraphrase database for simplification, in: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics</article-title>
          (Volume
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <year>2016</year>
          , pp.
          <fpage>143ś148</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Thomas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <article-title>Wordnet-based lexical simplification of a document.</article-title>
          ,
          <source>in: KONVENS</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>80ś88</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>G.</given-names>
            <surname>Paetzold</surname>
          </string-name>
          , L. Specia,
          <article-title>Text simplification as tree transduction</article-title>
          ,
          <source>in: Proceedings of the 9th Brazilian Symposium in Information and Human Language Technology</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>G.</given-names>
            <surname>Paetzold</surname>
          </string-name>
          , L. Specia,
          <article-title>Lexenstein: A framework for lexical simplification</article-title>
          ,
          <source>in: Proceedings of ACL-IJCNLP 2015 System Demonstrations</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>85ś90</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>R.</given-names>
            <surname>Baeza-Yates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dembowski</surname>
          </string-name>
          ,
          <article-title>Cassa: A context-aware synonym simplification algorithm</article-title>
          ,
          <source>in: Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>1380ś1385</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>C.</given-names>
            <surname>Horn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Manduca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kauchak</surname>
          </string-name>
          ,
          <article-title>Learning a lexical simplifier using wikipedia, in: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics</article-title>
          (Volume
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <year>2014</year>
          , pp.
          <fpage>458ś463</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>G.</given-names>
            <surname>Paetzold</surname>
          </string-name>
          , L. Specia,
          <article-title>Lexical simplification with neural ranking</article-title>
          ,
          <source>in: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume</source>
          <volume>2</volume>
          ,
          <string-name>
            <given-names>Short</given-names>
            <surname>Papers</surname>
          </string-name>
          ,
          <year>2017</year>
          , pp.
          <fpage>34ś40</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>J.</given-names>
            <surname>Qiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>Lexical simplification with pretrained encoders</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <year>1907</year>
          .06226.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>R.</given-names>
            <surname>Alarcon</surname>
          </string-name>
          ,
          <article-title>Dataset of sentences annotated with complex words and their synonyms to support lexical simplification</article-title>
          ,
          <year>2021</year>
          . URL: https://data.mendeley.com/datasets/ywhmbnzvmx/2.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cañete</surname>
          </string-name>
          , G. Chaperon,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fuentes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-H.</given-names>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <article-title>Spanish pre-trained bert model and evaluation data</article-title>
          ,
          <source>in: PML4DC at ICLR</source>
          <year>2020</year>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>E. F.</given-names>
            <surname>Tjong Kim Sang</surname>
          </string-name>
          , F. De Meulder,
          <article-title>Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition</article-title>
          ,
          <source>in: Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL</source>
          <year>2003</year>
          ,
          <year>2003</year>
          , pp.
          <fpage>142ś147</fpage>
          . URL: https://www.aclweb.org/anthology/W03-0419.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>A.</given-names>
            <surname>Trask</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Michalak</surname>
          </string-name>
          , J. Liu, sense2vec
          <article-title>- a fast and accurate method for word sense disambiguation in neural word embeddings</article-title>
          ,
          <year>2015</year>
          . arXiv:
          <volume>1511</volume>
          .
          <fpage>06388</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>C.</given-names>
            <surname>Cardellino</surname>
          </string-name>
          ,
          <source>Spanish Billion Words Corpus and Embeddings</source>
          ,
          <year>2019</year>
          . URL: https:// crscardellino.github.io/SBWCE/.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>E.</given-names>
            <surname>Grave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          ,
          <source>Learning word vectors for 157 languages</source>
          ,
          <year>2018</year>
          . arXiv:
          <year>1802</year>
          .06893.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>R.</given-names>
            <surname>Navigli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Ponzetto</surname>
          </string-name>
          , Babelnet:
          <article-title>Building a very large multilingual semantic network,</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>