<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IRAZ: Easy-to-Read Content Generation via Automated Text Simplification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thierry Etchegoyhen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jesús Calleja Pérez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Ponce</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fundación Vicomtech, Basque Research and Technology Alliance (BRTA)</institution>
          ,
          <addr-line>Donostia-San Sebastián, 20009, Gipuzkoa</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Information complexity is a critical communication and integration barrier for large segments of the population. This situation is exacerbated by the ever increasing volumes of content generated in modern digital societies. The IRAZ project aims to develop a flexible solution for easy-to-read text generation, by means of automated text simplification, to support the production of accessible content. The project sets to produce new datasets in the field, created by professionals or via synthetic data generation. It also explores neural approaches to lexical, syntactic and end-to-end text simplification, ofering configurable simplification hypotheses that can be post-edited by professional easy-to-read content creators. The project aims to develop a generic solution, with special emphasis on Basque and Spanish to provide further technological support for these two relatively under-resourced languages in the field.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Accessibility</kwd>
        <kwd>Easy-to-read</kwd>
        <kwd>Text Simplification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>and present each sentence separately, with one
grammatical segment per line.2 This method is one of the
Digital transformation is rapidly impacting societies main tools to facilitate the communication of information
around the world, with ever increasing volumes of infor- across domains. However, current processes are mainly
mation being produced and shared in multiples languages performed manually, without technological support for
and domains. Being able to understand this information the most part. This state of afairs hinders the production
is a requirement in modern society. Unfortunately, in of new accessible content and the development of a more
many cases, the available information is not accessible to inclusive society.
large portions of the population, due to its intrinsic com- Project IRAZ3 aims to address the current limitations
plexity. Thus, a communication barrier exists for persons in easy-to-read content production, via the development
with cognitive dysfunction or impairment, older people, of an integrated solution based on automated text
simplimigrant populations and refugees, or people sufering ifcation (ATS) technology. Its main objective is thus the
from learning dificulties, among others. development of an application that will assist
easy-to</p>
      <p>Although estimates may vary depending on the coun- read content creators, by providing simplified versions
try and the included segments of the population, it is of texts and allowing experts to post-edit system
suggestypically estimated that more than 25% of the population tions. The collected post-edited data will further allow
faces reading and comprehension dificulties in a country the training of improved versions of the text
simplificalike Spain [1]. For such a large portion of the popula- tion models that will be developed within the project.
tion, the inability to properly access the information in
essential domains such as health, education, culture or
media, can result in social exclusion in many sectors and 2. Consortium and Funding Body
activities.</p>
      <p>The easy-to-read method1 aims to facilitate informa- IRAZ is partially funded by the Basque Government via
tion access by presenting it under a series of guidelines, the Hazitek 2022 program of the Spri Group, as an
inincluding: favour the use of simple words and short sen- dustrial research project under Grant Agreement
ZLtences with less than 15 words, employ direct language, 2022/00788. The project started in April 2022 and will
ifnalise in December 2024.</p>
      <p>
        The consortium includes the following participants:
Vicomtech4 as the research centre leading research and
development activities; Merkatu Digital5, as project
coordinator; Gureak Marketing6; Lantegi Batuak7; Lectura SVM-based CWI and use these language models (LMs)
Fácil Euskadi8; and Merkatu Interactiva5. It is worth for substitution generation and ranking. In recent results
noting that the project includes leading experts in easy- from the TSAR-2022 Shared Task [11], participants have
to-read content creation and dissemination in the Basque mostly used neural LMs for the LS task. The best results
Country. on this task, for English, were obtained via prompts fed to
the very large generative pretrained language model
GPT3 [12]. Current LS approaches mainly use task-specific
3. State of the Art corpora for system benchmarking [13, 14].
Most current ATS models beyond LS employ
end-toAs noted above, most easy-to-read content is typically end (E2E) architectures, which attempt to model a
comcreated in a manual fashion by experts in the field, with- plete transformation from complex to simple text.
Diferout relevant technological support. A limited number of ent architectures have been explored along these lines,
projects, e.g., the Easy Reading project9, have attempted e.g., LSTMs [15] or Semantic Encoders Vu et al. [
        <xref ref-type="bibr" rid="ref3">16</xref>
        ],
alto include this type of support for easy-to-read content though, in recent years, Transformer models [17] have
generation and dissemination. Easy-to-read content usu- become ubiquitous for ATS as well. Thus, Zhao et al.
ally involves adaptation beyond text simplification, in par- [18] combine Transformers and paraphrase rules for the
ticular via the provision of additional information, such task, whereas Martin et al. [19] use parameter tokens to
as simple explanations of complex concepts. Nonetheless, control the output sequence in terms of character length
automated text simplification, defined as the reduction ratio and Levenshtein distance in their ACCESS model.
of complexity of a given text while retaining the origi- Variants of the latter approach have been proposed with
nal content and information [
        <xref ref-type="bibr" rid="ref23">2</xref>
        ], is typically viewed as pretrained BART models [20] and fine-tuned T5
moda key enabling technology to facilitate access to textual els [21, 22]. Alternatively, Omelianchuk et al. [23] use
information for people with reading dificulties. Projects a RoBERTA model to tag the input sequence with keep,
sMuuchTeaSs1P0SaEnTd [S3i]mapnldeTSeixmtp11l,ehxtav[4e]s,eotr,tmooprreovreidceentthliys, Ctyopne- delete or append tokens, among others, moving from a
generative task into a sequence labelling task.
of support via ATS technology. An important limitation for ATS modelling is the
      </p>
      <p>Texts may be simplified at diferent levels: lexical, re- scarcity of training corpora, particularly for E2E
modplacing complex words with simpler synonyms; syntactic, els. Most datasets are available only for English and the
transforming complex sentences involving coordination largest are derived from alignments of Wikipedia and
or diferent types of modifiers into separate simple sen- Simple Wikipedia content [24, 25]. Xu et al. [26]
hightences, or transforming passive voice into active, among light critical issues with Simple Wikipedia for ATS and
other operations; and conceptual, tackling coreference introduced Newsela, a professionally-produced corpus
resolution, for example. Earlier approaches attempted for English and Spanish. For Basque, only two small
to model these transformations via computational rules, hand-crafted datasets are currently available, one in the
hand-crafted [5] or inferred from aligned corpora [6]. science popularisation domain [27], the other on news
Later data-driven techniques such as Statistical Machine from the Irekia Open Government portal12 [28].
Translation led to formulating the simplification prob- To overcome data scarcity, diferent approaches have
lem as an end-to-end monolingual translation task, using been recently proposed for ATS. Thus, Surya et al. [29]
corpora of aligned complex and simple sentences [7]. described an unsupervised method based on a shared
en</p>
      <p>As in most natural language processing fields, data- coder and two attentional-decoders with
discriminationdriven approaches based on artificial neural networks based losses and denoising, which can be trained on
unlaand deep learning have become the dominant paradigm beled text data or in a weakly supervised fashion with a
in ATS research, in recent years. For lexical simplification few aligned examples. The creation of synthetic datasets
(LS), for example, LS-BERT [8] is currently a standard is another alternative to address training data scarcity.
baseline, based on a BiLSTM model for complex word Along these lines, Lu et al. [30] exploit machine
transidentification (CWI) and a pretrained BERT model [ 9] lationese to build pseudo-parallel datasets for the task.
for substitute candidate generation. For Spanish, one In Kim et al. [31], existing parallel datasets are exploited
of the languages of interest in IRAZ, Alarcón et al. [10] to create one-to-many monolingual parallel corpora, via
use contextual vectors extracted from pretrained mBERT machine translation, allowing the training of models that
and BETO models, among other features, to perform learn to split complex sentences into simpler ones.
6https://www.gureakmarketing.com
7https://www.lantegibatuak.eus
8https://lecturafacileuskadi.net
9https://www.easyreading.eu/the-project/
10https://www.upf.edu/web/conmutes
11https://anr.fr/Project-ANR-22-CE23-0019
12https://www.irekia.euskadi.eus/en</p>
    </sec>
    <sec id="sec-2">
      <title>4. IRAZ</title>
      <sec id="sec-2-1">
        <title>4.1. Objectives and Challenges</title>
        <sec id="sec-2-1-1">
          <title>The main objective of project IRAZ is the development</title>
          <p>of a software solution to support the generation of
easyto-read textual content, by means of automated text
simplification. 4.2.2. Simplification Methods</p>
          <p>Although the goal is to develop a generic solution Instead of providing one-size-fits-all simplification,
irrewhich could support any language, in principle, a sec- spective of a given adaptation task, IRAZ aims to
investiondary objective of the project is the development of rele- gate diferent types of ATS methods separately, and allow
vant resources for two specific languages, which are crit- users to select and test diferent ATS models individually
ical to the participants of the project: Spanish, for which or in combination, depending on the content at hand. For
limited resources are currently available, and Basque, for example, with specific types of content, performing only
which there exist only two small datasets, as previously lexical simplification might be an optimal choice, if other
indicated. types of transformation do not provide accurate results,</p>
          <p>As described in the previous section, ATS is a challeng- or if the user is only looking for this type of
simplificaing field in many respects. First, there is a significant tion. Alternatively, for other types of content, end-to-end
lack of resources across languages and domains, which ATS models may provide suficient quality and be used
hinders the development of data-driven models on a par directly to generate simplification hypotheses.
with those obtained in neighbouring fields such as neural A significant part of the research activities within the
language modelling and machine translation. project will thus address the following three main types
of ATS methods:
synthetic data merger, paraphrase generation or artificial
one-to-many data generation along the lines of [31].</p>
          <p>All data are automatically extracted and aligned, using
standard tools such as CATS [32] or in-house processing
scripts developed within the project.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>4.2. Approach</title>
        <p>Given the current challenges and limitations in the field
of ATS, the project adopted a multi-pronged approach
for the development of a flexible solution which could
both (i) benefit easy-to-read content creators, and (ii)
help advance the state-of-the-art in the field. The main
components of our approach are summarised below.
4.2.1. Data creation
The project involves specific resource collection and
generation activities. Documents consisting of complex or
simplified data 13 will thus be collected and processed
within various cycles of the project. This activity will
lead to the generation of new datasets in diferent
domains, in the two selected languages. It includes both
the collection of proprietary data from project
participants’ repositories, for which exploitation right within
the project have been cleared, and the preparation of new
datasets from public sources, such as Irekia (op. cit.) for
Basque and Spanish news.</p>
        <p>Considering the current lack of resources across the
board to train end-to-end neural simplification models,
part of the project activities are focused on synthetic data
generation. We thus investigate diferent methods to
generate synthetic datasets of complex-simple pairs, via
13For convenience, we refer to both simplified and adapted
content as simplified content, although adaptation can diverge from
strict simplification (which, by definition, should not alter the
information in the original text), via additional explanations or complex
concept removal, for instance.
• Lexical simplification : A specific set of models
and methods will be developed for multilingual
lexical simplification, based on statistical
methods, embeddings and pre-trained neural language
models such as BERT or XLM-R [33]. We explore
in particular the impact of cascading lexical
replacements, whole-word vs. subword masking
strategies, for agglutinative languages like Basque
in particular, and contextual coherence of lexical
substitution.
• Syntactic simplification : Under this designation,
we include all methods and models that
specifically tackle the transformation of complex
sentences into separate simple sentences. This
includes the extraction of complex modifiers, such
as relative or appositive clauses, and splitting
coordinated or juxtaposed sentences, for instance.
The project will focus on neural models for this
task, with a strong emphasis on the creation of
artificial datasets from which syntactic
simplification may be modelled.
• End-to-end simplification : Under this banner are
all neural text simplification models that perform
the task in and end-to-end fashion. This includes
the development and evaluation of models based
on control parameter tokens, synthetic corpora,
or pre-trained language models fine-tuned for
ATS tasks. Note that the data generated via the
ifrst two approaches above are also exploited for
synthetic data generation to train the end-to-end
ATS models.</p>
        <p>The IRAZ application has been designed to support
lfexible on-the-fly access to diferent type of ATS models,
in isolation or in combination via model ensembling. In
addition to models centred on specific aspects of text
simplification, it is worth noting that the project also
addresses automated text segmentation methods to support
easy-to-read text adaptation.
4.2.3. Post-editing
Even more so than in companion fields such as machine
translation, where high quality may be achieved from
existing parallel training resources, ATS output requires
post-editing by experts prior to its publication as
easyto-read content, for two main reasons. First, the overall
quality achieved by state-of-the-art models, in
particular for languages like Spanish or Basque, is still too low
for most results to be directly exploited in professional
settings. Secondly, as previously noted, easy-to-read
content may require further adaptation of the simplification
suggestions, be they at the lexical or grammatical level.</p>
        <p>The project takes these current technological
limitations into account, by providing a Web-based user
interface, where users can post-edit the suggested
simplifications generated with the underlying ATS models. Each
sentence in the original text is automatically split and its
automatic simplification presented in a separate editable
window, the two being presented side by side. Users can
freely post-edit the simplified text and download the
complete text once the overall editing process is completed.
The professionally post-edited data will then be collected
to re-train or fine-tune ATS models.</p>
      </sec>
      <sec id="sec-2-3">
        <title>4.3. Initial Results</title>
        <sec id="sec-2-3-1">
          <title>The first phase of the project, in 2022, centred on the</title>
          <p>design of the solution, based on collected user
requirements, an initial data collection phase, and preliminary
research on core ATS tools and methods, with specific
emphasis on the main use-cases of the project for the
Basque and Spanish languages.</p>
          <p>In its second phase, which debuted in 2023, the project
has entered the core research and development cycles
defined in the work plan, where ATS modelling and
prototype development is taking place, along with further
data collection and preparation.</p>
          <p>The initial results obtained to date are summarised
below:
• New corpora for Basque and Spanish, based on
data collected from the Irekia portal. The data
were collected from Web crawls, with automatic
pairing of complex and corresponding simplified
documents, text extraction and many-to-many
sentence alignment.14 We also prepared initial
datasets from project participants’ repositories,
with further data to be collected in the remaining
phases of the project.
• Initial models for lexical simplification and
comparative results for Basque and Spanish, based on
both LS-BERT and our own BERT-based variant
which includes additional features and candidate
substitution methods.
• Initial models for one-to-many sentence
transformation, following the BISECT approach adapted
to Basque and Spanish via high-quality machine
translation of the original English datasets and
data selection.
• Initial models for end-to-end neural text
simpliifcation, based on Transformer encoder-decoder
architectures and training on mixtures of
monolingual and simplified data.
• Initial automated text segmentation methods,
based on pretrained mBERT and predictions of
chunk completion via token masking.
• Initial development of the core components, API
and UI of the IRAZ solution.</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>Current initial results provide a solid basis for the remaining planned research and development activities of the project.</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Conclusions</title>
      <p>We have described the IRAZ project, whose main goal is
the development of a flexible solution to support
easyto-read content generation, by means of automated text
simplification. The project targets professional
easy-toread content creators, who lack technological support in
actual practice. The project aims to help optimise current
content generation processes in the field, thus enhancing
the creation of accessible content for the large segments
of the population in critical need of this type of content.
Research and development goals and activities have been
defined in close collaboration with the professional
endusers in the field who actively participate in the project.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>IRAZ is partially funded by the Basque Business
Development Agency, SPRI, under Grant Agreement
ZL2022/00788. We wish to thanks all participants of the
project for their contributions and insights. We also wish
to thank the anonymous SEPLN reviewer for their
comments and suggestions.</p>
      <p>14We aim to share the prepared datasets with the research
community in the near future.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <year>2021</year>
          , pp.
          <fpage>341</fpage>
          -
          <lpage>352</lpage>
          . [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Štajner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. C.</given-names>
            <surname>Sheang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          , Sentence sim-
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Intelligence</surname>
          </string-name>
          , volume
          <volume>36</volume>
          ,
          <year>2022</year>
          , pp.
          <fpage>12172</fpage>
          -
          <lpage>12180</lpage>
          . [23]
          <string-name>
            <given-names>K.</given-names>
            <surname>Omelianchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Raheja</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Skurzhanskyi</surname>
          </string-name>
          , Text
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>16th Workshop on Innovative Use of NLP for Build-</mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>ing Educational Applications</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>25</lpage>
          . [24]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bernhard</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          , A monolingual
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          cation,
          <source>in: Proceedings of the 23rd International</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <year>2010</year>
          ),
          <year>2010</year>
          , pp.
          <fpage>1353</fpage>
          -
          <lpage>1361</lpage>
          . [25]
          <string-name>
            <given-names>W.</given-names>
            <surname>Coster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kauchak</surname>
          </string-name>
          , Simple english wikipedia:
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>the 49th Annual Meeting of the Association for</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>nologies</surname>
          </string-name>
          ,
          <year>2011</year>
          , pp.
          <fpage>665</fpage>
          -
          <lpage>669</lpage>
          . [26]
          <string-name>
            <given-names>W.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Callison-Burch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoles</surname>
          </string-name>
          , Problems in
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>tional Linguistics</source>
          <volume>3</volume>
          (
          <year>2015</year>
          )
          <fpage>283</fpage>
          -
          <lpage>297</lpage>
          . [27]
          <string-name>
            <given-names>I.</given-names>
            <surname>Gonzalez-Dios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Aranzabe</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Díaz de Ilar-
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>Language Resources and Evaluation</source>
          <volume>52</volume>
          (
          <year>2018</year>
          )
          <fpage>217</fpage>
          -
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          247. [28]
          <string-name>
            <given-names>I.</given-names>
            <surname>Gonzalez-Dios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Gutiérrez-Fandiño</surname>
          </string-name>
          ,
          <string-name>
            <surname>O. M.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Readability (TSAR-2022)</surname>
          </string-name>
          ,
          <year>2022</year>
          , pp.
          <fpage>86</fpage>
          -
          <lpage>97</lpage>
          . [29]
          <string-name>
            <given-names>S.</given-names>
            <surname>Surya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Laha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Jain</surname>
          </string-name>
          , K. Sankara-
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>in: Proceedings of the 57th Annual Meeting of the</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Association for Computational Linguistics</surname>
          </string-name>
          ,
          <year>2019</year>
          , pp.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          2058-
          <fpage>2068</fpage>
          . [30]
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Qiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , An unsuper-
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>2021, Association for Computational Linguistics,</mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Punta</given-names>
            <surname>Cana</surname>
          </string-name>
          , Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>227</fpage>
          -
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          237. [31]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Maddela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kriz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xu</surname>
          </string-name>
          , C. Callison-
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <article-title>tences with bitexts</article-title>
          ,
          <source>in: Proceedings of the 2021</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>guage Processing</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>6193</fpage>
          -
          <lpage>6209</lpage>
          . [32]
          <string-name>
            <given-names>S.</given-names>
            <surname>Stajner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Franco-Salvador</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Ponzetto</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <source>in: Proceedings of the 55th Annual Meeting of the</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          2017, Vancouver, Canada,
          <source>July 30 - August 4</source>
          , Vol-
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          ume 2:
          <string-name>
            <surname>Short</surname>
            <given-names>Papers</given-names>
          </string-name>
          ,
          <year>2017</year>
          , pp.
          <fpage>97</fpage>
          -
          <lpage>102</lpage>
          . [33]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          , V. Chaud-
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <article-title>ceedings of the 58th Annual Meeting of the Asso-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>for Computational</surname>
            <given-names>Linguistics</given-names>
          </string-name>
          , Online,
          <year>2020</year>
          , pp.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>