<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Automatic Text Classification using Readability Levels in Galician and Spanish</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sandra Rodríguez Rey</string-name>
          <email>sandrarodriguez.rey@usc.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centro Singular de Investigación en Tecnoloxías Intelixentes (CITIUS), Universidade de Santiago de Compostela</institution>
          ,
          <addr-line>15782</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Santiago de Compostela</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>A text can be more or less complex to read depending on various parameters, such as sentence length, vocabulary variety, or the presence or absence of certain syntactic structures and linguistic phenomena. This concept of the level of reading complexity is called readability. This research aims to study automatic text classification in Galician and Spanish, focusing on the creation of corpora and the development of an automatic text classifier based on readability levels for Galician. First, the state of the art in text readability and automatic text classification is reviewed. Then, corpora of Galician and Spanish texts are compiled and classified by readability levels, and linguistic phenomena of diferent complexity levels are studied. Finally, diferent computational strategies for automatic text classification are evaluated.</p>
      </abstract>
      <kwd-group>
        <kwd>Text classification</kwd>
        <kwd>text complexity</kwd>
        <kwd>readability</kwd>
        <kwd>Automatic Readability Assessment</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Readability has long been studied, from the 19th century until now [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Knowing the degree
of reading complexity of a text is important in several domains, such as language learning or
automatic readability assessment [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Readability formulas have been the traditional method of
calculating reading complexity for decades. These formulas extract metrics such as word and
sentence length [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Nowadays, the most common methods to measure the readability level
of a text are based on linguistic features or on the use of deep learning models [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The latter
models are trained on corpora labeled based on readability levels.
      </p>
      <p>
        Automatic text classifiers based on complexity levels are useful tools for determining the
readability level of a text. FABRA [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is a good example of such a tool. To train automatic
classifiers, high-quality corpora labeled by complexity levels are needed. Some corpora are
available for Spanish, such as Coh-Metrix-Esp [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Newsela [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], CAES [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], or Simplext [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Nevertheless, these corpora contain texts from only three domains (literature, journalism, and
education) and are designed for text simplification or language teaching as a foreign language.
To the best of our knowledge, there are no corpora or automatic classifiers available for Galician.
      </p>
      <p>The main objective of this PhD project is the investigation of automatic methods for the
assessment of the readability of texts for adults in Galician and Spanish. To achieve this, three
specific objectives have been defined. First, to obtain Galician and Spanish corpora for testing
automatic text classifiers. Second, to test the performance of available automatic text classifiers
by readability levels for Galician and Spanish. And thirdly, to explore data augmentation and
transfer learning strategies to develop new models for Galician.</p>
      <p>This research project aims to contribute to the field by answering three research questions
(RQs). RQ1: Are reading complexity descriptors similar for Galician, Spanish and other
languages? The hypothesis for this research question is that they are similar, but that each language
has specific linguistic phenomena that afect text complexity. RQ2: Is it possible to adapt reliable
text classifiers designed for other languages with minor modifications? The hypothesis is that
it depends on the size of the corpus and the linguistic resources available, so the situation is
diferent for Galician and Spanish. RQ3: Since Galician is considered a language with few
resources, is it possible to use cross-linguistic strategies? The hypothesis is that it is possible by
using multilingual models and adapting resources from other languages to Galician.</p>
      <p>This paper gives a description of the research project and some results. First, some background
on the topic is given and some related work is highlighted. Then, a description including research
questions, hypotheses and objectives is provided, followed by the methods and techniques used.
Finally, a short introduction to the research work carried out during part of the research, and
some discussion questions are presented. Regarding the first results of this research project,
two corpora for automatic text classifier testing are presented: a Spanish and a Galician corpus
of texts for adult population labeled by levels of complexity.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Background and related work</title>
      <p>
        Readability is the ease or dificulty with which a text can be read and understood [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The first
studies on the characteristics that define the level of complexity of a text to be understood date
back to the 19th century and took place in the United States and Russia. These studies refer to
the lexicon and sentence length [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Vajjala [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] refers to Thorndike, Lively and Pressey, and
Vogel and Wash-burne as the first studies on how to measure the level of dificulty of reading a
text in the 1920s decade.
      </p>
      <p>
        The traditional calculation method is the application of readability formulas. These formulas
measure some superficial characteristics of the text [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], such as sentence length, number of
syllables in words or readers’ knowledge or unfamiliarity with the lexicon and extract a value
indicating the reading complexity of the text [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. An example is the Flesh-Lincaid metric (1975),
which calculates mainly the number of words per sentence and the number of syllables per
word [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Following Campos [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Du Bay states that other well-known formulas are the Flesch’s
Reading Ease Score (RES), the Dale-Chall, the SMOG, the Gunning FOG test or the Fry’s Chart.
      </p>
      <p>
        More recently, various methods and techniques have been used to automatically evaluate
the readability of a text. The most common ones are based on linguistic features or deep
learning models. The former focuses on word and sentence length, syntactic complexity, or the
percentage of occurrence of words included in various types of word lists [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. There are two
types of linguistic feature models that can be highlighted: from the statistical computation of
features in the text, or from trained machine learning models. Neural networks and classification
algorithms are some of the deep learning methods used [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. These models are trained on corpora
that are labeled based on readability levels.
      </p>
      <p>
        In the last 10 years, however, significant progress has been made using sophisticated NLP
techniques, such as automatic parsing and statistical language modeling [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This makes it
possible to evaluate a wide list of factors that afect the readability of a text. An example is
FABRA, a tool developed for French, which measures several factors, such as word and sentence
length, lexical diversity, orthographic neighborhood, lexical frequency, syntactic dependency,
syntactic coherence, the use of anaphoric elements, etc. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. However, to the best of our
knowledge, there are no resources or tools available for our target languages that aim to classify
texts of diferent genres according to readability levels.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Justification</title>
      <p>
        Information about the readability level of texts is useful in diverse fields: language learning,
automatic readability assessment, content creation, accessibility, etc. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Thus, tasks such
as appropriate reading materials selection or text adaptation to make texts clearer and more
accessible can be facilitated.
      </p>
      <p>To obtain text readability information, automatic text classifiers are an optimal tool. To
develop automatic text classifiers, high-quality corpora are needed. Concerning automatic text
classification models, we are not aware of the existence of this type of tool for Galician.</p>
      <p>
        Several corpora on text complexity already exist for various languages: Weekly Reader [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ],
WeeBit [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], (CLEAR) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]) or OneStopEnglish [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] for English; FLE-CORP [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], FLM-CORP [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
or FSW [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] for French; READ-IT [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or CELI [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] for Italian; a dataset from Instituto Camões
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] for Portuguese; Slovenian SB for Slovenian [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]; LBSPC [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] for Basque, or VikiWiki [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]
for Basque and Catalan.
      </p>
      <p>
        For Galician, we are not aware of any corpus of this type that is available. For Spanish, some
existing corpora are Coh-Metrix-Esp [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Newsela [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], CAES [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Simplext [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and kwiziq and
HablaCultura corpora [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Although more Spanish corpora exist, they are unavailable or they
have data protection licenses that prevent their use [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Spanish available resources include a
limited variety of textual genres and domains: literature, journalism, and teaching. Moreover,
some of them are designed for text simplification (Newsela or Simplext), and some others for
language teaching as a foreign language (CAES, kwiziq, or HablaCultura).
      </p>
      <p>Therefore, this research work aims to contribute to the field by creating new corpora for
Galician and Spanish and by exploring reliable automatic text classifiers for Galician texts.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Research description</title>
      <p>
        Text readability, understood as a predictor of text complexity for comprehension, afects reading
performance [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Automatic readability assessment tools make it possible to know the complexity
level a text represents for a reader and even certain characteristics that make texts more complex
to understand. This information can be valuable for diferent purposes, such as selecting and
simplifying texts for foreign language learners. Despite the number of Spanish speakers, the
development of this type of tool for Spanish is scarce [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. To the best of our knowledge,
Galician, which is considered a low-resource language, does not have a linguistic tool with this
functionality available.
      </p>
      <p>The main objective of this thesis is to study automatic methods to evaluate the readability of
texts for the adult population in Galician and Spanish. To achieve it, three specific objectives
have been defined:
• To obtain Galician and Spanish corpora that represent the wide variety of text genres and
topics an adult reader might encounter in his or her lifetime classified using readability
levels for testing automatic text classifiers.
• To evaluate the performance of several automatic text classifiers by readability levels for</p>
      <p>Galician and Spanish.
• To study data augmentation strategies by adapting existing resources from other languages
to Galician and transfer learning methods using multilingual models to develop reliable
text classifiers for Galician.</p>
      <p>The following research questions (RQs) and corresponding hypotheses (Hs) are formulated
in this thesis:
• RQ1: Are readability complexity descriptors similar for Galician, Spanish and related
languages?
H1: Complexity descriptors of text readability are similar for Galician and Spanish, and
also for related languages such as Portuguese or French. However, each language has
specific linguistic phenomena that afect the complexity of the text and must be taken
into account when determining the readability complexity descriptors. For Galician,
some examples can be the use of the dative of interest or the solidarity pronouns, some
periphrases (such as the construction “dar” + participle) or the inflected infinitive (although
this last one also exists in Portuguese).
• RQ2: Is it possible to quickly adapt reliable text classifiers designed for other languages
to Galician and Spanish?
H2: This depends on the size of the corpus and the linguistic resources available. For
Spanish, it may be possible, since there are resources available for measuring text
complexity (for example, word lists of concrete and abstract nouns, age of acquisition or
graded lexicons) and a corpus has already been created. For Galician it may be possible,
but the classifier will probably have a lower performance because the computational
resources for Galician are limited.
• RQ3: Since Galician is considered a language with few resources, is it possible to use
cross-linguistic strategies to obtain new data and develop an automatic text classifier?
H3: This is possible by adapting resources from other languages (e.g., Spanish, Portuguese
and French) to Galician using techniques such as transliteration and machine translation
and by using multilingual models and transfer learning methods.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Methodology</title>
      <p>To meet the aforementioned objectives, the following methodology is proposed.</p>
      <p>First, a review of the state of the art in text readability and automatic text classification
according to text complexity levels is carried out, including general parameters for diferent
languages and specific linguistic phenomena for Galician and Spanish. Second, a corpus of about
2000 Spanish texts classified by complexity levels will be created. Thirdly, a similar corpus of
about 400 texts in Galician will be compiled. Then, available Transformer models for Galician and
Spanish (both monolingual and multilingual) will be evaluated on text classification. Depending
on the performance, additional data for Galician may be created by exploring various data
augmentation methods. Subsequently, the evaluation of the resulting automatic text classifiers
for Galician will be performed.</p>
      <p>
        The main research techniques will be documentary and experimental. For the experimental
techniques, both quantitative and qualitative results will be obtained. Regarding the development
of linguistic tools for Galician texts, since Galician is considered a language with few resources,
the performance of classifiers designed for related languages will be studied [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Spanish,
Portuguese and French tools will be the main options to study, as these are the languages
involved in the project the author is engaged in and similar tools are being developed. Both
quantitative and qualitative analyses will be carried out using experimental techniques.
      </p>
      <p>
        After obtaining these results, we can explore methods such as synthetic data augmentation
using generative models, machine translation, or other transformations from related languages
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] if needed. Documentary techniques and questionnaires addressed to language experts and
students will be used to determine the characteristics that influence text readability in Spanish
and Galician. On the one hand, literature on readability and similar concepts, such as reading
comprehension and text simplification for easy reading will be reviewed. On the other hand,
specific linguistic aspects afecting readability in Galician, Spanish, Portuguese and French
will be studied. Since, to the best of our knowledge, there are no studies on Galician linguistic
phenomena that afect readability, specific aspects of Galician-related languages that may also
occur in Galician will be analyzed.
      </p>
      <p>Diferent types of automatic text classification models will be tested, from classical machine
learning models like Decision Stress and SVM, as well as the fine-tuning of language models
(such as BERT) for classification tasks using open-access libraries, such as Transformers by
HuggingFace. Both monolingual and multilingual models will be explored.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Results</title>
      <p>The research progress made includes the creation of two corpora based on readability levels.
Both resources include texts from a wide variety of genres and three levels of complexity defined
by experts.</p>
      <p>The first corpus is being developed within the iRead4Skills project (Intelligent Reading
Improvement System for Fundamental and Transversal Skills Development), an ongoing European
project involving researchers from diferent institutions, such as the NOVA University of Lisbon,
the INESC-ID Research Center (Lisbon), the Catholic University of Louvain (Belgium), the
University of Santiago de Compostela, the UAB (Universitat Autònoma de Barcelona) and the
Luxembourg Institute of Socio-Economic Research. This project aims to improve the reading
skills of the adult population with low literacy levels by creating an intelligent system that
analyzes text complexity and provides appropriate reading materials, thus facilitating their
access to information and culture1.</p>
      <p>
        This Spanish corpus, to which the author of this article is one of the two main contributors,
is part of a multilingual dataset that includes three corpora with similar characteristics in three
languages: Spanish, Portuguese and French [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. This corpus contains 2563 Spanish texts
classified into three levels of complexity: level 1 (very easy), level 2 (easy), and level 3 (plain).
A fourth level of more complex texts is also included, although it is not a consolidated level.
This fourth level is intended to represent the type of text in terms of complexity that should not
be considered level 3. The texts are also classified by text domains, genres, and subgenres, as
shown in Table 1. These categories are intended to represent the most common textual and
thematic genres that an adult reader might encounter, focusing on the types of texts that are of
most interest to an adult reader.
      </p>
      <p>The second resource to be presented is a Galician corpus to which the doctoral candidate is
the main contributor. Its design was based on the aforementioned multilingual dataset. This
corpus contains 424 texts classified into three levels of complexity and 11 domains. A fourth
level will also be included, but the work is still in progress. Although this corpus is inspired by
the multilingual dataset, some categories and subcategories may vary. This is due to the fact
that in certain genres and subjects no texts written in Galician have been found. In addition, the
corpus is smaller in size. The levels of dificulty of this corpus have been defined on the basis of
the corresponding levels of dificulty established for Spanish, Portuguese and French. However,
an adaptation was necessary to consider the specific linguistic aspects of Galician that afect the
readability of the texts. This adaptation was done by taking into account the Celga2 and CEFR3
classifications for Galician. This corpus has not been published yet, but will be available soon.
1https://iread4skills.com/
2https://www.lingua.gal/o-galego/aprendelo/celga
3https://www.lingua.gal/c/document_library/get_file?folderId=1647060&amp;name=DLFE-8921.pdf</p>
    </sec>
    <sec id="sec-8">
      <title>7. Discussion</title>
      <p>As research progresses, new questions arise about readability and complexity levels designed
for text simplification or educational purposes, general or language-specific readability features,
automated data transformation or validation of the classification. Some questions to be addressed
in the future may include the following:
• Texts adapted for educational purposes and classified following the CEFR are commonly
used to train and test automatic text classifiers based on complexity levels. Is there a
correlation between readability levels and CEFR levels?
• Some linguistic phenomena, such as specific verb tenses or syntactic structures, afect the
reading comprehension of a text. Focusing on the Galician case, which Galician-specific
linguistic phenomena afect readability?
• In a readability scenario, what is the best data augmentation method for Galician?
• Considering that we are dealing with readability, is it possible to obtain high-quality
resources by automatically transforming Spanish, Portuguese or French data into Galician?
• Regarding the classification of texts by readability levels, how can these classifications be
validated? Is it possible to use generative systems to classify texts? If we use more than
one annotator to validate the classification, how can we interpret the agreement between
them?</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>This work has received financial support from the European Commission (Project: iRead4Skills,
Grant number: 1010094837, Topic: HORIZON-CL2-2022-TRANSFORMATIONS-01-07, DOI:
10.3030/101094837), and from Xunta de Galicia - Consellería de Cultura, Educación, Formación
Profesional e Universidades (Centro de investigación de Galicia accreditation 2024-2027
ED431G2023/04), and the European Union (European Regional Development Fund - ERDF).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Campos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Contreras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Rifo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Véliz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reyes</surname>
          </string-name>
          ,
          <article-title>Complejidad textual, lecturabilidad y rendimiento lector en una prueba de comprensión en escolares adolescentes</article-title>
          ,
          <source>Universitas Psychologica</source>
          <volume>13</volume>
          (
          <year>2014</year>
          )
          <fpage>1135</fpage>
          -
          <lpage>1146</lpage>
          . URL: https://doi.org/10.11144/Javeriana.UPSY13-3.ctlr. doi:
          <volume>10</volume>
          .11144/Javeriana.UPSY13- 3.ctlr.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Vajjala</surname>
          </string-name>
          ,
          <article-title>Trends, limitations and open challenges in automatic readability assessment research</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Béchet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Blache</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Cieri</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Goggi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Isahara</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Maegaard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mariani</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Mazo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Odijk</surname>
          </string-name>
          , S. Piperidis (Eds.),
          <source>Proceedings of the Thirteenth Language Resources and Evaluation Conference</source>
          , European Language Resources Association, Marseille, France,
          <year>2022</year>
          , pp.
          <fpage>5366</fpage>
          -
          <lpage>5377</lpage>
          . URL: https://aclanthology. org/
          <year>2022</year>
          .lrec-
          <volume>1</volume>
          .
          <fpage>574</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Montemagni</surname>
          </string-name>
          , G. Venturi,
          <article-title>Assessing document and sentence readability in less resourced languages and across textual genres</article-title>
          , ITL - International
          <source>Journal of Applied Linguistics</source>
          <volume>165</volume>
          (
          <year>2014</year>
          )
          <fpage>163</fpage>
          -
          <lpage>193</lpage>
          . doi:
          <volume>10</volume>
          .1075/itl.165.2.03del.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Wilkens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Alfter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pintard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. P.</given-names>
            <surname>Yancey</surname>
          </string-name>
          , T. François, FABRA:
          <article-title>French aggregator-based readability assessment toolkit</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Béchet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Blache</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Cieri</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Goggi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Isahara</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Maegaard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mariani</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Mazo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Odijk</surname>
          </string-name>
          , S. Piperidis (Eds.),
          <source>Proceedings of the Thirteenth Language Resources and Evaluation Conference</source>
          , European Language Resources Association, Marseille, France,
          <year>2022</year>
          , pp.
          <fpage>1217</fpage>
          -
          <lpage>1233</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .lrec-
          <volume>1</volume>
          .
          <fpage>130</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Quispesaravia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Perez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Sobrevilla</given-names>
            <surname>Cabezudo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Alva-Manchego</surname>
          </string-name>
          ,
          <article-title>Coh-Metrix-Esp: A complexity analysis tool for documents written in Spanish</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Goggi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Grobelnik</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Maegaard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mariani</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Mazo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Odijk</surname>
          </string-name>
          , S. Piperidis (Eds.),
          <source>Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)</source>
          ,
          <source>European Language Resources Association (ELRA)</source>
          , Portorož, Slovenia,
          <year>2016</year>
          , pp.
          <fpage>4694</fpage>
          -
          <lpage>4698</lpage>
          . URL: https://aclanthology.org/L16-1745.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>W.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Callison-Burch</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Napoles, Problems in current text simplification research: New data can help, Transactions of the Association for Computational Linguistics 3 (</article-title>
          <year>2015</year>
          )
          <fpage>283</fpage>
          -
          <lpage>297</lpage>
          . URL: https://aclanthology.org/Q15-1021. doi:
          <volume>10</volume>
          .1162/tacl_a_
          <fpage>00139</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G.</given-names>
            <surname>Parodi</surname>
          </string-name>
          , Corpus de aprendices de español (caes),
          <source>Journal of Spanish Language Teaching</source>
          <volume>2</volume>
          (
          <year>2015</year>
          )
          <fpage>194</fpage>
          -
          <lpage>200</lpage>
          . doi:
          <volume>10</volume>
          .1080/23247797.
          <year>2015</year>
          .
          <volume>1084685</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Gómez-Martínez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Etayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Anula</surname>
          </string-name>
          , L. Bourg,
          <article-title>Text simplification in simplext: Making texts more accessible</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>47</volume>
          (
          <year>2011</year>
          )
          <fpage>341</fpage>
          -
          <lpage>342</lpage>
          . URL: https://www.researchgate.net/publication/277193726_Text_Simplification_ in_Simplext_Making_Text_More_Accessible.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Schwarm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ostendorf</surname>
          </string-name>
          ,
          <article-title>Reading level assessment using support vector machines and statistical language models</article-title>
          , in: K. Knight,
          <string-name>
            <given-names>H. T.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          Oflazer (Eds.),
          <article-title>Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL'05), Association for Computational Linguistics</article-title>
          , Ann Arbor, Michigan,
          <year>2005</year>
          , pp.
          <fpage>523</fpage>
          -
          <lpage>530</lpage>
          . URL: https://aclanthology.org/P05-1065. doi:
          <volume>10</volume>
          .3115/1219840.1219905.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Vajjala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Meurers</surname>
          </string-name>
          ,
          <article-title>On improving the accuracy of readability classification using insights from second language acquisition</article-title>
          ,
          <source>in: Proceedings of the Seventh Workshop on Building Educational Applications Using NLP</source>
          ,
          <string-name>
            <surname>NAACL</surname>
            <given-names>HLT</given-names>
          </string-name>
          '
          <volume>12</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, USA,
          <year>2012</year>
          , p.
          <fpage>163</fpage>
          -
          <lpage>173</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Heintz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Batchelor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Karimi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Malatinszky</surname>
          </string-name>
          ,
          <article-title>A large-scaled corpus for assessing text readability</article-title>
          ,
          <source>Behavior Research Methods</source>
          <volume>55</volume>
          (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .3758/ s13428- 022- 01802- x.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Vajjala</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Lučić</surname>
          </string-name>
          ,
          <article-title>OneStopEnglish corpus: A new corpus for automatic readability assessment and text simplification</article-title>
          , in: J.
          <string-name>
            <surname>Tetreault</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Burstein</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Kochmar</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Leacock</surname>
          </string-name>
          , H. Yannakoudakis (Eds.),
          <source>Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications</source>
          , Association for Computational Linguistics, New Orleans, Louisiana,
          <year>2018</year>
          , pp.
          <fpage>297</fpage>
          -
          <lpage>304</lpage>
          . URL: https://aclanthology.org/W18-0535. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W18</fpage>
          - 0535.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Ngo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Parmentier</surname>
          </string-name>
          ,
          <article-title>Towards sentence-level text readability assessment for French</article-title>
          , in: S. Štajner,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shardlow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Alva-Manchego</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the Second Workshop on Text Simplification, Accessibility and Readability</source>
          , INCOMA Ltd.,
          <string-name>
            <surname>Shoumen</surname>
          </string-name>
          , Bulgaria, Varna, Bulgaria,
          <year>2023</year>
          , pp.
          <fpage>78</fpage>
          -
          <lpage>84</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .tsar-
          <volume>1</volume>
          .8.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Grego Bolli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <article-title>Predicting readability of texts for italian l2 students: A preliminary study</article-title>
          ,
          <source>in: ALTE</source>
          (
          <year>2017</year>
          ).
          <source>Learning and Assessment: Making the Connections - Proceedings of the ALTE 6th International Conference, 3-5 May</source>
          <year>2017</year>
          , Association of Language Testers in Europe,
          <year>2017</year>
          , pp.
          <fpage>272</fpage>
          -
          <lpage>278</lpage>
          . URL: https://www.researchgate.net/publication/320532498_Predicting_Readability_of_ Texts_for_Italian_L2_Students_
          <string-name>
            <surname>A</surname>
          </string-name>
          _Preliminary_Study.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Mamede</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Baptista</surname>
          </string-name>
          ,
          <article-title>Automatic text readability assessment in European Portuguese</article-title>
          , in: P.
          <string-name>
            <surname>Gamallo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Claro</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Teixeira</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Real</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>H. G.</given-names>
          </string-name>
          <string-name>
            <surname>Oliveira</surname>
          </string-name>
          , R. Amaro (Eds.),
          <source>Proceedings of the 16th International Conference on Computational Processing of Portuguese, Association for Computational Lingustics</source>
          , Santiago de Compostela, Galicia/Spain,
          <year>2024</year>
          , pp.
          <fpage>97</fpage>
          -
          <lpage>107</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .propor-
          <volume>1</volume>
          .
          <fpage>10</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Martinc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pollak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Robnik-Šikonja</surname>
          </string-name>
          ,
          <article-title>Supervised and unsupervised neural approaches to text readability</article-title>
          ,
          <source>Computational Linguistics</source>
          <volume>47</volume>
          (
          <year>2021</year>
          )
          <fpage>141</fpage>
          -
          <lpage>179</lpage>
          . URL: https://aclanthology. org/
          <year>2021</year>
          .cl-
          <volume>1</volume>
          .6. doi:
          <volume>10</volume>
          .1162/coli_a_
          <fpage>00398</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>I.</given-names>
            <surname>Gonzalez-Dios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Aranzabe</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Díaz de Ilarraza, H. Salaberri,
          <article-title>Simple or complex? assessing the readability of Basque texts</article-title>
          , in: J.
          <string-name>
            <surname>Tsujii</surname>
          </string-name>
          , J. Hajic (Eds.),
          <source>Proceedings of COLING</source>
          <year>2014</year>
          ,
          <source>the 25th International Conference on Computational Linguistics: Technical Papers</source>
          , Dublin City University and Association for Computational Linguistics, Dublin, Ireland,
          <year>2014</year>
          , pp.
          <fpage>334</fpage>
          -
          <lpage>344</lpage>
          . URL: https://aclanthology.org/C14-1033.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>I. Madrazo</given-names>
            <surname>Azpiazu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Pera</surname>
          </string-name>
          ,
          <article-title>Is cross-lingual readability assessment possible?</article-title>
          ,
          <source>Journal of the Association for Information Science and Technology</source>
          <volume>71</volume>
          (
          <year>2020</year>
          )
          <fpage>644</fpage>
          -
          <lpage>656</lpage>
          . URL: https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/asi.24293. doi:https://doi.org/ 10.1002/asi.24293.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>L.</given-names>
            <surname>Vásquez-Rodríguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.-M.</given-names>
            <surname>Cuenca-Jiménez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Morales-Esquivel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Alva-Manchego</surname>
          </string-name>
          ,
          <article-title>A benchmark for neural readability assessment of texts in Spanish</article-title>
          , in: S. Štajner,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ferrés</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shardlow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. C.</given-names>
            <surname>Sheang</surname>
          </string-name>
          , K. North,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          , W. Xu (Eds.),
          <source>Proceedings of the Workshop on Text Simplification</source>
          , Accessibility, and
          <string-name>
            <surname>Readability</surname>
          </string-name>
          (TSAR-
          <year>2022</year>
          ),
          <article-title>Association for Computational Linguistics</article-title>
          , Abu Dhabi,
          <source>United Arab Emirates (Virtual)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>188</fpage>
          -
          <lpage>198</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .tsar-
          <volume>1</volume>
          .18. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          .tsar-
          <volume>1</volume>
          .
          <fpage>18</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bhargava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Vijayan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Anand</surname>
          </string-name>
          , G. Raina,
          <article-title>Exploration of transfer learning capability of multilingual models for text classification</article-title>
          ,
          <source>in: Proceedings of the 2023 5th International Conference on Pattern Recognition and Intelligent Systems</source>
          , PRIS '23,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2023</year>
          , p.
          <fpage>45</fpage>
          -
          <lpage>50</lpage>
          . doi:
          <volume>10</volume>
          .1145/3609703.3609711.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rehan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S. I.</given-names>
            <surname>Malik</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Jamjoom</surname>
          </string-name>
          ,
          <article-title>Fine-tuning transformer models using transfer learning for multilingual threatening text identification</article-title>
          ,
          <source>IEEE Access 11</source>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          . 1109/ACCESS.
          <year>2023</year>
          .
          <volume>3320062</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Pintard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>François</surname>
          </string-name>
          , J. Nagant de Deuxchaisnes, S. Barbosa,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Reis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Moutinho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Monteiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Amaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Correia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Rodríguez</given-names>
            <surname>Rey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Garcia</given-names>
            <surname>González</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Mu</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          <article-title>Blanco Escoda, iread4skills dataset 1: corpora by complexity level for fr, pt</article-title>
          and sp,
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.10889888.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>