<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ODESIA: Space for Observing the Development of Spanish in Artificial Intelligence</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Enrique Amigó</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Carrillo-de-Albornoz</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrés Fernández</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julio Gonzalo</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miguel Lucas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guillermo Marco</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roser Morante</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jacobo Pedrosa</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laura Plaza</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ibo Sanz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Beñat San-Sebastián</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Llorente y Cuenca Madrid</institution>
          ,
          <addr-line>S.L., Calle Lagasca 88, 28001 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>RMIT University</institution>
          ,
          <addr-line>124 La Trobe St, Melbourne VIC 3000</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universidad Nacional de Educación a Distancia, Calle Juan del Rosal</institution>
          ,
          <addr-line>16, 28040 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we present the ODESIA project (Space for Observing the Development of Spanish in Artificial Intelligence). ODESIA is a research collaborative project between the UNED University and the public institution RED.es, financed by the UE through the NextGenerationEU funds. The main objective of the project is the development of an annual index that quantitatively and qualitatively measures the gap between the language technologies in Spanish and English in terms of the state of the art, market solutions, the level of adoption of technologies and user experience. Other goals of the project are the production of Natural Language Processing resources in Spanish, such as datasets, a benchmark for evaluating the performance of language models, a website that provides information related to the progress of NLP in Spanish, and a tool for facilitating the evaluation of NLP systems.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;ODESIA</kwd>
        <kwd>NLP gap</kwd>
        <kwd>market solutions</kwd>
        <kwd>adoption level</kwd>
        <kwd>benchmarks</kwd>
        <kwd>evaluation benchmark</kwd>
        <kwd>NLP portal</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction and Motivation</title>
      <sec id="sec-1-1">
        <title>National Artificial Intelligence Strategy (ENIA) has iden</title>
        <p>tified seven objectives, one of which is to position the
ODESIA (Espacio de Observación del Desarrollo del Es- Spanish language as a leader in the development of tools,
pañol en la Inteligencia Artificial) is a collaborative re- technologies, and applications for the use of AI in various
search project between UNED University and the public ifelds, both in Spain and globally.
institution RED.es. It is funded by the European Union In this context, it is essential to establish an observation
through the NextGenerationEU funds. The project con- space that can analyze and measure the distance between
sortium includes several partners: the National Obser- the level of development of the AI in Spanish and the
vatory of Technology and Society (ONTSI),1 the NLP &amp; level of development of the AI in English, which is the
IR UNED research group at UNED,2 and the company dominant digital language. Specifically, the project will
Llorente &amp; Cuenca.3 focus on Natural Language Processing (NLP). ODESIA</p>
        <p>Currently, Artificial Intelligence (AI) plays a signifi- will provide the ENIA with the necessary infrastructure
cant role in all economic sectors and social activities, to measure the gap in NLP and monitor the evolution of
particularly in English-speaking countries. The Spanish the gap throughout the execution of the plan. This project
language is crucial in making Spain a significant player can be a vital contribution to help Spanish become one
in AI. In turn, AI is essential for defending the competi- of the leading languages in the deployment of AI and can
tive advantage that Spain has due to its language. The contribute to develop a more diverse, plural, and ethical
AI.
and English. This index will consider the state of the
art, market solutions, adoption of technologies, and user
experience as its key factors. The creation and periodic
update of this index should contribute to raise
awareness among major decision-makers and citizens about
the relevance of promoting the development of language
technologies in Spanish, which is crucial for Spain to
become a prominent player in the AI landscape. Figure 1
summarizes the main goals of the ODESIA project.</p>
        <p>ODESIA also pursuits the following specific objectives:</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Results</title>
      <sec id="sec-2-1">
        <title>3.1. Language Gap Index</title>
        <p>The main output of ODESIA is an index that measures
the gap in the development and adoption of language
technologies between English and Spanish. This index
is updated annually, providing a dynamic assessment of
the progress made in language technologies in Spanish
compared to English. The index is composed of four main
areas: (1) state of the art, (2) availability and functionality
of NLP market solutions, (3) adoption of NLP market
solutions, and (4) user experience and satisfaction. By
computing the gap in these areas, it will be possible to
quantify the diferences between the two languages and
monitor progress over time.
3.1.1. Gap in the State of the Art
The goal is to evaluate the state of language technologies
in Spanish and English by examining the dissemination
of research results in scientific media, the availability of
resources for NLP research and their efectiveness.
1. Developing a leaderboard that allows for compar- To assess the gap in the state of the art, three indexes
ison between diferent pre-trained language mod- are used: The dissemination index measures the gap in
els in Spanish and English for various NLP tasks. the publication of scientific works in top NLP conferences
The information provided by this leaderboard can as well as in funded NLP projects in both languages. The
be utilized by policymakers, researchers, and the resource index measures the diference in the
availabilindustry to monitor and comprehend the progress ity of pre-trained language models, annotated data, and
of language technologies in Spanish. language processing tools in English and Spanish. The
2. Developing bilingual datasets and evaluation efectiveness index quantifies the gap in performance
methodologies to compare pre-trained language between models and systems in both languages,
commodels in Spanish and English. This will also pared to baseline systems that do not use any linguistic
enable the comparison of NLP systems for tasks knowledge.</p>
        <p>with the top practical applications.
3. Creating a website that contains information 3.1.2. Gap in the Availability and Functionality of
about and allows to track the progress of the state NLP market Solutions
of the art in NLP for the Spanish language.
4. Developing a web application for evaluation that
allows to perform a comprehensive and
comparable evaluation of NLP systems for diferent tasks
and evaluation contexts.</p>
        <sec id="sec-2-1-1">
          <title>The gap is calculated by identifying strategic families of</title>
          <p>NLP applications (e.g., chat bots, reputation/sentiment
analysis) and selecting NLP products and services within
such applications, both in English and Spanish. The map
of functionalities associated with each family of
applications and the coverage of existing products in both
languages will be studied. The functionality index
quantifies the language gap in terms of the coverage of
the diferent functionalities ofered by the products found
in each application family.
3.1.3. Gap in the Adoption of NLP market</p>
          <p>Solutions</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>The objective is to quantify the level of adoption and value ofered by language technologies in the market, providing added value to both companies and end users.</title>
          <p>Two diferent indexes are calculated: The adoption in- 3.3.1. EXIST 2022
dex estimate the level of implementation/adoption of
language technologies in Spanish/English by companies
and by citizens, through the analysis of mentions in
corporate reports and (social) media. The impact index
estimates the reduction of costs or the increase in income
for companies that are related to the implementation of
NLP systems in any of their processes.</p>
          <p>
            The EXIST (EXism Identification in Social Networks) 2022
dataset aims to facilitate research on automatic detection
of sexism, containing texts from Twitter and Gab that are
labeled based on whether they express or describe sexist
attitudes or behaviors, and the type of sexism expressed
or described. A detailed overview of the dataset is
available at [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. The development and training datasets were
initially created for the EXIST 2021 evaluation campaign
3.1.4. User experience and satisfaction [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ] held as part of IberLEF 2021. However, the test set has
In this area, we aim to measure the user experience re- since been labeled as part of the ODESIA project and will
garding NLP products and services in both English and be kept confidential as part of the ODESIA leaderboard.
Spanish, using semi-automatic analysis of the
reputation of the products in social networks and user surveys. 3.3.2. DIPROMATS
To estimate this gap, we compute two types of indexes:
the reputational polarity index and the user satisfaction The DIPROMATS dataset is designed to train and test
index. models for identifying and characterizing propaganda
          </p>
          <p>The reputational polarity index estimates the over- techniques used by diplomats. It consists of labeled
all polarity of mentions of diferent NLP services and tweets, indicating the presence or absence of propaganda
products in both English and Spanish across various so- techniques, and if present, the type of propaganda used.
cial networks, as well as the polarity of various attributes The DIPROMAT dataset was used in the IberLef 2023
of the products/services. The user satisfaction index task on Automatic Detection and Characterization of
measures the degree of adoption of the products/services, Propaganda Techniques from Diplomats.4
both for personal (citizen adoption) and professional
(business adoption) use, the level of user satisfaction, 3.3.3. DIANN
and the limitations encountered during use.</p>
          <p>
            The DIANN (Disability Annotation) dataset was
developed at UNED to be used in the DIANN task at IberLEF
2018.5 Its main objective is to annotate mentions of
disabilities in scientific biomedical documents. A complete
description of the dataset can be found in [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ]. A new test
set has been labeled as part of the ODESIA project and
will be kept private as part of the ODESIA leaderboard.
          </p>
        </sec>
        <sec id="sec-2-1-3">
          <title>The ODESIA leaderboard aims to provide the AI research community with a tool to evaluate and compare the performance of NLP models, in an easy and homogeneous manner.</title>
          <p>The ODESIA leaderboard consists of two distinct
benchmarks: one for Spanish and another for English, 3.4. EvALL 2.0
where identical tasks are proposed and models can be EvALL 2.0 (Evaluate ALL) is an evaluation tool designed
applied to compare their performance in both languages. for information systems, ofering an extensive set of
met</p>
          <p>Currently, the ODESIA leaderboard is under develop- rics that cover various evaluation contexts, including
ment. It will include tasks such as named entity recog- classification, ranking, and clustering, with or without
nition, sexism detection and categorization, and propa- disagreement. The tool is designed with three core
conganda detection and categorization. However, it will be cepts in mind: persistence, replicability, and efectiveness.
frequently updated with new tasks, models, and results, Persistence enables users to store and retrieve past
evalumaking it a valuable resource for tracking progress in ations, while replicability ensures that all evaluations are
NLP research in both Spanish and English. carried out using the same methodology, making them
strictly comparable. Efectiveness means that all
met3.3. Datasets rics are based on measurement theory, and have been
double-implemented and evaluated.</p>
          <p>The following datasets have been developed (completely One of EvALL’s key features is its universality. It
ofor partially) during the first year of the ODESIA project, fers a single format that allows for the evaluation of
and are part of the ODESIA benchmark. multiple metrics, even from diferent evaluation contexts.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>3.2. The ODESIA Leaderboard</title>
        <sec id="sec-2-2-1">
          <title>4https://sites.google.com/view/dipromats2023/home 5http://nlp.uned.es/diann/</title>
        </sec>
        <sec id="sec-2-2-2">
          <title>Additionally, EvALL 2.0 is easy to use, thanks to its user</title>
          <p>friendly interface that enables users to generate
personalized or guided evaluations with just a few clicks.</p>
          <p>EvALL 2.0 provides a range of features to its users,
allowing them to: (i) evaluate their information systems
by providing the gold standard and their systems’
predictions; (ii) access past evaluations to validate the
efectiveness of new models; (iii) customize the selection of
metrics and adjust their parameters; (iv) choose from a
variety of input formats; (v) add new datasets to the EvALL
2.0 repository; and (vi) publish results related to datasets
included in the repository, making them available to the
wider scientific community.</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>3.5. ODESIA Portal</title>
        <p>art, as well as to understand the evolution of the
state of the NLP in Spanish over time.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Future work</title>
      <sec id="sec-3-1">
        <title>Our future work will concentrate on improving all the</title>
        <p>components that are necessary to carry out the project.</p>
        <p>We will compute various indexes in order to measure
the gap between the development of natural language
processing in Spanish and English. We plan to recalculate
and release these indexes yearly in the next two years, at
least.</p>
        <p>In addition, we aim to expand the data and
functionalities available in our ODESIA tools (Leaderboard, Portal,
and EvALL), and create more datasets in Spanish.</p>
        <p>The objective of the ODESIA portal6 (see Figure 2) is to
provide information on various NLP tasks in Spanish that Acknowledgments
have been tackled by the research community, including
the results obtained for these tasks and the datasets avail- ODESIA has been financed by the European Union
able for training and evaluating NLP systems. Version (NextGenerationEU funds) through the “Plan de
Recu1.0 of the portal contains information on 128 tasks and 95 peración, Transformación y Resiliencia”, by the Ministry
datasets from 125 competitions held between 2013 and of Economic Afairs and Digital Transformation and by
2022. the UNED University (C038/21-OT). However, the points</p>
        <p>The information has been gathered from the main of view and opinions expressed in this document are
evaluation forums in the NLP area: SemEval7, CLEF8, solely those of the author(s) and do not necessarily reflect
CoNLL9, COLING10, IberLEF11, IberEval12 and PAN13. those of the European Union or European Commission.</p>
        <p>The information in the ODESIA Portal is intended to Neither the European Union nor the European
Commissatisfy the needs of diferent stakeholders: sion can be considered responsible for them. Laura Plaza
and Jorge Carrillo-de-Albornoz are also financed by the
Ministry of Universities and the European Union through
the EuropeaNextGenerationUE funds and the “Plan de</p>
        <p>Recuperación, Transformación y Resiliencia”.
• Researchers seeking information on the diferent</p>
        <p>NLP tasks that have been tackled in Spanish, their
datasets and evalualtion measures; the existing
data sets, with information about the domain,
linguistic variety of texts, type of texts, annotations,
etc.; and the main evaluation forums or
competitions in the Spanish NLP scene.
• Companies that want quick access to information
about the variety of NLP tasks that may
incorporate into their systems, the performance that they
may expect to obtain for a task, the existing data
sets and their characteristics.
• Public entities and decision-makers financing
activities related to NLP and need to detect research
gaps in NLP in Spanish, evaluate the novelty and
relevance of project proposals and compare the
results of finaced projects with the state of the
6http://portal.odesia.uned.es
7https://semeval.github.io
8http://www.clef-initiative.eu
9https://www.signll.org/conll
10https://coling2022.org
11https://sites.google.com/view/iberlef2022
12https://sites.google.com/view/ibereval-2018
13https://pan.webis.de</p>
        <p>Figure 2: ODESIA Portal main page.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rodríguez-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mendieta-Aragón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Marco-Remón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Makeienko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2022:
          <article-title>sexism identification in social networks</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>69</volume>
          (
          <year>2022</year>
          )
          <fpage>229</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rodríguez-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          , L. Plaza,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Comet</surname>
          </string-name>
          , T. Donoso, Overview of EXIST 2021:
          <article-title>sexism identification in social networks</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <year>2021</year>
          )
          <fpage>195</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Fabregat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Martínez-Romo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Araujo</surname>
          </string-name>
          ,
          <article-title>Overview of the DIANN Task: Disability annotation task</article-title>
          ,
          <source>Proceedings of the Third Workshop on Evaluation of Human Language Technologies for Iberian Languages (IberEval</source>
          <year>2018</year>
          )
          <article-title>co-located with 34th Conference of the Spanish Society for Natural Language Processing (SEPLN</article-title>
          <year>2018</year>
          )
          <article-title>(</article-title>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>