<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>INFOTEC-NLP at HOMO-LAT 2025: Testing a Novel Multi-region Spanish Model to Monitor Opinion in Latin American LGBTQI+ Social Media</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Guillermo Ruiz</string-name>
          <email>luis.ruiz@infotec.mx</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rogelio Campos</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tania Ramirez-delreal</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniela Moctezuma</string-name>
          <email>dmoctezuma@centrogeo.edu.mx</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mario Graf</string-name>
          <email>mario.graff@infotec.edu.mx</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eric Sadit Tellez</string-name>
          <email>eric.tellez@infotec.mx</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centro de Investigación en Ciencias de Información Geoespacial (CentroGeo)</institution>
          ,
          <addr-line>Aguascalientes, Ags., 20213</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>INFOTEC Centro Público de Investigación en Tecnologías de la Información y Comunicación</institution>
          ,
          <addr-line>112 Circuito Tecnopolo Sur, Parque</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Independent researcher</institution>
          ,
          <addr-line>Puebla</addr-line>
          ,
          <country country="MX">México</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Industrial Tecnopolo 2</institution>
          ,
          <addr-line>Aguascalientes, 20326</addr-line>
          ,
          <country country="MX">México</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Secretaría de Ciencia</institution>
          ,
          <addr-line>Humanidades, Tecnología e Innovación (SECIHTI), 1582 Insurgentes Sur 1582, Crédito Constructor, Ciudad</addr-line>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>de México</institution>
          ,
          <addr-line>03940</addr-line>
          <country country="MX">México</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>This paper presents our solution notebook for the HOMO-LAT@IberLEF2025 challenge, which consists of two tasks that require models that take advantage of subtle diferences in language between Spanish-speaking countries to identify the polarity of LGBTQI+ discourse in Reddit posts coming from Latin American Spanish-speaking countries. Our approach uses our BILMALAT model, a novel region-aware Large Language Model based on RoBERTa that is designed to understand and process Spanish variants from several countries, making it uniquely suited for capturing the diverse linguistic and cultural nuances present in the HOMO-LAT dataset. We show the high performance of BILMALAT in this challenge, demonstrating its eficacy in solution achieved the third position in both HOMO-LAT Tasks 1 and 2, solely fine-tuning the model for the task without extra strategies to balance the data and tackling large messages beyond the capacity of our model. hate and hope speech identification, multi-dialect Spanish language models, text classification models ∗Corresponding author.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Sentiment analysis, also known as opinion mining, applied to the online discourse on the LGBTQI+
community in Latin America, provides an indispensable framework for comprehending societal
perceptions and attitudes. Social media platforms, e.g., Reddit, which facilitate discussions, represent
extensive repositories of public sentiment, encompassing a spectrum from supportive to explicitly
harmful. Examining this content is essential to identify the predominant biases, monitor the
dissemination of discriminatory language, and evaluate the efectiveness of inclusion initiatives in various
Spanish-speaking contexts. This paper outlines our participation in the HOMO-LAT 2025 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] shared
task from IberLEF [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], specifically addressing this significant need by focusing on the polarity analysis
of LGBTQI+-related terms in Latin American Spanish.
      </p>
      <sec id="sec-1-1">
        <title>1.1. Related work</title>
        <p>
          To promote safer online environments, Bel-Enguix et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] introduced HOMOMEX, a shared task that
seeks computational models for automatically identifying hate speech against LGBTQI+ people on
social networks, particularly focusing on Mexican Spanish tweets. Subsequently, Gómez-Adorno et al.
        </p>
        <p>CEUR
Workshop</p>
        <p>
          ISSN1613-0073
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] presented a second edition of the challenge that maintained its focus on the Mexican Spanish dialect
while expanding the variety of message sources by incorporating song lyrics. The present iteration
has evolved into HOMO-LAT, a task centered on identifying polarity across various Latin American
countries using Reddit posts instead of tweets. This new version presents distinct challenges, regarding
multi-region data, where the training and test sets are sourced from the same countries, and cross-region
data, where the training and test sets comprise diferent countries. In HOMO-LAT, regular messages
indicate the country of origin, which allows the prediction model to take advantage of this information.
        </p>
        <p>
          When trained on extensive text datasets, it is acknowledged that transformer-based language models
(LMs) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] are highly efective for linguistic modeling; BERT [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and RoBERTa [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] are notable example
of such models. Although the initial training of these models is expensive, they can be eficiently
ifne-tuned for specific applications once pre-trained. Original BERT focuses on the English language
and provides a multilingual model that covers Spanish. Since BERT models are highly dependent on
how they were trained, including the data presented, several Spanish models have been introduced,
improving performance in most tasks. For example, BETO [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], BERTIN [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], and RoBERTuito [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
These models are highly performant and work for tasks in generic Spanish. However, experiments
like HOMO-LAT show that regional information is important for several domains, and we need more
sophisticated models, as noted in [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], to be able to consider regional diversity.
        </p>
        <p>
          In addition, when considering the study of specific language variants, academic investigations such as
those cited in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] have examined regional variations in the Spanish language among diferent
countries. The conclusions of these studies underscore the need for further initiatives to improve the
the Spanish language resources. Tellez et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] present several lexical and regional word embeddings
and BERT-based models tailored to Spanish-speaking countries. The study emphasizes that training
models in the same language variant as the dataset defining a task results in improved performance.
Jorge Ortiz-Fuentes1 has fine-tuned BETO models specifically for the Chilean Spanish, reinforcing the
idea that models designed for specific regions perform better in regional tasks.
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>1.2. Task description</title>
        <p>The HOMO-LAT of IberLEF 2025 asks systems to accurately predict the polarity of a Reddit post,
which involves determining whether the post is positive, negative, or neutral regarding the LGBTQI+
community. These posts are composed in Spanish by users from various Latin American nations.
Solutions should consider the regional context to account for local references and regionalisms that
might ofer diferent interpretations, making it crucial to identify the intent of these communications.
The posts are categorized into three types: identification of hate speech, hopeful messages, and sentiment
analysis. The task comprises two subtasks: Task 1 was assessed using a test dataset from the same
countries as in the training phase. Task 2 used a test consisting of messages from Bolivia, Costa
Rica, Cuba, the Dominican Republic, Ecuador, El Salvador, Guatemala, Honduras, Nicaragua, Panama,
Paraguay, Peru, Puerto Rico, Uruguay, and Venezuela.</p>
        <p>The training set contains 5,700 Reddit posts and the development set includes 1,400 messages. Each
post has at least one keyword from the following: trans, lgbt, gay/gays, lesbiana/as, bisexual/es, asexual/es,
transexual/les, travesti/is, queer/s, transgénero, pansexual/es, intersexual/es. The training has 3174 posts
classified as neutral, 1960 as negative, and 633 as positive. For the development partition, 888 are neutral,
475 negative, and 80 positive.</p>
      </sec>
      <sec id="sec-1-3">
        <title>Roadmap</title>
        <p>This section introduces the underlying opinion mining problem behind HOMO-LAT@IberLEF2025 and
presents a review of the related work. Section 2 briefly describes our BILMALAT model and describes
how we use it to solve the HOMO-LAT tasks. Section 3 details our experimental methodology and lists
our results. Finally, conclusions are given in Section 4.
1Models and comparisons can be accessed at https://huggingface.co/dccuchile/patana-chilean-spanish-bert and https://
huggingface.co/dccuchile/tulio-chilean-spanish-bert.</p>
        <sec id="sec-1-3-1">
          <title>Country</title>
        </sec>
        <sec id="sec-1-3-2">
          <title>Token</title>
        </sec>
        <sec id="sec-1-3-3">
          <title>Argentina</title>
        </sec>
        <sec id="sec-1-3-4">
          <title>Colombia</title>
        </sec>
        <sec id="sec-1-3-5">
          <title>Chile</title>
        </sec>
        <sec id="sec-1-3-6">
          <title>Mexico</title>
        </sec>
        <sec id="sec-1-3-7">
          <title>Uruguay</title>
        </sec>
        <sec id="sec-1-3-8">
          <title>Panama</title>
        </sec>
        <sec id="sec-1-3-9">
          <title>Peru</title>
        </sec>
        <sec id="sec-1-3-10">
          <title>Paraguay</title>
        </sec>
        <sec id="sec-1-3-11">
          <title>Guatemala</title>
        </sec>
        <sec id="sec-1-3-12">
          <title>Venezuela</title>
        </sec>
        <sec id="sec-1-3-13">
          <title>Honduras</title>
        </sec>
        <sec id="sec-1-3-14">
          <title>El Salvador</title>
        </sec>
        <sec id="sec-1-3-15">
          <title>Republica Dominicana</title>
        </sec>
        <sec id="sec-1-3-16">
          <title>Bolivia</title>
        </sec>
        <sec id="sec-1-3-17">
          <title>Ecuador</title>
        </sec>
        <sec id="sec-1-3-18">
          <title>Nicaragua</title>
        </sec>
        <sec id="sec-1-3-19">
          <title>Costa Rica</title>
          <p>Cuba
_ar
_co
_cl
_mx
_uy
_pa
_pe
_py
_gt
_ve
_hn
_sv
_do
_bo
_ec
_ni
_cr
_cu</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Our system solution</title>
      <p>Our approach was to fine-tune an early bird version of our BILMALAT 2 model. BILMALAT uses a
RoBERTa-based training strategy on 600 million tweets from Latin America countries, trained from
scratch. The model was designed such that the where and when of a given text is captured, i.e. it learns
about regions and periods of time. Our first real-world application is to use it in HOMO-LAT; the model
can be accessed and tested with the Huggingface framework 3.</p>
      <p>For BILMALAT to be region- and period-aware, we use additional tokens to indicate the region.
The first token corresponds to the country, the exact list is shown in Table 1. The second and third
tokens correspond to the year and month of publication. The tokens can be _2015, _2016, … , _2023 and
_01, … , _12 for the year and month, respectively. The rest of the tokens constitute the actual post.</p>
      <p>Here, we describe our solution. We first used the BETO 4 model as a baseline. We fine-tunned the
model in the training set for 6 epochs with a learning rate of 10−5. We conducted various experiments
to choose these hyperparameters.</p>
      <p>We then fine-tunned the BILMALAT with the messages provided. Since the training data included
the country of origin, we could include the first token, for the second and third we simply used a generic
x and _01 to force the model to predict without temporal bias 5. We fine-tuned the model using diferent
hyperparameters.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>Our experiments were carried out on Google Colab6. We employed the Huggingface framework with a
Pytorch back-end. Our BETO baseline obtained a best accuracy of 0.5385 and an F1-score of 0.4487.
2https://huggingface.co/guillermoruiz/bilmaLAT
3https://huggingface.co/
4https://huggingface.co/dccuchile/bert-base-spanish-wwm-cased
5For instance, our internal tests show that our model is aware of diferent meanings for terms like infección in COVID times
and other periods; we need much more research to take advantage of this extra-information.
6Google Colab site https://colab.research.google.com/.</p>
      <sec id="sec-3-1">
        <title>Test Learning rate</title>
      </sec>
      <sec id="sec-3-2">
        <title>Epoch</title>
      </sec>
      <sec id="sec-3-3">
        <title>Best epoch Validation loss 1 2</title>
        <p>Our model selection procedure was quite simple, focused on fine-tuning the HOMO-LAT base-model;
using the given training and development sets, we varied the learning rate in 10−4, 10−5, and 10−6; and
looking for the best model among the computed epochs. This process was not systematic and was
performed manually due to the lack of time. Table 2 displays these results for Task 1. We found that
models with 10−5 show a higher performance, notoriously a learning rate of 10−6 produced very low
scores, which means that we should give more epochs, but we were unable to do it due to our late
joining the challenge. We did not fine-tune for task 2 and use the same model.</p>
        <sec id="sec-3-3-1">
          <title>Oficial Results</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions</title>
      <p>Our BILMALAT achieved a third place of seven in the gold standard with an F1 score of 0.5137 vs.
0.5296 of the best system. Our system also achieved third place (of seven) with an F1 score of 0.4639,
where an F1 score 0.5086 was the best achieved.</p>
      <p>This challenge was a perfect scenario to test our BILMALAT and compare with the popular alternative
BETO model. Our approach, a language model enriched with geographical and cultural context,
demonstrated a clear advantage in navigating the subtle expressions of sentiment prevalent in the
dataset due to the message’s origin. Our model supports up to 128 input tokens, which might be enough
for tweets, but the Reddit posts used in this task were much longer, which gives them greater complexity.</p>
      <p>This result underscores the need for fine-grained regional awareness in natural language processing,
moving beyond generic linguistic models to include local information in the training data. The insights
gained from this competition provide a valuable contribution to the field of sentiment analysis and
open the way for more sophisticated and culturally sensitive AI applications.</p>
      <sec id="sec-4-1">
        <title>Future work</title>
        <p>Due to our delayed awareness of the challenge, we had limited time to explore advanced strategies for
task management. Consequently, we could not test sophisticated techniques like data augmentation
or other methods to address the dataset’s inherent imbalance and the uneven distribution between
training and development data. Additionally, similar to other BERT-based approaches, our model is
constrained by sentence size, meaning it cannot handle huge messages without additional strategies, i.e.,
BILMALAT requires additional strategies to process large messages. With further focus on addressing
the unique features of the challenge and overcoming the limitations of BERT-based models, our results
can be enhanced.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, we applied the Writefull’s model for grammar and spelling checks.
After using these services, we reviewed and edited the content as needed and assume full responsibility
for the content of the publication.
ume 1: Long Papers), Association for Computational Linguistics, Mexico City, Mexico, 2024, pp.
3689–3711. URL: https://aclanthology.org/2024.naacl-long.204. doi:10.18653/v1/2024.naacl-long.
204.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Bel-Enguix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gómez-Adorno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ojeda-Trueba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sierra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Barco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dunstan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Manrique</surname>
          </string-name>
          ,
          <article-title>Overview of HOMO-LAT at IberLEF 2025: Human-centric polarity detection in Online Messages Oriented to the Latin American-speaking lgbtq+ populaTion</article-title>
          ,
          <source>Procesamiento del lenguaje natural 75</source>
          (
          <year>2025</year>
          )
          <article-title>-</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Á</surname>
          </string-name>
          .
          <string-name>
            <surname>González-Barba</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Chiruzzo</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <article-title>Overview of IberLEF 2025: Natural Language Processing Challenges for Spanish and other Iberian Languages, in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2025), co-located with the 41st Conference of the Spanish Society for Natural Language Processing (SEPLN 2025), CEUR-WS</article-title>
          .org,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G.</given-names>
            <surname>Bel-Enguix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gómez-Adorno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sierra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vásquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Andersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ojeda-Trueba</surname>
          </string-name>
          ,
          <article-title>Overview of homo-mex at iberlef 2023: Hate speech detection in online messages directed towards the mexican spanish speaking lgbtq+ population</article-title>
          ,
          <source>Procesamiento del lenguaje natural 71</source>
          (
          <year>2023</year>
          )
          <fpage>361</fpage>
          -
          <lpage>370</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Gómez-Adorno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Bel-Enguix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Calvo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ojeda-Trueba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Andersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vásquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Alcántara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Soto</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Macias, Overview of homo-mex at iberlef 2024: Hate speech detection towards the mexican spanish speaking lgbt+ population</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>73</volume>
          (
          <year>2024</year>
          )
          <fpage>393</fpage>
          -
          <lpage>405</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          ,
          <year>2019</year>
          . URL: https://arxiv.org/abs/
          <year>1907</year>
          . 11692. arXiv:
          <year>1907</year>
          .11692.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cañete</surname>
          </string-name>
          , G. Chaperon,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fuentes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-H.</given-names>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <article-title>Spanish pre-trained bert model and evaluation data</article-title>
          ,
          <source>in: PML4DC at ICLR</source>
          <year>2020</year>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>J. D.</surname>
          </string-name>
          la Rosa y Eduardo G. Ponferrada y Manu Romero y Paulo Villegas y Pablo González de Prado Salas y María Grandury,
          <article-title>Bertin: Eficient pre-training of a spanish language model using perplexity sampling</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>68</volume>
          (
          <year>2022</year>
          )
          <fpage>13</fpage>
          -
          <lpage>23</lpage>
          . URL: http://journal. sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6403.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>J. M. Pérez</surname>
            ,
            <given-names>D. A.</given-names>
          </string-name>
          <string-name>
            <surname>Furman</surname>
            ,
            <given-names>L. Alonso</given-names>
          </string-name>
          <string-name>
            <surname>Alemany</surname>
            ,
            <given-names>F. M.</given-names>
          </string-name>
          <string-name>
            <surname>Luque</surname>
          </string-name>
          ,
          <article-title>RoBERTuito: a pre-trained language model for social media text in Spanish, in: Proceedings of the Thirteenth Language Resources</article-title>
          and Evaluation Conference, European Language Resources Association, Marseille, France,
          <year>2022</year>
          , pp.
          <fpage>7235</fpage>
          -
          <lpage>7243</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .lrec-
          <volume>1</volume>
          .
          <fpage>785</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E. S.</given-names>
            <surname>Tellez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Moctezuma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Miranda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Graf</surname>
          </string-name>
          , G. Ruiz,
          <article-title>Regionalized models for spanish language variations based on twitter</article-title>
          ,
          <source>Language Resources and Evaluation</source>
          <volume>57</volume>
          (
          <year>2023</year>
          ). URL: https: //doi.org/10.1007/s10579-023-09640-9. doi:
          <volume>10</volume>
          .1007/s10579-023-09640-9.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Jimenez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Dueñas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Rodriguez-Diaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mancera</surname>
          </string-name>
          ,
          <article-title>Automatic detection of regional words for pan-hispanic spanish on twitter</article-title>
          , in: G. R.
          <string-name>
            <surname>Simari</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Fermé</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Gutiérrez Segura</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          <string-name>
            <surname>Rodríguez Melquiades</surname>
          </string-name>
          (Eds.),
          <source>Advances in Artificial Intelligence - IBERAMIA 2018</source>
          , Springer International Publishing, Cham,
          <year>2018</year>
          , pp.
          <fpage>404</fpage>
          -
          <lpage>416</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>España-Bonet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Barrón-Cedeño</surname>
          </string-name>
          ,
          <article-title>Elote, choclo and mazorca: on the varieties of Spanish</article-title>
          , in: K. Duh,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , S. Bethard (Eds.),
          <source>Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          (Vol-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>