<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>ZaRa-IU-NLP at EXIST 2023 - Sexism Identification: Specialized or Generalized?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Zackary Leech</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ravi Regulagedda</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandra Kübler</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Linguistics, Indiana University</institution>
          ,
          <addr-line>Bloomington, IN</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Luddy School of Informatics, Computing, and Engineering, Indiana University</institution>
          ,
          <addr-line>Bloomington, IN</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>8</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>We present our approach to EXIST (sEXism Identification in Social neTworks) Task 1, at CLEF 2023, comprising automatic sexism detection in both English and Spanish on Twitter data. We compare two methods, the first being a bi-ensemble method that combines two pre-trained BERT architectures, BETO and RoBERTa, for each specific language, Spanish and English, respectively. The second method utilizes the larger multilingual transformer, RoBERTa-XLM-base, and considers the entire dataset despite language diferences. We show that the language-specific ensemble performs better than the generalized model and is a better choice when looking at sexism detection in mixed Spanish and English data.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;sexism detection</kwd>
        <kwd>RoBERTa</kwd>
        <kwd>BETO</kwd>
        <kwd>language-specific classification</kwd>
        <kwd>language-agnostic classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        violence against women” with “the overall prevalence rate by region at 76% in North America
and 91% in Latin America and Caribbean” [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>From outwardly explicit misogyny to more subtle implicit misogyny, automatic classification
of sexism will create a more eficient process of creating a safer environment online. With rising
hate speech towards women online and little research into the detection of sexism, contributions
to this issue are urgent.</p>
      <p>We present the approach by team ZaRa-IU-NLP. This approach was developed during a
course on machine learning in NLP. Our work focuses on comparing a multilingual neural
model with an ensemble of language specific models.</p>
      <p>The remainder of this report will proceed as follows: Section 2 outlines previous research
for prior EXIST shared tasks, Sections 3 and 4 provide details on the task description, the
dataset, and the preprocessing methods, respectively. Section 5 describes the two rival model
architectures. In section 6 we analyze the results of the experiments before concluding in
section 7.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Team avacaondata [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] provided the winning approach to the 2022 EXIST Task 1; they utilized an
ensemble of transformer models using BERTweet-large, RoBERTa, and DeBERTa v3 for English,
and BETO, BERTIN, MarIA-base, and Robertuito for Spanish. With this combination, the team
achieved an overall F1 of 0.7996. To reach optimal performance given the computational load of
the models and to avoid overgeneration, the training was carried out in 2 phases, choosing to
optimize parameters with smaller amounts of data before expanding to the entire dataset [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
      </p>
      <p>
        Team CIMATCOLMEX ranked first in the evaluation of Spanish tweets, with an accuracy of
0.7801 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This team utilized an ensemble of 10 RoBERTuito and 10 BERT models. They reached
the highest scores for Spanish, surpassing avacaondata in F1 score by 2.27% absolute. Using a
bi-ensemble method, Villa-Cueva et. al [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] merged two transformers ensembles, one for Spanish
and one for English. Though there is a high computational cost, this model scored second in
Task 1 with a diference of 0.0038 in F1 to the winning team [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In regard to preprocessing,
team CIMATCOLMEX used the following normalization steps: lowercasing, removing emojis,
replacing usernames with “@user”, replacing any URL with the token “&lt;URL&gt;”, and removing
any whitespace at the beginning or end of the tweet [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Task Description and Dataset</title>
      <p>EXIST provides an opportunity to address the issue of online sexism with a wider reach, by
providing participants with data containing bilingual, sexist speech, in both Spanish and English.
Task 1 focuses on a binary classification of whether the text is sexist or not for both explicit
or implicit examples. The data is constructed out of tweets with any form of oppression or
prejudice against women because of their sex, explicit or implicit. It is worth noting that while
the dataset labels every tweet as either English or Spanish, it includes tweets mixing both of
the languages.
Data</p>
      <p>English
Spanish</p>
      <p>RoBERTa Tokenizer</p>
      <p>BETO Tokenizer</p>
      <p>RoBERTa</p>
      <p>BETO</p>
      <p>The EXIST 2023 dataset consists of 6 920 tweets for training, 1 038 tweets for validation, and
2 076 tweets for testing, where all sets are randomly selected from the 9 000 and 4 000 sets
sampled and created by the CLEF 2023 organizers, for balance. The annotation process was
carried out by a balanced group of 3 women and 3 men, in order to avoid gender bias. When
choosing the label for a given tweet from these 6 labels, we used a simple majority. In the case
of ties, we randomly selected either of YES or NO.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Preprocessing</title>
      <p>Our focus is on a comparison of two model architectures, a generalized model, processing
both English and Spanish tweets, and a specialized model using language specific classifiers
for each language. As such, preprocessing was kept to a minimum. Tweets were used with
URLs, emojis, duplications of characters, etc in place. Further, the dataset included a set of
tweets containing varying levels of both languages. No alterations were made to the content
of these individual bilingual tweets. Rather, each architecture was given an opportunity to
make judgement on tweets based on the language information provided by the shared task. I.e.,
tweets labeled English were passed through RoBERTa and tweets labeled Spanish were passed
through BETO. The fluidity of language, specifically in the informal context of twitter, creates
ample opportunity for bilingual users to hide implicit sexism through language switching and
as such, should be included in each model’s ability to detect implicit bias.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Model Architecture</title>
      <p>In our work, we compare two diferent approaches, a multilingual model, and a language
specific model. The multilingual model uses XLM-RoBERTa, while the language specific model,
RoBERTa for English, and BETO for Spanish. Figure 1 shows an overview of the data flow in
the two models.</p>
      <sec id="sec-5-1">
        <title>5.1. XLM-RoBERTa</title>
        <p>
          One of the first approach to sexism detection in a Spanish-English context was using
XLMRoBERTa [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. This is a multi-lingual version of RoBERTa [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] trained on a database of 100
languages. It was initially pre-trained for masked language modeling (MLM).
        </p>
        <p>We also use this model for our multilingual approach and finetune it on the EXIST dataset on
the binary classification task. The data is tokenized using XLM-RoBERTa’s tokenizer before
passing into the model that was finetuned in 3 epochs. The finetuned model is used for inference.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. RoBERTa + BETO</title>
        <p>This model is a pipeline composed of two models pre-trained on English and Spanish texts
respectively. For English, we chose RoBERTa, a model trained on English texts using an MLM
objective. RoBERTa is an improved model over the core BERT architecture with more parameters
and trained on a larger corpus to provide a more robust performance than the base BERT model.</p>
        <p>
          We chose BETO [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], a BERT-based Spanish language model for the Spanish language tweets
in the dataset. BETO was trained on a large Spanish corpus [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. Similar to RoBERTa, it was
pre-trained on an MLM objective.
        </p>
        <p>We fine-tuned both these models using the data for their respective languages. Since the data
was split fairly evenly between English and Spanish texts, RoBERTa and BETO had about half
the training data to finetune in comparison to XLM-RoBERTa.</p>
        <p>For this ensemble, we performed tokenization based on the specific language using the
pre-trained tokenizer. The data is then passed into the corresponding model.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Evaluation</title>
        <p>
          We participated in the HARD-HARD evaluation, i.e., we provided a single label per tweet, which
was evaluated against the gold label. For the oficial evaluation on the test set, we report ICM
and F1 for the positive class (sexist). For our internal results on the development set, we report
the macro-averaged F1 score, the F1 score per class, and ICM. ICM is a score developed to
measure the similarity between two datapoints more accurately [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. It is a generalization of
Pointwise Mutual Accuracy (PMA) and compares the closeness between two diferent outputs
by comparing them to a ground truth value.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Results</title>
      <sec id="sec-6-1">
        <title>6.1. Oficial Results</title>
        <p>Table 1 shows the oficial scores on the test set. These results show the ICM metric scores and
the F1 scores of the positive class for the two models in the HARD-HARD evaluation
ZaRaIU-NLP_1 is based on the multilingual XLM-RoBERTa, and ZaRa-IU-NLP_2 on a combination
of RoBERTa and BETO. Our scores show that the combination of language specific models
outperforms the generalized model in each metric, even though the individual language models
were trained on only half the data. The models produce consistent results across languages for
sexist tweets, with the language specific models producing an F1 score of 0.7332 for Spanish
tweets and 0.7263 for English tweets, while the single, large model reached 0.6956 and 0.6954,
respectively.</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Results on the Validation Set</title>
        <p>Table 2 compares the performance of our two models on the validation set, including language
specific results for both architectures. The results are in line with the oficial results,
ZaRa-IUNLP_2, the combination of language specific models, outperforms the multilingual
ZaRa-IUNLP_1 with an ICM of 0.4444 as compared to 0.3168. This shows that it is possible to obtain
solid results given a small set of data and that the quality of the data (wrt. the language) is more
important than the size of the training data. A look at the F1-scores per class shows a balanced
performance, the score for the sexist class is only slightly lower than the one for the majority
class of non-sexist tweets.</p>
        <p>We then had a closer look at precision and recall per language and per class. These results
are shown in Table 3. These results show that the multilingual model ZaRa-IU-NLP_1 shows
the same preference for the non-sexist class given the higher recall for this class. However,
this trend is much more pronounced for Spanish where the recall for non-sexist reaches 81.08%
while the recall for the sexist class is at 63.24%. In the combined model ZaRa-IU-NLP_2, the
Spanish classifier has a preference for the non-sexist class (recall: 85.81) while the English
classifier has a preference for the sexist class (recall: 86.08). These results suggest that we might
be able to gain better performance for Spanish if we upsample the sexist class. We leave this for
future research.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>Overall, the study addresses the dificulties of online sexism detection in social networks in a
bilingual setting. We carried out a comparison between two distinct architectures: the first,
a large, multilingual model that processes the entire data set, and the second, a combination
of language specific, smaller models that implement a split in language categorization for
processing. Given our minimal pre-processing, the higher accuracy of the language specific
model can be attributed to the more specialized language models. Our results, when tested
on the developmental set, show that the combination of two specialized models outperforms
the single, generalized one by a diference of 3 percent points in macro-averaged F1 score,
and a diference in ICM of 0.13. Overall, this study contributes to the process of automatic
sexism detection in social networks in highlighting the greater eficiency and accuracy of two
specialized models, rather than a multi-lingual model.</p>
      <p>As described above, this system was developed as a project in a course on machine learning.
For this reason, the system was intentionally kept simple, so that it could be carried out in a
short amount of time. This is the reason, for example, why we focused on comparing the two
systems without preprocessing the data, and without investigating other language models, etc.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Santos-Rios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vilares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Alonso</surname>
          </string-name>
          ,
          <article-title>Some experiments on the use of natural language processing for sexism detection and classification in social media</article-title>
          , in: A.
          <string-name>
            <surname>Leitao</surname>
          </string-name>
          , L. Ramos (Eds.),
          <source>Proceedings of V XoveTIC Conference (XoveTIC)</source>
          , volume
          <volume>14</volume>
          of Kalpa Publications in Computing, EasyChair,
          <year>2023</year>
          , pp.
          <fpage>24</fpage>
          -
          <lpage>27</lpage>
          . URL: https://easychair.org/publications/paper/ rdrm. doi:
          <volume>10</volume>
          .29007/8z6l.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>U. N.</given-names>
            <surname>Women</surname>
          </string-name>
          ,
          <article-title>Accelerating eforts to tackle online and technology-facilitated violence against women and girls</article-title>
          , https: //www.unwomen.org/en/digital-library/publications/2022/10/ accelerating-eforts
          <article-title>-to-tackle-online-and-technology-facilitated-violence-against-women-and-</article-title>
          <string-name>
            <surname>girls</surname>
          </string-name>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>The</given-names>
            <surname>Economist Intelligence Unit</surname>
          </string-name>
          ,
          <article-title>Measuring the prevalence of online violence against women, The Economist (</article-title>
          <year>2021</year>
          ). URL: https://onlineviolencewomen.eiu.com/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de-Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2023 -
          <article-title>Learning with Disagreement for Sexism Identification and Characterization</article-title>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Giachanou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Aliannejadi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Vlachos</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality</source>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          , Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de-Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2023 -
          <article-title>Learning with Disagreement for Sexism Identification and Characterization (Extended Overview)</article-title>
          , in: M.
          <string-name>
            <surname>Aliannejadi</surname>
            , G. Faggioli,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
          </string-name>
          , M. Vlachos (Eds.),
          <source>Working Notes of CLEF 2023 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaca-Serrano</surname>
          </string-name>
          ,
          <article-title>Detecting and classifying sexism by ensembling transformers models</article-title>
          ,
          <source>in: Proceedings of IberLEF, A Coruña</source>
          ,
          <year>Spain</year>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E.</given-names>
            <surname>Villa-Cueva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sánchez-Vega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. P.</given-names>
            <surname>López-Monroy</surname>
          </string-name>
          ,
          <article-title>Bi-ensembles of transformer for online bilingual sexism detection</article-title>
          ,
          <source>in: Proceedings of IberLEF, A Coruña</source>
          ,
          <year>Spain</year>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chaudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wenzek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Guzmán</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Unsupervised cross-lingual representation learning at scale, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>8440</fpage>
          -
          <lpage>8451</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .acl-main.
          <volume>747</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .acl-main.
          <volume>747</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized BERT pretraining approach</article-title>
          , CoRR abs/
          <year>1907</year>
          .11692 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1907</year>
          .11692. arXiv:
          <year>1907</year>
          .11692.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cañete</surname>
          </string-name>
          , G. Chaperon,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fuentes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-H.</given-names>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <article-title>Spanish pre-trained BERT model and evaluation data</article-title>
          ,
          <source>in: Practical ML for Developing Countries Workshop @ ICLR</source>
          <year>2020</year>
          ,
          <string-name>
            <given-names>Addis</given-names>
            <surname>Ababa</surname>
          </string-name>
          , Ethiopia,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cañete</surname>
          </string-name>
          ,
          <source>Compilation of large Spanish unannotated corpora</source>
          ,
          <year>2019</year>
          . URL: https://doi.org/ 10.5281/zenodo.3247731. doi:
          <volume>10</volume>
          .5281/zenodo.3247731.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Delgado</surname>
          </string-name>
          ,
          <article-title>Evaluating extreme hierarchical multi-label classification</article-title>
          ,
          <source>in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>5809</fpage>
          -
          <lpage>5819</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>