<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluation of Intermediate Pre-training for the Detection of O ensive Language</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Segun Taofeek Aroyehun</string-name>
          <email>aroyehun.segun@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Gelbukh</string-name>
          <email>gelbukh@gelbukh.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CIC, Instituto Politecnico Nacional Mexico City</institution>
          ,
          <country country="MX">Mexico</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents an evaluation of intermediate pretraining for the task of o ensive language identi cation. We leverage recent advances in multilingual contextual representation and ne-tuning of pre-trained language models. We compare the performance of a pretrained language model adapted for the social media domain and another that was further trained on multilingual sentiment analysis data. We found that the intermediate pre-training steps prior to ne-tuning on the target task yield performance gains. The best submissions by our team, NLP-CIC, achieved rst and second place on the non-contextual Spanish (Subtask 1) and Mexican Spanish (Subtask 3) subtasks of the MeO endEs-IberLEF 2021 shared task respectively.</p>
      </abstract>
      <kwd-group>
        <kwd>XLM-RoBERTa</kwd>
        <kwd>Social Media</kwd>
        <kwd>Spanish</kwd>
        <kwd>Mexican Spanish</kwd>
        <kwd>O ensive Language Identi cation</kwd>
        <kwd>Sentiment Analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The purpose of social media is for information exchange. This involves
interactions among users on the various social media platforms. During these
interactions, users often show unhealthy and anti-social behaviour such as insults and
personal attacks. This kind of behaviour hampers meaningful conversations at
the least and can cause harm to individuals, groups, and the society at large.
Natural language processing research can help in identifying o ensive language
to help reduce incidences of unacceptable behaviour. Research into this problem
has gained attention especially in English language. This is attributable to
availability of labeled data and pre-trained word embeddings and language models.
Recently, there has been a number of shared tasks with a focus on languages
other than English. One example is the IberLEF shared task series on o ensive
language identi cation in Mexican Spanish. For the 2021 edition [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] the tasks
include MeO endEs [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], a track on o ensive language identi cation on several
social media platforms. The challenge includes datasets for Spanish and Mexican
Spanish.
      </p>
      <p>
        Multilingual language models (LM) are becoming a popular area of focus
with the transformer architecture [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] which makes it possible to combine text
written in di erent languages to learn a single multilingual representation. There
are several successful instances of this approach in multilingual BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
XLMRoBERTa [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and recently multilingual T5 [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. However, these pre-trained
multilingual LMs mainly cover domains with text written in consistent and formal
style in contrast to social media text which are noisy, irregular and informal. To
adapt LMs to a speci c domain, the authors of [
        <xref ref-type="bibr" rid="ref4 ref7">4, 7</xref>
        ] showed that using the
transformer architecture and its pre-trained weights, a domain-speci c model can be
derived by continuing pre-training on text speci c to the domain of interest. For
the social media domain, [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] are examples of this adaptation in the
monolingual English setting. For the multilingual case, the authors of [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] introduced
an adaptation of XLM-RoBERTa to multilingual twitter text. XLM-RoBERTa
was further trained with the masked language modeling objective on twitter
text (about 12GB) in over 30 languages. Furthermore, this LM was trained on a
uni ed collection of sentiment analysis data in eight languages with the goal of
demonstrating the e ectiveness of the multilingual LM trained on twitter text.
      </p>
      <p>
        In this paper, our focus is to evaluate the e ectiveness of the pre-trained
LMs on the identi cation of o ensive language in tweets written in Spanish
and Mexican Spanish. We examine the e ect that intermediate pre-training on
sentiment analysis, a la [
        <xref ref-type="bibr" rid="ref10 ref12">10, 12</xref>
        ], has on o ensive language identi cation.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <p>
        Task. We address the non-contextual classi cation of o ensive language in tweets
written in Spanish (Subtask 1) and Mexican Spanish (Subtask 3). For Subtask
1, the task is to classify comments written in Spanish using only the textual
content into one of four categories: O ensive where the target is a person (OFP);
O ensive where the target is a group of people (OFG); non-o ensive, but with
inadequate language (NOM); non-o ensive (NO). This subtask also assess the
agreement between the con dence of model predictions and the con dence of
human annotators. Subtask 3 is a binary classi cation of tweets written in Mexican
Spanish. It requires predicting whether a comment is o ensive or not.
Data. The MeO endEs 2021 shared task [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] provides two corpora, O endEs
and O endMEX, which are collections of messages on social media platforms
in Spanish and Mexican Spanish annotated with labels indicating o ensiveness.
The generic Spanish data consist of labeled comments focusing on popular young
Spanish in uencers collected from di erent social media platforms (YouTube,
Instagram, and Twitter). The Mexican Spanish dataset was collected from Twitter
and manually labeled for o ensiveness. In addition, metadata for each comment
is provided for the classi cation in the contextual tracks of the competition. We
only participate in the non-contextual tracks in both languages. Table 1 provides
details of the Mexican Spanish dataset. Also, the details of the Spanish dataset
is in Table 2.
      </p>
      <p>
        We perform minimal pre-processing of the data in our experiments as it was
reported in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] that extensive pre-processing tends to hurt performance of
pretrained LMs. Hence, we normalize the text by converting user mentions and web
links to @USER and URL. We also replace multiple consecutive whitespaces
with a single one and punctuation marks are surrounded by a single whitespace
character on both sides. Then, the sequence of text is tokenized with the subword
tokenizer provided with the XLM-RoBERTa model, which is a Sentence Piece
model (using a unigram language model) with a vocabulary size of 250K [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. We
set the maximum sequence length to 128 subword tokens.
      </p>
      <p>
        Fine-tuning. We use the huggingface transformers library [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] for the
experiments. We add a linear prediction layer on top of the pooled output of the last
transformer layer and optimize this layer jointly with the pre-trained layers. We
optimize the model using Adam (without bias correction) with a batch size of
128 on a single Nvidia V100 GPU (32GB) and a maximum learning rate within
the range of 1e-5 to 5e-5. We use a warmup ratio of 0.1 and set the maximum
number of epochs to 10 with earlystopping on the validation performance metric
(micro F1) using a patience of 2 evaluation runs. We evaluate the performance
of the model every 20 steps on the validation set. Furthermore, we employ three
regularization techniques: weight decay with a factor of 0.01, dropout applied
to the pooled output of the last transformer layer with a probability of 0.2,
and label smoothing with a factor of 0.1. Our submissions vary in their use of
the regularization approaches. Details of the settings for each submission are
in Table 3. The con gurations are based on XLM-twitter 1 and
XLM-twittersentiment 2 pre-trained models introduced in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. For all experiments, we set
the random seed to 42. On average, the ne-tuning procedure takes about 900
seconds (wall time) for the Spanish task and approximately 600 seconds (wall
time) for the Mexican Spanish task.
1 https://huggingface.co/cardi nlp/twitter-xlm-roberta-base
2 https://huggingface.co/cardi nlp/twitter-xlm-roberta-base-sentiment
The scores of our submissions on the development and test sets for Subtask 1
(generic Spanish) are in Table 4. Three submissions are allowed for this subtask.
Submission-II which is based on the XLM-RoBERTa model trained on both
multilingual twitter text and sentiment analysis dataset achieved our best
submission out of the three. The model also has the least mean squared error, an
indication of greater agreement with the con dence of human annotators. The
overall ranking showed that this system is the best on the competition for the
non-contextual classi cation in Spanish. Figure 1 shows the confusion matrix of
the best model (Submission II) on the validation dataset for Subtask 1. It can
be observed that most of the mistakes on the validation data occurs where the
model predicts non-o ensive (NO) when the comments are actually o ensive to
a person (OFP). Also, the model performs poorly on the o ensive to a group
category (OFG). It makes the correct prediction on 1 out of 4 examples in the
validation set.
      </p>
      <p>Table 5 presents the results received by the NLP-CIC team on the
leaderboard on the unseen test set for Subtask 3 (Mexican Spanish). The maximum
number of submissions for this task is ve. It can be observed that the
XLMRoberta model that has been further trained on twitter data and a collection
of sentiment analysis datasets in eight languages (Submission-I) that we
netune on Subtask 3 dataset has the highest score out of our four submissions.
Also, the results show that label smoothing was bene cial for this task. On the
overall ranking for the competition, this system is in second place. In Figure
2, the confusion matrix provides an overview into the best model performance
(Submission I) across the two classes. The model predicts the non-o ensive label
(NO) on 8 examples when the true label is o ensive (OFF) compared to the
converse where the model predicts the o ensive label on 1 example when the true
label is non-o ensive. It shows that the model is relatively better at identifying
the non-o ensive category on the validation dataset. This can be linked to the
number of examples for the non-o ensive category which is about three times
more than the o ensive category in the training set.</p>
      <p>We observed that overall, the scores on the Spanish dataset is far higher
than the Mexican Spanish dataset even though it's a binary classi cation task.
We think that the amount of data available for the Spanish task is a factor
for this di erence in performance. The consistent performance of the model that
includes sentiment analysis as part of the pre-training for both tasks con rms our
hypothesis that sentiment analysis can be bene cial for detecting o ensiveness.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>We address the task of o ensive language identi cation in Spanish and Mexican
Spanish using a pre-trained language model adapted for the twitter domain.
We found that a further training on multilingual sentiment analysis is bene cial
to the task. In addition, label smoothing proved useful on the Mexican Spanish
dataset. The best systems submitted by our team, NLP-CIC, achieved rst place
on the non-contextual Spanish task and second place on the non-contextual
Mexican Spanish task.</p>
      <p>In the future, we will like to examine whether a model trained on Spanish
data can be seamlessly transferred to Mexican Spanish for this task and vice
versa. Our models only use textual content, it is very likely that the addition of
metadata can improve their performance.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>Thanks to the competition organizers for their support. The authors thank
CONACYT for the computer resources provided through the INAOE
Supercomputing Laboratory's Deep Learning Platform for Language Technologies.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aroyehun</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelbukh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>NLP-CIC at HASOC 2020: Multilingual O ensive Language Detection using All-in-one Model</article-title>
          .
          <source>In: FIRE (Working Notes)</source>
          . pp.
          <volume>331</volume>
          {
          <issue>335</issue>
          (
          <year>2020</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2826</volume>
          /
          <fpage>T2</fpage>
          -31.pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Barbieri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anke</surname>
            ,
            <given-names>L.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Camacho-Collados</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <string-name>
            <surname>XLM-T: A Multilingual Language</surname>
          </string-name>
          <article-title>Model Toolkit for Twitter</article-title>
          .
          <source>arXiv preprint arXiv:2104.12250</source>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Barbieri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Camacho-Collados</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Espinosa</surname>
            <given-names>Anke</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Neves</surname>
          </string-name>
          , L.:
          <article-title>TweetEval: Uni ed benchmark and comparative evaluation for tweet classi - cation</article-title>
          . In:
          <article-title>Findings of the Association for Computational Linguistics: EMNLP 2020</article-title>
          . pp.
          <volume>1644</volume>
          {
          <fpage>1650</fpage>
          . Association for Computational Linguistics,
          <source>Online (Nov</source>
          <year>2020</year>
          ). https://doi.org/10.18653/v1/
          <year>2020</year>
          . ndings-emnlp.
          <volume>148</volume>
          , https://www.aclweb.org/anthology/2020. ndings-emnlp.
          <fpage>148</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Beltagy</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lo</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>SciBERT: A pretrained language model for scienti c text</article-title>
          .
          <source>In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          . pp.
          <volume>3615</volume>
          {
          <fpage>3620</fpage>
          . Association for Computational Linguistics, Hong Kong,
          <source>China (Nov</source>
          <year>2019</year>
          ). https://doi.org/10.18653/v1/
          <fpage>D19</fpage>
          -1371, https://www.aclweb.org/anthology/D19-1371
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Conneau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khandelwal</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaudhary</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wenzek</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guzman</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoyanov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Unsupervised cross-lingual representation learning at scale</article-title>
          .
          <source>In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          . pp.
          <volume>8440</volume>
          {
          <fpage>8451</fpage>
          . Association for Computational Linguistics,
          <source>Online (Jul</source>
          <year>2020</year>
          ). https://doi.org/10.18653/v1/
          <year>2020</year>
          .aclmain.
          <volume>747</volume>
          , https://www.aclweb.org/anthology/2020.acl-main.
          <fpage>747</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          : BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers). pp.
          <volume>4171</volume>
          {
          <fpage>4186</fpage>
          . Association for Computational Linguistics, Minneapolis,
          <source>Minnesota (Jun</source>
          <year>2019</year>
          ). https://doi.org/10.18653/v1/
          <fpage>N19</fpage>
          -1423, https://www.aclweb.org/anthology/N19-1423
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gururangan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marasovic</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swayamdipta</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lo</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beltagy</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Downey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          :
          <article-title>Don't stop pretraining: Adapt language models to domains and tasks</article-title>
          .
          <source>In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          . pp.
          <volume>8342</volume>
          {
          <fpage>8360</fpage>
          . Association for Computational Linguistics,
          <source>Online (Jul</source>
          <year>2020</year>
          ). https://doi.org/10.18653/v1/
          <year>2020</year>
          .acl-main.
          <volume>740</volume>
          , https://www.aclweb.org/anthology/2020.acl-main.
          <fpage>740</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Montes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aragon</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agerri</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alvarez-Carmona</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alvarez Mellado</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Chiruzzo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freitas</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            <given-names>Adorno</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Gutierrez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Jimenez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.M.</given-names>
            ,
            <surname>Lima</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Plaza-de Arco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.M.</given-names>
            ,
            <surname>Taule</surname>
          </string-name>
          , M. (eds.):
          <source>Proceedings of the Iberian Languages Evaluation Forum (IberLEF</source>
          <year>2021</year>
          ) (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>D.Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vu</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tuan</surname>
            <given-names>Nguyen</given-names>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>BERTweet: A pre-trained language model for English tweets</article-title>
          .
          <source>In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source>
          . pp.
          <volume>9</volume>
          {
          <fpage>14</fpage>
          . Association for Computational Linguistics,
          <source>Online (Oct</source>
          <year>2020</year>
          ). https://doi.org/10.18653/v1/
          <year>2020</year>
          .emnlp-demos.2, https://www.aclweb.org/anthology/2020.emnlp-demos.
          <fpage>2</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Phang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calixto</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Htut</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pruksachatkun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vania</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kann</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          :
          <article-title>English intermediate-task training improves zeroshot cross-lingual transfer too</article-title>
          .
          <source>In: Proceedings of the 1st Conference of the Asia-Paci</source>
          c
          <article-title>Chapter of the Association for Computational Linguistics and the 10th</article-title>
          <source>International Joint Conference on Natural Language Processing</source>
          . pp.
          <volume>557</volume>
          {
          <fpage>575</fpage>
          . Association for Computational Linguistics, Suzhou,
          <source>China (Dec</source>
          <year>2020</year>
          ), https://www.aclweb.org/anthology/2020.aacl-main.
          <fpage>56</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <article-title>Plaza-del-</article-title>
          <string-name>
            <surname>Arco</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Casavantes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Jair</given-names>
            <surname>Escalante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Mart</surname>
          </string-name>
          n-Valdivia,
          <string-name>
            <given-names>M.T.</given-names>
            ,
            <surname>Montejo-Raez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Montes-</surname>
          </string-name>
          y-Gomez,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Jarqu</surname>
          </string-name>
          n-Vasquez,
          <string-name>
            <surname>H.</surname>
          </string-name>
          ,
          <article-title>Villasen~or-</article-title>
          <string-name>
            <surname>Pineda</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Overview of the MeO endEs task on o ensive text detection at IberLEF 2021</article-title>
          .
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <issue>0</issue>
          ) (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Pruksachatkun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Phang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Htut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.M.</given-names>
            ,
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            ,
            <surname>Pang</surname>
          </string-name>
          , R.Y.,
          <string-name>
            <surname>Vania</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kann</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          :
          <article-title>Intermediate-task transfer learning with pretrained language models: When and why does it work? In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</article-title>
          . pp.
          <volume>5231</volume>
          {
          <fpage>5247</fpage>
          . Association for Computational Linguistics,
          <source>Online (Jul</source>
          <year>2020</year>
          ). https://doi.org/10.18653/v1/
          <year>2020</year>
          .acl-main.
          <volume>467</volume>
          , https://www.aclweb.org/anthology/2020.acl-main.
          <fpage>467</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaiser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Attention is all you need</article-title>
          . In: Guyon,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.V.</given-names>
            ,
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Vishwanathan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Garnett</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds.)
          <source>Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          , pp.
          <volume>5998</volume>
          {
          <fpage>6008</fpage>
          . Curran Associates, Inc. (
          <year>2017</year>
          ), http://papers.nips.cc/paper/7181-attention
          <article-title>-is-all-you-need</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Wolf</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Debut</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanh</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaumond</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delangue</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cistac</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rault</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Louf</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Funtowicz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davison</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shleifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , von Platen,
          <string-name>
            <surname>P.</surname>
          </string-name>
          , Ma,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Jernite</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Plu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Le Scao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Gugger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Drame</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Lhoest</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            ,
            <surname>Rush</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Transformers: State-of-the-Art Natural Language Processing</article-title>
          .
          <source>In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source>
          . pp.
          <volume>38</volume>
          {
          <fpage>45</fpage>
          . Association for Computational Linguistics,
          <source>Online (Oct</source>
          <year>2020</year>
          ). https://doi.org/10.18653/v1/
          <year>2020</year>
          .emnlp-demos.6, https://www.aclweb.org/anthology/2020.emnlp-demos.
          <fpage>6</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Xue</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Constant</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kale</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Rfou</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siddhant</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barua</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ra el</surname>
          </string-name>
          , C.: mt5:
          <article-title>A massively multilingual pre-trained text-to-text transformer</article-title>
          . arXiv preprint arXiv:
          <year>2010</year>
          .
          <volume>11934</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>