<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>GuillemGSubies at IberLEF-2021 DETOXIS task: Detecting Toxicity with Spanish BERT</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Guillem Garc a Subies</string-name>
          <email>guillem.garcia@iic.uam.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Instituto de Ingenier a del Conocimiento</institution>
          ,
          <addr-line>Francisco Tomas y Valiente st., 11 EPS, B Building, 5th oor UAM Cantoblanco. 28049 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes a system created for the DETOXIS 2021 shared task, framed within the IberLEF 2021 workshop. We present an approach mainly based in ne-tuned BERT models using a GridSearch and Data Augmentation with MLM substitution. This approach only takes into account the textual data from the dataset to prove the power of language models. Our models far outperform the baselines and achieve results close to the state-of-the-art.</p>
      </abstract>
      <kwd-group>
        <kwd>Toxicity Detection</kwd>
        <kwd>BERT</kwd>
        <kwd>Transformers</kwd>
        <kwd>Data Augmentation</kwd>
        <kwd>BETO</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Polarization can be a very problematic issue in society, especially on social media.
There are manual mechanisms to report these behaviors, however they can be
slow and ine cient. To address this, we can use NLP to detect automatically
these undesirable toxic behaviors. The DETOXIS (DEtection of TOxicity in
comments In Spanish) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] shared task proposes, during this third edition of the
IberLEF [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] workshop, a corpus to detect toxicity level in comments on internet
forums and newspapers discussions.
      </p>
      <p>
        This article summarizes our participation in all the DETOXIS tasks. Given
the success of Transformer-inspired language models [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], both in academia and
industry [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], we decided to use already pre-trained BERT [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] models.
Specifically, we will use BETO [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] with some extra transfer learning techniques for
ordinal classi cation problems and a hyperparameters Grid-Search. To address
the problem of small data, we will use Data Augmentation techniques.
      </p>
      <p>In the next section, we will brie y see some previous work related to this
topic. In Section 3 we will go through a brief description of the tasks and the
corpus. Then, in Section 4, we will explain the main ideas behind the proposed
models. In Section 5 we will present a summary of the experiments we carried out
and the results we got. Finally, in Section 6 we will expose the main conclusions
of our work and results and we will also propose some ideas for future work.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>There is an extensive bibliography on Sentiment Analysis and text classi cation
in social networks, however not that much work has been done about identifying
and classifying toxic behaviors until 2019.</p>
      <p>
        Most of the toxicity detection datasets are focused on classifying what kind
of toxicity is present in the text, instead of the level of toxicity. Basile et al.
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] propose a task for the identi cation of toxicity against women and migrant
people in Spanish and English Tweets. Stru et al., [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] also focus on Twitter, but
now with German tweets. Other corpora focus in a more multilingual emphasis
like Kumar et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and Zampieri et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        However, all have a common denominator. The best results have been
obtained with some kind of Transformers or BERT-based model. For instance, the
best models in Stru et al., [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] are ne-tuned BERTs with German pretraining
on general data or ne-tuned BERTs pre-trained with speci c German tweets.
The best results in Kumar et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] also point at BERT and Transformers,
speci cally a bootstrap aggregation of BERT models. Finally, the best results
in Zampieri et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] are also some kind of Transformer based models and
ensembles of them.
      </p>
      <p>This is a clear indicator of the trends in the state-of-the-art for this topic.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Tasks Description</title>
      <p>The main corpus consists of 3463 comments posted in Spanish online newspapers
and forums for the train split and 890 for the test one. They were collected from
August 2017 to July 2020. Furthermore, the articles were selected taking into
account their potential toxicity and the number of comments in them (more than
50 comments).</p>
      <p>For the rst task, the comments are annotated into two categories; toxic
and not toxic. The second task consists of further classifying that toxicity into
four levels of toxicity; toxicity level 0=not toxic, toxicity level 1=mildly
toxic, toxicity level 2=toxic and toxicity level 3=very toxic.</p>
      <p>In addition to the classi cation labels, for every sample there are also
annotation about the argumentation, constructiveness, stance, target, stereotype,
sarcasm, mockery, insult, improper language, aggressiveness and intolerance.
However none of these are public in the test set, so they will be ignored in this study.
Finally, for every comment, there is also a label indicating if the comment is a
response to another comment or not. This information will not be used either
because we will only focus on the textual data.</p>
      <p>
        The metrics used to evaluate the results are the F-measure for the task1 and
the Closeness Evaluation Metric (CEM) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for the second task. This last metric
is very useful for ordinal classi cation problems given that it takes into account
the order of the classes using concepts from Measurement Theory.
      </p>
      <p>In the table above we can see the distribution of the samples in the train split.
The most notable fact is that 67% of the samples are not toxic, so the dataset
is unbalanced. For the second task, we can also see a very notable unbalance.</p>
      <p>In the table below, we can see some illustrative examples of the data and
their labels:
We performed a simple preprocessing where we substituted some expressions
with a more normalized form:
{ Every URL was replaced with the token \[URL]" so we don't get strange
tokens when the tokenizer tries to process and URL. Furthermore, no semantic
information about toxicity can be inferred from a URL, the only information
relevant for the model is that there is a URL in that token.
{ Finally we normalized every laugh (\jasjajajajj" ! \haha") so we minimize
the noise of the misspellings, common in social networks.
4.2</p>
      <sec id="sec-3-1">
        <title>Baselines</title>
        <p>We created some baselines so we can compare our models properly. We selected
a HashingVectorizer + RandomForest. This way, we can compare our models to
a classic feature extraction model.
4.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Language Models</title>
        <p>
          We used BETO [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], a BERT model trained with the Spanish Unannotated
Corpora (SUC) [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] that has proven to be much better than the multilingual BERT
model.
        </p>
        <p>
          We tried di erent training strategies given that the classes are related to each
other:
{ The simplest approach we tried is treating both tasks as the same one, with
a multiclass classi cation model. For the rst task, everything di erent from
not toxic would be considered toxic.
{ For the second approach, we rst trained a binary classi cation model to
distinguish between not toxic and toxic for the rst task. Then we trained
a multiclass model to classify between the three levels of toxicity.
{ Similarly to the last approach, we then tried to transform the second task
into three di erent binary classi cation problems; classifying between not
toxic and the rest of the classes, mildly toxic and toxic or very toxic,
and between toxic and very toxic. With this, we tried to have very speci c
models that can di erentiate slight changes in toxicity.
{ As there are not too many samples for the last models in the previous
approach to learn correctly, we also tried with a transfer learning approach,
similar to the one presented by Sun et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Instead of using always the
same BETO pretrained model for every netuned model, we used the
netuned model from the step before, i.e. the model that classi es mildly toxic
and toxic has as base model the one that classi es not toxic and mildly
toxic.
        </p>
        <p>In addition, for the ne-tuning process, we carried out a Grid-search
optimization over the main parameters of the neural network: learning rate, batch
size and dropout rate. The search was performed with a 5-fold strati ed
crossvalidation with the following grid: Learning rate, (1e 6; 1e 5; 3e 5; 5e 5; 1e
4); batch size, (8; 16; 32) and dropout rate, (0:08; 0:1; 0:12). The best parameters
for both models were: learning rate, 1e 5; batch size, 16 and dropout rate, 0:1.
4.4</p>
      </sec>
      <sec id="sec-3-3">
        <title>Data Augmentation</title>
        <p>As the dataset is relatively small, we decided to run Data Augmentation
techniques. The selected strategy was the Data Augmentation through the masking
of words with a Masked Language Model, BETO.</p>
        <p>For every sample in the dataset, we randomly masked 15% of the tokens and
used BETO to predict them, creating a modi ed sample. With this method, we
obtained double the amount of the original samples.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments and Results</title>
      <sec id="sec-4-1">
        <title>Experimental Setup</title>
        <p>We trained all the models with a NVIDIA Tesla P100-PCIE-16GB GPU and
a Intel(R) Xeon(R) CPU E5-2640 v4 @ 2.40GHz CPU with 500GB of RAM
memory.</p>
        <p>
          The software we used was Python3.8, transformers 4.5.1 [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], pytorch 1.8.1
[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], scikit-learn 0.24.1 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and nlpaug [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] 1.1.3.
5.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Results</title>
        <p>
          In the Table 3 we can see the results for our models in the test set of the
rst task. Note that the ChainBOW baseline, Word2VecSpacy baseline and the
SINAI team (winner of the task) results are taken from the task Overview [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
Our runs for this task are BETO-multiclass and BETO-binary both with and
without data augmentation as explained in Section 4.3. Note that some of the
results presented here were obtained after the labeled test set was published so
we could analyze in depth our models.
        </p>
        <p>We can see that there is almost no di erence between the binary and
multiclass models. This might happen because there is a great di erence in the
amount of samples and toxicity between the not toxic comments and the rest,
which makes them easy to indentify in every situation. Finally we can see the the
Data Augmentation strategy obtained around 0:02 points more than the models
without augmentation. Our result was the second best in the competition, which
proves that the simplicity of using BETO with some Grid-Search can yield really
good results.</p>
        <p>
          For the second task, the results were similar to the ones obtained in the rst
task. In the Table 4 we can look at them in more detail. Again, ChainBOW
baseline, Word2VecSpacy baseline and the SINAI team results are taken from
the task Overview [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. We can see that our transfer learning approach
(BETOtransfer ) obtains better results than the other approaches and that there is
almost no di erence between the simple multiclass approach and the one that
rst detects the not toxic comments (BETO-2models-aug ). These results are in
line with the ones in the rst task, showing that adding the not toxic class to
the models, will not make them worse.
        </p>
        <p>This results are placed fth among all the participating teams (24), which
proves that our approach, given it's simplicity and the lack of any linguistic
analysis, is very good.</p>
        <p>Model CEM
Word2VecSpacy 0.6116
HV+RF 0.6214
ChainBOW 0.6535
BETO-2 models-aug 0.6891
BETO-multiclass-aug 0.6913
BETO-3 models-aug 0.704
BETO-transfer 0.7172
BETO-transfer-aug 0.7189
SINAI team 0.7495</p>
        <p>Table 4. Results for task2
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>Through this shared task, we have seen that NLP can be of great help in
detecting and classifying unwanted toxic behavior in social networks and there is
still a long way to go.</p>
      <p>The results obtained by our systems are very promising given their great
performance and their simplicity. This compilation of methods is very signi
cant because it could lead to much better results when combined with other
improvements from the state-of-the-art.</p>
      <p>
        We believe that our results could improve a lot using speci c language models
trained with corpora from social networks like TWilBert [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Another interesting
approach would be to use a general language model and further pre-train it with
corpora from the same domain [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] as the nal task. Finally, we have proven
that good hyperparameters are also key for a good neural network so a better
search, like the Population Based Training [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], would further improve the model.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work has been partially funded by the Instituto de Ingenier a del Conocimiento
(IIC) and the hardware used was also provided by the IIC.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Amigo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mizzaro</surname>
            , S., de Albornoz,
            <given-names>J.C.</given-names>
          </string-name>
          :
          <article-title>An e ectiveness metric for ordinal classi cation: Formal properties and experimental results (</article-title>
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bosco</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fersini</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nozza</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Rangel</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.M.</given-names>
            ,
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Sanguinetti</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.:</surname>
          </string-name>
          <article-title>SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter</article-title>
          .
          <source>In: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          . pp.
          <volume>54</volume>
          {
          <fpage>63</fpage>
          . Association for Computational Linguistics, Minneapolis, Minnesota, USA (Jun
          <year>2019</year>
          ). https://doi.org/10.18653/v1/
          <fpage>S19</fpage>
          -2007, https://www.aclweb.org/anthology/S19- 2007
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Can~ete, J.:
          <article-title>Compilation of large spanish unannotated corpora</article-title>
          (May
          <year>2019</year>
          ). https://doi.org/10.5281/zenodo.3247731, https://doi.org/10.5281/zenodo.3247731
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Can~ete, J.,
          <string-name>
            <surname>Chaperon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuentes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perez</surname>
          </string-name>
          , J.:
          <article-title>Spanish pre-trained bert model and evaluation data</article-title>
          . In: to appear
          <source>in PML4DC at ICLR</source>
          <year>2020</year>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding (</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Angel</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Hurtado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.F.</given-names>
            ,
            <surname>Pla</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          :
          <article-title>Twilbert: Pre-trained deep bidirectional transformers for spanish twitter</article-title>
          .
          <source>Neurocomputing</source>
          (
          <year>2020</year>
          ). https://doi.org/https://doi.org/10.1016/j.neucom.
          <year>2020</year>
          .
          <volume>09</volume>
          .078, http://www.sciencedirect.com/science/article/pii/S0925231220316180
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jaderberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dalibard</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osindero</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Czarnecki</surname>
            ,
            <given-names>W.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Donahue</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Razavi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinyals</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Green</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunning</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simonyan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fernando</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kavukcuoglu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Population based training of neural networks (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ojha</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malmasi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zampieri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Evaluating aggression identi cation in social media</article-title>
          .
          <source>In: Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying</source>
          . pp.
          <volume>1</volume>
          {
          <issue>5</issue>
          .
          <string-name>
            <given-names>European</given-names>
            <surname>Language Resources Association</surname>
          </string-name>
          (ELRA), Marseille, France (May
          <year>2020</year>
          ), https://www.aclweb.org/anthology/2020.trac-
          <volume>1</volume>
          .
          <fpage>1</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Ma</surname>
          </string-name>
          , E.:
          <article-title>Nlp augmentation</article-title>
          . https://github.com/makcedward/nlpaug (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Montes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aragon</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agerri</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Angel Alvarez Carmona,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Alvarez</given-names>
            <surname>Mellado</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>de Albornoz</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chiruzzo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freitas</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adorno</surname>
            ,
            <given-names>H.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gutierrez</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zafra</surname>
            ,
            <given-names>S.M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lima</surname>
          </string-name>
          , S.,
          <string-name>
            <surname>de Arco</surname>
            ,
            <given-names>F.M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taule</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021)</article-title>
          .
          <source>In: CEUR Workshop Proceedings</source>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Paszke</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gross</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Massa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lerer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bradbury</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chanan</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Killeen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gimelshein</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antiga</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Desmaison</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kopf</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>DeVito</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raison</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tejani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chilamkurthy</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steiner</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bai</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chintala</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Pytorch:
          <article-title>An imperative style, high-performance deep learning library</article-title>
          . In: Wallach,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Larochelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Beygelzimer</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>d'</surname>
            Alche-Buc,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fox</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garnett</surname>
            ,
            <given-names>R</given-names>
          </string-name>
          . (eds.)
          <source>Advances in Neural Information Processing Systems</source>
          <volume>32</volume>
          , pp.
          <volume>8024</volume>
          {
          <fpage>8035</fpage>
          . Curran Associates, Inc. (
          <year>2019</year>
          ), http://papers.neurips.cc/paper/9015-pytorch
          <article-title>-animperative-style-high-performance-deep-learning-library</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanderplas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Passos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cournapeau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brucher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perrot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duchesnay</surname>
          </string-name>
          , E.:
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          ,
          <volume>2825</volume>
          {
          <fpage>2830</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Stru</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siegel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruppenhofer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiegand</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klenner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Overview of germeval task 2, 2019 shared task on the identi cation of o ensive language</article-title>
          .
          <source>In: Preliminary proceedings of the 15th Conference on Natural Language Processing (KONVENS</source>
          <year>2019</year>
          ),
          <source>October</source>
          <volume>9</volume>
          {
          <fpage>11</fpage>
          , 2019 at Friedrich-Alexander-Universita
          <article-title>t Erlangen-Nurnberg</article-title>
          . pp.
          <volume>352</volume>
          {
          <fpage>363</fpage>
          .
          <article-title>German Society for Computational Linguistics &amp; Language Technology und Friedrich-Alexander-Universitat Erlangen-Nurnberg</article-title>
          , Munchen [u.a.] (
          <year>2019</year>
          ), https://nbn-resolving.org/urn:nbn:de:bsz:
          <fpage>mh39</fpage>
          -
          <lpage>93197</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qiu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>How to ne-tune bert for text classi cation? (</article-title>
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Taule</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ariza</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nofre</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amigo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the detoxis task at iberlef-2021: Detection of toxicity in comments in spanish</article-title>
          .
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaiser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Attention is all you need (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Wolf</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Debut</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanh</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaumond</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delangue</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cistac</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rault</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Louf</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Funtowicz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davison</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shleifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , von Platen,
          <string-name>
            <surname>P.</surname>
          </string-name>
          , Ma,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Jernite</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Plu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Scao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.L.</given-names>
            ,
            <surname>Gugger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Drame</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Lhoest</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            ,
            <surname>Rush</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.M.</surname>
          </string-name>
          : Transformers:
          <article-title>State-of-the-art natural language processing</article-title>
          .
          <source>In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source>
          . pp.
          <volume>38</volume>
          {
          <fpage>45</fpage>
          . Association for Computational Linguistics,
          <source>Online (Oct</source>
          <year>2020</year>
          ), https://www.aclweb.org/anthology/2020.emnlp-demos.
          <fpage>6</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Zampieri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenthal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanasova</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karadzhov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mubarak</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derczynski</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pitenis</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , Coltekin, C.: SemEval-2020 task 12:
          <article-title>Multilingual o ensive language identi cation in social media (O ensEval 2020)</article-title>
          .
          <source>In: Proceedings of the Fourteenth Workshop on Semantic Evaluation</source>
          . pp.
          <volume>1425</volume>
          {
          <fpage>1447</fpage>
          . International Committee for Computational Linguistics,
          <source>Barcelona (online) (Dec</source>
          <year>2020</year>
          ), https://www.aclweb.org/anthology/2020.semeval-
          <volume>1</volume>
          .
          <fpage>188</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>