<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AI-UPV at IberLEF-2021 DETOXIS task: Toxicity Detection in Immigration-Related Web News Comments Using Transformers and Statistical Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Angel Felipe Magnoss~ao de Paula[</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ipek Baris S</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universitat Politecnica de Valencia</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes our participation in the DEtection of TOXicity in comments In Spanish (DETOXIS) shared task 2021 at the 3rd Workshop on Iberian Languages Evaluation Forum. The shared task is divided into two related classi cation tasks: (i) Task 1: toxicity detection and; (ii) Task 2: toxicity level detection. They focus on the xenophobic problem exacerbated by the spread of toxic comments posted in di erent online news articles related to immigration. One of the necessary e orts towards mitigating this problem is to detect toxicity in the comments. Our main objective was to implement an accurate model to detect xenophobia in comments about web news articles within the DETOXIS shared task 2021, based on the competition's o cial metrics: the F1-score for Task 1 and the Closeness Evaluation Metric (CEM) for Task 2. To solve the tasks, we worked with two types of machine learning models: (i) statistical models and (ii) Deep Bidirectional Transformers for Language Understanding (BERT) models. We obtained our best results in both tasks using BETO, a BERT model trained on a big Spanish corpus. We obtained the 3rd place in Task 1 o cial ranking with the F1-score of 0.5996, and we achieved the 6th place in Task 2 o cial ranking with the CEM of 0.7142. Our results suggest: (i) BERT models obtain better results than statistical models for toxicity detection in text comments; (ii) Monolingual BERT models have an advantage over multilingual BERT models in toxicity detection in text comments in their pre-trained language.</p>
      </abstract>
      <kwd-group>
        <kwd>Spanish text classi cation</kwd>
        <kwd>Toxicity detection</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>Transformers</kwd>
        <kwd>BERT</kwd>
        <kwd>Statistical models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The increase in the number of news pages where the reader can openly discuss
the articles has driven the dissemination of internet users' opinions through
social media [
        <xref ref-type="bibr" rid="ref10">18, 10</xref>
        ]. A survey carried out in the US by The Center for Media
Engagement at the University of Texas at Austin states that most of the
comments on news articles are posted by internet users who we call active-users or
in uencers [16]. They are highly active and generate huge amounts of data.
      </p>
      <p>
        The imbalance in the amount of data generated by in uencers and non-active
users creates a distorted reality where in uencers' opinions end up representing
the opinion of all internet users to society [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This distorted reality can aggravate
the existing social problems, as is the case with xenophobia, a heavy sense of
aversion, or dread of people from other countries [19].
      </p>
      <p>
        In recent years, the problem with xenophobia has been exacerbated by the
increase in the spread of toxic comments posted in di erent online news articles
related to immigration [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. One of the rst steps to mitigate the problem is to
detect toxic comments regarding news articles [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. For this reason, the Iberian
Languages Evaluation Forum proposed the DEtection of TOXicity in comments
In Spanish (DETOXIS) shared task 2021 [17].
      </p>
      <p>The DETOXIS shared task comprises Task 1 and Task 2, which are
respectively toxicity detection and toxicity level detection. The two tasks are performed
on comments posted in Spanish in response to di erent online news articles
related to immigration. Task 1 is a binary classi cation problem where the
objective is to classify a Spanish text comment as `toxic' or `not toxic'. Task 2 aims
to classify the same comment but among four classes: `not toxic', `mildly toxic',
`toxic', or `very toxic'. Table 1 displays examples of comments classi ed across
all classes.
asi me gusta, que se maten entre ellos y en alta mar. Mas
inmigrantes asi porfavor</p>
    </sec>
    <sec id="sec-2">
      <title>A esosmoros hay que echarlos pero ya.O los politicos hacen algo o la gente tendra que "actuar"</title>
      <p>
        The detection of toxicity in comments is mostly done with Machine Learning
(ML) models, especially deep learning models, which require large amounts of
annotated datasets for robust predictions [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. However, labeling toxicity is a
challenging and time-consuming task that requires many annotators to avoid
bias, and the annotators should be aware of social and cultural contexts [
        <xref ref-type="bibr" rid="ref11">15, 11</xref>
        ].
      </p>
      <p>
        Our main goal was to implement an accurate model to detect xenophobic
comments on web news articles within the DETOXIS shared task 2021, using
the competition's o cial metrics. We decided to solve the problem by applying
models that can learn using only a small amount of data, which can be done
with statistical models and most advanced pre-trained deep learning models.
Roughly speaking, there are two types of statistical models: Generative and
Discriminative [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. We chose to use one of each type. Thus, we tried a Naive Bayes
(Generative) and a Maximum Entropy (Discriminative) model. Among the most
advanced and highly e ective deep learning models is Deep Bidirectional
Transformers for Language Understanding (BERT), which comes with its parameters
pre-trained in an unsupervised manner in a large corpus [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Therefore, it only
needs a tuned train that can be run on a small set of data, which suits our
problem. Our source code is publicly available1
      </p>
      <p>The work's main contribution is to help in the e ort to improve the results
in the identi cation of toxic comments in news articles related to immigration.
Unlike the vast majority of works [14], we use ML models that can tackle the
xenophobia detection problem having only little data available. The second
contribution is to build ML models and nd their best con guration to deal not
only with the classi cation of news articles as `toxic' and `not toxic', but also to
infer the toxicity level of the comments into `not toxic', `mildly toxic', `toxic', or
`very toxic'. As far as we know, there are few works in the literature in which
the solution model tries to infer the toxicity level of the comments posted in the
news related to immigration. On the DETOXIS o cial ranking, we obtained the
3rd place in Task 1 with the F1-score of 0.5996, and we achieved the 6th place
in Task 2 with the CEM of 0.7142.</p>
      <p>The article is organized as follows: Section 2 contains the methodology with
fundamental concepts; Section 3 describes the experiments; Section 4 contains
the results and discussions, and and Section 5 draws some conclusion and future
work.
2</p>
      <sec id="sec-2-1">
        <title>Methodology</title>
        <p>This section explains the data structure, the evaluation metrics, and the ML
models applied to solve our classi cation problems. In addition, the text
representation used to encode the text comments.
2.1</p>
        <sec id="sec-2-1-1">
          <title>Dataset</title>
          <p>The DETOXIS shared task organization granted its participants the
NewsComTOX dataset [17] divided into train set and test set where text data are in
Spanish. The train set consists of 3463 instances, and the test set consists of 891
instances. Both sets have as main labels: (i) `Comment id' and (ii) `Comment';
but only the train set has the labels: (iii) `Toxicity' and (iv) `Toxicity level',
respectively for Task 1 and Task 2. The `Comment id' is a unique reference number
assigned to each instance within the NewsCom-TOX dataset. The `Comment'
label is a text message posted in response to a Spanish online news article from
di erent sources such as El Mundo, NIUS, ABC, etc., or discussion forums like
1 https://github.com/AngelFelipeMP/Machine-Learning-Tweets-Classification
Meneame. Moreover, `Toxicity' labels the comment for a particular instance
between `toxic' or `not toxic' and the `Toxicity level' label classi es the same
comment as `not toxic', `mildly toxic', `toxic', or `very toxic'. Table 2 shows the
label's distribution for `Toxicity' and `Toxicity level'. We can see that the labels
are unbalanced in both cases.</p>
          <p>The data annotation process was carried out by four annotators where two
were linguists experts, and two were trained linguistic students. Three of them
labeled all news article comments in parallel. Once they nished, an
interannotator agreement test was executed. When a disagreement happens, the three
annotators plus the senior annotator reviewed it in order to achieve accordance
with the nal label [17]. In Table 1, we can see examples from the DETOXIS
train set of comments and its labels attributed by the annotators for `Toxicity'
and `Toxicity level'.</p>
          <p>Next we explain how we used the data during the project development in
both tasks. First, we applied 10-fold cross-validation in the train set to nd the
best ML model. After that, we trained the selected model in the whole train
set. Subsequently, we applied the selected model to make predictions on the
o cial test set, as shown in Figure 1. These predictions were submitted to the
DETOXIS shared task 2021.
2.2</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Evaluation metrics</title>
          <p>
            Because the train set is imbalanced, as we can see in Table 2, we selected
evaluation metrics that are able to fairly evaluate ML models in this circumstance.
For Task 1, we adopted Accuracy, Recall, Precision, and F1-score, which was
the DETOXIS o cial evaluation metric for Task 1. For Task 2, we adopted
Accuracy, F1-macro, F1-weighted, Recall, Precision, and CEM [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ], the DETOXIS
o cial evaluation metric for Task 2. We used the DETOXIS o cial metrics as
performance measures to rank and select the best ML models during the
crossvalidation process for Task 1 and Task 2.
There are two types of statistical models: the Generative models and the
Discriminative models [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]. We used one model from each kind. We adopted the
Naive Bayes (Generative) model and the Maximum Entropy (Discriminative)
model. Among the Transformers models, we decided to use the BERT models,
one of the most advanced and highly e ective Transformer models. They come
with their parameters pre-trained in an unsupervised manner in a large corpus
[
            <xref ref-type="bibr" rid="ref7">7</xref>
            ]. Therefore, they only need a supervised ne-tune train in the downstream
task that can be run on a small set of data. We adopted: (i) the BETO model,
a BERT model trained on a big unannotated Spanish corpus composed of three
billion tokens [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ]; and (ii) the mBERT, a BERT model pre-trained on the top
102 languages with the most extensive Wikipedia corpus. However, the balance
among the language in the corpus was not perfect. For example, the English
partition of the corpus was 1000 bigger than the Icelandic partition [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ].
2.4
          </p>
        </sec>
        <sec id="sec-2-1-3">
          <title>Text representation</title>
          <p>To represent our text data in a way that the statistical models could handle
it, we used two encode methods: (i) Bag of Words (BOW) [20]; and (ii) Term
Frequency - Inverse Document Frequency (TF-IDF) [13]. The BOW represents a
text comment by a unidimensional vector whose length is the size of the training
vocabulary. In this case, each column of this vector contains the number of times
a particular word from the vocabulary appears in the speci c comment. The
TFIDF representation for each text comment is also a at vector with the size of
the training vocabulary. However, the value for each word on the vector follows
the well-known TF-IDF calculation [13].</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Experiments</title>
        <p>This section explains the environment setup, the data preprocessing, and
statistical models' feature extraction. Furthermore, the section also contains
explanations for the 10-fold cross-validation process and how we selected the model
to make predictions on the DETOXIS test set, which we submitted as our nal
results to the competition.
3.1</p>
        <sec id="sec-2-2-1">
          <title>Environment setup</title>
          <p>For code purposes, we used python 3.7.10. As a code editor/machine, we used
Google collaborator2. The main python libraries that we used were: (i) NumPy
1.19.5 to work with matrix, (ii) Pandas 1.1.5 to handle and visualize data, (iii)
Spacy 2.2.4 and (iv) the Natural Language Toolkit (NLTK) 3.2.5 for natural
language transformations, (v) Pytorch 1.6.0, and (vi) Transformers 3.0.0 to
actually implement the BERT models. In addition, we used (vii) Sklearn 0.22.2 to
implement the statistical models.
3.2</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>Preprocessing</title>
          <p>For both tasks, we only preprocess the data for the statistical models. The
preprocessing step was carried out on the text data from the train and test sets.
We used the built-in python model for Regular Expression (RegEx) and the
NLTK python library. Applying RegEx, we removed stock market tickers,
oldstyle retweet text, hashtags, hyperlinks and changed the numbers to the tag
\&lt;number&gt;". We employed the NLTK on the text comments to remove
stopwords, stem and tokenize the words.
3.3</p>
        </sec>
        <sec id="sec-2-2-3">
          <title>Feature extraction</title>
          <p>
            The feature extraction process was executed to focus on achieving good results
with the statistical models. These models' performance is susceptible to their
input features [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]. Hence, after preprocessing the datasets for the statistical
models, we executed the feature extraction process to create good input features.
We encode the text comments in two di erent manners: (i) BOW [20]; and (ii)
TF-IDF [13].
          </p>
          <p>The two proposed encode methods are based on word occurrences, and
unfortunately, they completely ignore the relative position information of the words
in the comments. Therefore, we lose the information about the local ordering
of the words. In order to mitigate this problem and preserve some of the local
word ordering information, we increase the vocabulary by extracting 2-grams
and 3-grams from the text comments despite dimensional increasing.
2 https://colab.research.google.com/
3.4</p>
        </sec>
        <sec id="sec-2-2-4">
          <title>Cross-validation</title>
          <p>The cross-validation process was performed on the train set aiming to nd the
best ML model to make the prediction on the DETOXIS test set. We can see the
summary of our cross-validation process in Figure 2. During the cross-validation,
each statistical model received di erent input features, and the BERT models
tried di erent hyper-parameters.</p>
          <p>In Figure 3, we can see the cross-validation process that focuses on the BERT
models. We tried di erent combinations for the Output BERT, Learning Rate,
Batch Size, and Epochs. The BERT models are composed of a pre-trained model
plus a linear layer at the top which receives the output of the pre-trained BERT
model. We have two di erent options for the Output BERT: (i) the sequence of
hidden states at the output of the last layer which we performed a mean pooling
and max pooling operation and concatenated them into a uni ed unidimensional
vector that we called `hidden'; (ii) the pooler of the last layer's hidden state of
the rst token of the sequence further processed by a linear layer and a tanh
activation function that we called `pooler'. For the Learning Rate, we tried 1E-5,
3E-5, and 5E-5. For the Batch Size, we tried 8, 16, 32, and 64. The number of
Epochs was from 1 to 20.</p>
          <p>Figure 4 illustrates the cross-validation process for the statistical models. We
can see that we tried four algorithm versions of the Naive Bayes (NB) model:
the Multinomial, the Bernoulli, the Gaussian, and the Complement ones. On the
other hand, we tried only the original version of the Maximum Entropy (ME)
model but with di erent solvers: the liblinear, newton, sag, saga, and lbfgs. We
call solvers the algorithms used in the optimization problem. We tried di erent
vocabulary sizes for all statistical models using di erent n-grams combinations
and the two encode methods: BOW and TF-IDF.
parameters. For Task 1, the evaluation metrics are Accuracy, F1-score, Recall,
and Precision, and for Task 2, the evaluation metrics are Accuracy, F1-macro,
F1-weighted, Recall, Precision, and CEM.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Newton</title>
      <p>Model OBuEtRpTut LeRaranteing BSaitzceh Epochs Accuracy F1-score Recall Precision
Model OBuEtRpTut LeRaranteing BSaitzceh Epochs Accuracy F1-score Recall Precision
mBERT
pooler
hidden
hidden
hidden
pooler
3E-05
5E-05
5E-05
3E-05
3E-05
BETO
pooler
pooler
pooler
pooler
pooler
1E-05
1E-05
1E-05
5E-05
At the end of the cross-validation, we selected the best model for each task
accordingly with the DETOXIS o cial metric for the speci ed task, as shown
in Figure 5.</p>
      <p>OBuEtRpTut LeRaranteing BSaitzceh Epochs Accuracy mFac1ro weiFgh1ted Recall Precision CEM</p>
      <p>After the cross-validation, we chose the best model for Task 1, which following
Table 12 is BETO with the respective parameters: (i) pooler as Output BERT;
(ii) 1E-05 Learning Rate; (iii) Batch Size equal 32; and (iv) 4 training Epochs.
We also selected the best model for Task 2 that following Table 14 is BETO with
the respective parameters: (i) hidden Output BERT; (ii) 1E-05 Learning Rate;
(iii) Batch Size equal 16; and (iv) 4 training Epochs. Having the best models
and their parameters, we trained the models on the train set.</p>
      <p>Once the best models are trained, we use those models to make the
predictions on the DETOXIS test set. These predictions afterward were submitted to
the DETOXIS shared task organization as our nal results.
4</p>
      <sec id="sec-3-1">
        <title>Results and Discussion</title>
        <p>We discovered important information on the cross-validation results. Looking at
Table 3, we can see that the ME model achieves its best results on Task 1 with
the BOW encode based on the F1-score evaluation metric, which is 0.4679. The
highest Accuracy 0.7126 and Recall 0.4019 are also performed with the BOW
encode. The only performance metric in which the TF-IDF encode obtains a
higher score is Precision that is 0.8928. Thus, we can conclude that BOW is
the best encoding for the ME model on Task 1 in the DETOXIS training set.
Moreover, employing the Sag solver, the ME model achieved a higher F1-score,
Recall, and Precision. Hence, it seems to us that Sag was the best solver for the
ME model on Task 1 in the DETOXIS training set. We do not have a de nitive
conclusion about the vocabulary size because the ME model achieved its highest
results with di erent numbers of n-grams for each metric.</p>
        <p>Observing Table 4, we see that the NB model achieves its best results on
Task 1 with the BOW encode based on the F1-score evaluation metric, which
is 0.5355. The NB model also obtained the higest Recall 0.8004 with the BOW
encode, but its highest results for Accuracy 0.6933 and Precision 0.7282 were
with the TF-IDF encode. Therefore, we can not conclude which encode method
is the best for the NB model on Task 1 in the DETOXIS training set. A similar
case occurs with the vocabulary size where the NB model that employed 1-grams,
2-grams, and 3-grams achieved the highest F1-score and Recall. However, the
NB model with a 1-grams vocabulary size obtained the highest Accuracy and
Precision. The di erent NB algorithms obtained a similar performance based
on the F1-score. In most cases, they achieved their best results with the BOW
encode.</p>
        <p>We can see in Table 5 that the ME model achieved its best results on Task
2 with the BOW encode based on the CEM evaluation metric, which is 0.7080.
The highest F1-macro 0.3587, F1-weighted 0.6125, Recall 0.3367, and Precision
0.4942 are also obtained with the BOW encode. The only performance metric in
which the TF-IDF encode obtains a higher score is the Accuracy, which is 0.6846.
Thus, we can conclude that BOW is the best encoding for the ME model on Task
2 in the DETOXIS training set. Moreover, employing the Newton solver, the ME
model achieved a higher Accuracy, F1-macro, F1-weighted, Recall, and Precision.
Hence, it seems to us that Newton was the best solver for the ME model on
Task 2 in the DETOXIS training set. We concluded that the vocabulary size of
1-grams is the best for the ME model on Task 2 in the DETOXIS training set
because the ME model achieved its highest Accuracy, F1-macro, F1-weighted,
Recall, and Precision.</p>
        <p>Table 6 shows that the NB model achieved its best results on Task 2 with
the TF-IDF encode based on the CEM evaluation metric, which is 0.6882. The
NB model also obtained its highest Accuracy 0.6769, F1-weighted 0.5845, and
Precision 0.3171, with the TF-IDF encode, but its highest results for F1-macro
0.2916 and Recall 0.3365 were obtained with the BOW encode. Therefore, we
can conclude that the TF-IDF encode best suits the NB model on Task 2 in the
DETOXIS training set. We see indications in Table 6 that the ideal vocabulary
for the NB model on Task 2 in the DETOXIS training set is composed of 1-grams
and 2 grams. Once with this vocabulary, the model obtained its highest Accuracy,
Recall, Precision, and CEM results. The di erent NB algorithms obtained similar
performance based on the CEM ranged from 0.49 to 0.68.</p>
        <p>Based on the F1-score, the mBERT model achieved its best performance on
Task 1 with a value of 0.6010, as we can see in Table 7. The model parameters are
the following: (i) pooler as Output BERT; (ii) 3E-05 Learning Rate; (iii) Batch
Size equal 32; and (iv) 11 training Epochs. Table 8 shows that the BETO model
obtained its best performance on Task 1 also based on the F1-score with the
following parameters: (i) pooler as Output BERT; (ii) 1E-05 Learning Rate; (iii)
Batch Size equal 32; and (iv) 4 training Epochs. The BETO model obtained a
F1score value of 0.6314, which was also the highest among all the ML models in the
cross-validation process. For this reason, the BETO model with the mentioned
parameters was used for our Task 1 o cial prediction on the DETOXIS test set.
These predictions afterward were submitted as our o cial Task 1 results.</p>
        <p>Observing Table 9, we can conclude that based on the CEM, the mBERT
model achieved its best performance on Task 1 with the following parameters: (i)
pooler as Output BERT; (ii) 1E-05 Learning Rate; (iii) Batch Size equal to 16;
and (iv) 12 training Epochs. This model achieved the CEM of 0.7599. Table 10
shows that the BETO model obtained its best performance on Task 2 also based
on the CEM with the following parameters: (i) hidden as Output BERT; (ii)
1E05 Learning Rate; (iii) Batch Size equal to 16; and (iv) 4 training Epochs. The
BETO model obtained CEM value of 0.7769, which was also the highest among
all the ML models in the cross-validation process. For this reason, the BETO
model with the mentioned parameters was used for our Task 2 o cial prediction
on the DETOXIS test set. These predictions afterward were submitted as our
o cial Task 2 results.</p>
        <p>To sum up the comments about the cross-validation results, looking at Tables
12 and 14, we can see that the BETO model with di erent combinations of
parameters obtained the ve rst positions on the ranking for the best ML
model for Task 1 and Task 2.</p>
        <p>The DETOXIS organization provided us with the results of the test set.
Table 15 shows our result on Task 1 plus the three o cial DETOXIS baselines:
Random Classi er, Chain BOW, and BOW Classi er. Our model obtained an
F1-score around 59% greater than the results obtained by the best baseline on
Task 1.</p>
        <p>Table 16 shows the results of our model and the three DETOXIS baselines
on Task 2. Our BETO model was able to achieve a CEM of 9% higher than the
best DETOXIS baseline result obtained by the Random Classi er.</p>
        <p>On the DETOXIS o cial ranking, we obtained 3rd place on Task 1 with
F1-score 0.5996, and we achieved 6th place on Task 2 with CEM 0.7142.
5</p>
      </sec>
      <sec id="sec-3-2">
        <title>Conclusion and Future Work</title>
        <p>Xenophobia is a problem which is aggravated by the increase in the spread of
toxic comments posted in di erent online news articles related to immigration.
In this paper, to address this problem within the DETOXIS 2021 shared task,
we tried two types of ML models: (i) statistical models and (ii) BERT models.
We obtained the best results in both tasks using BETO, a BERT model
pretrained with a big Spanish corpus. Our contributions are as follows: (i) help in
the e ort to improve the results in the identi cation of toxic comments in news
articles related to immigration. Unlike the vast majority of works, we use ML
models that can tackle the xenophobia detection problem having only little data
available; (ii) We build an ML model and nd its best con guration to deal not
only with the classi cation of news articles as `toxic' and `not toxic', but also to
infer the toxicity level of the comments into `not toxic', `mildly toxic', `toxic', or
`very toxic'.</p>
        <p>Based on the DETOXIS o cial metrics, we concluded that our results
indicate that: (i) BERT models obtain better results than statistical models for
toxicity and toxicity level detection in text comments; and (ii) Monolingual BERT
models achieve higher results in comparison with the multilingual BERT models
in toxicity detection and toxicity level detection in their pre-trained language.</p>
        <p>After all, our BETO model obtained the 3rd position on Task 1 o cial
ranking with the F1-score of 0.5996, and it achieved the 6th position on Task 2 o cial
ranking with the CEM of 0.7142. As future work, we aim to include sentiment
lexicons on the model's input to boost its performance.
13. Pimpalkar, A.P., Raj, R.J.R.: In uence of pre-processing strategies on the
performance of ml classi ers exploiting tf-idf and bow features. ADCAIJ: Advances in
Distributed Computing and Arti cial Intelligence Journal 9(2), 49{68 (2020)
14. Plaza-Del-Arco, F.M., Molina-Gonzalez, M.D., Uren~a Lopez, L.A., Mart
nValdivia, M.T.: Detecting misogyny and xenophobia in spanish tweets
using language technologies. ACM Trans. Internet Technol. 20(2) (Mar 2020).
https://doi.org/10.1145/3369869
15. Risch, J., Krestel, R.: Delete or not delete? semi-automatic comment moderation
for the newsroom. In: Proceedings of the rst workshop on trolling, aggression and
cyberbullying (TRAC-2018). pp. 166{176 (2018)
16. Stroud, N.J., Van Duyn, E., Peacock, C.: News commenters and news comment
readers. Engaging News Project pp. 1{21 (2016)
17. Taule, M., Ariza, A., Nofre, M., Amigo, E., Rosso, P.: Overview of the detoxis task
at iberlef-2021: Detection of toxicity in comments in spanish. Procesamiento del
Lenguaje Natural 67 (2021)
18. Winter, S., Bruckner, C., Kramer, N.C.: They came, they liked, they commented:
Social in uence on facebook news channels. Cyberpsychology, Behavior, and Social
Networking 18(8), 431{436 (2015)
19. Xenophobia: Retrieved from https://en.oxforddictionaries.com/de nition/money.</p>
        <p>Oxford Online Dictionary (2021)
20. Zhang, Y., Jin, R., Zhou, Z.H.: Understanding bag-of-words model: a statistical
framework. International Journal of Machine Learning and Cybernetics 1(1-4),
43{52 (2010)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Amigo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mizzaro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carrillo-de Albornoz</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>An e ectiveness metric for ordinal classi cation: Formal properties and experimental results</article-title>
          . arXiv preprint arXiv:
          <year>2006</year>
          .
          <volume>01245</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Baeza-Yates</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>Biases on social media data: (keynote extended abstract)</article-title>
          .
          <source>In: Companion Proceedings of the Web Conference</source>
          <year>2020</year>
          . p.
          <volume>782</volume>
          {
          <fpage>783</fpage>
          . WWW '
          <volume>20</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA (
          <year>2020</year>
          ). https://doi.org/10.1145/3366424.3383564
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Blaya</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Audrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Toward an understanding of the characteristics of secondary school cyberhate perpetrators</article-title>
          .
          <source>In: Frontiers in Education</source>
          . vol.
          <volume>4</volume>
          , p.
          <fpage>46</fpage>
          .
          <string-name>
            <surname>Frontiers</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Canete</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaperon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuentes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perez</surname>
          </string-name>
          , J.:
          <article-title>Spanish pre-trained bert model and evaluation data</article-title>
          .
          <source>PML4DC at ICLR</source>
          <year>2020</year>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Davidson</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warmsley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Macy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Automated hate speech detection and the problem of o ensive language</article-title>
          .
          <source>In: Proceedings of the International AAAI Conference on Web and Social Media</source>
          . vol.
          <volume>11</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Devlin</surname>
          </string-name>
          , J.:
          <article-title>Multilingual bert readme document</article-title>
          . https://github.com/ google-research/bert/blob/a9ba4b8d7704c1ae18d1b28c56c0430d41407eb1/ multilingual.md (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Gheisari</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhuiyan</surname>
            ,
            <given-names>M.Z.A.</given-names>
          </string-name>
          :
          <article-title>A survey on deep learning in big data</article-title>
          .
          <source>In: 2017 IEEE International Conference on Computational Science and Engineering (CSE) and IEEE International Conference on Embedded and Ubiquitous Computing (EUC)</source>
          .
          <source>vol. 2</source>
          , pp.
          <volume>173</volume>
          {
          <issue>180</issue>
          (
          <year>2017</year>
          ). https://doi.org/10.1109/CSE-EUC.
          <year>2017</year>
          .215
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jebara</surname>
          </string-name>
          , T.:
          <article-title>Machine learning: discriminative and generative</article-title>
          , vol.
          <volume>755</volume>
          . Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <article-title>M.: Exploring the in uence of comment tone and content in response to misinformation in social media news</article-title>
          .
          <source>Journalism Practice</source>
          <volume>15</volume>
          (
          <issue>4</issue>
          ),
          <volume>456</volume>
          {
          <fpage>470</fpage>
          (
          <year>2021</year>
          ). https://doi.org/10.1080/17512786.
          <year>2020</year>
          .1739550
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Korencic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baris</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fernandez</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leuschel</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanchez Salido</surname>
          </string-name>
          , E.:
          <article-title>To block or not to block: Experiments with machine learning for news comment moderation</article-title>
          .
          <source>In: Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation</source>
          . pp.
          <volume>127</volume>
          {
          <fpage>133</fpage>
          . Association for Computational Linguistics,
          <source>Online (Apr</source>
          <year>2021</year>
          ), https://www.aclweb.org/anthology/2021.hackashop-
          <volume>1</volume>
          .
          <fpage>18</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Nelson</surname>
          </string-name>
          , W.B.:
          <article-title>Accelerated testing: statistical models, test plans, and data analysis</article-title>
          , vol.
          <volume>344</volume>
          . John Wiley &amp; Sons (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>