<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>TheNorth @ HaSpeeDe 2: BERT-based Language Model Fine-tuning for Italian Hate Speech Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eric Lavergne</string-name>
          <email>eric.lavergne@gmx.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rajkumar Saini</string-name>
          <email>rajkumar.saini@ltu.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gy o¨rgy Kov a´cs</string-name>
          <email>gyorgy.kovacs@ltu.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Killian Murphy</string-name>
          <email>killian.murphy@telecom-sudparis.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lulea ̊ Tekniska Universitet</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. This report was written to describe the systems that were submitted by the team “TheNorth” for the HaSpeeDe 2 shared task organised within EVALITA 2020. To address the main task which is hate speech detection, we fine-tuned BERT-based models. We evaluated both multilingual and Italian language models trained with the data provided and additional data. We also studied the contributions of multitask learning considering both hate speech detection and stereotype detection tasks.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Organised as part of the 7th EVALITA
evaluation campaign
        <xref ref-type="bibr" rid="ref3">(Basile et al., 2020)</xref>
        , the HaSpeeDe
2 shared task focuses on the detection of online
hate speech
        <xref ref-type="bibr" rid="ref19">(Sanguinetti et al., 2020)</xref>
        in
ItalianHate speech occurs frequently on social media. It
is defined as “any communication that disparages
a person or a group on the basis of some
characteristics such as race, colour, ethnicity, gender,
sexual orientation, nationality, religion, or other
characteristics”
        <xref ref-type="bibr" rid="ref14">(Nockleby, 2000)</xref>
        . Regulating all
user messages is very time-consuming for a
human, and this is one of the reasons why automatic
methods are important.
      </p>
      <p>Beside the main task of binary hate speech
classification - aimed at deciding whether a message
contains hate speech or not - the HaSpeeDe 2
shared task has two more sub-tasks. One being
stereotype detection, and the other the
identification of nominal utterances. All tasks being
evaluated both on in-domain (tweets) data, and
outof-domain (newspaper headlines) data. Here, we</p>
      <p>
        Copyright c 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
tackle both the main task, and the first sub-task
of Stereotype Detection that is potentially useful
for the main task. For this sub-task the
organisers use the following definition of Stereotype: “a
standardized mental picture that is held in
common by members of a group and that represents
an oversimplified opinion, prejudiced attitude, or
uncritical judgment”
        <xref ref-type="bibr" rid="ref13">(Merriam-Webster, 2020)</xref>
        .
      </p>
      <p>
        Here, we have two binary classification tasks. A
simple way to perform text classification is based
on bag-of-words representation counting the
number of occurrences of each word within text. It is
often combined with term frequency-inverse
document frequency
        <xref ref-type="bibr" rid="ref22">(Sparck Jones, 1988)</xref>
        (TF-IDF)
representation. TF-IDF allows the frequencies to
be normalized according to how often the words
appear in all documents. With the rise of
neural networks, word vectors have provided useful
features for text classification tasks. Recurrent
Neural Networks as the Bidirectional Long-Short
Term Memory (BiLSTM) network
        <xref ref-type="bibr" rid="ref20">(Schuster and
Paliwal, 1997)</xref>
        have then be used to encode the
long-term dependencies between the words. These
systems were the most successful in the previous
HaSpeeDe campaign
        <xref ref-type="bibr" rid="ref5 ref8">(Bosco et al., 2018)</xref>
        .
      </p>
      <p>
        In
        <xref ref-type="bibr" rid="ref1">(Aluru et al., 2020)</xref>
        , the authors showed
that when dealing with very low monolingual
resources, multilingual approaches can be
interesting for hate speech. In
        <xref ref-type="bibr" rid="ref16 ref17">(Polignano et al., 2019b)</xref>
        ,
the AlBERTo monolingual Italian BERT-based
language model was trained that outperformed the
state-of-the-art on the HaSpeeDe 2018 evaluation
task
        <xref ref-type="bibr" rid="ref16 ref17">(Polignano et al., 2019a)</xref>
        .
      </p>
      <p>We have chosen to deepen the approach of
finetuning a BERT based language model, comparing
multilingual and monolingual settings. We also
assessed the contribution of additional hate speech
data from different online sources. We finally
submitted the results of the same model fine-tuned
with and without multitask learning between hate
speech and stereotype detection tasks.
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>System Description</title>
      <sec id="sec-2-1">
        <title>Fine-tuning process</title>
        <p>
          The chosen classification approach is to fine-tune
a BERT-based language model. This kind of
approach is the state-of-the-art for many text
classifications tasks today
          <xref ref-type="bibr" rid="ref21 ref23">(Sun et al., 2019; Seganti et
al., 2019)</xref>
          . BERT is a language model which aims
to learn the distribution of language
          <xref ref-type="bibr" rid="ref7">(Devlin et al.,
2018)</xref>
          . It is trained with the prediction of masked
tokens in a text. The next sentence prediction task
that was used simultaneously for training has been
removed for some later BERT-based models such
as RoBERTa
          <xref ref-type="bibr" rid="ref12">(Liu et al., 2019)</xref>
          . BERT is a
Transformer. In a Transformer, the recurrence of
Recurrent Neural Networks is replaced by the
mechanism of attention
          <xref ref-type="bibr" rid="ref24">(Vaswani et al., 2017)</xref>
          .
        </p>
        <p>It has been shown that it is possible to fine-tune
these models for many downstream natural
language processing tasks, including the one we are
interested in, which is text classification. This can
be achieved by removing the language modelling
head and replacing it by a head appropriate for
the target task. The designers of BERT prepared
this by adding a token at the beginning of each
text sequence, named CLS for classification. The
purpose of this token is to contain the information
useful for the classification task at the end of the
forwarding process. Then a classifier head can just
take this CLS token as input to classify the whole
text sequence. In our case we decided to add a
simple linear layer with a softmax on top of it, for
simplicity and because it is efficient enough since
the other layers are fine-tuned.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Layer-wise learning rate</title>
        <p>
          An important consideration of fine-tuning
described in
          <xref ref-type="bibr" rid="ref23">(Sun et al., 2019)</xref>
          is the choice of the
learning rate. Besides being as usual the most
important hyper-parameter in the gradient descent
learning algorithm, it could also be responsible
here for some catastrophic forgetting if it were too
high. Catastrophic forgetting refers to the fact of
erasing the information of the weights of the
pretrained model and can happen when the gradient
updates are too high.
        </p>
        <p>Moreover, the learning rate can be gradually
decreased in the first layers of the models. It aims at
limiting the update in these first layers that have
been showed to contain the most primal
information about the language. One can think of the
classical example in computer vision neural networks
where the basics shapes features are extracted by
the first layers and the task-specific combinations
are processed in the last ones. Thus we applied
layer-wise learning rate with the following
geometric equation: the learning rate in a layer is the
one of the following multiplied by a decay factor
between 0 and 1.</p>
        <p>LRk 1 =</p>
        <p>LRk
where LRk is the learning rate of the k-th layer.</p>
        <p>Then the case when is one is the case of
classic fine-tuning with the same learning rate
everywhere, and the case when is zero is the case of
feature extraction with the whole language model
weights that are frozen and only the parameters of
the classification head are trainable. This
hyperparameter was learned with the others during the
hyper-parameters tuning process.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Monolingual and multilingual language models</title>
        <p>
          We compared the use of several language
models. Many models similar to BERT have been
trained since 2018, and a lot are available for use.
Although the models are often first and foremost
trained for English, multilingual models have been
trained on data of several languages in order to
counteract the lack of data for some languages. It
is the case of mBERT and XLM-Roberta
          <xref ref-type="bibr" rid="ref6">(Conneau et al., 2020)</xref>
          . Also machine learning
researchers trained monolingual models for their
own language, as CamemBERT for French and
AlBERTo or UmBERTo for Italian. Multilingual
models have the advantage that they are trainable
on data in different languages; it is very useful for
low-resources tasks. However, they are expected
to perform in dozens of languages while
monolingual models focus on just one, with the same
number of parameters. For this reason,
monolingual models often perform better when sufficient
data is available, as we show here.
        </p>
        <p>
          We evaluated two multilingual models, mBERT
and XLM-RoBERTa, and three Italian
monolingual models, AlBERTo, UmBERTo, and
PoliBERT. AlBERTo was pretrained on TWITA, that
is a collection of Italian tweets
          <xref ref-type="bibr" rid="ref16 ref17">(Polignano et al.,
2019b)</xref>
          . UmBERTo was pretrained on
Commoncrawl ITA exploiting OSCAR Italian large corpus
          <xref ref-type="bibr" rid="ref15">(Parisi et al., 2020)</xref>
          . Finally, PoliBERT was
finetuned for sentiment analysis on Italian tweets by
its creators
          <xref ref-type="bibr" rid="ref2">(Barone, 2020)</xref>
          .
        </p>
        <p>We tried to use more data, with different
settings. For the multilingual models, we could use
all type of hate speech data. For the monolingual
models, we used the little data available for
Italian but we tried also to use translated multilingual
data. These additions were not conclusive, so we
stuck to the HaSpeeDe 2 data for the submissions.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Random search hyper-parameters tuning</title>
        <p>
          The tuning of the hyper-parameters is relevant in
order to get good results, and that is especially
the case for the learning rate and the layer-wise
decay factor . We tuned hyper-parameters with
random search which has been shown to be
often more efficient than grid-search
          <xref ref-type="bibr" rid="ref4">(Bergstra and
Bengio, 2012)</xref>
          . The hyper-parameters to be tuned
are the batch size, the learning rate, the layer-wise
multiplier and the length of the model (maximum
number of tokens). We did ten trials for each
language model. The number of epochs is selected
with early stopping on the validation macro
F1score with a split of 80/20. Table 1 shows the best
hyper-parameters obtained that have been used for
the systems submitted.
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Hyper-parameter</title>
        <p>Learning rate
Layer-wise
Batch Size
Max Length
Language Model</p>
      </sec>
      <sec id="sec-2-6">
        <title>Value</title>
        <p>2.10-4
0.35</p>
        <p>32
100
UmBERTo</p>
        <p>It is very important that the learning rate and the
layer-wise multiplier are tuned simultaneously
because the choice of the multiplier strongly
modifies the amplitude of the gradient.
2.5</p>
      </sec>
      <sec id="sec-2-7">
        <title>Multitask Learning</title>
        <p>
          We evaluated the usage of multitask learning
between the two classification tasks of the
competition that are hate speech detection and stereotype
detection. Multitask learning consists of learning
to perform several tasks. It can be done by
learning the tasks simultaneously with common first
layers but task-specific heads
          <xref ref-type="bibr" rid="ref18">(Ruder, 2017)</xref>
          . In
our case each task has its own output linear layer.
When the tasks should be based on similar
representations, it is supposed to do a good
regularization with useful shared representations. It is
then a kind of transfer learning. The error
analysis conducted on HaSpeeDe 2018 evaluation
suggests a significant correlation between the usage
of stereotype and hate speech
          <xref ref-type="bibr" rid="ref8">(Francesconi et al.,
2019)</xref>
          . Moreover, they showed that the false
positive rate of hate speech tweets is slightly bigger
for tweets with stereotype.
        </p>
        <p>
          A question that arises when doing multitasking
is the way to combine the loss of the tasks in one.
The simple solution is to sum them uniformly. It
might not be the best solution when there is
imbalance between the tasks, for instance when the scale
of the outputs of one is much higher than the
others. A solution brought by
          <xref ref-type="bibr" rid="ref10">(Kendall et al., 2017)</xref>
          is to use trainable weights based on uncertainty.
          <xref ref-type="bibr" rid="ref11 ref8">(Liebel and Ko¨rner, 2018)</xref>
          upgrades the
regularisation term of this solution and
          <xref ref-type="bibr" rid="ref9">(Gong et al., 2019)</xref>
          shows in a benchmark that this last solution is
often the best. We evaluated this solution and we
compared with the single-task setting.
2.6
        </p>
      </sec>
      <sec id="sec-2-8">
        <title>Cross-validation ensembling and submitted models</title>
        <p>Two submissions are allowed during the
HaSpeeDe 2 test phase. We chose to submit
a fine-tuned UmBERTo trained separately for
each of the two tasks and a fined-tuned UmBERTo
with multitasking on both Stereotype and Hate
Speech detection. The hyper-parameters used to
train these models were presented in Table 1.</p>
        <p>Since we compared the different language
models with 5-fold cross-validation, we then
ensembled the 5 models obtained for each fold in order to
get the final model. The ensembling was done by
considering the mean of the probabilities returned
by each model.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data Description</title>
      <p>The organisers provided a train dataset of 6,839
tweets, annotated with Hate Speech and
Stereotype labels (as described in Table 2).</p>
      <sec id="sec-3-1">
        <title>Dataset HS</title>
        <p>Development Data (Tweets) 0.404
Test Data (Tweets) 0.492
Test Data (News) 0.362</p>
      </sec>
      <sec id="sec-3-2">
        <title>Ster</title>
        <p>0.445
0.450
0.350</p>
        <p>The test data of HaSpeeDe 2 consists of two
subsets: an in-domain set (1,263 tweets) and an
out-of-domain set (500 newspaper headlines).</p>
        <p>The hate speech labels are slightly unbalanced
towards non-hate speech. Thus we tried to use
adapted losses to prevent tendency towards
nonhate speech predictions. We used class-weighted
loss, which assigns a higher weight to the
observations from the minority class in the computing
of the loss. We also tried to use a smoothed
F1score – a differentiable loss in phase with the F1.
Neither approach improved the results in a
significant way.</p>
        <p>The pre-processing was simple. We removed
emoticons and hashtags and we replaced urls and
user names with associated tags as done in the
evaluation data. Each tweet was padded with a
size of 100. Then we used the pre-processing and
tokenization pipeline specific to each language
model as provided by the authors of the models.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <sec id="sec-4-1">
        <title>Macro F1-score</title>
        <p>The metric used for the evaluation is the macro
F1-score. The F1-score of a class is computed by
calculating the harmonic mean between the
precision and recall for this class. The macro F1-score
is the mean between the F1-scores for each class.
It is less sensitive to the imbalance between the
classes.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Baselines</title>
        <p>We used several baselines to evaluate our results
during the development process. The first ones
are those obtained by dummy classifiers, one that
always predicts the most frequent class and the
other one that makes a random stratified
prediction according to the distribution of the classes in
the training data. We also computed the results of
more developed systems, that are a TF-IDF bag of
words and a BiLSTM with trainable word vectors
inputs.</p>
        <p>The HaSpeeDe 2 organisers provided two
baseline systems after the results were submitted. The
first is a most frequent class predictor and the
second is a linear SVM with unigrams, char-grams
and TF-IDF representation.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Validation Results</title>
        <p>We tuned the hyper-parameters for each evaluated
language model as described in Section 2.4. For
each language model, we then computed 5-fold
cross-validation results on HaSpeeDe 2 training
data. The averages of the 5 macro F1-scores are
shown in Table 3.</p>
        <p>HS</p>
      </sec>
      <sec id="sec-4-4">
        <title>System</title>
        <sec id="sec-4-4-1">
          <title>Baselines</title>
          <p>Most Frequent Class 0.374
TF-IDF Bag-of-words 0.703
Word vectors + BiLSTM 0.721</p>
          <p>Multilingual language models
mBERT 0.757
XLM-RoBERTa 0.761</p>
          <p>Italian language models
AlBERTo 0.773
PoliBERT 0.795
UmBERTo 0.799</p>
          <p>Ster
The scores of our two systems evaluated on the
HaSpeeDe 2 test data are summarized in Table 4.
These systems are 5 UmBERTo models trained on
each of the 5 training folds and ensembled. The
second system is the same as the first with the use
of multitask learning.</p>
        </sec>
      </sec>
      <sec id="sec-4-5">
        <title>System Tweets</title>
        <p>Hate Speech Detection
Most Frequent Class 0.337
Classic Features + SVM 0.721
UmBERTo 0.790
UmBERTo + Multitasking 0.809
Best HaSpeeDe 2 0.809</p>
        <p>Stereotype Detection
Most Frequent Class 0.355
Classic Features + SVM 0.715
UmBERTo 0.772
UmBERTo + Multitasking 0.768
Best HaSpeeDe 2 0.772
News
0.389
0.621
0.671
0.660
0.774
0.394
0.669
0.685
0.647
0.720
According to Table 3, multilingual models
performed worse than monolingual models based on
HaSpeeDe 2 data alone, although they achieved
respectable results.</p>
        <p>Moreover, even when we used additional data
from other languages to train the multilingual
models, they still did not manage to outperform
the monolingual models, as we were hoping they
would.</p>
        <p>Within the Italian models, UmBERTo and
PoliBERT performed better than AlBERTo on
these tasks. While the good performance of
PoliBERT can be linked to its pre-training for a tweet
classification task (sentiment analysis) potentially
useful for hate speech detection, it is more
difficult to explain the competitiveness of UmBERTo,
which was trained on data not coming from
Twitter and less numerous than for AlBERTo. One
explanation could be the better quality of this data,
or a better optimisation by its creators.
5.2</p>
      </sec>
      <sec id="sec-4-6">
        <title>Out-of-domain data and in-domain data</title>
        <p>Our results on the HaSpeeDe 2 test dataset are
summarized in the Table 4. The results obtained
on in-domain data correspond to what we
expected from our cross-validation results. Our
systems achieved the best macro F1-scores on the
indomain test set (Tweets) for both hate speech and
stereotype detection. However, the results on
outof-domain data (News) are far from being as good.
This can be explained by the different distribution
of this data compared to the training data.</p>
        <p>Table 5 shows the confusion matrix for our first
system evaluated on out-of-domain data. The
error is mostly due to the high number of false
negatives. The classifier predicts too many sequences
as non-hate speech. This suggests that this
classifier trained with hate speech on Twitter is
struggling to detect hate speech in newspaper headlines.
It can be assumed that hate speech in newspapers
is more subtle, with less coarseness and
aggressiveness that make it easier to detect on Twitter.</p>
        <sec id="sec-4-6-1">
          <title>False True</title>
        </sec>
        <sec id="sec-4-6-2">
          <title>Predicted False</title>
          <p>312
117</p>
          <p>Predicted True</p>
          <p>7
64
We have chosen to submit a system with multitask
learning on both Stereotype and Hate Speech
detection and an other one without, in order to study
the benefits of it. Indeed, the system with
multitasking learning performed much better on the
indomain data for the hate speech detection task. It
is not the case however for the out-of-domain data,
neither for the stereotype detection task.</p>
          <p>Table 6 describes in more detail the differences
between the predictions of the two systems for
data containing stereotypes and data not
containing stereotypes. We observed that the
improvement linked to multitask learning consists mainly
in a reduction in the number of false positives in
favour of the number of true negatives in data not
labeled as Stereotype. Assuming that hate speech
makes significant use of stereotype, one could
suppose that the multitask model has learned to
discard some data that do not have the characteristics
of stereotypes and are therefore unlikely to contain
hate speech.</p>
        </sec>
        <sec id="sec-4-6-3">
          <title>False True</title>
        </sec>
        <sec id="sec-4-6-4">
          <title>False True</title>
        </sec>
      </sec>
      <sec id="sec-4-7">
        <title>Data labeled as Stereotype</title>
        <p>Predicted False Predicted True
+3 -3
+7 -7</p>
      </sec>
      <sec id="sec-4-8">
        <title>Data not labeled as Stereotype</title>
        <p>Predicted False Predicted True
+28 -28
+1 -1
In this work, we compared the fine-tuning of
multilingual and monolingual BERT-based
language models for hate speech detection. We
also investigated the addition of multitask learning
with the Stereotype detection task linked to hate
speech. We obtained the best macro F1-scores of
HaSpeeDe 2 on the in-domain test data. However,
the results were worse for out-of-domain test data,
and further research could be conducted to better
understand the reasons for this and address it.</p>
        <p>AuxCoRR,</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Sai</given-names>
            <surname>Saketh</surname>
          </string-name>
          <string-name>
            <surname>Aluru</surname>
          </string-name>
          , Binny Mathew, Punyajoy Saha, and
          <string-name>
            <given-names>Animesh</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Deep Learning Models for Multilingual Hate Speech Detection</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Gianfranco</given-names>
            <surname>Barone</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Politic BERT based Sentiment Analysis</article-title>
          . https://huggingface.co/ unideeplearning/polibert_sa.
          <source>accessed on Sept 18</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>EVALITA 2020: Overview of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>James</given-names>
            <surname>Bergstra</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Random Search for Hyper-Parameter Optimization</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          ,
          <volume>13</volume>
          :
          <fpage>281</fpage>
          -
          <lpage>305</lpage>
          ,
          <fpage>03</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          , Felice Dell'Orletta,
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Poletto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Tesconi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of the EVALITA 2018 Hate Speech Detection Task</article-title>
          . In EVALITA@CLiC-it.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Alexis</given-names>
            <surname>Conneau</surname>
          </string-name>
          , Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzma´n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Unsupervised Cross-lingual Representation Learning at Scale</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          . CoRR, abs/
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Chiara</given-names>
            <surname>Francesconi</surname>
          </string-name>
          , Cristina Bosco, Fabio Poletto, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Error Analysis in a Hate Speech Detection Task: The Case of HaSpeeDe-TW at EVALITA 2018</article-title>
          .
          <article-title>In CLiC-it</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Ting</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Tyler</given-names>
            <surname>Lee</surname>
          </string-name>
          , Cory Stephenson, Venkata Renduchintala, Suchismita Padhy, Anthony Ndirango, Gokce Keskin, and
          <string-name>
            <given-names>Oguz</given-names>
            <surname>Elibol</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A comparison of loss weighting strategies for multi-task learning in deep neural networks</article-title>
          .
          <source>IEEE Access</source>
          , PP:
          <fpage>1</fpage>
          -
          <lpage>1</lpage>
          ,
          <fpage>09</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Alex</given-names>
            <surname>Kendall</surname>
          </string-name>
          , Yarin Gal, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Cipolla</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Multi-Task Learning Using Uncertainty to weigh Losses for Scene Geometry and Semantics</article-title>
          . CoRR, abs/1705.07115.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Lukas</given-names>
            <surname>Liebel</surname>
          </string-name>
          and Marco Ko¨rner.
          <year>2018</year>
          .
          <article-title>iliary Tasks in Multi-task Learning</article-title>
          . abs/
          <year>1805</year>
          .06334.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Yinhan</given-names>
            <surname>Liu</surname>
          </string-name>
          , Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen,
          <string-name>
            <surname>Omer Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mike</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>RoBERTa: A Robustly Optimized BERT Pretraining Approach</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Merriam-Webster</surname>
          </string-name>
          .
          <year>2020</year>
          . stereotype, noun. https://www.merriam-webster.com/ dictionary/stereotype. Accessed on 2020-
          <volume>11</volume>
          -05.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>John T. Nockleby. 2000. Hate</given-names>
            <surname>Speech</surname>
          </string-name>
          . Macmillan, New York.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Loreto</given-names>
            <surname>Parisi</surname>
          </string-name>
          , Simone Francia, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Magnani</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>UmBERTo: an Italian Language Model trained with whole word Masking</article-title>
          . https://github.com/ musixmatchresearch/umberto. accessed
          <source>on Sept 18</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Polignano</surname>
          </string-name>
          , Pierpaolo Basile, Marco De Gemmis, and
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Semeraro</surname>
          </string-name>
          . 2019a.
          <article-title>Hate Speech Detection through AlBERTo Italian Language Understanding Model</article-title>
          . In NL4AI@AI*IA.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Polignano</surname>
          </string-name>
          , Pierpaolo Basile, Marco De Gemmis, Giovanni Semeraro, and
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          .
          <year>2019b</year>
          .
          <article-title>AlBERTo: Italian BERT Language Understanding Model for NLP Challenging Tasks Based on Tweets</article-title>
          .
          <source>In Proceedings of the Sixth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2019</year>
          ), volume
          <volume>2481</volume>
          . CEUR.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Ruder</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>An Overview of MultiTask Learning in Deep Neural Networks</article-title>
          . CoRR, abs/1706.05098.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Gloria Comandini, Elisa Di Nuovo, Simona Frenda, Marco Stranisci, Cristina Bosco, Tommaso Caselli, Viviana Patti, and
          <string-name>
            <given-names>Irene</given-names>
            <surname>Russo</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Overview of the EVALITA 2020 Second Hate Speech Detection Task (HaSpeeDe 2)</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA 2020), Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Schuster</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.K.</given-names>
            <surname>Paliwal</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Bidirectional recurrent neural networks</article-title>
          .
          <source>Trans. Sig. Proc.</source>
          ,
          <volume>45</volume>
          (
          <issue>11</issue>
          ):
          <fpage>2673</fpage>
          -
          <lpage>2681</lpage>
          , November.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Seganti</surname>
          </string-name>
          , Helena Sobol, Iryna Orlova, Hannam Kim, Jakub Staniszewski, Tymoteusz Krumholc, and
          <string-name>
            <given-names>Krystian</given-names>
            <surname>Koziel</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>NLPR@SRPOL at SemEval-2019 Task 6 and Task 5: Linguistically enhanced deep learning offensive sentence classifier</article-title>
          . In SemEval@NAACL-HLT.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Karen</given-names>
            <surname>Sparck Jones</surname>
          </string-name>
          ,
          <year>1988</year>
          .
          <article-title>A Statistical Interpretation of Term Specificity and Its Application in Retrieval</article-title>
          , page
          <volume>132</volume>
          -
          <fpage>142</fpage>
          . Taylor Graham Publishing, GBR.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Chi</given-names>
            <surname>Sun</surname>
          </string-name>
          , Xipeng Qiu, Yige Xu,
          <string-name>
            <given-names>and Xuanjing</given-names>
            <surname>Huang</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>How to Fine-Tune BERT for Text Classification</article-title>
          ? In Maosong Sun, Xuanjing Huang, Heng Ji, Zhiyuan Liu, and Yang Liu, editors,
          <source>Chinese Computational Linguistics</source>
          , pages
          <fpage>194</fpage>
          -
          <lpage>206</lpage>
          , Cham. Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Ashish</given-names>
            <surname>Vaswani</surname>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,
          <article-title>Ł ukasz Kaiser, and</article-title>
          <string-name>
            <given-names>Illia</given-names>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Attention is All you Need</article-title>
          . In I. Guyon,
          <string-name>
            <given-names>U. V.</given-names>
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          , and R. Garnett, editors,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          , pages
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          . Curran Associates, Inc.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>