<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Svandiela @ HaSpeeDe: Detecting Hate Speech in Italian Twitter Data with BERT</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Svea Klaus Anna-Sophie Bartle Daniela Rossmann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Eberhard Karls Universita ̈t Tu ̈ bingen</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. This paper explains the system developed for the Hate Speech Detection (HaSpeeDe) shared task within the 7th evaluation campaign EVALITA 2020 (Basile et al., 2020). The task solution proposed in this work is based on a fine-tuned BERT model. In cross-corpus evaluation, our model reached an F1 score of 77,56% on the tweets test set, and 60,31% on the news headlines test set.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>The detection of Hate Speech has been a popular
task in Natural Language Processing. Because
there is no universal definition of the term ’hate
speech’, we follow the EVALITA 2018 organizers
in defining it as any expression ”that is abusive,
insulting, intimidating, harassing, and/or incites
to violence, hatred, or discrimination. It is
directed against people on the basis of their race,
ethnic origin, religion, gender, age, physical</p>
      <p>
        Copyright © 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
condition, disability, sexual orientation, political
conviction, and so forth”
        <xref ref-type="bibr" rid="ref6">(Erjavec and Kovacˇicˇ,
2012)</xref>
        .
      </p>
      <p>
        Apart from being hurtful to the person or group
that the hateful message is aimed at, its
systematic usage can be the cause of hate crime and other
criminal acts towards these groups. Mass and
social media help to spread hate speech a lot faster
than traditional communication channels
        <xref ref-type="bibr" rid="ref14">(Sponholz, 2018)</xref>
        . However, social media platforms
like Twitter, YouTube and Facebook lack
systematic control in monitoring and removing hateful
comments. Although these platforms discourage
hateful content, its removal depends on
individual users and trusted reports
        <xref ref-type="bibr" rid="ref6">(Erjavec and Kovacˇicˇ,
2012)</xref>
        , thus indicating that automated detection of
such utterances is a crucial problem to solve. Our
goal within the HaSpeeDe task was to develop a
system for automated detection of hateful
messages against muslims, roma, and immigrants. The
first section introduces related works on the topic.
In the second section, we explain the task setup,
followed by the description of our approach.
Finally, we show our results and discuss them with
regards to possible future work on hate speech
detection.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        In previous work, automated detection of hateful
messages has been approached in various ways,
starting from simpler lexicon-based approaches
and Naive Bayes classifiers to more state of the art
Convolutional Neural Networks
        <xref ref-type="bibr" rid="ref16 ref2 ref4 ref8">(Zhang and Luo,
2018)</xref>
        . The EVALITA 2020 shared task follows
SemEval 2019
        <xref ref-type="bibr" rid="ref10">(May et al., 2019)</xref>
        and EVALITA
2018
        <xref ref-type="bibr" rid="ref3">(Bosco et al., 2018)</xref>
        , where the automated
detection of hateful speech has also been among
the core topics.
      </p>
      <p>
        Early work in this area includes Spertus’
automatic recognition of hostile messages with
the Smokey system. She found that only 12%
of such messages contained explicit keywords.
Therefore, she compiled a set of rules resulting
in a 47-element feature vector per sentence to
capture semantic and syntactic information.
For instance, imperative statements have higher
chances of containing insulting content than
indicative utterances. The same applies to sentences
starting with you. For evaluation, decision trees
were trained on the vectors and the results were
compared to human assessments. Overall, in 36%
of the cases the instances labeled as insulting
matched with the human classification.
        <xref ref-type="bibr" rid="ref13">(Spertus,
1997)</xref>
        .
      </p>
      <p>
        Another approach introduced by Greevy and
Smeaton in 2004 involved support vector
machines for classifying racist texts. In their work,
they compared part-of-speech distributions across
racist and non-racist documents as well as
different feature representations like bag-of words and
bigrams. The bag-of-words model was found to be
more useful than the bigram model (accuracy of
87.77% vs. 84.77%)
        <xref ref-type="bibr" rid="ref7">(Greevy and Smeaton, 2004)</xref>
        .
      </p>
      <p>
        Since around 2015 and with the gaining
popularity of deep learning, various methods
involving neural networks have been proposed.
For instance, Kamble and Joshi compared a CNN,
LSTM, and BiLSTM to one another for detecting
code-mixed Hindi-English hate speech within the
context of ICON 2018. The CNN was fed with
domain-specific embeddings and showed the best
performance (F1 score of 80.85%)
        <xref ref-type="bibr" rid="ref16 ref2 ref4 ref8">(Kamble and
Joshi, 2018)</xref>
        . The growing interest in hate speech
detection is further reflected in other shared
tasks, workshops, and data mining competitions
on Abusive Language, Trolling, Aggression,
Cyberbyullying, Misogyny detection and so forth
        <xref ref-type="bibr" rid="ref16 ref2 ref4 ref8">(Zhang and Luo, 2018)</xref>
        . For the most part, these
models are trained on English text data, paying
little attention to other languages. Therefore,
Italian hate speech detection has been introduced
within the context of EVALITA
        <xref ref-type="bibr" rid="ref11 ref12">(Sanguinetti et
al., 2020a)</xref>
        .
      </p>
      <p>
        In 2018, the EVALITA organizers presented
three subtasks: In the first task, Facebook data
was used to classify a message as not hateful (0)
or hateful (1) and in Task 2, the same challenge
was conducted on Twitter data. Task 3 asked the
participants to train on the Facebook data and
test on the Twitter data, and vice versa. With an
F1 score of 0.82, the best performance on the
Facebook task was achieved by a team that used
polarity and subjectivity lexicons as well as two
word-embedding lexicons as external resources
together with a 2-layer BiLSTM. The same team
reached the best performance for the Twitter data
(F1 score of 0.79). However, systems that were
cross-corpus tested performed significantly worse
with an F1 score of 0.65% with the Facebook
training set and 0.69% with the Twitter train
data. The former score was achieved with a
neural network with three hidden layers involving
word embeddings that were trained on previously
extracted Facebook comments; the latter was
once again achieved by the team with the 2-layer
BiLSTM
        <xref ref-type="bibr" rid="ref4">(Cimino et al., 2018)</xref>
        .
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Task Description and Dataset</title>
      <p>
        We participated in subtask A of HaSpeeDe – a
binary classification task to predict the presence
or absence of hate speech in Italian Twitter
messages
        <xref ref-type="bibr" rid="ref11 ref12">(Sanguinetti et al., 2020b)</xref>
        . The training
dataset provided by the task organizers consists
of 6837 text samples collected from Twitter and
corresponding binary labels: 1 if the text sample
contains hate speech and 0 otherwise. Among the
tweets, 4071 are labeled as not containing hate
speech, 2766 are labeled as hate speech. Table 1
shows two examples with their labels.
      </p>
      <p>
        id
1940
6777
text
Ma quindi solo io sono
preoccupato che il terrorista stava in
Italia?
Cacciamo tutti gli immigrati visto
che sono un pericolo
hs
0
1
To solve the task, we fine-tuned the language
model Bidirectional Encoder Representations
from Transformers (BERT). BERT was developed
by Google and offers great possibilities not only
for hate speech detection, but for all kinds of
tasks that involve processing natural language
        <xref ref-type="bibr" rid="ref5">(Devlin et al., 2019)</xref>
        . Since BERT is available for
multiple languages, we were interested in which
version of BERT – the multilingual BERT
(bertbase-multilingual-cased) or the Italian version of
BERT (dbmdz/bert-base-italian-cased)
        <xref ref-type="bibr" rid="ref15">(Wolf et
al., 2019)</xref>
        – would perform best for the task at
hand to determine Italian hate speech in tweets
and news headlines. The multilingual BERT
cased is a language model that has been trained
on 104 languages whereas the latter version has
been pretrained solely on Italian language.
      </p>
      <p>For faster and more efficient processing while
fine-tuning the model, we used Google
Colab (https://colab.research.google.
com) in all experiments as it provides free GPU.
We further experimented with the training data
by comparing model performance on the data as
it was provided by the event organizers and after
cleaning it. Leaving data as is could have several
advantages: On the one hand, it can be helpful to
leave in junk characters that appear in tweets as
well as trailing white spaces. For instance, a tweet
written in all capital letters might indicate an
insult and therefore contain useful information for
the classifier. On the other hand, the task at hand
did not solely require hate speech detection on
social media but was evaluated on newspaper
articles. Therefore, the model might adapt too much
to the specific style of the Twitter genre and lower
classifier performance when trying to generalize
to another domain (like newspaper articles where
these kinds of characters do not occur). For both
our runs of the final model we cleaned the data as
previous test runs showed better performance.
4.1</p>
      <p>System Description
To solve the task, we fine-tuned a BERT model.
After experimenting with the different language
models as described in the previous section, we
found the bert-base-italian-cased model to be
the best fit. The data was split into training and
validation set during the first phase of the training.
Cross-validation was used on the training set to
prevent overfitting, and the validation set was used
to assess how the model will generalize to unseen
data. In the second training phase, the whole
training data was used for training purposes.</p>
      <p>Before experimenting with different
estimators, the data was cleaned from @user-marks,
trailing whitespaces, and we corrected errors like
”&amp;amp” to ”&amp;”. Since BERT is an already trained
language model, extensive preprocessing of the
data is not unnecessary. However, we assume
that some preprocessing will be useful for
crossdomain evaluation. After preprocessing, the text
data was tokenized by the Italian BERT tokenizer
(AutoTokenizer) that splits texts into tokens. It
adds special [CLS] and [SEP] tokens to mark that
the sentences can now be used for classification
purposes and to separate sentences so that each
token within a sentence can be assigned a segment
token. Afterwards, the tokens are converted
into token ids using the pre-defined indices of
BERT’s tokenizer vocabulary. Additionally, those
embeddings are also assigned attention masks that
specify how much attention the system should pay
to each of the words.</p>
      <p>
        Since we implemented BERT with PyTorch,
we used the optimization module AdamW for
finetuning. Finding a good learning rate can be
difficult. AdamW takes care of this issue by
adapting the learning rates for different
parameters which makes the training process more
efficient
        <xref ref-type="bibr" rid="ref9">(Kingma and Ba, 2015)</xref>
        . Following the
recommendations of the developers of AdamW,
we set the learning rate to 5e-5 as default which
also achieved the best results overall. Moreover,
we tried various epochs, again using the
recommended number of epochs, to see whether the
performance of the model would improve. The
best F1-score and overall accuracy was achieved
with only two epochs. During each epoch the
model is trained and evaluated on the validation
set. The batch size was set to 16 and we set the
random seed to 42 to ensure reproducibility.
      </p>
      <p>Even though we are dealing with binary
classification, the model makes predictions by
calculating probabilities using the softmax function.
Moreover, we used a threshold of 0.9% to reduce
prediction errors; 90% certainty is very high when
we compare the default threshold of 50% that is
typically used for this purpose. However, after
manually going through some of the test data, it is
sometimes fairly difficult even for a human to
uncover hate speech, especially for the news dataset.
Therefore, our goal was to produce realiable
predictions. For both our runs we used the same
system playing around with some of its parameters
according to the results received from the first run.
Therefore, our second run performs slightly better.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>
        When evaluating our model with the two test
sets provided by the EVALITA organizers, we
received the scores shown in Table 2. Our model
performed 17% better on test data containing
Tweets
        <xref ref-type="bibr" rid="ref1">(Basile et al., 2020)</xref>
        compared to the news
data with overall F1 macro-scores of 77.56% (on
tweets) and 60% (on news).
      </p>
      <p>The organizers provided two baseline models
(see Table 3 – most frequent class (MFC) and
Linear SVM with unigrams, char-grams and
TFIDF representation. Our model achieved higher
scores for the news headlines and the twitter test
set compared to the MFC baseline that achieved
Macro-F1 scores of 38.94% and 33.66%
respectively. However, our model failed to beat the
baseline of the Linear SVM for the news test set which
scored 62.1%. Nevertheless, it performed better
on the tweets test set compared to the Linear SVM
(72.12%).</p>
      <sec id="sec-4-1">
        <title>Test Data</title>
      </sec>
      <sec id="sec-4-2">
        <title>News</title>
        <p>Tweets</p>
        <p>non-hate hate
F1 P R F1 P R
0.82 0.70 0.98 0.39 0.25 0.9
0.79 0.75 0.83 0.76 0.81 0.72
As expected, model performance decreases in
cross-corpus evaluation, especially in the news
headlines test data. We assume that our model
learned characteristics of the Twitter data
alongside the characteristics of hate speech. Therefore,
the model performs worse when applied to
domains that entail different linguistic surface
structures. The F1 macro-scores in Table 2 show that
the scores for the two labels are evenly distributed
(79% for non-hate and 76% for hate). Contrary to
this, the model tested on the news data is a lot more
likely to detect non-hate items (with 82%) whereas
its performance on finding hate items only lies at
39%. The confusion matrices for both test sets for
the second run can be seen in Table 4 and Table 5.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Predicted</title>
        <p>Positive Negative
314 5
136 45
la Positive
u
tc Negative</p>
        <p>A
Table 4: Confusion Matrix of news headlines test
set</p>
      </sec>
      <sec id="sec-4-4">
        <title>Predicted</title>
        <p>Positive Negative
534 107
175 447
la Positive
u
tc Negative
A</p>
        <p>Table 5: Confusion Matrix of tweets test set
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Error Analysis</title>
      <p>Identifying hate speech in Twitter data was
obviously easier for our model because it had been
trained on similar data. However, the model had
more difficulties in making predictions on the
news headlines as hints towards hate speech were
much more subtle and harder to grasp. This
became especially clear when we tried to identify
hate speech in the test data ourselves. For the
tweets test data, the use of hate speech was more
obvious and direct. Another and bigger problem
might have been missing context information as
we were limited to the headlines, thereby
missing the content of the article. Since we had
difficulties identifying especially hate speech for the
news headlines test data it is only reasonable that
our model had similar difficulties and performed
worse compared to the tweets test set. Table 6 and
7 show some examples where our system failed to
detect hate speech correctly. Table 6 contains
examples with upper-cased words which are used to
highlight strong ideas and opinions. In this
context, the upper-cased language is used to highlight
the rage of the user. Therefore, our model should
have been made more sensible towards the
intentional use of capital letters to classify content
containing hate speech more accurately. Nevertheless,
none of these examples, including Table 7 were
correctly classified as hate speech.
text
@user A me pare una scelta
politica suicida puntare tutto su una
battaglia sicuramente perdente in
favore dell’immigrazione
incontrollata...Meglio cos`ı, spariranno piu`
velocemente!
Rosarno, le case popolari? Solo agli
immigrati Hanno avuto bisogno di
governi non eletti, di gente imposta ad
un popolo disarmato. Una volta messi
li, i VIGLIACCHI hanno dato inizio
alla ns fine! Se e quando si scatenera`
la rabbia vera, ne faro` parte!!URL
I CRISTIANI ATTACCATI DAL
MONDO ISLAMICO: IRAQ SIRIA
SRI LANKA E ED EUROPA.E LA
CHIESA DIVISA TRA DUE PAPI,
BENEDETTO AUTOREVOLE
RINTUZZA LA RIVOLUZIONE
TRASGRESSIVA DEI COSTUMI,
FRANCESCO LASCIA FARE.</p>
      <p>CRISTIANI PERSEGUITATI MA IL
PROBLEMA SONO I MIGRANTI</p>
      <p>
        URL
Our goal was to develop a system for Hate Speech
Detection in Italian Twitter data. After cleaning
the data, we fine-tuned a BERT model with a batch
size of 16 and a learning rate of 5e-5. Overall, our
model reached an F1 score of 77.56% on the
Twitter test data, and 60% on the news data. Ideas for
future work include adding training data that has
been collected from other sources apart from
Twitter, incorporating a lexicon of hate words, such as
Hurtlex
        <xref ref-type="bibr" rid="ref2">(Bassignana et al., 2018)</xref>
        , or using topic
modelling techniques to extract information about
topics that are likely to be involved in hate speech
on social media.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>EVALITA 2020: Overview of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Bassignana</surname>
          </string-name>
          , Valerio Basile, and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hurtlex: A Multilingual Lexicon of Words to Hurt</article-title>
          . In Elena Cabrio, Alessandro Mazzei, and Fabio Tamburini, editors,
          <source>Proceedings of the Fifth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2018</year>
          ), Torino, Italy,
          <source>December 10-12</source>
          ,
          <year>2018</year>
          , volume
          <volume>2253</volume>
          <source>of CEUR Workshop Proceedings. CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Christina</given-names>
            <surname>Bosco</surname>
          </string-name>
          , Felice Dell'Orletta, Fabio Poletto, Manuela Sanguinetti, and
          <string-name>
            <given-names>Maurizio</given-names>
            <surname>Tesconi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of the EVALITA 2018 Hate Speech Detection Task</article-title>
          . EVALITA@
          <article-title>CLiC-it</article-title>
          , pages
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Cimino</surname>
          </string-name>
          , Lorenzo De Mattei, and Felice Dell'Orletta.
          <year>2018</year>
          .
          <article-title>Multi-task learning in deep neural networks at EVALITA 2018</article-title>
          . In Tommaso Caselli, Nicole Novielli, Viviana Patti, and Paolo Rosso, editors,
          <source>Proceedings of the Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          )
          <article-title>co-located with the Fifth Italian Conference on Computational Linguistics (CLiC-it</article-title>
          <year>2018</year>
          ), Turin, Italy,
          <source>December 12-13</source>
          ,
          <year>2018</year>
          , volume
          <volume>2263</volume>
          <source>of CEUR Workshop Proceedings. CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>Proceedings of NAACL-HLT</source>
          , pages
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Karmen</given-names>
            <surname>Erjavec</surname>
          </string-name>
          and Melita Poler Kovacˇicˇ.
          <year>2012</year>
          . ”You
          <string-name>
            <surname>Don't Understand</surname>
          </string-name>
          ,
          <article-title>This Is a New War!” Analysis of Hate Speech in News Web Sites' Comments</article-title>
          . Mass Communication and Society.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Edel</given-names>
            <surname>Greevy and Alan F. Smeaton</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Classifying racist texts using a support vector machine</article-title>
          .
          <source>In SIGIR 2004 - the 27th Annual International ACM SIGIR Conference</source>
          ,
          <volume>25</volume>
          -
          <issue>29</issue>
          <year>July 2004</year>
          , Sheffield, UK.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Satyajit</given-names>
            <surname>Kamble</surname>
          </string-name>
          and
          <string-name>
            <given-names>Aditya</given-names>
            <surname>Joshi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hate speech detection from code-mixed hindi-english tweets using deep learning models</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Diederik P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jimmy</given-names>
            <surname>Ba</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Adam: A Method for Stochastic Optimization</article-title>
          . CoRR, abs/1412.6980.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Jonathan</given-names>
            <surname>May</surname>
          </string-name>
          , Ekaterina Shutova, Aurelie Herbelot, Xiaodan Zhu, Marianna Apidianaki, and Saif M. Mohammad, editors.
          <source>2019. Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          , Minneapolis, Minnesota, USA, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Gloria Comandini, Elisa Di Nuovo, Simona Frenda, Marco Stranisci, Cristina Bosco, Tommaso Caselli, Viviana Patti, and
          <string-name>
            <given-names>Irene</given-names>
            <surname>Russo</surname>
          </string-name>
          .
          <year>2020a</year>
          .
          <article-title>HaSpeeDe 2@EVALITA2020: Overview of the EVALITA 2020 Hate Speech Detection Task</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Gloria Comandini, Elisa Di Nuovo, Simona Frenda, Marco Stranisci, Cristina Bosco, Tommaso Caselli, Viviana Patti, and
          <string-name>
            <given-names>Irene</given-names>
            <surname>Russo</surname>
          </string-name>
          . 2020b.
          <article-title>Overview of the EVALITA 2020 Hate Speech Detection (HaSpeeDe 2) Task</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA 2020), Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Ellen</given-names>
            <surname>Spertus</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Smokey: Automatic Recognition of Hostile Messages</article-title>
          .
          <source>In Proceedings of the Fourteenth National Conference on Artificial Intelligence and Ninth Conference on Innovative Applications of Artificial Intelligence</source>
          , pages
          <fpage>1058</fpage>
          -
          <lpage>1065</lpage>
          . AAAI Press.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Liriam</given-names>
            <surname>Sponholz</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hate Speech in den Massenmedien: Theoretische Grundlagen und empirische Umsetzung</article-title>
          .
          <source>VS Verlag fu¨r Sozialwissenschaften.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Wolf</surname>
          </string-name>
          , Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Re´mi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and
          <string-name>
            <surname>Alexander</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Rush</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>HuggingFace's Transformers: State-of-the-art Natural Language Processing</article-title>
          . ArXiv, abs/
          <year>1910</year>
          .03771.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Ziqi</given-names>
            <surname>Zhang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lei</given-names>
            <surname>Luo</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hate speech detection: A solved problem? the challenging case of long tail on twitter</article-title>
          .
          <source>Semantic Web</source>
          ,
          <volume>10</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>