<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>HaSpeeDe 2 @ EVALITA2020: Overview of the EVALITA 2020 Hate Speech Detection Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Manuela Sanguinetti?</string-name>
          <email>manuela.sanguinetti@unica.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gloria Comandini</string-name>
          <email>gloria.comandini@unitn.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elisa di Nuovo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simona Frenda</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Stranisci</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cristina Bosco</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tommaso Caselli</string-name>
          <email>t.caselli@rug.nl</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Viviana Patti</string-name>
          <email>pattig@di.unito.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Irene Russoy</string-name>
          <email>irene.russo@ilc.cnr.it</email>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>The Hate Speech Detection (HaSpeeDe 2) task is the second edition of a shared task on the detection of hateful content in Italian Twitter messages. HaSpeeDe 2 is composed of a Main task (hate speech detection) and two Pilot tasks, (stereotype and nominal utterance detection). Systems were challenged along two dimensions: (i) time, with test data coming from a different time period than the training data, and (ii) domain, with test data coming from the news domain (i.e., news headlines). Overall, 14 teams participated in the Main task, the best systems achieved a macro F1-score of 0.8088 and 0.7744 on the indomain in the out-of-domain test sets, respectively; 6 teams submitted their results for Pilot task 1 (stereotype detection), the best systems achieved a macro F1-score of 0.7719 and 0.7203 on in-domain and outof-domain test sets. We did not receive any submission for Pilot task 2.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction and Motivations</title>
      <p>
        From a NLP perspective, much attention has been
paid to the automatic detection of Hate Speech
(HS) and related phenomena (e.g., offensive or
abusive language among others) and behaviors
(e.g., harassment and cyberbullying). This has led
to the recent proliferation of contributions on this
topic
        <xref ref-type="bibr" rid="ref21 ref29">(Nobata et al., 2016; Waseem et al., 2017;
Fortuna et al., 2019)</xref>
        , corpora and lexica1,
ded
      </p>
      <p>Copyright © 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).</p>
      <p>1More details and an overview of available HS resources
have been recently presented in Poletto et al. (2020).
icated workshops2, and shared tasks within
national3 and international4 evaluation campaigns.</p>
      <p>
        As for Italian, the first edition of HaSpeeDe
        <xref ref-type="bibr" rid="ref12 ref27 ref6">(Bosco et al., 2018)</xref>
        , a task specifically focused
on HS detection, was proposed at EVALITA
2018
        <xref ref-type="bibr" rid="ref8">(Caselli et al., 2018)</xref>
        . The task consisted of
the binary classification (HS vs not-HS) of texts
from Twitter and Facebook. For each social
media platform, training and test data were provided.
Furthermore, two cross-platform sub-tasks were
introduced to test the systems’ ability to generalize
across platforms.
      </p>
      <p>
        The ultimate goal of HaSpeeDe 2 at EVALITA
2020
        <xref ref-type="bibr" rid="ref24 ref3">(Basile et al., 2020)</xref>
        is to take a step further
in state-of-the-art HS detection for Italian. By
doing this, we also intend to explore other side
phenomena and see the extent to which they can be
automatically distinguished from HS.
      </p>
      <p>
        We propose a single training set made of tweets,
but two separate test sets within two different
domains: tweets and news headlines. While social
media are still one of the main channels used to
spread hateful content online
        <xref ref-type="bibr" rid="ref1 ref30">(Alkiviadou, 2019;
Wodak, 2018)</xref>
        , an important role in this respect is
also played by traditional media, and newspapers
in particular.
      </p>
      <p>
        Furthermore, we chose to include another
HSrelated phenomenon, namely the presence of
stereotypes referring to one of the targets
identified within our dataset (i.e., muslims, Roma and
immigrants). With the term stereotype we mean
any explicit or implicit reference to typical beliefs
and attitudes about a given target
        <xref ref-type="bibr" rid="ref27">(Sanguinetti et
al., 2018)</xref>
        . An error analysis of the main systems
on the HaSpeeDe 2018 dataset itself (Francesconi
2More detailed informations in: https://www.
workshopononlineabuse.com/
      </p>
      <p>
        3HASOC
        <xref ref-type="bibr" rid="ref19">(Mandl et al., 2019)</xref>
        , Poleval
        <xref ref-type="bibr" rid="ref25">(Ptaszynski et al.,
2019)</xref>
        or VLSD
        <xref ref-type="bibr" rid="ref28">(Vu et al., 2019)</xref>
        .
      </p>
      <p>
        4Hateval task at Semeval 2019
        <xref ref-type="bibr" rid="ref2 ref7">(Basile et al., 2019)</xref>
        .
et al., 2019) showed that the occurrence of these
elements constitutes a common source of error in
HS identification.
      </p>
      <p>
        Finally, it has been observed that in social media
and newspapers’ headlines, the most hateful parts
are often verbless sentences or a verbless
fragments, also known as Nominal Utterances (NUs)
        <xref ref-type="bibr" rid="ref11">(Comandini et al., 2018)</xref>
        . The relevant presence of
NUs has been investigated in the POP-HS-IT
corpus
        <xref ref-type="bibr" rid="ref10 ref12 ref25">(Comandini and Patti, 2019)</xref>
        . In order to have
a better understanding of the syntactic strategies
used in HS, we include the recognition of NUs in
hateful tweets and news headlines.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Task Description</title>
      <p>HaSpeeDe 25 consists of a Main task and two Pilot
tasks and is based on two datasets, one containing
messages from a social media platform, namely
Twitter, and the other one news headlines. The
three tasks are shortly described as follows:
• Task A - Hate Speech Detection (Main
Task): binary classification task aimed at
determining the presence or the absence of
hateful content in the text towards a given target
(among immigrants, Muslims and Roma)
• Task B - Stereotype Detection (Pilot Task
1): binary classification task aimed at
determining the presence or the absence of a
stereotype towards the same targets as Task
A
• Task C - Identification of Nominal
Utterances (Pilot Task 2): sequence labeling task
aimed at recognizing NUs in data previously
labeled as hateful.</p>
      <p>This edition of the task presents several
distinguishing features with respect to the first one.
Besides including new and more-richly annotated
data, news headlines were introduced as
crossdomain test data. Furthermore, two additional
tasks are proposed. Finally, the Twitter test set
intentionally contains tweets published in a different
time frame than those in the training set to verify
the systems’ ability to detect HS forms
independently of biases. These biases result from
contextrelated features, such as events – regarding one of
our HS targets – that can be controversial or be
subject to heated and polarized debates.</p>
      <p>5Task repository:
https://github.com/msang/haspeede/tree/
master/2020.</p>
    </sec>
    <sec id="sec-3">
      <title>Datasets and Formats</title>
      <p>In this section we describe the datasets and
formats used in the three tasks.
3.1</p>
      <sec id="sec-3-1">
        <title>Twitter Dataset</title>
        <p>
          Task A: The Twitter portion of the data of
HaSpeeDe 2018 was included in the training set
(4,000 tweets posted from October 2016 to April
2017). Moreover, new Twitter data were included
for this competition, a subset of the data
gathered for the Italian hate speech monitoring project
“Contro l’Odio”
          <xref ref-type="bibr" rid="ref7">(Capozzi et al., 2019)</xref>
          . The data
were retrieved using the Twitter Stream API and
filtered using the set of keywords described in
Poletto et al. (2017). The newly annotated tweets
were posted between September 2018 and May
2019 and were annotated by Figure Eight (now
Appen) contributors for hate speech and by the
task organizers for the stereotype category. In
particular, only data posted between January and May
2019 were included in the test set.
        </p>
        <p>Task B: The HaSpeeDe Twitter corpus – used
in the first edition of the task – was already
annotated for stereotype since it was part of the Italian
Hate Speech corpus described in Sanguinetti et al.
(2018). We then used the same guidelines to
enrich the new data from “Contro l’Odio” with this
annotation layer. The annotation was carried out
by the task organizers.</p>
        <p>
          Task C: The HaSpeeDe Twitter corpus was also
annotated for the presence of Nominal Utterances
(NUs) within a side project
          <xref ref-type="bibr" rid="ref10 ref12 ref25">(Comandini and Patti,
2019)</xref>
          . We used an updated version of its
guidelines (available in the task repository) to enrich
the new hateful data introduced in the campaign.
Similarly to the stereotype level, the annotation of
NUs was carried out by the task organizers
specifically for this task’s purposes.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>News Dataset</title>
        <p>Task A: For task A a new test corpus
composed of newspapers’ headlines about immigrants
was made available. The data were retrieved
between October 2017 and February 2018 from
online newspapers (La Stampa, La Repubblica, Il
Giornale, Liberoquotidiano) and annotated within
the context of a Master’s degree thesis discussed in
2018 at the Department of Foreign Languages at
the University of Turin. Data annotation includes
the same categories annotated in the Twitter
corpus.</p>
        <p>Task B: The News corpus also includes
stereotype annotation, performed according to the same
guidelines used for developing the Twitter corpus.
Task C: Similarly to the Twitter dataset, the
third annotation level was added in the News
corpus from scratch and specifically for the present
task.</p>
        <p>Tables 1, 2 and 3 show the data distribution for
each task.</p>
        <p>TASK A
Train
Test Tweets
Test News</p>
        <p>HS</p>
        <p>NOT HS
4073
641
319</p>
        <p>TOT.
6839
1263
500
The whole dataset consists of 8,012 tweets and
500 news headlines for Task A and B, and 3,388
tweets and 181 news (i.e., the sub-set with hateful
data only) for Task C.</p>
        <p>In Task A and B, HS and stereotype represent
the 41.8% and 44.6%, respectively, of the
Twitter dataset. In contrast, in the News dataset, the
portion of hateful content and stereotype lowers to
36% and 35%.</p>
        <p>Table 3 shows statistics about the total number of
texts with or without NUs in Task C. The
percentage of hateful tweets featuring at least one NU is
57.4%; the percentage of news headlines having at
least one NU is 83.4%. This distribution is in line
with the one found in Comandini and Patti (2019).
Task A and B: For both tasks A and B data are
provided in a tab-separated values (TSV) file
including ID, text, HS and stereotype class (0 or 1).
Mentions and URLs were replaced with @user
and URL placeholders. Table 4 shows some
annotation examples.</p>
        <p>Task C: The dataset provided for Task C was
annotated using WebAnno and converted into a IOB
(Inside-Outside-Beginning) format. The resulting
IOB2 alphabet consists of I-NU-CGA, O and
BNU-CGA.</p>
        <p>The annotation includes the ID, followed by an
hyphen to mark the token number, the token, and the
IOB2 annotation of the NUs.</p>
        <p>Below an example taken from the training set.
#Text= E` UNA PROVOCAZIONE...ORA BASTA..
NESSUNO SBARCHI IN #ITALIA6
9602-23 E`
9602-24 UNA
9602-25 PROVOCAZIONE
9602-26 .
9602-27 .
9602-28 .
9602-29 ORA
9602-30 BASTA
9602-31 .
9602-32 .
9602-33 NESSUNO
9602-34 SBARCHI
9602-35 IN
9602-36 #
9602-37 ITALIA
O
O
O
O
O
O
B-NU-CGA
I-NU-CGA
I-NU-CGA
I-NU-CGA
O
O
O
O
O
To prevent participants from cheating, the released
test set for Task C also contains non-hateful
messages. However, the evaluation of the systems is
conducted only on the hateful messages since we
are interested in investigating the relationship
between these two phenomena.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>For each task, participants were allowed to submit
up to 2 runs. A separate official ranking was
provided, and the evaluation was performed
according to the standard metrics, i.e, Precision, Recall
and F-score.</p>
      <p>For Task A and Task B, the scores were computed
for each class separately, and finally the F-score
was macro-averaged to get the overall results.</p>
      <p>6“IT’S A PROVOCATION...THAT’S ENOUGH...NO
LANDINGS IN #ITALY”
1
1
0
0
1
0
1
0</p>
      <p>For Task C, token-wise scores were computed,
and a NU was considered correct only in case of
exact match, i.e., if all tokens that compose it were
correctly identified.</p>
      <p>
        Different baseline systems were built according
to the task type:
• For Task A and B, besides a typical
classifier based on the most frequent class
(Baseline MFC in Tables 5–8), a Linear SVM with
TF-IDF of unigrams and 2–5 char-grams was
used (Baseline SVC).
• For Task C, the baseline replicates the one
presented for the COSMIANU corpus
        <xref ref-type="bibr" rid="ref11">(Comandini et al., 2018)</xref>
        , which identifies as
correct in the test the NUs that appear in the
training set (memory-based approach);
baseline results in Table 9.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Task Overview: Participation and</title>
    </sec>
    <sec id="sec-6">
      <title>Results</title>
      <sec id="sec-6-1">
        <title>5.1 Participants</title>
        <p>
          A total amount of 14 teams participated in the
Main task on HS detection, 6 teams also
submitted their results for the Pilot task 1 (i.e. Task B)
on stereotype detection, while we did not receive
any submission for the Pilot task 2 (i.e. Task C)
on NUs identification. Except for one case, all
teams submitted 2 runs for their tasks.
Furthermore, 4 teams used the same systems to
participate in other (and partly related) tasks within the
EVALITA 2020 campaign: YNU OXZ and
Jigsaw participated in the task on Automatic
Misogyny Identification (AMI) (Fersini et al., 2020),
while TextWiller and Venses also participated in
the task on Stance Detection in Italian Tweets
(SardiStance)
          <xref ref-type="bibr" rid="ref9">(Cignarella et al., 2020)</xref>
          . It is worth
pointing out that in this second edition we
registered a higher participation of non-Italian and
nonacademic teams, and that HaSpeeDe 2 has been
one of the most participated EVALITA 2020 tasks.
5.2
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>Systems Overview</title>
        <p>
          Approaches The participating models are
characterized by different architectures that exploit
principally BERT-based models and linguistic
features. Transformers are a popular choice in this
edition. Jigsaw
          <xref ref-type="bibr" rid="ref17">(Lees et al., 2020)</xref>
          , Svandiela
          <xref ref-type="bibr" rid="ref15">(Klaus et al., 2020)</xref>
          , DH-FBK
          <xref ref-type="bibr" rid="ref18">(Leonardelli et al.,
2020)</xref>
          , TheNorth
          <xref ref-type="bibr" rid="ref16">(Lavergne et al., 2020)</xref>
          finetuned BERT, AlBERTo7 and UmBERTo8
language models for both runs. YNU OXZ
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref17 ref18 ref26 ref5">(Ou
and Li, 2020)</xref>
          exploited the pre-trained
XLMRoBERTa9 multi-language model as input of
Neural Networks architecture. Fontana-Unipi
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref17 ref18 ref26 ref5">(Fontana and Attardi, 2020)</xref>
          developed a model
that is an ensemble of fixed number of instances
of two principal transformers (AlBERTo and
DBMDZ10) and a combination of DBMDZ input and
a dense layer. The DBMDZ is used also by
By1510 (Deng et al., 2020) in a transfer learning
approach. UO team
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref17 ref18 ref26 ref5">(Rodriguez Cisnero and
Ortega Bueno, 2020)</xref>
          , on the other hand, used a
BiLSTM with the addition of linguistic features in
7https://github.com/marcopoli/
AlBERTo-it
        </p>
        <p>8https://github.com/
musixmatchresearch/umberto</p>
        <p>9https://huggingface.co/transformers/
model_doc/xlmroberta.html</p>
        <p>
          10https://huggingface.co/dbmdz/
bert-base-italian-uncased
the first run, while using the pre-trained DBMDZ
model in the second one. CHILab
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref17 ref18 ref26 ref5">(Gambino and
Pirrone, 2020)</xref>
          experimented transformer encoders
in the first run and depth-wise Separable
Convolution techniques in the second one. Moreover,
some teams explored classical machine learning
approaches such as No Place For Hate Speech
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref17 ref18 ref26 ref5">(dos
S. R. da Silva and T. Roman, 2020)</xref>
          , TextWiller
(Ferraccioli et al., 2020), UR NLP
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref17 ref18 ref26 ref5">(Hoffmann and
Kruschwitz, 2020)</xref>
          and Montanti
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref17 ref18 ref26 ref5">(Bisconti and
Montagnani, 2020)</xref>
          . Finally, Venses (Delmonte,
2020), based on the parser for Italian ItGetaruns,
applied six different rule-based classifiers.
        </p>
      </sec>
      <sec id="sec-6-3">
        <title>Features and Lexical Resources Various fea</title>
        <p>
          tures are tested and explored by participants.
Morphosyntactic features are exploited by
CHILab, using Part-of-Speech tags as additional
input. To adapt the POS tagging model provided by
Python’s spaCy library to social media language,
they added emoticons, emojis, hashtags and URLs
to vocabulary. In addition, to preprocess the texts,
they used sentiment lexicon for replacing
emoticons with appropriate labels about the expressed
sentiment. Semantic and lexical features are
exploited by Venses and UO teams. In particular,
UO team used WordNet to catch lexical
ambiguity, syntactic patterns and similarity among words;
calculated information gain to capture the most
relevant words; used lexicons such as HurtLex
          <xref ref-type="bibr" rid="ref4">(Bassignana et al., 2018)</xref>
          and SenticNet11 to
feature words with hateful categories and sentiment
information. Finally, different types of
representation of tweets are tested by Montanti: TF-IDF,
DistilBert12 and GloVe
          <xref ref-type="bibr" rid="ref22">(Pennington et al., 2014)</xref>
          vectors as well as their combination.
        </p>
        <p>Additional data Some teams preferred to use
additional data to improve the knowledge of their
classifiers. To extend the provided training set,
YNU OXZ exploited Facebook data provided in
the first edition of HaSpeeDe and DH-FBK used
a set of Italian tweets that covers similar topics.
Jigsaw, for one of the submissions, used
additional user-generated comments to fine-tune their
model. CHILab used additional tweets taken from
TWITA 201813 by means of some keywords
extracted from the provided training set to extend the
11https://www.sentic.net/
12https://huggingface.co/transformers/
model_doc/distilbert.html
13http://twita.di.unito.it/
embedding layer of their model. Finally, the
SENTIPOLC 2016 dataset was exploited by UO team.</p>
      </sec>
      <sec id="sec-6-4">
        <title>Interaction between Task A and B Except for</title>
        <p>TheNorth team, most of the participants did not
consider the interaction between Task A and B.
Taking into account the possible correlation
between texts containing hate speech and texts
expressing stereotyped ideas about targets, TheNorth
tested the performance of multitasking approach
for both tasks (second run) against a fine-tuned
UmBERTo model (first run). In particular,
observing competition results we can notice the efficacy
of multitasking in hate speech identification and
not in stereotype detection.
5.3</p>
      </sec>
      <sec id="sec-6-5">
        <title>Results</title>
        <p>In Table 5, 6, 7 and 8, we report the official results
of HaSpeeDe 2 for Task A and B, ranked by the
macro-F1 score. In case of multiple runs, a suffix
has been appended to each team name, in order to
distinguish the run ID of the submitted file.</p>
        <p>Team
TheNorth 2
TheNorth 1
CHILab 1
Fontana-Unipi
CHILab 2
By1510 1
Svandiela 2
YNU OXZ 1
Jigsaw al
UR NLP 2
DHFBK 2
DHFBK 1
No Place For Hate Speech STT
Svandiela 1
Montanti 1
UR NLP 1
YNU OXZ 2
Montanti 2
UO 2
Baseline SVC
Jigsaw js
By1510 2
No Place For Hate Speech LRT
UO 1
Venses 1
Venses 2
TextWiller 1
Baseline MFC
TextWiller 2
Macro-F1
0.8088
0.7897
0.7893
0.7803
0.7782
0.7766
0.7756
0.7717
0.7681
0.7598
0.7534
0.7495
0.7491
0.7452
0.7432
0.7399
0.7345
0.7279
0.7214
0.7212
0.717
0.7065
0.7057
0.6878
0.5054
0.4726
0.3604
0.3366
0.3317</p>
        <p>As a general remark, we can observe that the
indomain Main task registered better results
(macroF1=0:8088) both compared to the cross-domain
counter-part (0:7744) and the Pilot task 1; in turn,
CHILab 1
UO 2
Montanti 1
CHILab 2
DHFBK 2
UR NLP 2
YNU OXZ 2
Montanti 2
Jigsaw js
DHFBK 1
TheNorth 1
UR NLP 1
UO 1
By1510 2
YNU OXZ 1
TheNorth 2
Fontana-Unipi
Jigsaw al
No Place For Hate Speech STN
No Place For Hate Speech LRN
Baseline SVC
By1510 1
Svandiela 2
Svandiela 1
Venses 1
Baseline MFC
Venses 2
TextWiller 1
TextWiller 2
Macro-F1
better results were obtained in the latter with the
in-domain data compared to the News set (0:7744
and 0:7203, respectively). The best performances
overall provided by the systems used for Task A on
Twitter data is also reflected in the average value
of the macro-F1 scores of each ranking: 0:6899
for the latter, 0:6306 for Task B on Twitter data,
0:6144 for Task A on News data and 0:5972 for
Task B on News data.</p>
        <p>We also considered the overall results achieved by
all participating teams and observed that, as
regards Task A, 12 and 13 teams (in the Twitter
and News test set, respectively) obtained higher
scores than the SVM-based baseline with at least
one of the submitted runs, and 13 teams, on both
domains, outperformed the one based on the most
frequent class. For Task B, and with respect to the
SVM baseline, the same is true for 4 teams out of
6 in the Twitter set and for 3 teams in the News
set, while all teams beat the majority-class
baseline with at least one run.</p>
        <p>Regarding Task C, since the training set is
composed of tweets, we first investigated the macro
F-score value on a validation set created by
splitting the training set in 80%-20%. We then tested
the memory-based baseline described in Section
TheNorth 1
TheNorth 2
CHILab 1
Jigsaw al
CHILab 2
Baseline SVC
Montanti 1
Montanti 2
Jigsaw js
TextWiller 2
Venses 1
Venses 2
Baseline MFC
TextWiller 1
Team
CHILab 1
CHILab 2
Montanti 1
TheNorth 1
Jigsaw al
Montanti 2
Baseline SVC
TheNorth 2
Jigsaw js
TextWiller 2
Venses 1
Baseline MFC
Venses 2
TextWiller 1
4 on the two test sets released for the task.
Table 9 shows the macro-F1 values obtained in the
validation set, in the Twitter test set as well as in
the News test set. As mentioned earlier, no
submissions were made for this task, but the
baselines’ values for both domains are reported in this
overview as reference points for further works.</p>
        <p>Baseline
Baseline validation
Baseline test Tweets
Baseline test News</p>
        <p>Macro-F
0.1459
0.0706
0.0087
A discussion of results, especially those
regarding the Main task, necessarily involves a
preliminary comparison with the ones obtained in the
first edition of HaSpeeDe, in particular in the two
tasks where Twitter data were used for training,
i.e. HaSpeeDe TW and Cross-HaSpeeDe TW.
The best systems attained macro-F1=0:7993 in the
former task and 0:6985 in the latter. While these
results are in line with those reported for Task A
on the in-domain data, the results obtained in this
edition on News data are better than the part
crossdomain task, where the test set was made up of
Facebook comments. We hypothesize that the
homogeneity of hate target in News and Twitter
corpora (immigrants) has meant more than the similar
linguistic features in Twitter and Facebook data,
stemming from the fact that they are both social
media texts.</p>
        <p>
          Participants achieved promising results in the
detection of stereotypes, a new pilot task proposed
at HaSpeeDe this year for the first time. In our
view, stereotype and HS are meant as
orthogonal dimensions of abusive language, which do not
necessarily coexist. This influenced the design of
HaSpeeDe 2, where we proposed two independent
tasks for the detection of such categories.
However, a first analysis of systems participating in
both tasks suggests that most teams did not
design a dedicated system for stereotype recognition,
but focused on developing a HS detection model,
adapting the same model to stereotype
recognition, reducing de facto stereotypes to
characteristics of HS. We hypothesize that this could be one
of the factors that led the systems to not
generalize well when applied to the stereotype
detection task, especially in texts that are not hateful
but contain stereotypes. This hypothesis is
confirmed by the high percentage of false negatives
(21% in tweets and 35% in news headlines) of
the stereotype class in non-hateful texts, with
respect to false negatives (5% in tweets and 28% in
news headlines) in hateful ones. It is possible to
notice the same increase also in false positives in
hateful texts. These values suggest that stereotype
appears as a more subtle phenomenon that could
not give rise to hurtful message. The percentages
have been computed taking into account the set
of common incorrect predictions of the three best
runs in Task B, and calculated in relation to the
actual distribution of HS and stereotype in the test
set. Analyzing the predictions of the three best
runs in Task A, similar influence of stereotype is
observed in false negative and positive, but to a
minor extent. These results are in line with the
observations about emerged from the error analysis
of HaSpeeDe 2018
          <xref ref-type="bibr" rid="ref12">(Francesconi et al., 2019)</xref>
          .
        </p>
        <p>To conclude the discussion on this edition’s
results, we comment on the baseline scores obtained
for Task C. As it can be noticed from Table 9, the
value obtained on the validation set is higher than
the ones obtained on both test sets. This variation
can be explained by the main characteristics of the
data at hand: on the Twitter side, this is due to
the different time frames of tweet’s publication
included in training and test set, while on the News
side, such low value is expected by virtue of the
different text domain. Since this baseline uses a
memory-based approach, such a low performance
is to be expected in datasets from different time
frames, since the discussion topics are different
and Twitter users change their hashtags and
slogans, which are the main repeated items.
7</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>In its second edition, the HaSpeeDe task proposed
the detection of hateful content in Italian, by
challenging systems along two dimensions, time and
domain, and taking into account also the category
of stereotype, which often co-occurs with HS. This
paves the way for further investigations also about
the relationships linking stereotype and HS.</p>
      <p>
        In order to take a step further in
state-of-theart HS detection, the task provided novel
benchmarks for exploring different facets of the
phenomenon and laying the foundations for deeper
studies about the impact of bias, topic and text
domain. In this line, also a pilot task about
recognition of NUs was proposed, devoted to study
this kind of linguistic form in hateful messages
in tweets and newspaper headlines, as it has been
proved that both headlines in journalistic writings
        <xref ref-type="bibr" rid="ref20">(Mortara Garavelli, 1971)</xref>
        and social media texts
        <xref ref-type="bibr" rid="ref11">(Ferrari, 2011; Comandini et al., 2018)</xref>
        are a fertile
ground for NUs. Even though we did not receive
any submission for Pilot task 2, our hope is that the
fine-grained annotation of hateful data concerning
these aspects can be the subject of deeper studies
to shed light on the syntax of hate, a topic still
understudied.
      </p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>
        The work of Cristina Bosco, Simona Frenda,
Viviana Patti and Marco Stranisci is partially
funded by Progetto di Ateneo/CSP 2016
(Immigrants, Hate and Prejudice in Social Media,
S1618.L2.BOSC.01) and by the project “Be
Positive!”
        <xref ref-type="bibr" rid="ref1">(under the 2019 “Google.org Impact
Challenge on Safety” call)</xref>
        .
      </p>
      <p>Tao Deng, Yang Bai, and Hongbing Dai. 2020.</p>
      <p>By1510 @ HaSpeeDe 2: Identification of Hate
Speech for Italian Language in Social Media Data.
In Proceedings of the 7th Evaluation Campaign of
Natural Language Processing and Speech Tools for
Italian (EVALITA 2020).</p>
      <p>Adriano dos S. R. da Silva and Norton T. Roman. 2020.</p>
      <p>No Place For Hate Speech @ HaSpeeDe 2:
Ensemble to identify hate speech in Italian. In
Proceedings of the 7th Evaluation Campaign of Natural
Language Processing and Speech Tools for Italian
(EVALITA 2020).</p>
      <p>Federico Ferraccioli, Andrea Sciandra, Mattia Da Pont,
Paolo Girardi, Dario Solari, and Livio Finos. 2020.
TextWiller @ SardiStance, HaSpeede2: Text or
Con-text? A smart use of social network data in
predicting polarization. In Proceedings of the 7th
Evaluation Campaign of Natural Language Processing
and Speech Tools for Italian (EVALITA 2020).
Angela Ferrari. 2011. Enunciati
nominali. Enciclopedia dell’Italiano. http:
//www.treccani.it/enciclopedia/
enunciati-nominali_(Enciclopedia_
dell’Italiano)/.</p>
      <p>Elisabetta Fersini, Debora Nozza, and Paolo Rosso.
2020. AMI @ EVALITA2020: Automatic
Misogyny Identification. In Proceedings of the 7th
evaluation campaign of Natural Language Processing and
Speech tools for Italian (EVALITA 2020), Online.
Michele Fontana and Giuseppe Attardi. 2020.</p>
      <p>Fontana-Unipi @ HaSpeeDe2: Ensemble of
transformers for the Hate Speech task at Evalita. In
Proceedings of the 7th Evaluation Campaign of Natural
Language Processing and Speech Tools for Italian
(EVALITA 2020).</p>
      <p>Paula Fortuna, Joa˜o Rocha da Silva, Juan
SolerCompany, Leo Wanner, and Se´rgio Nunes. 2019.
A Hierarchically-Labeled Portuguese Hate Speech
Dataset. In Proceedings of the Third Workshop on
Abusive Language Online.
guage Detection in Online User Content. In
Proceedings of the 25th International Conference on
World Wide Web (WWW’16).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Natalie</given-names>
            <surname>Alkiviadou</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Hate speech on social media networks: towards a regulatory framework?</article-title>
          <source>Information &amp; Communications Technology Law</source>
          ,
          <volume>28</volume>
          (
          <issue>1</issue>
          ):
          <fpage>19</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Rangel, Paolo Rosso, and
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>SemEval2019 Task 5: Multilingual Detection of Hate Speech Against Immigrants and Women in Twitter</article-title>
          .
          <source>In Proceedings of SemEval</source>
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Evalita 2020: Overview of the 7th evaluation campaign of natural language processing and speech tools for italian</article-title>
          .
          <source>In Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ), Online.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Bassignana</surname>
          </string-name>
          , Valerio Basile, and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hurtlex: A Multilingual Lexicon of Words to Hurt</article-title>
          .
          <source>In Proceedings of the Fifth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Elia</given-names>
            <surname>Bisconti</surname>
          </string-name>
          and
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Montagnani</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Montanti @ HaSpeeDe2 EVALITA 2020: Hate Speech Detection in online contents</article-title>
          .
          <source>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <surname>Dell'Orletta Felice</surname>
            , Fabio Poletto, Manuela Sanguinetti, and
            <given-names>Tesconi</given-names>
          </string-name>
          <string-name>
            <surname>Maurizio</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of the EVALITA 2018 hate speech detection task</article-title>
          .
          <source>In Proceedings of the Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Arthur TE Capozzi</surname>
            , Mirko Lai, Valerio Basile, Fabio Poletto, Manuela Sanguinetti, Cristina Bosco, Viviana Patti, Giancarlo Ruffo, Cataldo Musto,
            <given-names>Marco</given-names>
            Polignano, Giovanni Semeraro, and Marco
          </string-name>
          <string-name>
            <surname>Stranisci</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Computational linguistics against hate: Hate speech detection and visualization on social media in the ”Contro L'Odio” project</article-title>
          .
          <source>In Proceedings of the Sixth Italian Conference on Computational Linguistics</source>
          , CLiC-it
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Tommaso</given-names>
            <surname>Caselli</surname>
          </string-name>
          , Nicole Novielli, Viviana Patti, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>EVALITA 2018: Overview of the 6th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          .
          <source>In Proceedings of Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Alessandra</given-names>
            <surname>Teresa</surname>
          </string-name>
          <string-name>
            <surname>Cignarella</surname>
          </string-name>
          , Mirko Lai, Cristina Bosco, Viviana Patti, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>SardiStance@EVALITA2020: Overview of the Task on Stance Detection in Italian Tweets</article-title>
          .
          <source>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Gloria</given-names>
            <surname>Comandini</surname>
          </string-name>
          and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>An Impossible Dialogue! Nominal Utterances and Populist Rhetoric in an Italian Twitter Corpus of Hate Speech against Immigrants</article-title>
          .
          <source>In Proceedings of the Third Workshop on Abusive Language Online.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Gloria</given-names>
            <surname>Comandini</surname>
          </string-name>
          , Manuela Speranza, and
          <string-name>
            <given-names>Bernardo</given-names>
            <surname>Magnini</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Effective Communication without Verbs? Sure! Identification of Nominal Utterances in Italian Social Media Texts</article-title>
          .
          <source>In Proceedings of the Fifth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2018</year>
          ), volume
          <volume>2253</volume>
          .
          <article-title>CEUR-WS.org</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Chiara</given-names>
            <surname>Francesconi</surname>
          </string-name>
          , Cristina Bosco, Fabio Poletto, and
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Error Analysis in a Hate Speech Detection Task: The case of HaSpeeDe-TW at EVALITA 2018</article-title>
          .
          <source>In Proceedings of the Sixth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Gambino</surname>
          </string-name>
          and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Pirrone</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>CHILab @ HaSpeeDe 2: Enhancing Hate Speech Detection with Part-of-Speech Tagging</article-title>
          .
          <source>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Julia</given-names>
            <surname>Hoffmann</surname>
          </string-name>
          and
          <string-name>
            <given-names>Udo</given-names>
            <surname>Kruschwitz</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>UR NLP @ HaSpeeDe 2 at EVALITA 2020: Towards Robust Hate Speech Detection with Contextual Embeddings</article-title>
          .
          <source>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Svea</given-names>
            <surname>Klaus</surname>
          </string-name>
          ,
          <string-name>
            <surname>Anna-Sophie Bartle</surname>
            , and
            <given-names>Daniela</given-names>
          </string-name>
          <string-name>
            <surname>Rossmann</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Svandiela @ HaSpeeDe: Detecting Hate Speech in Italian Twitter Data with BERT</article-title>
          .
          <source>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Eric</given-names>
            <surname>Lavergne</surname>
          </string-name>
          , Rajkumar Saini,
          <source>Gyo¨rgy Kova´cs, and Killian Murphy</source>
          .
          <year>2020</year>
          .
          <article-title>TheNorth @ HaSpeeDe 2: BERT-based Language Model Fine-tuning for Italian Hate Speech Detection</article-title>
          .
          <source>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Alyssa</given-names>
            <surname>Lees</surname>
          </string-name>
          , Jeffrey Sorensen, and
          <string-name>
            <given-names>Ian</given-names>
            <surname>Kivlichan</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Jigsaw @ AMI and HaSpeeDe2: Fine-Tuning a Pre-Trained Comment-Domain BERT Model</article-title>
          .
          <source>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Leonardelli</surname>
          </string-name>
          , Stefano Menini, and
          <string-name>
            <given-names>Sara</given-names>
            <surname>Tonelli</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>DH-FBK @ HaSpeeDe2: Italian Hate Speech Detection via Self-Training and Oversampling</article-title>
          .
          <source>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Mandl</surname>
          </string-name>
          , Sandip Modha, Prasenjit Majumder, Daksh Patel, Mohana Dave, Chintak Mandlia, and
          <string-name>
            <given-names>Aditya</given-names>
            <surname>Patel</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Overview of the hasoc track at fire 2019: Hate speech and offensive content identification in Indo-European languages</article-title>
          .
          <source>In Proceedings of the 11th Forum for Information Retrieval Evaluation.</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Bice</given-names>
            <surname>Mortara Garavelli</surname>
          </string-name>
          .
          <year>1971</year>
          .
          <article-title>Fra norma e invenzione: lo stile nominale</article-title>
          . In Accademia della Crusca, editor,
          <source>Studi di grammatica italiana</source>
          , volume
          <volume>1</volume>
          , pages
          <fpage>271</fpage>
          -
          <lpage>315</lpage>
          . G. C. Sansoni Editore, Firenze, Italia.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Chikashi</given-names>
            <surname>Nobata</surname>
          </string-name>
          , Joel Tetreault, Achint Thomas,
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Mehdad</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Yi</given-names>
            <surname>Chang</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Abusive LanXiaozhi Ou</article-title>
          and
          <string-name>
            <given-names>Hongling</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>YNU OXZ @ HaSpeeDe 2 and AMI : XLM-RoBERTa with Ordered Neurons LSTM for classification task at EVALITA 2020</article-title>
          .
          <article-title>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey</surname>
            <given-names>Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>GloVe: Global vectors for word representation</article-title>
          .
          <source>In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP).</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Poletto</surname>
          </string-name>
          , Marco Stranisci, Manuela Sanguinetti, Viviana Patti, and
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Hate Speech Annotation: Analysis of an Italian Twitter Corpus</article-title>
          .
          <source>In Proceedings of the Fourth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Poletto</surname>
          </string-name>
          , Valerio Basile, Manuela Sanguinetti, Cristina Bosco, and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Resources and Benchmark Corpora for Hate Speech Detection: a Systematic Review</article-title>
          .
          <source>Language Resources and Evaluation.</source>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Michal</given-names>
            <surname>Ptaszynski</surname>
          </string-name>
          , Agata Pieciukiewicz, and
          <string-name>
            <given-names>Paweł</given-names>
            <surname>Dybała</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Results of the PolEval 2019 Shared Task 6: First Dataset and Open Shared Task for Automatic Cyberbullying Detection in Polish Twitter</article-title>
          .
          <source>In Proceedings of the PolEval 2019 Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <source>Mariano Jason Rodriguez Cisnero and Reynier Ortega Bueno</source>
          .
          <year>2020</year>
          .
          <article-title>UO@HaSpeeDe2: Ensemble Model for Italian Hate Speech Detection</article-title>
          .
          <source>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Fabio Poletto, Cristina Bosco, Viviana Patti, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Stranisci</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>An Italian Twitter Corpus of Hate Speech against Immigrants</article-title>
          .
          <source>In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC'18).</source>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Xuan-Son</surname>
            <given-names>Vu</given-names>
          </string-name>
          , Thanh Vu,
          <string-name>
            <surname>Mai-Vu</surname>
            <given-names>Tran</given-names>
          </string-name>
          , Thanh LeCong, and
          <string-name>
            <surname>Huyen T M Nguyen</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>HSD Shared Task in VLSP Campaign 2019: Hate Speech Detection for Social Good</article-title>
          .
          <source>In Proceedings of VLSP</source>
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Zeerak</given-names>
            <surname>Waseem</surname>
          </string-name>
          , Thomas Davidson, Dana Warmsley, and
          <string-name>
            <given-names>Ingmar</given-names>
            <surname>Weber</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Understanding abuse: A typology of abusive language detection subtasks</article-title>
          .
          <source>In Proceedings of the First Workshop on Abusive Language Online</source>
          , pages
          <fpage>78</fpage>
          -
          <lpage>84</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>Ruth E.</given-names>
            <surname>Wodak</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Introductory remarks from 'hate speech' to 'hate tweets'</article-title>
          .
          <source>In Mojca Pajnik and Birgit Sauer</source>
          , editors,
          <article-title>Populism and the web: communicative practices of parties and movements in Europe, pages xvii-xxiii</article-title>
          . Rourledge.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>