<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Policycorpus XL: An Italian Corpus for the Detection of Hate Speech Against Politics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fabio Celli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mirko Lai</string-name>
          <email>mirko.lai@unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Armend Duzha</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cristina Bosco</string-name>
          <email>bosco@di.unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Viviana Patti</string-name>
          <email>patti@di.unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>. Research</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Development</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gruppo Maggioli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>. Dept. of Informatics, University of Turin</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we describe the largest corpus annotated with hate speech in the political domain in Italian. Policycorpus XL has 7000 tweets, manually annotated, and a presence of hate labels above 40%, while in other corpora of the same type is usually below 30%. Here we describe the collection of data and test some baseline with simple classification algorithms, obtaining promising results. We suggest that the high amount of hate labels boosts the performance of classifiers, and we plan to release the dataset in a future evaluation campaign.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In recent years, computer mediated
communication on social media and microblogging websites
has become more and more aggressive
        <xref ref-type="bibr" rid="ref33">(Watanabe
et al., 2018)</xref>
        . It is well known that people use
social media like Twitter for a variety of purposes
like keeping in touch with friends, raising the
visibility of their interests, gathering useful
information, seeking help and release stress
        <xref ref-type="bibr" rid="ref34">(Zhao and
Rosson, 2009)</xref>
        , but the spread of fake news
        <xref ref-type="bibr" rid="ref3 ref31">(Shu
et al., 2019; Alam et al., 2016)</xref>
        has exacerbated a
cultural clash between social classes that emerged
at least since after the debate about Brexit
        <xref ref-type="bibr" rid="ref2 ref3 ref9">(Celli
et al., 2016)</xref>
        and more recently during the
pandemics
        <xref ref-type="bibr" rid="ref23">(Oliver et al., 2020)</xref>
        . Despite the fact that
the behavior online is different from the
behavior offline
        <xref ref-type="bibr" rid="ref7">(Celli and Polonio, 2015)</xref>
        , we observe
more and more hate speech in social media, to the
point where it has become a serious problem for
free speech and social cohesion.
      </p>
      <p>
        Copyright © 2021 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0)
Hate speech is defined as any expression that is
abusive, insulting, intimidating, harassing, and/or
incites, supports and facilitates violence, hatred,
or discrimination. It is directed against people
(individuals or groups) on the basis of their race,
ethnic origin, religion, gender, age, physical
condition, disability, sexual orientation, political
conviction, and so forth
        <xref ref-type="bibr" rid="ref13">(Erjavec and Kovacˇicˇ, 2012)</xref>
        .
In response to the growing number of hate
messages, the Natural language Processing (NLP)
community focused on the classification of hate
speech
        <xref ref-type="bibr" rid="ref4">(Badjatiya et al., 2017)</xref>
        and the analysis
of online debates
        <xref ref-type="bibr" rid="ref8">(Celli et al., 2014)</xref>
        . In
particular, many worked on systems to detect offensive
language against specific vulnerable groups (e.g.,
immigrants, LGBTQ communities among others)
        <xref ref-type="bibr" rid="ref24">(Poletto et al., 2017)</xref>
        <xref ref-type="bibr" rid="ref25">(Poletto et al., 2021)</xref>
        , as well
as aggressive language against women
        <xref ref-type="bibr" rid="ref27">(Saha et
al., 2018)</xref>
        . An under-researched - yet important
area of investigation is anti-politics hate: the hate
speech against politicians, policy makers and laws
at any level (national, regional and local). While
anti-policy hate speech has been addressed in
Arabic
        <xref ref-type="bibr" rid="ref17">(Guellil et al., 2020)</xref>
        and German
        <xref ref-type="bibr" rid="ref16 ref19">(Jaki and
De Smedt, 2019)</xref>
        , most European languages have
been under-researched. The bottleneck in this field
of research is the availability of data to train good
hate speech detection models. In recent years,
scientific research contributed to the automatic
detection of hate speech from text with datasets
annotated with hate labels, aggressiveness,
offensiveness, and other related dimensions
        <xref ref-type="bibr" rid="ref28">(Sanguinetti et
al., 2018)</xref>
        . Scholars have presented systems for the
detection of hate speech in social media focused
on specific targets, such as immigrants
        <xref ref-type="bibr" rid="ref11">(Del
Vigna et al., 2017)</xref>
        , and language domains, such as
racism
        <xref ref-type="bibr" rid="ref20">(Kwok and Wang, 2013)</xref>
        , misogyny
        <xref ref-type="bibr" rid="ref5">(Basile
et al., 2019)</xref>
        or cyberbullying
        <xref ref-type="bibr" rid="ref22">(Menini et al., 2019)</xref>
        .
Each type of hate speech has its own vocabulary
and its own dynamics, thus the selection of a
specific domain is crucial to obtain clean data and
to restrict the scope of experiments and learning
tasks.
      </p>
      <p>
        In this paper we present a new corpus, called
Policycorpus XL, for hate speech detection from
Twitter in Italian. This corpus is an extension of the
Policycorpus
        <xref ref-type="bibr" rid="ref12">(Duzha et al., 2021)</xref>
        . We selected
Twitter as the source of data and Italian as the
target language because Italy has, at least since the
elections in 2018, a large audience that pays
attention to hyper-partisan sources on Twitter that
are prone to produce and retweet messages of hate
against policy making
        <xref ref-type="bibr" rid="ref16">(Giglietto et al., 2019)</xref>
        .
The paper is structured as follows: after a
literature review (Section 2), we describe how we
collected and annotated the data (Section 3), we
evaluate some baselines (Section 4), and we pave the
way for future work (Section 5).
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Hate Speech in social media is a complex
phenomenon, whose detection has recently gained
significant traction in the Natural Language
Processing community, as attested by several recent
review works
        <xref ref-type="bibr" rid="ref25">(Poletto et al., 2021)</xref>
        . High-quality
annotated corpora and benchmarks are key
resources for hate speech detection and haters
proifling in general
        <xref ref-type="bibr" rid="ref18">(Jain et al., 2021)</xref>
        , considering the
vast number of supervised approaches that have
been proposed
        <xref ref-type="bibr" rid="ref21">(MacAvaney et al., 2019)</xref>
        .
      </p>
      <p>
        Early datasets on Hate Speech, especially in
English, were produced outside any evaluation
campaigns
        <xref ref-type="bibr" rid="ref32">(Waseem and Hovy, 2016)</xref>
        ,
        <xref ref-type="bibr" rid="ref15">(Founta et al.,
2018)</xref>
        as well as inside such competitions. These
include SemEval 2019, where a multilingual hate
speech corpus against immigrants and women in
English and Spanish
        <xref ref-type="bibr" rid="ref5">(Basile et al., 2019)</xref>
        was
released, and PAN 2021, that provided a dataset for
the detection of hate spreader authors in English
and Spanish
        <xref ref-type="bibr" rid="ref26">(Rangel et al., 2021)</xref>
        . Most Italian
datasets in the field of hate speech have been
released during competitions and evaluation
campaigns. There are:
• the Italian HS corpus
        <xref ref-type="bibr" rid="ref24">(Poletto et al., 2017)</xref>
        ,
• HaSpeeDe-tw2018 and HaSpeeDe-tw2020,
the datasets released during the EVALITA
campaigns
        <xref ref-type="bibr" rid="ref29">(Sanguinetti et al., 2020)</xref>
        ,
• the Policycorpus
        <xref ref-type="bibr" rid="ref12">(Duzha et al., 2021)</xref>
        , the
only dataset in Italian that is annotated with
hate speech in the political domain.
      </p>
      <p>
        The Italian HS corpus is a collection of more
than 5700 tweets manually annotated with hate
speech, aggressiveness, irony and other forms
of potentially harassing communication. The
HaSpeeDe-tw corpora are two collections of 4000
and 8100 tweets respectively, manually annotated
with hate speech labels and containing mainly
anti-immigration hate
        <xref ref-type="bibr" rid="ref28 ref6">(Bosco et al., 2018)</xref>
        . The
Policycorpus is a collection of 1260 tweets
manually annotated with hate speech labels against
politics and politicians. We decided to expand it and
produce a new dataset.
      </p>
      <p>
        Hate speech is hard to annotate and hard to
model, with the risk of creating data that is
biased and making the models prone to overfitting.
In addition to this, literature also reports cases
of annotators’ insensitivity to differences in
dialect that can lead to racial bias in automatic hate
speech detection models, potentially amplifying
harm against minority populations. It is the case of
African American English
        <xref ref-type="bibr" rid="ref30">(Sap et al., 2019)</xref>
        but it
potentially applies to Italian as well, as it is a
language full of dialects and regional offenses.
      </p>
      <p>
        Hate speech is intrinsically associated to
relationships between groups, and also relying in
language nuances. There are many definitions of hate
speech from different sources, such as European
Union Commission, International minorities
associations (ILGA) and social media policies
        <xref ref-type="bibr" rid="ref14 ref16 ref27 ref33">(Fortuna and Nunes, 2018)</xref>
        . In most definitions, hate
speech has specicfi targets based on specific
characteristics of groups. Hate speech is to incite
violence, usually towards a minority. Moreover, hate
speech is to attack or diminish. Additionally,
humour has a specific status in hate speech, and it
makes more difficult to understand the boundaries
about what is hate and what is not.
      </p>
      <p>In the political domain we find all of these
aspects, especially messages against a minority
(politicians) to attack or diminish. We think that
more resources are needed for the classification
of hate speech in Italian in the political domain,
hence we decided to collect and annotate more
data for this task.</p>
      <p>In the next section, we describe how we created
the dataset and annotated it with hate speech
labels.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Data Collection and Annotation</title>
      <p>Starting from the Policycorpus, we expanded it
from 1260 to 7000 tweets in Italian, collected
using snowball sampling from Twitter APIs. As
initial seeds, we used the same set of hashtags used
for the Policycorpus, for instance: #dpcm (decree
of the president of the council of ministers), #legge
(law) and #leggedibilancio (budget law). We
removed duplicates, retweets and tweets containing
only hashtags and urls. At the end of the
sampling process, the list of seeds included about 6000
hashtags that co-occurred with the initial ones.
We grouped the hashtags into the following
categories:
• Laws, such as #decretorilancio
(#relaunchdecree), #leggelettorale (#electorallaw),
#decretosicurezza (#securitydecree)
• Politicians and policy makers, such as
#Salvini, #decretoSalvini (#Salvinidecree),
#Renzi, #Meloni, #DraghiPremier
• Political parties, such as #lega (#league), #pd
(#Democratic Party)
• Political tv shows, such as #ottoemezzo,
#nonelarena, #noneladurso, #Piazzapulita
• Topics of the public debate, such as #COVID,
#precari (#precariousworkers), #sicurezza
(#security), #giustizia (#justice), #ItalExit
• Hyper-partisan slogans, such as
#vergognaConte (#shameonConte),
#contedimettiti (#ConteResign) or #noicontrosalvini
(#WeareagainstSalvini)
Examples of collected hashtags are reported in
Figure 1</p>
      <p>
        Recent shared tasks
        <xref ref-type="bibr" rid="ref1 ref10 ref2">(Agerri et al., 2021;
Cignarella et al., 2020; Aker et al., 2016)</xref>
        promoted the use of contextual information about the
tweet and its author (including his/her social
media network) for improving the performance of
stance detection. Here, with the aim to
stimulate the exploration of data augmentation on hate
speech detection, we shared additional contextual
information based on the post such as: the number
of retweets and the number of favours (the number
of tweets that given user has marked as favorite
favours count field) the tweet received, the device
used for posting it (e.g. iOS or Android), the
posting date and location, and an attribute that states if
the post is a tweet, a retweet, a reply, or a quote.
Furthermore, we collected contextual information
related to the authors of these posts such as: the
number of tweets ever posted, the user’s
description and location, the number of her/his followers
and of her/his friends, the number of public lists
that this user is a member of and the date her/his
account has been created.
      </p>
      <p>All these contextual information are
respectively part of the “root-level” attributes of the
Tweets and Users objects that Twitter returns in
JSON format through its APIs. Additionally, we
planned to explore the interests of the author
collecting the list of her/his following (the users
she/he follows) employing the following API
endpoint. Moreover, for exploring the author’s social
interactions, we used the Academic Full Search
API for recovering the list of the users that she/he
has retweeted to and replied to in the last two
years.</p>
      <p>
        The enhanced Policycorpus has been finally
anonymised mapping each tweet id, users id, and
mention with a randomly generated ID. To
produce gold standard labels, we asked two Italian
native speakers, experts of communication, to
manually label the tweets in the corpus, distinguishing
between hate and normal tweets according to the
following guidelines: By definition, hate speech
is any expression that is abusive, insulting,
intimidating, harassing, and/or incites to violence,
hatred, or discrimination. It is directed against
people on the basis of their race, ethnic origin,
religion, gender, age, physical condition,
disability, sexual orientation, political conviction, and
so forth.
        <xref ref-type="bibr" rid="ref13">(Erjavec and Kovacˇicˇ, 2012)</xref>
        . Below
We provide some examples with translation in
English:
1. “Un chiaro #NO all #Olanda che ci
vorrebbe s`ı utilizzatori delle risorse economiche
del #MES ma in cambio della rinuncia dell
Italia alla propria autonomia di bilancio. All
Olanda diciamo: grazie e arrivederci NON
CI INTERESSA!!”1
The first example is normal because it does not
contain hate, insults, intimidation, violence or
discrimination.
      </p>
      <p>2. “...Sta settimanale passerella dello
#sciacallo #no #proprioNo! Ascoltare un
#pagliaccio padano dopo un vero PATRIOTA un
medico di #Bergamo non si puo` reggere
ne vedere ne ascoltare. Giletti dovrebbe
smetterla di invitare certi CAZZARIPADANI!
#COVID-19 #NonelArena”2
The second example contains hate speech,
including insults like #clown and #jackal.</p>
      <p>3. “Dico la mia... #Draghi e` un grande
economista ma a noi non serve un
economista stile #Monti... A noi non
serve un altro #governo tecnico per ubbidire
alla lobby delle banche! A noi serve un
leader politico! A noi serve un #ItalExit! A
noi serve la #Lira! #No a #DraghiPremier”3
The last example is a normal case, despite the
strong negative sentiment. It might be
controversial for the presence of the term lobby, often
used in abusive contexts, but in this case, it is
1a clear #NO to the #Netherlands that would like us to be
users of the #MES economic resources but in exchange for
Italy’s renunciation of its budgetary autonomy. To
Netherlands we say: thank you and goodbye, WE ARE NOT
INTERESTED !!</p>
      <p>2... There is a weekly catwalk of the #jackal #no
#notAtAll! Listening to a Padanian #clown after a true PATRIOT
a doctor from #Bergamo cannot be held, seen or heard. Giletti
should stop inviting certain SLACKERS FROM THE PO
VALLEY! #COVID-19 #NonelArena</p>
      <p>3I have my say ... #Draghi is a great economist but we
don’t need a #Monti-style economist ... We don’t need
another technical #government to obey the banking lobby! We
need a political leader! We need a #ItalExit! We need the
#Lira! #No to #DraghiPremier
not directed against people on the basis of their
race, ethnic origin, religion, gender, age, physical
condition, disability, sexual orientation or political
conviction.</p>
      <p>
        The Inter-Annotator Agreement is k=0.53.
Although this score is not high, it is in line with
the score reported in the literature for hate speech
against immigrants (k=0.54)
        <xref ref-type="bibr" rid="ref24">(Poletto et al., 2017)</xref>
        and indicates that the detection of hate speech is a
hard task for humans.
      </p>
      <p>All the examples in disagreement were
discussed and an agreement was reached between the
annotators, with the help of a third supervisor. The
cases of disagreements occurred more often when
the sentiment of the tweet was negative, this was
mainly due to:
• The use of vulgar expressions not explicitly
directed against specific people but
generically against political choices.
• The negative interpretation of hyper-partisan
hashtags, such as #contedimettiti
(#ConteResign) or #noicontrosalvini
(#WeareagainstSalvini), in tweets without explicit insults or
abusive language.
• The substitution of explicit insults with
derogatory words, such as the word “circus”
instead of “clowns”.</p>
      <p>The amount of hate labels in the original
Policycorpus was 11% (1124 normal and 140 hate
tweets), strongly unbalanced like the Italian HS
corpus (17% of hate tweets), because it reflects
the raw distribution of hate tweets in Twitter. The
HaSpeeDe-tw corpus (32% of hate tweets) instead
has a distribution that oversamples hate tweets and
it is better for training hate speech models.
Following the HaSpeeDe-tw example, in
Policycorpus XL we collected more tweets of hate,
randomly discarding normal tweets to reach at least
40% of hate tweets in the corpus. In the end we
have 40.6% of hate labels and 59.4% of normal
labels, distributed between training and test set as
shown in figure 2.</p>
      <p>We note in the style of these tweets that there
is a substantial overlap among the top unigrams in
the two classes, as shown in Figure 3. We suggest
that weak signals, like less frequent words, are key
features for the classification task.</p>
      <p>In the next section, we report and discuss the
results of classification experiments.
In order to set the baselines for the hate speech
classification task on Policycorpus-XL, we tested
different classification algorithms. We are using
a 70 train and 30 test percentage split, the
training set shape is 4900 instances and 300 features,
while the test set shape is 2100 instances and 300
features. The 300 features are the normalized
frequencies of the 300 most frequent words extracted
from tweets without removing the stopwords.
Table 1 reports the result of classification.</p>
      <p>algorithm
majority baseline
naive bayes
decision trees
SVMs
balanced acc
0.500
0.783
0.763
0.788
macro F1
0.37
0.78
0.76
0.79</p>
      <p>We used Scikit-Learn to compute a majority
baseline with a dummy classifier, that assigns all
the instances to the most frequent class (normal
tweets), a naive bayes classifier, a decision tree
and Support Vector Machines (SVMs). The best
performance for the classification of hate speech
has been achieved with the SVM classifier, that
has a very high precision (0.94) and poor recall
(0.60). All the algorithms a The results are in line
with the scores obtained by the systems on the
HaSpeeDe-tw 2020 dataset at EVALITA, and we
believe that there is still great room for
improvement with the Policycorpus-XL, as we exploited
very simple and limited features.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Work</title>
      <p>We presented a large corpus of Twitter data in
Italian, manually annotated with hate speech labels.
The corpus is an extension of a previous one, the
ifrst corpus annotated with hate speech in the
political domain in Italian.</p>
      <p>Given the rising amount of hate messages
online, not just against minorities but more and more
against policies and policymakers, it is urgent to
understand the phenomenon and train classifiers
that could prevent people to disseminate hate in
the public debate. This is very important to keep
democracies alive and grant a free speech that is
respectful of other people’s freedom.</p>
      <p>We plan to distribute the corpus in the next
edition of EVALITA for a specific HaSpeeDe-tw task.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>The research leading to the results presented in
this paper has received funding from the
PolicyCLOUD project, supported by the European
Union’s Horizon 2020 research and innovation
programme under Grant Agreement no 870675.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Rodrigo</given-names>
            <surname>Agerri</surname>
          </string-name>
          , Roberto Centeno, Mar´ıa Espinosa, Joseba Fernandez de Landa, and
          <string-name>
            <given-names>Alvaro</given-names>
            <surname>Rodrigo</surname>
          </string-name>
          .
          <year>2021</year>
          . VaxxStance@IberLEF 2021:
          <article-title>Going Beyond Text in Crosslingual Stance Detection</article-title>
          .
          <source>In Proceedings of the Iberian Languages Evaluation Forum (IberLEF</source>
          <year>2021</year>
          ).
          <article-title>CEUR-WS.org</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Ahmet</given-names>
            <surname>Aker</surname>
          </string-name>
          , Fabio Celli, Adam Funk, Emina Kurtic,
          <string-name>
            <given-names>Mark</given-names>
            <surname>Hepple</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Rob</given-names>
            <surname>Gaizauskas</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Sheffield-trento system for sentiment and argument structure enhanced comment-to-article linking in the online news domain</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Firoj</given-names>
            <surname>Alam</surname>
          </string-name>
          , Fabio Celli, Evgeny Stepanov, Arindam Ghosh, and
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Riccardi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>The social mood of news: self-reported annotations to design automatic mood detection systems</article-title>
          .
          <source>In Proceedings of the Workshop on Computational Modeling of People's Opinions, Personality, and Emotions in Social Media (PEOPLES)</source>
          , pages
          <fpage>143</fpage>
          -
          <lpage>152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Pinkesh</given-names>
            <surname>Badjatiya</surname>
          </string-name>
          , Shashank Gupta, Manish Gupta, and
          <string-name>
            <given-names>Vasudeva</given-names>
            <surname>Varma</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Deep learning for hate speech detection in tweets</article-title>
          .
          <source>In Proceedings of the 26th International Conference on World Wide Web Companion</source>
          , pages
          <fpage>759</fpage>
          -
          <lpage>760</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Cristina Bosco, Elisabetta Fersini, Nozza Debora, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso,
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , et al.
          <year>2019</year>
          .
          <article-title>Semeval-2019 task 5: Multilingual detection of hate speech against immigrants and women in twitter</article-title>
          .
          <source>In 13th International Workshop on Semantic Evaluation</source>
          , pages
          <fpage>54</fpage>
          -
          <lpage>63</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          , Felice Dell'Orletta, Fabio Poletto, Manuela Sanguinetti, and
          <string-name>
            <given-names>Maurizio</given-names>
            <surname>Tesconi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of the evalita 2018 hate speech detection task</article-title>
          .
          <source>In EVALITA 2018-Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</source>
          , volume
          <volume>2263</volume>
          , pages
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          . CEUR.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Celli</surname>
          </string-name>
          and
          <string-name>
            <given-names>Luca</given-names>
            <surname>Polonio</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Facebook and the real world: Correlations between online and offline conversations</article-title>
          .
          <source>CLiC it</source>
          , page
          <volume>82</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Celli</surname>
          </string-name>
          , Giuseppe Riccardi, and
          <string-name>
            <given-names>Arindam</given-names>
            <surname>Ghosh</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Corea: Italian news corpus with emotions and agreement</article-title>
          .
          <source>In Proceedings of CLIC-it 2014</source>
          , pages
          <fpage>98</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Celli</surname>
          </string-name>
          , Evgeny A Stepanov,
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Poesio</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Riccardi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Predicting brexit: Classifying agreement is better than sentiment and pollsters</article-title>
          .
          <source>In PEOPLES@ COLING</source>
          , pages
          <fpage>110</fpage>
          -
          <lpage>118</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Alessandra</given-names>
            <surname>Teresa</surname>
          </string-name>
          <string-name>
            <surname>Cignarella</surname>
          </string-name>
          , Mirko Lai, Cristina Bosco, Viviana Patti, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Sardistance@evalita2020: Overview of the task on stance detection in italian tweets</article-title>
          .
          <source>In Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ), volume
          <volume>2765</volume>
          <source>of CEUR Workshop Proceedings</source>
          , Aachen, Germany, December.
          <source>CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Fabio Del Vigna</surname>
            ,
            <given-names>Andrea</given-names>
          </string-name>
          <string-name>
            <surname>Cimino</surname>
            , Felice Dell'Orletta,
            <given-names>Marinella</given-names>
          </string-name>
          <string-name>
            <surname>Petrocchi</surname>
            , and
            <given-names>Maurizio</given-names>
          </string-name>
          <string-name>
            <surname>Tesconi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Hate me, hate me not: Hate speech detection on facebook</article-title>
          .
          <source>In Proceedings of the First Italian Conference on Cybersecurity (ITASEC17)</source>
          , pages
          <fpage>86</fpage>
          -
          <lpage>95</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Armend</given-names>
            <surname>Duzha</surname>
          </string-name>
          , Cristiano Casadei,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Tosi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Celli</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Hate versus politics: detection of hate against policy makers in italian tweets</article-title>
          .
          <source>SN Social Sciences</source>
          ,
          <volume>1</volume>
          (
          <issue>9</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Karmen</given-names>
            <surname>Erjavec</surname>
          </string-name>
          and Melita Poler Kovacˇicˇ.
          <year>2012</year>
          .
          <article-title>“you don't understand, this is a new war!” analysis of hate speech in news web sites' comments</article-title>
          .
          <source>Mass Communication and Society</source>
          ,
          <volume>15</volume>
          (
          <issue>6</issue>
          ):
          <fpage>899</fpage>
          -
          <lpage>920</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Paula</given-names>
            <surname>Fortuna</surname>
          </string-name>
          and Se´rgio Nunes.
          <year>2018</year>
          .
          <article-title>A survey on automatic detection of hate speech in text</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>51</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Antigoni</given-names>
            <surname>Maria</surname>
          </string-name>
          <string-name>
            <surname>Founta</surname>
          </string-name>
          , Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Sirivianos</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Nicolas</given-names>
            <surname>Kourtellis</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Large scale crowdsourcing and characterization of twitter abusive behavior</article-title>
          .
          <source>In Twelfth International AAAI Conference on Web and Social Media.</source>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Giglietto</surname>
          </string-name>
          , Nicola Righetti, Giada Marino, and
          <string-name>
            <given-names>Luca</given-names>
            <surname>Rossi</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Multi-party media partisanship attention score. estimating partisan attention of news media sources using twitter data in the leadup to 2018 italian election</article-title>
          .
          <source>Comunicazione politica</source>
          ,
          <volume>20</volume>
          (
          <issue>1</issue>
          ):
          <fpage>85</fpage>
          -
          <lpage>108</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Imane</given-names>
            <surname>Guellil</surname>
          </string-name>
          , Ahsan Adeel, Faical Azouaou, Sara Chennoufi, Hanene Maafi, and
          <string-name>
            <given-names>Thinhinane</given-names>
            <surname>Hamitouche</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Detecting hate speech against politicians in arabic community on social media</article-title>
          .
          <source>International Journal of Web Information Systems.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Rakshita</given-names>
            <surname>Jain</surname>
          </string-name>
          , Devanshi Goel, Prashant Sahu, Abhinav Kumar, and Jyoti Prakash Singh.
          <year>2021</year>
          .
          <article-title>Profiling hate speech spreaders on twitter</article-title>
          .
          <source>In CLEF.</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Sylvia</given-names>
            <surname>Jaki</surname>
          </string-name>
          and Tom De Smedt.
          <year>2019</year>
          .
          <article-title>Right-wing german hate speech on twitter: Analysis and automatic detection</article-title>
          . arXiv preprint arXiv:
          <year>1910</year>
          .07518.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Irene</given-names>
            <surname>Kwok</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yuzhou</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Locate the hate: Detecting tweets against blacks</article-title>
          .
          <source>In Proceedings of the twenty-seventh AAAI conference on artificial intelligence</source>
          , pages
          <fpage>1621</fpage>
          -
          <lpage>1622</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Sean</surname>
            <given-names>MacAvaney</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hao-Ren</surname>
            <given-names>Yao</given-names>
          </string-name>
          , Eugene Yang, Katina Russell, Nazli Goharian, and
          <string-name>
            <given-names>Ophir</given-names>
            <surname>Frieder</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Hate speech detection: Challenges and solutions</article-title>
          .
          <source>PloS one</source>
          ,
          <volume>14</volume>
          (
          <issue>8</issue>
          ):
          <fpage>e0221152</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Menini</surname>
          </string-name>
          , Giovanni Moretti, Michele Corazza, Elena Cabrio, Sara Tonelli, and
          <string-name>
            <given-names>Serena</given-names>
            <surname>Villata</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A system to monitor cyberbullying based on message classification and social network analysis</article-title>
          .
          <source>In Proceedings of the Third Workshop on Abusive Language Online</source>
          , pages
          <fpage>105</fpage>
          -
          <lpage>110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Nuria</given-names>
            <surname>Oliver</surname>
          </string-name>
          , Bruno Lepri, Harald Sterly, Renaud Lambiotte, Se´bastien Deletaille, Marco De Nadai, Emmanuel Letouze´,
          <string-name>
            <surname>Albert Ali Salah</surname>
          </string-name>
          , Richard Benjamins,
          <string-name>
            <given-names>Ciro</given-names>
            <surname>Cattuto</surname>
          </string-name>
          , et al.
          <year>2020</year>
          .
          <article-title>Mobile phone data for informing public health actions across the covid-19 pandemic life cycle</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Poletto</surname>
          </string-name>
          , Marco Stranisci, Manuela Sanguinetti, Viviana Patti, and
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Hate speech annotation: Analysis of an italian twitter corpus</article-title>
          .
          <source>In 4th Italian Conference on Computational Linguistics</source>
          , CLiC-it
          <year>2017</year>
          , volume
          <year>2006</year>
          , pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . CEUR-WS.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Poletto</surname>
          </string-name>
          , Valerio Basile, Manuela Sanguinetti, Cristina Bosco, and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Resources and benchmark corpora for hate speech detection: a systematic review</article-title>
          .
          <source>Language Resources &amp; Evaluation</source>
          ,
          <volume>55</volume>
          :
          <fpage>477</fpage>
          -
          <lpage>523</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Francisco</surname>
            <given-names>Rangel</given-names>
          </string-name>
          , GLDLP Sarrace´n, BERTa Chulvi, Elisabetta Fersini, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Profiling hate speech spreaders on twitter task at pan 2021</article-title>
          . In CLEF.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Punyajoy</given-names>
            <surname>Saha</surname>
          </string-name>
          , Binny Mathew, Pawan Goyal, and
          <string-name>
            <given-names>Animesh</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hateminers: detecting hate speech against women</article-title>
          . arXiv preprint arXiv:
          <year>1812</year>
          .06700.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Fabio Poletto, Cristina Bosco, Viviana Patti, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Stranisci</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>An italian twitter corpus of hate speech against immigrants</article-title>
          .
          <source>In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Gloria Comandini, Elisa Di Nuovo, Simona Frenda, Marco Stranisci, Cristina Bosco, Tommaso Caselli, Viviana Patti, and
          <string-name>
            <given-names>Irene</given-names>
            <surname>Russo</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Overview of the evalita 2020 second hate speech detection task (haspeede 2)</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA 2020), Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>Maarten</given-names>
            <surname>Sap</surname>
          </string-name>
          , Dallas Card, Saadia Gabriel, Yejin Choi,
          <source>and Noah A Smith</source>
          .
          <year>2019</year>
          .
          <article-title>The risk of racial bias in hate speech detection</article-title>
          .
          <source>In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>1668</fpage>
          -
          <lpage>1678</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <given-names>Kai</given-names>
            <surname>Shu</surname>
          </string-name>
          , Xinyi Zhou, Suhang Wang,
          <string-name>
            <surname>Reza Zafarani</surname>
          </string-name>
          , and Huan Liu.
          <year>2019</year>
          .
          <article-title>The role of user profiles for fake news detection</article-title>
          .
          <source>In Proceedings of the 2019 IEEE/ACM international conference on advances in social networks analysis and mining</source>
          , pages
          <fpage>436</fpage>
          -
          <lpage>439</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <given-names>Zeerak</given-names>
            <surname>Waseem</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dirk</given-names>
            <surname>Hovy</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Hateful symbols or hateful people? predictive features for hate speech detection on twitter</article-title>
          .
          <source>In Proceedings of the NAACL student research workshop</source>
          , pages
          <fpage>88</fpage>
          -
          <lpage>93</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <given-names>Hajime</given-names>
            <surname>Watanabe</surname>
          </string-name>
          , Mondher Bouazizi, and
          <string-name>
            <given-names>Tomoaki</given-names>
            <surname>Ohtsuki</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hate speech on twitter: A pragmatic approach to collect hateful and offensive expressions and perform hate speech detection</article-title>
          .
          <source>IEEE access</source>
          ,
          <volume>6</volume>
          :
          <fpage>13825</fpage>
          -
          <lpage>13835</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <given-names>Dejin</given-names>
            <surname>Zhao</surname>
          </string-name>
          and Mary Beth Rosson.
          <year>2009</year>
          .
          <article-title>How and why people twitter: the role that micro-blogging plays in informal communication at work</article-title>
          .
          <source>In Proceedings of the ACM 2009 international conference on Supporting group work</source>
          , pages
          <fpage>243</fpage>
          -
          <lpage>252</lpage>
          . ACM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>