<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Hate Speech and Topic Shift in the Covid-19 Public Discourse on Social Media in Italy</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Komal Florio</string-name>
          <email>komal.florio@unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valerio Basile</string-name>
          <email>valerio.basile@unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Viviana Patti</string-name>
          <email>viviana.patti@unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Turin</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The availability of large annotated corpora from social media and the development of powerful classification approaches have contributed in an unprecedented way to tackle the challenge of monitoring users' opinions and sentiments in online social platforms across time but also arose the challenge of temporal robustness of such detection and monitoring systems. We used as case study a dataset of tweets in Italian related to the COVID-19 induced lockdown in Italy to measure how quickly the most debated topic online shifted in time. We concluded that it is a promising approach but dedicated corpora and fine tuning of algorithms are crucial for more insightful results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        The task of abusive message detection is a very
challenging one and from multiple perspective.
From the computational point of view, despite
the increasing interest and effort of the
community on developing automatic systems abusive
language detection and related tasks for different
languages
        <xref ref-type="bibr" rid="ref16 ref19">(Poletto et al., 2021; Vidgen and
Derczynski, 2021)</xref>
        , the robustness of detection and
monitoring systems emerges as a crucial factor to be
addressed, where one of the main limitations
observed is to consider the Natural Language
Processing (NLP) task of detecting abusive language
in isolation, without taking into account the
intersection with the contextual or social dimensions,
that could contribute to a more holistic
comprehension of the abusive phenomena in language. In
fact, it is becoming increasingly evident that the
      </p>
      <p>
        Copyright © 2021 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
goodness of hate speech prediction systems, and
of NLP algorithms in general, is rooted in how
well they capture and model all the relevant
characteristics of language in the context of a specific
phenomenon and its evolution over time
        <xref ref-type="bibr" rid="ref11 ref13 ref14 ref15 ref18 ref2 ref4">(Jurafsky
and Martin, 2000; Nadkarni et al., 2011; Feldman,
2013; Schmidt and Wiegand, 2017; Fortuna and
Nunes, 2018)</xref>
        . This brought us to intersect our
NLP research with the field of Computational
Social Science.
      </p>
      <p>The recent availability of long-term and
largescale digital corpora and the effectiveness of
methods for representing words over time can play a
crucial role in the recent advances in this field. In
particular, social media have recently become one
of the predominant sources of linguistic data,
being the venue for noticeable phenomena in the
domain of NLP tasks. They represent the ideal
communication context to address the challenges we
have outlined.</p>
      <p>
        This paper aims to characterize how the
online conversation on the Italian Twitter around the
ifrst Covid-19 lockdown, imposed in Italy in 2020,
shifted very quickly from one heated debate to
another one, following the quick succession of news
reports on both news cases and institutional
advice and rules on how to navigate everyday life
as the crisis was unfolding in the entirety of the
world. At first we tried to identify the most
polarizing conversation by analyzing the presence
of hate speech using AlBERTo
        <xref ref-type="bibr" rid="ref17">(Polignano et al.,
2019)</xref>
        but we found that this BERT
        <xref ref-type="bibr" rid="ref10">(Devlin et
al., 2019)</xref>
        based algorithm, trained on Italian
Social Media language, seemed to under-perform, in
comparison with similar case studies
        <xref ref-type="bibr" rid="ref7">(Capozzi et
al., 2019)</xref>
        . We hence performed the same task
using an abusive language computational lexicon,
Hurtlex
        <xref ref-type="bibr" rid="ref4">(Bassignana et al., 2018)</xref>
        . We
discovered the most recurrent types of abusive language,
their distribution over time and correlation with
real life events regarding the ongoing pandemic.
To identify the most debated topics we resorted
to topic modeling and in particular the Dynamic
Topic Modeling allowed us to describe how the
most frequent topics evolved over time and shed
lights on the interplay with the governmental
measures that sparked the most debated conversations.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Abusive Speech Prediction</title>
      <p>
        In this work we use as case study a dataset of
tweets related to the COVID-19 induced lockdown
in Italy, as this was the perfect example of
government measures that deeply affected everyday life
of citizens and hence had the potential to spark
very heated debates online. We rely on a recently
developed resource, named 40wita 1
        <xref ref-type="bibr" rid="ref1 ref12 ref8">(Basile and
Caselli, 2020)</xref>
        , created by means of filtering with
a set of dedicated keywords the publicly available
TWITA dataset
        <xref ref-type="bibr" rid="ref2 ref4">(Basile et al., 2018)</xref>
        , a long term
collection of tweets in Italian. The filtering was
run from 1st February 2020 to 30th April 2020 and
resulted in the collection of 3309704 tweets.
AlBERTo Our first experiment to detect the
most debated conversation consisted in a hate
speech prediction with AlBERTo, using the same
set of hyper-parameters as in
        <xref ref-type="bibr" rid="ref12">(Florio et al., 2020)</xref>
        .
The findings show a peak of 6% of daily abusive
messages around mid February 2020 and at the
end of April 2020, while for the rest of the
timestamps the rates were much lower (in some cases
almost close to zero) than those found in other
Twitter-based datasets (see for example
        <xref ref-type="bibr" rid="ref8">(Capozzi
et al., 2020)</xref>
        ).
      </p>
      <p>
        Even allowing for the influence of a different
context, this finding induced us to conclude that
an unknown but not negligible percentage of
hateful messages were left undetected. We believe
that increasing the training dataset size and
quality could lead to better results. For this
experiment the data were annotated using guidelines
developed for an hate speech detection task, while
a set of new guidelines developed specifically for
this context could be a significant improvement in
the quality of the labelled data. Another possible
adjustment relies on the number of annotators and
the exploration of the best metric to compute their
disagreement, following the latest published work
on annotating subjective tasks
        <xref ref-type="bibr" rid="ref16 ref3">(Basile et al., 2021)</xref>
        .
Hurtlex In order to get a broader insight of the
hateful messages in this dataset that were
poten1https://osf.io/n39ks/
tially left out by AlBERTo, we performed the
same task by means of Hurtlex
        <xref ref-type="bibr" rid="ref4">(Bassignana et
al., 2018)</xref>
        2, a multilingual computational lexicon
that contains 17 different categories of abusive
language, each of them consisting of a list of
characterising words.
      </p>
      <p>The predominant categories of hate speech
are represented by tweets containing derogatory
words, abusive terms related to moral and
behavioural defects, and words indicating cognitive
disabilities and diversity. To gain a deeper insight
on how this classification has unfolded we
analysed which were the most common words that
classified a tweet into a specific category. Quite
often the words that determine whether a tweet
falls or not into a category, and independently
on the category, are very generic (e.g.,
“problema”=“problem”, “storia”=“history”) or can
assume very different meaning depending on the
context (e.g.: “cane”=“dog” can be used as a
derogatory term or with a neutral meaning), and
this contributes in creating a noisy tweets
classiifcation. This insight is meaningful in showing
why HurtLex presents some struggles in the
accuracy of this task. For this reason, the division
into pre-defined categories turned out to be not as
informative as we were hoping at the beginning.
An improvement on this would encompass a
manual revision of the list of words for each category,
in order to exclude the most generic ones and
retain only those which can potentially improve the
accuracy of the result. We also conducted a
manual revision of all the tweets belonging to the
categories with less than 30 tweets, while for the
other categories we choose a random sample of
30 tweets, for consistency with the previous case.
One of the most interesting findings was that in
the category “rci - locations and demonyms”, in
contrast to the global dimension of the pandemic,
our data counter-intuitively showed that the debate
was centered strictly around the measures taken in
Italy and the differences between national and
local rules.</p>
      <p>This lexicon-based approach, even though it did
not lead to the desired outcome, was nevertheless
important to gain more information on our corpus
and experience for future directions. In the next
sections we will focus on the most powerful
classiifcation tool that we employed on this dataset: two
2http://hatespeech.di.unito.it/resourc
es.html
different algorithms for unsupervised topic
modeling.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Topic Modeling</title>
      <p>We implemented two different classification
algorithms. At first we run an exploratory topics
analysis with a Latent Dirichlet Allocation (or LDA)
and then a Dynamic Topic Modeling (or DTM) to
better capture the temporal evolution of topics in
the discourse.</p>
      <p>
        Latent Dirichlet Allocation The first topic
model algorithm that we applied to our dataset is
the Latent Dirichlet Allocation, which was first
introduced by Blei
        <xref ref-type="bibr" rid="ref6">(Blei et al., 2003)</xref>
        . The
popularity and versatility of such algorithm relies on the
human-interpretable form of the extracted topics
and on being, by construction, very robust when
deployed on unseen documents.
      </p>
      <p>This model was able to correctly and precisely
identify the conversations around the first
relevant news around the incoming pandemic.
Examples of this include the first restrictions on
movements following the first Covid-19 outbreak in
Lombardy and Veneto, the national lockdown
issued in March and the consequent gradual shift of
the conversation towards the difficulties of normal
life in such a new context.</p>
      <p>As powerful as this model is, it showed a
fundamental limit for our perspective and purpose.
The relevant topics were punctual but, as expected,
not consistent over time because the model was
completely re-trained on data from every single
week, hence the results for each single time slice
were agnostic of the result for every other time
slices, and therefore not time-consistent, or
comparable, by design. To overcome this issue we
implemented a Dynamic Topic Modeling.</p>
      <sec id="sec-3-1">
        <title>Dynamic Topic Modeling The Dynamic Topic</title>
        <p>
          Modeling
          <xref ref-type="bibr" rid="ref5">(Blei and Lafferty, 2006)</xref>
          allows to split
the datasets into custom time slices and extracts
the same exact topics over all of them, thus
enabling an analysis on how topics evolve over time.
        </p>
        <p>
          At rfist we fine tuned the model by optimizing
the perplexity and the coherence score. The first
score captures the behaviour of the model towards
data which were previously unknown by means of
a normalised log-likelihood of a held-out test set.
However there are relevant studies (for example
          <xref ref-type="bibr" rid="ref9">(Chang et al., 2009)</xref>
          ) proving that perplexity and
human judgement not only often do not correlate,
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Topic No. Italian</title>
        <p>Topic 0
Topic 1
Topic 2
Topic 3
Topic 4
quarantena
altro
lavoro
governo
sanita`</p>
      </sec>
      <sec id="sec-3-3">
        <title>English</title>
        <p>quarantine
other
work
government
healthcare
but sometimes they even anti-correlate. For this
reason a second metric was elaborated: the
coherence score, to better model human judgement.
This measure captures the degree of semantic
similarity between the words related to each single
topic ( i.e., a measure of the likeness of their
meaning). We did not have an annotated corpus that
can serve as a training set, hence we only explored
the trend of the coherence score with reference
to changes in the number of topics, chunksize of
data, number of passes and evaluation score. We
then concluded for 5 topics and 20 words per
topics, as listed in the following Table 1. We chose to
leave one topic undetermined (“Topic 1 - Other”)
to label all the messages that the algorithm
struggled to correctly assign to a specific topic.</p>
        <p>The DTM outputs each unlabelled topic as a list
of words with a relevance value. This value,
between 0 and 1, represents the probability of a
single word to be affiliated with a specific topic. The
rationale behind the decision of choosing only 5
topics is that a higher number did not improve the
understanding of the corpus as it led to a noisier
classification. Each additional topic consisted of a
list of words that were either very general in their
meaning, or not very close semantically, or both,
which made it very difficult to find a topic label
that properly represented all the listed tokens.</p>
        <p>The most powerful feature of the DTM is that,
for each topic, it is possible to rank the most
relevant words based on their attached probability
value (of referring to the specific topic) and see
how they evolve over time. In the following
Figure 1, the change in ranking for all the 20 words
involved is presented as a coloured heatmap, where
the blue values represents words with higher
ranking while the red ones are at the lower end of
ranking.</p>
        <p>There are two main insights we can gain from
this visualization. The first one is that topics
Topic 0 "Quarantine" Words Ranking
are lists of pretty common words, which proves
how hard of a task topic detection is, because of
the complexity and versatility of human language,
where general words can be used in different
contexts with different meanings. The second insight
is that the biggest changes in the word ranking
happen within the first time slices. A possible
explanation may be traced back to how this dataset
was created. The list of hashtags and trends used
to filter the tweets was compiled in February and
was fixed in time. This means that potentially
interesting tweets were left out because they
contained hashtags that emerged as relevant later in
time but hence were not captured by the keywords
used for selecting relevant tweets.</p>
        <p>In order to measure the temporal trend of
predominance for each topics, we computed, for each
of the 13 time slices, the ratio of documents
labeled as predominantly referring to each of the
topic.</p>
        <p>We plotted in Figure 2 the normalized share of
documents classified as containing each of the
topics in each time slices, to highlight the relative
trends over time.</p>
        <p>Topics Distribution Sorted by Mean/Max Value
topic 0 = quarantine
topic 1 = other
topic 2 = employment
topic 3 = government
topic 4 = healthcare
1.0
0.8
is that the discourse on Twitter does not only
follow closely the most recent and relevant news but
it quickly shifts from one topic to the other. In
fact, all major peaks in Fig. 2 are followed by a
sharply decreasing trend, indicating an immediate
loss of predominance and hence an alternation of
the dominant arguments of debates.</p>
        <p>We explored in a similar way also the temporal
evolution of the share of tweets labelled with the
Hurtlex categories.</p>
        <p>For each of the time slices we computed the
relative frequency of tweets labeled with every
categories and then created a stacked plot of their
maximum values (shown in Figure 3) and the
normalized mean values (shown in Figure 4) of their
frequencies, to identify both peaks and categories
that were consistently predominant through the
time.</p>
        <p>The relevance of the Hurtlex category related to
derogatory words detected over the whole dataset,
as described in Section 2, confirms its validity also
at a weekly time granularity, as shown by Figure
3. Looking at the chart as a whole it is important to
notice that, as we have already highlighted before,
the peaks occur in time slices 3 and 5, which
respectively correspond the the issue of the first red
zones in Italy and two major public health news
regarding Lombardy, the hardest hit region of Italy
in the first phases of the pandemic (see Table 2 for
details).</p>
        <p>It is relevant to notice that these peaks occur
exactly in the same time slices as the peaks in
Figure 2 for the topics ”quarantine” and
”healthcare”, showing that the most heated debates
happened around public measures that affected
directly and immediately on both the collectivity
(”healthcare”) and personal life (”quarantine”).
Analysing the mean value of the frequencies, in
Figure 4, we can see that categories rank
differently from Figure 3. More specifically we see that
for example ”ddf - physical disabilities and
diversity” is by far the most consistent over time but
it represents somehow a generic type of offensive
language, not correlated with the pandemic, and
to some extent this is as a noisy classification of
tweets and it would be interesting to investigate
further how to improve on this result.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Final Remarks</title>
      <p>In this work we tried to tackle the challenge of
measuring and quantifying the topic shift in the
public discourse on Social Media, using as a case
study the online debate on Twitter following the
Covid-19 related lockdown in Italy in 2020, by
means of a dedicated dataset. By combining
multiple classification methods we gathered insights
into which governmental measures generated the
most debated online conversation but we also
concluded for the need of deeper investigation on how
to build ad hoc corpora and methods to investigate
specific linguistic phenomena as online
conversation with rapid topic shift following the flow of
news coming from both online and traditional
media outlets. We also tried to inform AlBERTo with
information extracted from topic modeling but the
results were far from satisfying. This is a
promising way to enhance the accuracy of hate speech
prediction, but we concluded that a further
investigation on size and characteristics of datasets is
essential to gain better results.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          and
          <string-name>
            <given-names>Tommaso</given-names>
            <surname>Caselli</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>40twita 1.0: A collection of Italian Tweets during the COVID-</article-title>
          19 Pandemic.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Mirko Lai, and
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Long-term Social Media Data Collection at the University of Turin</article-title>
          .
          <source>In Proceedings of the Fifth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2018</year>
          ), volume
          <volume>2253</volume>
          <source>of CEUR Workshop Proceedings</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          , Torino,Italy. CEURWS.org.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio,
          <string-name>
            <given-names>Alexandra</given-names>
            <surname>Uma</surname>
          </string-name>
          , et al.
          <year>2021</year>
          .
          <article-title>We need to consider disagreement in evaluation</article-title>
          .
          <source>In 1st Workshop on Benchmarking: Past, Present and Future</source>
          , pages
          <fpage>15</fpage>
          -
          <lpage>21</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Bassignana</surname>
          </string-name>
          , Valerio Basile, and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hurtlex: A multilingual lexicon of words to hurt</article-title>
          .
          <source>In Proceedings of the Fifth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2018</year>
          ), Torino, Italy,
          <source>December 10-12</source>
          ,
          <year>2018</year>
          , volume
          <volume>2253</volume>
          <source>of CEUR Workshop Proceedings</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . CEURWS.org.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>David M Blei and John D Lafferty</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Dynamic topic models</article-title>
          .
          <source>In Proceedings of the 23rd international conference on Machine learning</source>
          , pages
          <fpage>113</fpage>
          -
          <lpage>120</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>David M Blei</surname>
            , Andrew Y Ng, and
            <given-names>Michael I</given-names>
          </string-name>
          <string-name>
            <surname>Jordan</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>3</volume>
          :
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          ,
          <year>March</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Arthur TE Capozzi</surname>
            , Mirko Lai, Valerio Basile, Fabio Poletto, Manuela Sanguinetti, Cristina Bosco, Viviana Patti, Giancarlo Ruffo, Cataldo Musto,
            <given-names>Marco</given-names>
          </string-name>
          <string-name>
            <surname>Polignano</surname>
          </string-name>
          , et al.
          <year>2019</year>
          .
          <article-title>Computational linguistics against hate: Hate speech detection and visualization on social media in the” contro l'odio” project</article-title>
          .
          <source>In 6th Italian Conference on Computational Linguistics</source>
          , CLiC-it
          <year>2019</year>
          , volume
          <volume>2481</volume>
          , pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . CEUR-WS.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Arthur TE Capozzi</surname>
            , Mirko Lai, Valerio Basile, Fabio Poletto, Manuela Sanguinetti, Cristina Bosco, Viviana Patti, Giancarlo Ruffo, Cataldo Musto,
            <given-names>Marco</given-names>
          </string-name>
          <string-name>
            <surname>Polignano</surname>
          </string-name>
          , et al.
          <year>2020</year>
          .
          <article-title>“contro l'odio”: A platform for detecting, monitoring and visualizing hate speech against immigrants in italian social media</article-title>
          .
          <source>IJCoL</source>
          .
          <source>Italian Journal of Computational Linguistics</source>
          ,
          <volume>6</volume>
          (
          <issue>6</issue>
          -1):
          <fpage>77</fpage>
          -
          <lpage>97</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Jonathan</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            Boyd-Graber,
            <given-names>Chong</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Sean Gerrish</surname>
          </string-name>
          , and David M Blei.
          <year>2009</year>
          .
          <article-title>Reading tea leaves: How humans interpret topic models</article-title>
          .
          <source>In Proceedings of the 22nd International Conference on Neural Information Processing Systems</source>
          , NIPS'
          <volume>09</volume>
          , page 288-296,
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA. Curran Associates Inc.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In Proceedings of the</source>
          <year>2019</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis</article-title>
          , MN, USA, June 2-7,
          <year>2019</year>
          , Volume
          <volume>1</volume>
          (Long and Short Papers), pages
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          , Minneapolis, MN, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Ronen</given-names>
            <surname>Feldman</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Techniques and applications for sentiment analysis</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>56</volume>
          (
          <issue>4</issue>
          ):
          <fpage>82</fpage>
          -
          <lpage>89</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Komal</given-names>
            <surname>Florio</surname>
          </string-name>
          , Valerio Basile, Marco Polignano, Pierpaolo Basile, and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Time of your hate: The challenge of time in hate speech detection on social media</article-title>
          .
          <source>Applied Sciences</source>
          ,
          <volume>10</volume>
          (
          <issue>12</issue>
          ):
          <fpage>4180</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Paula</given-names>
            <surname>Fortuna</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sergio</given-names>
            <surname>Nunes</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A survey on automatic detection of hate speech in text</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>51</volume>
          (
          <issue>4</issue>
          ):
          <fpage>85</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Jurafsky and James H. Martin</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Speech and Language Processing: An Introduction to Natural Language Processing</article-title>
          , Computational Linguistics, and
          <string-name>
            <given-names>Speech</given-names>
            <surname>Recognition. Prentice Hall</surname>
          </string-name>
          <string-name>
            <surname>PTR</surname>
          </string-name>
          , USA, 1st edition.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Prakash M Nadkarni</surname>
            ,
            <given-names>Lucila</given-names>
          </string-name>
          <string-name>
            <surname>Ohno-Machado</surname>
          </string-name>
          , and Wendy W Chapman.
          <year>2011</year>
          .
          <article-title>Natural language processing: an introduction</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          ,
          <volume>18</volume>
          (
          <issue>5</issue>
          ):
          <fpage>544</fpage>
          -
          <lpage>551</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Poletto</surname>
          </string-name>
          , Valerio Basile, Manuela Sanguinetti, Cristina Bosco, and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Resources and benchmark corpora for hate speech detection: a systematic review</article-title>
          .
          <source>Lang. Resour. Evaluation</source>
          ,
          <volume>55</volume>
          (
          <issue>2</issue>
          ):
          <fpage>477</fpage>
          -
          <lpage>523</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Polignano</surname>
          </string-name>
          , Valerio Basile, Pierpaolo Basile, Marco de Gemmis, and
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Semeraro</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>AlBERTo: Modeling Italian Social Media Language with BERT</article-title>
          .
          <source>Italian Journal of Computational Linguistics - IJCOL</source>
          , -
          <volume>2</volume>
          , n.2.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Anna</given-names>
            <surname>Schmidt</surname>
          </string-name>
          and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Wiegand</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A survey on hate speech detection using natural language processing</article-title>
          .
          <source>In Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          , Valencia, Spain, April. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Bertie</given-names>
            <surname>Vidgen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Leon</given-names>
            <surname>Derczynski</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Directions in abusive language training data, a systematic review: Garbage in, garbage out</article-title>
          .
          <source>PLOS ONE</source>
          ,
          <volume>15</volume>
          (
          <issue>12</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>32</lpage>
          ,
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>