<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Effective Detection of Hate Speech Spreaders on Twitter</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julian Höllig</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yeong Su Lee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nina Seemann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michaela Geierhos</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Research Institute CODE, Bundeswehr University Munich</institution>
          ,
          <addr-line>Neubiberg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <abstract>
        <p>In this paper, we summarize our participation in the task of “Profiling Hate Speech Spreaders on Twitter” at the PAN@CLEF Conference 2021. Our models obtained an average accuracy of 76% (79% for Spanish and 73% for English). For English, we used a Linear Support Vector Machine with tf-idf features on noun chunk level, while for Spanish we used a Ridge Classifier with simple counts on noun chunk level. Both classifiers were fed with additional features obtained from a Convolutional Neural Network.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;author profiling</kwd>
        <kwd>hate speech</kwd>
        <kwd>noun chunks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In recent years, there has been growing awareness that hate speech became an increasing issue in
social media, which offers anonymity and virality to authors of hateful posts. On average, Twitter
had 199 million daily active users in the first quarter of 2021, compared to 166 million active
users counted the year before, which is an increase of almost 20 percent [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. These developments
caused political forces and social media providers to act. For example, the EU initiated measures
such as the European Council’s “No Hate Speech Movement”, which aims at mobilizing online
consumers to take action against hate [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The EU, with the involvement of YouTube, Twitter,
Facebook, and Microsoft, has also drafted a “Code of conduct on countering illegal hate speech
online” [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], in which these companies commit to check hate speech notifications within 24
hours [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. However, the most effective and efcfiient way to combat hate speech is through its
automatic detection, where machine learning applications can play a crucial role.
      </p>
      <p>
        In the PAN shared task [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], international researchers focus on modeling such applications to
identify hate speech spreaders on Twitter. Compared to other hate speech detection tasks [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ],
the data for this task consist of tweet collections belonging to the same author, each representing
a sample in the dataset. One profound challenge in detecting hate speech is its unclear definition.
Table 1 compares the attitudes of political, social media, and scientific representatives on four
important questions about the definition of hate speech [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The contents of the table were
retrieved from ofcfiial sources of the institutions and from scientific papers [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. While there
is agreement that hate speech is directed to specific targets, the attitudes differ regarding the
influence of humor, whether or not hate speech is intended to incite hate, and whether or not hate
speech is intended to attack and disparage. This leads to critically different definitions of hate
speech and how to combat it. For example, according to Table 1, YouTube would not define a
verbal attack on a person as hate speech, while inciting violence against the same person would
be considered as hate speech. Facebook would take the opposite position, according to Table 1.
Due to the broad definition of hate speech, it is difficult to find large harmonized data collections,
since annotators often have low agreement when building new collections [9, 10, as cited in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]].
Consequently, it is challenging to establish a common standard for modeling hate speech so far.
      </p>
      <p>The paper is organized as follows. In Section 2, we present relevant related work. In Section 3,
we explain our methods in detail and present the evaluation of the final models in Section 4.
Finally, we conclude in Section 5.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>In recent years, the detection of hate speech has been of great interest for many researches. As a
result, there is a vast literature on this topic. In the following, we will only focus on some other
shared tasks and their results.</p>
      <p>
        Mandl et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] describe the identification of hate speech and offensive content in
IndoEuropean languages1 at FIRE 2019. There were three subtasks: subtask A was the coarse-grained
binary classification into non Hate-Offensive and Hate &amp; Offensive (HOF). If a post was classified
as HOF, then it was handled by subtask B, which further classified it into either hate speech,
offensive, or profane. Subtask C addressed the targeting and non-targeting of individuals, groups,
or others when a post was classified as HOF. For example, the team with the best performance on
the English data achieved a macro F1 score of 78.82% and a weighted F1 score of 83.95% for
subtask A, a macro F1 of 54.46% and a weighted F1 of 72.77% for subtask B, and a macro F1 of
51.11% and a weighted F1 of 75.63% for subtask C.
      </p>
      <p>
        HatEval [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] consists of detecting hateful content in Twitter posts for English and Spanish.
There were two subtasks: subtask A focused on detecting hate speech against immigrants and
1The three languages provided were English, Hindi, and German.
women, i.e., a binary classification into hateful or not. In subtask B, a fine-grained classification
had to be performed. Hateful tweets had to be further classified in terms of (i) aggressive attitude,
i.e., is a tweet aggressive or non-aggressive, and (ii) target classification, i.e., is a specific target
harassed or a generic group. Both tasks in subtask B are binary. For subtask A, the best systems
obtained a macro-averaged F1 score of 0.651 for English and a macro-averaged F1 score of 0.73
for Spanish. For subtask B, the best systems achieved an Exact Match Ratio (EMR) of 0.570 for
English and 0.705 for Spanish.
      </p>
      <p>
        Bosco et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] describe the hate speech shared task at the Sixth Evaluation Campaign of
Natural Language Processing and Speech Tools for Italian (Evalita) in 2018. Both tasks, Evalita
and PAN, focused on detecting hate speech on social media (Facebook and Twitter). However,
Evalita targeted hate speech detection at the post level, i.e., tweets. The PAN task focused on
identifying an author as hate speech spreader based on a collection of his/her tweets, which
introduced additional fuzziness to the challenging task of defining and detecting hate speech.
Nine out of ten Evalita participants used external resources to improve their systems, such as
pre-trained embeddings, dictionaries, and datasets related to the task. The winning team achieved
an F1 score of almost 80% by using additional data from a subjectivity and polarity lexicon and
the SENTIPOLC dataset on sentiment analysis along with linear SVM and BiLSTM models. In
our work, we also successfully experimented with additional external data. However, the data was
not directly used to identify hate speech spreaders, but was used to score the tweets themselves to
create a ‘hate weight’ for each author.
      </p>
      <p>
        Struß et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] summarize the 2019 GermEval shared task on identifying offensive language
in Twitter data, where ‘offensive’ is defined as insulting, abusive, or profane language. The shared
task was divided in three subtasks: (1) a binary classification into offensive and non-offensive
tweets, (2) a multi-classification into insulting, abusive, profane, and non-offensive tweets, and
(3) a binary classification of offensive tweets into implicitly and explicitly offensive. The dataset
consisted of 7,000 tweets in total (4,000 train set, 3,000 test set). For subtask (2), the offensive
tweets were divided into the three indicated groups. For subtask (3), 2,900 offensive tweets were
split into 400 implicitly and 2,500 explicitly offensive tweets. The best performing system on all
subtasks was a BERT model, which achieved 77%, 54%, and 73% in macro average F1 score (on
subtasks (1), (2), and (3)). It was pre-trained on six million German tweets and fine-tuned on the
GermEval data. The average performances of all participants obtained on the subtasks were 72%,
47%, and 67%.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methods</title>
      <p>We experimented with different approaches and methods, using both deep learning, i.e. neural
networks, and machine learning, i.e. more traditional algorithms. In the following sections, we
provide an overview of the PAN dataset and its challenges before describing in detail the steps
taken to obtain our final results for the task.</p>
      <sec id="sec-3-1">
        <title>3.1. Dataset</title>
        <p>For both English and Spanish, the organizers provided us with a dataset containing 200 tweets
for 200 authors each, indicating whether or not the author is considered as hate speech spreader.
This classifies 100 authors as hate speech spreaders and 100 as legitimate users. In total, the
dataset contains 40,000 tweets for each language. An overview of the dataset can be found in
Table 2. What makes this author profiling task so challenging is the fact that the tweets collected
for an author are neither all hate speech nor all harmless. Hence, not all of the 200 tweets per hate
speech spreader are hateful per se. Information on the annotation scheme or threshold (i.e., the
minimum number of hateful tweets per author) for classifying an author as hate speech spreader
was not provided at the time of the competition.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Additional Features</title>
        <p>
          Inspired by the challenging mix of hate speech and harmless tweets per author, we developed
additional features to weigh the amount of hateful content produced by each author. Therefore,
we relied on external data, which we describe in Section 3.2.1. In recent years, many NLP
applications — including text classification tasks — have been significantly improved by the use
of deep learning algorithms. Hence, we decided to train a Convolutional Neural Network (CNN,
[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]) with external data and applied the resulting model on the PAN data at the tweet level to
generate additional features. In the following subsections, we describe this process in more detail.
        </p>
        <sec id="sec-3-2-1">
          <title>3.2.1. External data</title>
          <p>
            To create the additional features described in Section 3.2, we searched for external data containing
hate speech or hate speech related concepts such as offensive language. Since the PAN dataset
consists of tweets, we looked for external data also retrieved from Twitter. Our search resulted
in several good sources for English. We chose the data from (i) CONAN [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ], (ii) Davidson et
al. [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ], (iii) HASOC track [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ], and (iv) SemEval-2019 Task 5 HatEval [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]. Unfortunately,
there was not much data available for Spanish, so we only used the Spanish part of the HatEval
dataset [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]. An overview of the sizes and the percentage of offensive tweets in the datasets is
given in Table 3.
          </p>
        </sec>
        <sec id="sec-3-2-2">
          <title>3.2.2. Training of the CNN</title>
          <p>
            We used the Tensorflow/Keras API 2 for the implementation of the CNN. Since the input for neural
networks does not require much preprocessing, we simply lowercased the external dataset and
removed unwanted characters, symbols, and emojis. We implemented a character-level model
with 49 features for English and 57 features for Spanish. The number of features is determined
by the number of different characters present after preprocessing. After examining the character
length of all tweets, we set the maximum sequence length to 300 per tweet. Keras provides a
tokenizer that converts the input into a list of integers (similar to the bag-of-words approach) and
a method to convert these integers into a fixed-length vector. This means that tweets containing
more than 300 characters were reduced in size and shorter tweets were padded with zeros to
maximum length. We trained the CNN with the following setting:
• filter size: [
            <xref ref-type="bibr" rid="ref5 ref6 ref7">5,6,7</xref>
            ]
• number of filters: 100
• activation: ReLU
• output: sigmoid
We used the Adam optimizer [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] and trained the CNN for 80 epochs. After the last epoch,
the model had an accuracy of 91.01% on the English external data and 92.12% on the Spanish
external data.
          </p>
        </sec>
        <sec id="sec-3-2-3">
          <title>3.2.3. Applying the model to the PAN data</title>
          <p>For each author, we let the model obtained by the CNN predict the class for each of his/her tweets.
We recorded the values of the predictions as ‘HateCounts’ and ‘LoveCounts’. After classifying
all tweets, we used the majority vote on these counts to predict whether an author was a hate
speech spreader or not. In numbers, whenever ‘HateCount’ was &gt; 100, the author was classified
as hate speech spreader and vice versa. Unfortunately, this resulted in an accuracy of about 0.5
for both languages. So we decided to iteratively lower the threshold of ‘HateCount’ from 100 to 0
to get better accuracy. For English, a threshold of 48 gave the best accuracy of 67.58%, while we
obtained the best accuracy of 71.5% with a threshold of 33 for Spanish. Since these performances
were not convincing, we decided to move to more traditional machine learning algorithms (see
2https://www.tensorflow.org/api_docs/python/tf/keras
Section 3.3). Unlike deep learning models, traditional machine learning models have no sequence
length limit, so we could use them to process all tweets per author at once. However, since the
‘HateCounts’/‘LoveCounts’ showed at least some influence on the classification of hate speech
spreaders, we kept them as additional features for the subsequent experiments with the traditional
models. Furthermore, we obtained the class probability calculated by the CNN for each tweet
and also stored the mean probabilities for each author (’ProbMean’).</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Experiments</title>
        <p>In the following sections, we describe our experiments using traditional methods that led us to
our final models. All experiments were implemented in Python using the scikit-learn library 3.</p>
        <sec id="sec-3-3-1">
          <title>3.3.1. Experimental setup A</title>
          <p>In our first setup, we tested vfie algorithms for different features. As preprocessing, we combined
all tweets of an author into one document. To mark the beginning of a new tweet, we added a
special start-of-tweet token. Furthermore, the data was lowercased. As features, we used simple
counts of bag-of-words and tf-idf at the word and character level. To reduce the feature space,
we limited the maximum number of features to 5,000. Additional features were not considered
for these experiments. We split the data into 70% for training and 30% for testing, using three
different random states for splitting. The experimental results are presented in Table 4. The
values shown are the mean accuracy for the different data splits. The results were not sufficient,
but gave us a reasonable basis on which to improve further.</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>3.3.2. Experimental setup B</title>
          <p>For our second setup, we used the same setup as for experimental setup A, but fed the ‘LoveCounts’
and the ‘ProbMeans’ into the models as additional features to improve our performance. Since
the PAN dataset contains only 200 instances in total, we also applied oversampling to the training
3https://scikit-learn.org/stable/index.html
data using the BorderlineSMOTE library from imbalanced-learn4. In doing so, we performed
oversampling for both classes rather than just one class (as usually done by SMOTE). Preliminary
results showed that an oversampling of 500 gave the best results. Again, we ran experiments with
the same vfie algorithms, but this time with vfie different random states for splitting into training
and test sets. For English, the results obtained are shown in Table 5. The results show that both
the additional features and oversampling increased the model performances.</p>
        </sec>
        <sec id="sec-3-3-3">
          <title>3.3.3. Experimental setup C</title>
          <p>
            For our third setup and final experiments, we optimized our data preprocessing to achieve further
improvements. The following experiments were performed for both English and Spanish, but for
brevity we present only the English results. To this end, we first removed stop words by using the
stop word list provided by the spaCy library5. Second, we used an emoji sentiment recognizer6
based on the research by Novak et al. [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ]. We replaced positive emojis with ‘EMOJI-positive’,
negative emojis with ‘EMOJI-negative’, and the remaining ones with ‘EMOJI’7. Finally, we
extracted all noun chunks from the tweets by again using the spaCy library8. These noun chunks
were then used as input for our experiments. Each word of a noun chunk was lemmatized.
          </p>
          <p>As shown in Table 6, we added more algorithms in addition to those from the previous setups.
We used count and tf-idf vectors based on noun chunks as features, both in combination with
the ‘LoveCount’ and ‘ProbMean’ features, as these were the most promising from our previous
results. We did not limit the number of features for this setup. The mean accuracy for the English
test set can be found in Table 6.</p>
          <p>4https://imbalanced-learn.org/stable/
5https://spacy.io/usage/spacy-101#language-data
6https://github.com/FLAIST/emosent-py. This was adapted, further developed, and made available for our work
by D. Schwimmbeck from the Research Institute CODE, Bundeswehr University Munich, Germany.</p>
          <p>7We did not distinguish between one or more consecutive emojis: All consecutive occurrences of the same emoji
were replaced by a corresponding counterpart.</p>
          <p>8https://spacy.io/usage/linguistic-features#noun-chunks
word</p>
        </sec>
        <sec id="sec-3-3-4">
          <title>3.3.4. Beyond n-gram features</title>
          <p>As shown in Table 6, most models with tf-idf on noun chunks outperform their counterparts with
bag-of-words count. Moreover, the models of the Ridge Classifier and the Linear SVM show
remarkable improvements. Hence, we compared the feature importance for tf-idf vectors from
n-gram and noun chunks models. Table 7 shows the vfie most positive and negative features
with their weights9. While all n-gram models indicate the importance of the words president and
trump, they are not listed in the vfie most important features of the noun chunks model 10. On
the other hand, it should be also noted that the word white alone receives a high weighting in the
1-gram and 1- &amp; 2-gram models; the word sequence white people makes more sense in the 2- &amp;
3-gram and noun chunks models.
9n-grams were obtained by the standard word analyzer with stop words elimination for English, while the noun
chunks by spaCy noun chunks. The first word was removed from them if it is a stop word.</p>
          <p>10The word sequence president trump occurs only in the 18th position.</p>
        </sec>
        <sec id="sec-3-3-5">
          <title>3.3.5. spaCy noun chunks and feature extraction</title>
          <p>Table 7 shows some awkward noun chunks like illegal and y’. Therefore, we examined the posts
containing these words in more detail. The word illegal occurs in 185 different posts. However,
among these, the word illegals occurs in 54 different posts in 21 training documents. The noun
chunk illegal refers to this use of the word. Of these, only 4 documents were classified as non-hate
speech spreader. The other 21 documents were classified as hate speech spreader. From this fact,
it can be concluded that the word illegal could be weighted highly for the task. In the following,
some examples from the training documents are given:
• The illegals took most of the blacks jobs!
• This is why I don’t want more illegals in the USA until we take care of ALL our Vets!
• Why are you protecting illegals?</p>
          <p>For example, the spaCy noun chunk module tagged The illegals as noun chunks from the first
sentence. But as described in 3.3.4, the first word The is a stop word and was removed (see 3.3.3),
and the word illegals has been lemmatized to illegal11.</p>
          <p>The word y’ occurs only in the combination of y’all in 266 different posts, and the spaCy noun
chunk module separates the word y’ as noun chunk from the word all. For example, spaCy assigns
noun chunks to three units (I, y’, college) from the following sentence taken from a training post:
I can see why some of y’all ain’t go to college.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation on the PAN Test Sets</title>
      <p>For the final evaluation on the test set provided by the PAN2021 shared task organizers, we
applied the following algorithms and settings:
• English: Linear SVM with tf-idf vectors on noun chunks
• Spanish: Ridge Classifier with count vectors on noun chunks</p>
      <p>
        On the English test set, we achieved an accuracy of 73%, on the Spanish test set an accuracy
of 79%. The overall average accuracy for both languages was 76%. According to the PAN2021
Overview [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], this corresponds to rank 8. Due to technical issues, we could not use TIRA[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
We sent our results to the organizers by email.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>By retrieving noun chunks from the data, we employed a specific linguistic feature that proved its
suitability for the classification task. As shown, it outperforms most methods based on n-gram
features. In contrast to n-gram features, noun chunks form linguistically meaningful units and
are therefore more comprehensible. Furthermore, we have developed the additional features
‘LoveCount’ and ‘ProbMean’, which add a kind of a ‘hate weight’ to each author.</p>
      <p>In future work, we want to explore how linguistic units such as phrasal expressions, and
specifically predicate-argument structures, can be obtained and embedded, and to what extent
they can contribute to the classification and other tasks related to textual content.</p>
      <p>11https://spacy.io/usage/linguistic-features#lemmatization</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This research is partially funded by dtec.bw – Digitalization and Technology Research Center of
the Bundeswehr within the project MuQuaNet.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Yahoo</surname>
          </string-name>
          ! Finance,
          <source>Twitter Announces First Quarter 2021 Results</source>
          ,
          <year>2021</year>
          . URL: https://s22. q4cdn.com/826641620/files/doc_financials/
          <year>2021</year>
          /q1/Q1'
          <fpage>21</fpage>
          -
          <string-name>
            <surname>Earnings-Release</surname>
          </string-name>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] Council of Europe,
          <source>No Hate Speech Youth Campaign</source>
          ,
          <year>2017</year>
          . URL: https://www.coe.int/en/ web/no-hate-campaign.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>European</given-names>
            <surname>Commission</surname>
          </string-name>
          ,
          <source>The EU Code of conduct on countering illegal hate speech online</source>
          ,
          <year>2020</year>
          . URL: https://ec.europa.eu/info/policies/ justice-and
          <article-title>-fundamental-rights/combatting-discrimination/racism-and-xenophobia/ eu-code-conduct-countering-illegal-hate-speech-online_en.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hern</surname>
          </string-name>
          , Facebook, YouTube,
          <source>Twitter and Microsoft sign EU hate speech code</source>
          ,
          <year>2016</year>
          . URL: https://www.theguardian.com/technology/2016/may/31/ facebook-youtube
          <article-title>-twitter-microsoft-eu-hate-speech-code.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. L. D. L. P.</given-names>
            <surname>Sarracén</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chulvi</surname>
          </string-name>
          , E. Fersini,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <source>Profiling Hate Speech Spreaders on Twitter Task at PAN</source>
          <year>2021</year>
          ,
          <article-title>in: CLEF 2021 Labs and Workshops, Notebook Papers, CEUR-WS</article-title>
          .org,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Poletto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Tesconi, Overview of the EVALITA 2018 hate speech detection task</article-title>
          , in: T. Caselli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Novielli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          , P. Rosso (Eds.),
          <source>Proceedings of the Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          )
          <article-title>co-located with the Fifth Italian Conference on Computational Linguistics (CLiC-it</article-title>
          <year>2018</year>
          ), Turin, Italy,
          <source>December 12-13</source>
          ,
          <year>2018</year>
          , volume
          <volume>2263</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2018</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2263</volume>
          /paper010.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Al-Khalifa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Magdy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Darwish</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Elsayed</surname>
          </string-name>
          , H. Mubarak (Eds.),
          <source>Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools</source>
          ,
          <article-title>with a Shared Task on Offensive Language Detection</article-title>
          , European Language Resource Association, Marseille, France,
          <year>2020</year>
          . URL: https://www.aclweb.org/anthology/2020.osact-
          <volume>1</volume>
          .0.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Fortuna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nunes</surname>
          </string-name>
          ,
          <article-title>"a survey on automatic detection of hate speech in text"</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>51</volume>
          (
          <year>2018</year>
          )
          <fpage>85</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B.</given-names>
            <surname>Roß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rist</surname>
          </string-name>
          , G. Carbonell, B.
          <string-name>
            <surname>Cabrera</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Kurowsky</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>M. Wojatzki, Measuring the Reliability of Hate Speech Annotations: The Case of the European Refugee Crisis</article-title>
          ,
          <source>Originally published in Bochumer Linguistische Arbeitsberichte</source>
          <volume>17</volume>
          ,
          <string-name>
            <surname>NLP4CMC</surname>
            <given-names>III</given-names>
          </string-name>
          :
          <article-title>3rd Workshop on Natural Language Processing for Computer-Mediated Communication, by Michael Beißwenger, Michael Wojatzki</article-title>
          and Torsten Zesch (Eds.),
          <source>22 September 2016 (ISSN 2190-0949)</source>
          . (
          <year>2016</year>
          )
          <fpage>6</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Waseem</surname>
          </string-name>
          ,
          <article-title>Are You a Racist or Am I Seeing Things? Annotator Influence on Hate Speech Detection on Twitter</article-title>
          .,
          <source>in: Proceedings of the First Workshop on NLP and Computational Social Science</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>138</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Modha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mandlia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <article-title>Overview of the hasoc track at fire 2019: Hate speech and offensive content identification in indoeuropean languages</article-title>
          ,
          <source>in: Proceedings of the 11th Forum for Information Retrieval Evaluation</source>
          , FIRE '19,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2019</year>
          , p.
          <fpage>14</fpage>
          -
          <lpage>17</lpage>
          . doi:
          <volume>10</volume>
          .1145/3368567.3368584.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fersini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nozza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Rangel Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , M. Sanguinetti, SemEval
          <article-title>-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter</article-title>
          ,
          <source>in: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Minneapolis, Minnesota, USA,
          <year>2019</year>
          , pp.
          <fpage>54</fpage>
          -
          <lpage>63</lpage>
          . URL: https://www.aclweb.org/anthology/S19-2007. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>S19</fpage>
          -2007.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>J. M. Struß</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Siegel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Ruppenhofer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Wiegand</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Klenner, Overview of GermEval Task 2, 2019 Shared Task on the Identification of Offensive Language</article-title>
          , in: G. S. for Computational Linguistics (Ed.),
          <source>Proceedings of the 15th Conference on Natural Language Processing (KONVENS)</source>
          <year>2019</year>
          , s.a.,
          <source>Nürnberg/Erlangen</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>354</fpage>
          -
          <lpage>365</lpage>
          . doi:
          <volume>10</volume>
          .5167/ uzh-178687.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y.</given-names>
            <surname>LeCun</surname>
          </string-name>
          , L. Bottou,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Haffner</surname>
          </string-name>
          ,
          <article-title>Gradient-based learning applied to document recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE</source>
          ,
          <year>1998</year>
          , pp.
          <fpage>2278</fpage>
          -
          <lpage>2324</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Y.-L. Chung</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Kuzmenko</surname>
            ,
            <given-names>S. S.</given-names>
          </string-name>
          <string-name>
            <surname>Tekiroglu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Guerini, CONAN - COunter NArratives through nichesourcing: a multilingual dataset of responses to fight online hate speech, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>2819</fpage>
          -
          <lpage>2829</lpage>
          . URL: https://www.aclweb.org/anthology/P19-1271. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P19</fpage>
          -1271.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Davidson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Warmsley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Macy</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Weber</surname>
          </string-name>
          ,
          <article-title>Automated hate speech detection and the problem of offensive language</article-title>
          ,
          <source>in: Proceedings of the 11th International AAAI Conference on Web and Social Media</source>
          ,
          <source>ICWSM '17</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>512</fpage>
          -
          <lpage>515</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochastic optimization</article-title>
          , in: Y. Bengio, Y. LeCun (Eds.),
          <source>3rd International Conference on Learning Representations, ICLR</source>
          <year>2015</year>
          , San Diego, CA, USA, May 7-
          <issue>9</issue>
          ,
          <year>2015</year>
          , Conference Track Proceedings,
          <year>2015</year>
          . URL: http: //arxiv.org/abs/1412.6980.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kralj Novak</surname>
          </string-name>
          , J. Smailovic´,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sluban</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Mozeticˇ</surname>
          </string-name>
          , Sentiment of emojis,
          <source>PLOS ONE 10</source>
          (
          <year>2015</year>
          )
          <article-title>e0144296</article-title>
          . doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0144296</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chulvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. L. D. L. P.</given-names>
            <surname>Sarracén</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Manjavacas</surname>
          </string-name>
          , I. Markov,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wolska</surname>
          </string-name>
          , , E. Zangerle, Overview of PAN 2021:
          <article-title>Authorship Verification,Profiling Hate Speech Spreaders on Twitter,and Style Change Detection</article-title>
          ,
          <source>in: 12th International Conference of the CLEF Association (CLEF</source>
          <year>2021</year>
          ), Springer,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gollub</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          , TIRA Integrated Research Architecture, in: N.
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          Peters (Eds.),
          <source>Information Retrieval Evaluation in a Changing World, The Information Retrieval Series</source>
          , Springer, Berlin/Heidelberg/New York,
          <year>2019</year>
          , pp.
          <fpage>123</fpage>
          -
          <lpage>160</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -22948-
          <issue>1</issue>
          _
          <fpage>5</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>