<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mining Annotator Perspectives from Hate Speech Corpora</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>ª¨©Universita degli Studi di Torino</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Turin</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy ªmic.fell@gmail.com</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>firstname.lastnameg@unito.it</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Disagreement in annotation, traditionally treated mostly as noise, is now more and more often considered as a source of valuable information instead. We investigate a particular form of disagreement, occurring when the focus of an annotated dataset is a subjective and controversial phenomenon, therefore inducing a certain degree of polarization among the annotators' judgments. We argue that the polarization is indicative of the con icting perspectives held by di erent annotator groups, and propose a quantitative method to model this phenomenon. Moreover, we introduce a method to automatically identify shared perspectives stemming from a common background. We test our method on several corpora in English and Italian, manually annotated according to their hate speech content, validating prior knowledge about the groups of annotators, when available, and discovering characteristic traits among annotators with unknown background. We found several precisely dened perspectives, described in terms of increased sensitivity towards textual content expressing attitudes such as xenophobia, islamophobia, and homophobia.</p>
      </abstract>
      <kwd-group>
        <kwd>Linguistic Annotation</kwd>
        <kwd>Perspective Identi cation tator Bias</kwd>
        <kwd>Hate Speech</kwd>
        <kwd>Polarization of Opinions</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Most modern approaches to Natural Language Processing (NLP) tasks rely on
supervised Machine Learning. This is true, among other tasks, for text classi
cation tasks such as abusive language and hate speech detection [
        <xref ref-type="bibr" rid="ref35 ref6">35,6</xref>
        ]. However,
while bias in datasets has been investigated [
        <xref ref-type="bibr" rid="ref27 ref33">33,27</xref>
        ], the bias in the annotation
of the datasets used for training hate speech models is relatively less studied.
      </p>
      <p>
        Recent works highlight the importance of a \perspectivist turn", i.e., a change
of paradigm in supervised machine learning moving away from datasets
aggregated by majority vote, and towards frameworks that consider multiple
annotator perspectives in data creation, model training, and evaluation [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Taking
Copyright c 2021 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
an inclusive stance towards disagreement in data annotation does not only have
ethical implications, but rather it has practical impact on the performance of
predictive systems [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] and on the reliability of the evaluation [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        We focus on hate speech (HS), and in general abusive phenomena in
online verbal communication, for several reasons. Firstly, hateful discourse online
is growing at a worrying rate [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ], and it is linked to an increase of violence
and hatred towards vulnerable communities, with strong negative social
impact [
        <xref ref-type="bibr" rid="ref14 ref21 ref22">14,21,22</xref>
        ]. Moreover, hate speech is a highly subjective phenomenon. While
no phenomenon is neither totally subjective nor totally objective, the position of
hate speech on a hypothetical inter-subjectivity spectrum [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] is far from the
center, as its judgment is in uenced by factors such as socioeconomic background,
ethnicity, gender, among others [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]. Moreover, the hatred is typically directed
towards targets carrying speci c socio-economic, cultural, or demographic traits,
which are likely aligned to the factors in uencing the judgments of hateful
messages by human annotators. Indeed, messages containing hateful content are
often controversial, that is, they reference events, people, and issues that prompt
very di erent reactions depending on the recipient of the message [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
      </p>
      <p>
        In the area of hate speech detection, Akhtar et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] introduced a
quantitative measure of the polarization of the annotation induced by the controversiality
of the messages. They show how in presence of highly subjective phenomena like
hate speech, systematic patterns emerge that suggest a diversi cation of the
annotators' perspectives beyond the mere disagreement. In a follow-up work,
Akhtar et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] leveraged the polarization in the annotation to create multiple
perspective-encoding classi ers, boosting the classi cation performance in the
process. In this work, we further explore the polarization of annotation, and
particularly at providing a methodology to qualitatively study emerging groups
of annotators holding di erent, and sometimes con icting, perspectives.
      </p>
      <p>More speci cally, in this work we deal with shared perspectives, that is,
the set of factors that cause a certain annotation by a group of human
annotators (each holding an individual perspective). By analyzing the annotation
with computational methods, we aim at i) distinguishing groups of annotators
holding di erent shared perspectives, and ii) identifying the nature of the shared
perspectives, providing a human-readable description. First, we provide a formal
de nition of perspective in the context of the annotation of NLP datasets,
hinging on the di erence between label agreement and the novel concept of feature
agreement We then empirically demonstrate the emergence of perspectives in
real datasets of hate speech, computed with a straightforward yet e ective
procedure, and illustrated in the form of important words and selected examples.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Most research related to the identi cation of perspectives mainly focuses on the
perspective of the author of the messages. The literature is typically concerned
with subjective phenomena in natural language, such as abusive language, where
an abundance of expressions of emotions, opinions and sentiments is found.
Subjective language is considered a catalyst for multiple perspectives [
        <xref ref-type="bibr" rid="ref28 ref32">32,28</xref>
        ] and
varying opinions at sentence level [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ]. Political discourse analysis is an important
research area and many researchers worked in identifying di erent perspectives
on political topics including election campaigns as a qualitative analysis task
[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Lin et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] automatically identi ed perspectives at the document and
sentence level with high accuracy by developing statistical models and learning
algorithms on articles about the Israeli-Palestinian con ict.
      </p>
      <p>
        In NLP, the task of stance detection [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] aims at identifying points of view,
judgments or opinions on a given topic of interest in natural language. The
social and political issues on which individuals tend to express their opinions are
usually controversial in nature, causing polarization among people [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
BeigmanKlebanov et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] worked on perspective identi cation in public stance on
controversial topics such as abortion.
      </p>
      <p>
        Highly Controversial topics, such as hate speech, are a rich source to identify
and analyze con icting perspectives in online environments. When social media
users express di erent opinions on topics or social issues, the text depicts high
level of controversy due to varying perspectives [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. When such phenomena are
manually annotated by human judges, high controversy is bound to have an
impact on such annotations, in terms of agreement between the human judges.
      </p>
      <p>
        In the aforementioned work, Akhtar et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] developed a novel method to
measure the level of polarization in con icting annotations on social media
corpora. The authors developed a quantitative index, called polarization index, to
measure the level at which polarized opinions appear in individual messages.
The authors extended their work [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] by developing perspective-aware models
based on automatically clustered groups of annotators. State-of-the-art machine
learning models are trained on gold standard training sets based on this
division, successfully picking up the divergence of opinions in group-based test data.
The same authors recently developed a novel multi-perspective abusive language
dataset [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] on Brexit to identify and model perspectives expressed by
annotators with di erent ethnic, cultural and demographic background. In contrast to
traditionally published NLP corpora, this dataset provides a natural grouping
of the annotators into groups of similar backgrounds.
      </p>
      <p>
        It is noteworthy that disagreement in annotation is a relevant topic also in
more objective tasks such as POS tagging [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and even outside the scope of NLP;
for instance, Basile et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] describe the high disagreement in the annotation of
medical images by experts.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Mining Perspectives in Annotations</title>
      <p>
        We postulate that annotators and their individual perspectives in uence how
they annotate di erent items related to a given topic. This is particularly
relevant to annotation tasks that exhibit a high degree of subjectivity as here the
in uence of the perspective on the ratings may be higher. According to a
common de nition, a judgment is considered subjective when it is mainly \based
on, or in uenced by, personal feelings, tastes, or opinions."; we usually contrast
this concept with that of objective, a term that characterizes judgments that,
ideally, are not in uenced by personal feelings or idiosyncrasies and which, on a
practical level, the vast majority of people would see and label in the same way.
For instance, in hate speech detection, di erent annotators have been shown to
diverge highly in their ratings and are polarized [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Furthermore, the o
ensiveness of words depends on the context in which the words are uttered [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. For
example, consider the di erence between the use of the word \nigga" in a Rap
song, where it is considered as lowly o ensive, as opposed to using such words in
a political discourse, where it is understood as highly o ensive. We assume that
annotators implicitly or explicitly take perspectives on topics, and we model this
as described in the following.
      </p>
      <p>In order to mine shared annotator perspectives in a given dataset, we
postulate a two-step procedure. First, we detect perspectives that are shared among
annotators. To this end, we measure how much the annotators agree on item
labels, the label agreement. Second, we measure to what extent annotators agree
on the importance of linguistic features of the items, the feature agreement.
Combined, our method ensures that annotators in the same shared perspective
label items similarly and do so for similar reasons. In this work, we only use
unigrams as linguistic features to allow simpler explanations. For instance,
annotators holding the perspective that the word \fag" is especially hateful, tend
to always label items containing this word as hate speech, i.e., they exhibit a
high feature agreement on this unigram.1 We nally perform analysis on shared
perspectives consisting of annotators that are similar both in label agreement
and feature agreement. Such annotators tend to agree both on their item labels
and the importances they give to the item features (unigrams).</p>
      <p>Individual Perspectives Given a list of items, an annotator A judges these items,
according to their opinion on each of them. We call this labeling the individual
perspective of annotator A on the items. Formally, given n items, assume there
are possible opinions 0; 1; :::; c for each item. Then, an annotator A takes a
perspective pA by holding an opinion on each item. We call pA 2 f0; 1; :::; cgn the
perspective of A (on the items). By modelling annotator perspectives as vectors,
we can compare them quantitatively.</p>
      <p>
        In order to identify perspectives in annotations, we require items to have
disagreeing annotations. This is only possible in the case where annotations have
not previously been aggregated into a single label, what in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] has been called
diamond standard. This is in contrast with the usual gold standard paradigm
where multiple annotations are harmonized into one gold label, often
implemented by majority voting. Under the paradigm of annotator perspectives that
we have introduced above, the reduction of multiple labels (annotator opinions)
into a gold label by majority vote is equivalent to taking the majority perspective.
Shared Perspectives While each annotator takes their own perspective, we are
more interested in nding perspectives which are shared among annotators. We
      </p>
      <sec id="sec-3-1">
        <title>1 Note that our method is agnostic to the type of features extracted from the messages,</title>
        <p>and it could therefore be used in conjunction with other, more re ned features.
call perspectives pA; pB shared based on their similarity. We employ clustering
to nd clusters of annotators that share perspectives. While shared perspectives
arise from an agreement of annotators on item labels (label agreement), we
also aim to understand how shared perspectives are linguistically de ned. To
this end, we analyze the importance that di erent annotators give to di erent
linguistics features, i.e. which words annotators in a shared perspective agree to
be important (feature agreement).
3.1</p>
        <sec id="sec-3-1-1">
          <title>Label Agreement</title>
          <p>
            We measure label agreement in terms of inter-annotator agreement. We use
Krippendor 's alpha reliability [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ] and cluster the annotators based on the
label agreement. We proceed as follows:
{ Given n annotators that label the same k items, we obtain a label matrix
V 2 Rn k, where Vi;j is the rating of annotator i of item j. We compute
the similarity matrix A 2 Rn n, where Ai;j = (Vi;:; Vj;:) that encodes the
pairwise agreement between annotators i; j, where is Krippendor 's alpha
reliability, and Vi;: is the label vector of annotator i. Then, the distance
matrix D = 1 A induces a clustering of the annotators. The distances in
D are the pairwise disagreements between annotators.
{ We use an o -the-shelf clustering algorithm to cluster the annotators based
on their distances D to one another, into groups of annotators with high
intra-group label agreement and low inter-group label agreement.
{ A high label agreement (i; j) indicates that annotators i; j tend to give
similar labels on the items (texts).
          </p>
          <p>Note that Krippendor 's alpha is also de ned for incomplete annotation, i.e.,
where not all annotators covered all the instances. This is a typical scenario in
crowdsourcing, but could happen with other annotation procedures as well.
3.2</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Feature Agreement</title>
          <p>We want to measure whether annotators agree on the importance of
linguistic features of the textual items. In this paper, we use a simple bag of words
(BOW) to model the texts; the features are unigram counts. Feature agreement
between annotators i; j arises when i and j give similar importance to features.
We measure the importance of each feature to an annotator by computing the
chi-square ( 2) statistics between the feature distribution and the label
distribution in the annotator, following a univariate feature selection approach. This
measures how the annotator's label depends on the presence of a word in an
item. For instance, the presence of the unigram bitch often coincides with the
label hate speech, while this is not the case for the word sunny. The 2 statistics
captures this; it is much higher for bitch than it is for sunny. When annotators
tend to agree on the importance of words, they exhibit an overall high feature
agreement. Speci cally, we compute feature agreement as follows:
{ We extract k features from the corpus. Since we employ the BOW model,
given a corpus of n documents, we compute the term-document matrix F 2
Nn k, where Fi;j indicates the count of word j in document i. The columns
of F are the features, as F:;i is the word counts of word i over the documents.
{ For each feature fi = F:;i and annotator r, we compute the importance
imp of feature fi to annotator r as imp(fi; r) = 2(fi; VrT;:), where V is the
previously introduced label matrix.
{ We de ne the feature agreement between annotators i; j by comparing their
feature importances. To this end, let I 2 Rk n be the importance matrix,
where Ii;j = imp(fi; j). Then, the vector of all importances of annotator
j is given by I:;j . The feature agreement between annotators i; j is then
computed by the cosine similarity of their importances vectors: (i; j) =
cosine(I:;i; I:;j ), where cosine(x; y) = x y (kxk kyk) 1.
{ A high feature agreement (i; j) indicates that annotators i; j tend to give
similar labels when similar words are present.
{ Given n annotators, we compute the similarity matrix Bi;j = (i; j) that
encodes the pairwise feature agreement between annotators i; j. Analogously
to the label agreement case, we use the distance matrix D = 1 B to
cluster the annotators into groups of annotators with high intra-group feature
agreement and low inter-group feature agreement.</p>
          <p>Since the 2 statistics requires a dense label matrix, if an annotator has not
labelled an item, we insert the negative label (i.e., not hate speech). Truly
unimportant words then correctly get low importance, while truly important words
get assigned a somewhat diminished importance.
3.3</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Label Feature Agreement</title>
          <p>We consider two di erent ways of clustering the annotators: by label agreement
and by feature agreement. These two clusterings sometimes di er, for instance,
when two annotators agree on the item labels (label agreement), but do not
agree on the importance of words related to those labels (feature agreement).
Since our goal is to nd annotators that label similarly and do so for similar
reasons, we analyze all annotators that cluster in the exact same way in both
labels and features. Speci cally, we perform two clusterings for the annotators
ai. First, we cluster the ai into k di erent clusters f1, 2, ..., kg according to label
agreement, assigning each ai a cluster Lab(ai) 2 f1; 2; :::; kg. Analogously, each
ai is assigned a cluster F eat(ai) 2 f1; 2; :::; kg according to feature agreement.
Then, we only consider such annotators ai that cluster in the same way, i.e.,
label feature agreement is de ned as fai : Lab(ai) = F eat(ai)g.
3.4</p>
        </sec>
        <sec id="sec-3-1-4">
          <title>Cluster Analysis</title>
          <p>Given the clusters of annotators we obtain, we analyze certain cluster properties
statistics, how the clusters di er and which words are important to each cluster.
Speci cally, we perform the following analyses.</p>
          <p>Quantitative cluster description: the number of annotators in the cluster, the
positive label rate %, the label agreement , the number of features, and the
feature agreement . We compare the cluster numbers also with the numbers for
all annotators disregard of their cluster a liation.</p>
          <p>Qualitative cluster description: we inspect the most characteristic unigrams
for the clusters, i.e. the words with the highest relative importance R to the
cluster. We measure RC (w) of a word w to a cluster C as RC (w) = 11++mmeeddffiimmpp((ww;;ii))ggii22:CC ,
i.e., the median importance to all annotators inside the cluster vs. the median
importance to all annotators outside the cluster. We inspect examples that are
polarized between the clusters, i.e., they are annotated with disagreement
between the clusters. These examples often carry important words as vocabulary.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Datasets</title>
      <p>The experiments described in this paper are conducted on several hate speech
corpora, consisting of Twitter messages (tweets); they are published in various
research studies on hate speech. This section provides details about the datasets,
such as the annotation process with scheme and guidelines, and information on
the annotators.
4.1</p>
      <sec id="sec-4-1">
        <title>HS Dataset on Brexit</title>
        <p>
          The hate speech dataset on Brexit was recently published [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Originally, the
authors gathered the data from a study on stance detection in political debates [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]
where around 5 millions tweets were collected during the Brexit voting period,
June 2016. The authors developed a multi-perspective dataset to automatically
detect abusive language on social media with the intention to model annotator
perspectives and polarized opinions. The collected tweets are ltered with a list
of selected abusive keywords based on a previous study [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. 1,120 tweets were
randomly sampled and annotated for hate speech, Aggressiveness, O ensiveness
and Stereotype, following the scheme described in [
          <xref ref-type="bibr" rid="ref25 ref29">25,29</xref>
          ]. In total, six
annotators contributed to the dataset. Three of the annotators were researchers with
western background and experience in linguistic annotation who volunteered to
annotate the data. The other three volunteers were rst- or second-generation
immigrants and migrants as students from the developing countries to Europe
and the UK, of Muslim background. The group of migrants is named Target and
the locals are named Control. The dataset is unique in the way that it involves
migrants as the victims of abuse on social media. Personal details of all
annotators such as cultural and demographic background and ethnicity are known
and considered a valuable source of information for perspective-aware abusive
language detection. For the current study, we only used the hate speech label.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>HS Dataset in Italian</title>
        <p>
          The hate speech dataset in Italian language [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] (HS Italian) consists of 3,200
tweets collected from TWITA [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] in 2017, partially overlapping with the Italian
        </p>
        <p>
          Hate Speech Corpus [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]. The collection was ltered by the authors of the
original dataset with a list of handcrafted keywords related to migrants and ethnic
and religious minorities in Italy. The tweets were annotated on the Figure Eight
platform2. A minimum of three annotators annotated the whole corpus,
subsequently aggregated by the crowd-sourcing platform to create a gold standard
dataset. We requested and obtained the dataset from the authors.
4.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>HS Dataset in English</title>
        <p>
          Davidson et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] developed a hate speech dataset to perform automatic hate
speech detection as a multi-classi cation task. The authors gathered around
85.4 million tweets from a total of 33,458 Twitter users. A hate speech lexicon
containing hateful words and phrases was used to query the tweets. This lexicon
was compiled by Hatebase.org and the hateful words in the lexicon were identi ed
by internet users. The authors randomly selected about 27,000 tweets from the
dataset by using the keywords from the hate speech lexicon. CrowdFlower (now
Appen) workers were hired to manually annotate the tweets. The annotation
scheme comprises the labels hate speech, o ensive but not hate speech, and neither
o ensive nor hate speech. The authors developed detailed guidelines with their
own de nitions of di erent hate speech terms including the context in which
the words were used. Each tweet in the dataset was annotated by three or more
annotators. The Davidson dataset is only available for download in an aggregated
gold standard form3, therefore we requested and obtained the non-aggregated
dataset from the authors.4
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Perspective Mining Experiments</title>
      <p>We performed analysis on all the datasets that we introduced in the previous
section. An important factor for our experiments is what prior knowledge we
have about the annotators that annotated the datasets. Where such background
information is given, we can con rm or reject our ndings by comparing our
empirically found annotator clusters (shared perspectives) with groupings of human
annotators. As stated in the dataset section before, we have the following
information on the dataset annotators. On the Brexit dataset: the personal details
of all annotators such as cultural and demographic background and ethnicity
are known. On all other datasets: no background information on the annotators
is available. Note that the HS Italian and Davidson datasets are sparsely
annotated, as annotators have only labeled a fraction of the instances. This is in
opposition to the Brexit dataset which has a dense annotation matrix.</p>
      <sec id="sec-5-1">
        <title>2 https://www. gure-eight.com/, now Appen.</title>
      </sec>
      <sec id="sec-5-2">
        <title>3 https://github.com/t-davidson/hate-speech-and-o ensive-language</title>
      </sec>
      <sec id="sec-5-3">
        <title>4 As we needed non-aggregated data for our work, we only found aggregated gold</title>
        <p>
          standard data on author's GitHub repository. Therefore, we requested the authors
to provide us with pre-aggregated data and we are grateful to them for providing us
the required format of the dataset.
{ Preprocessing: we removed URLs and Twitter handles (@username) from the
tweets, tokenized them using the NLTK5 Tweet Tokenizer and lemmatized
them using spaCy6.
{ BOW features: we created the BOW feature space with the scikit-learn7
CountVectorizer, where we set the minimum document frequency to 10. This
number was decided based on the fact that some tweets occured duplicate
or near-duplicate, because of the dialog structure of Twitter, users will cite
each other. We alleviate the problem by setting a rather high minimum
document frequency of 10. Furthermore, we counted each word once per
document (\binary") and extracted solely unigrams.
{ Clustering: we used the KMeans algorithm with di erent numbers k of
clusters, we settled to k = 2 which appeared most reasonable based on the
inspection of 2D-PCA embeddings of the datasets. This parameter choice
makes us conform with the polarization paradigm, i.e. we analyze two
conicting/polarized perspectives.
{ Important words: for each cluster, the top 20 words with highest relative
importance are considered. The polarized examples are extracted using the
polarization index method [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
5.2
        </p>
        <sec id="sec-5-3-1">
          <title>Perspectives in the Brexit Corpus</title>
          <p>In the Brexit datasets, the inter-annotator agreement is measured as = 0:35.
The positive label rate is 12.9%. We extract 266 features from the corpus and
obtain = 0:7 as feature agreement between all annotators.</p>
        </sec>
      </sec>
      <sec id="sec-5-4">
        <title>5 https://www.nltk.org</title>
      </sec>
      <sec id="sec-5-5">
        <title>6 https://spacy.io</title>
      </sec>
      <sec id="sec-5-6">
        <title>7 https://scikit-learn.org</title>
        <p>After label feature agreement (see Section 3.3), we obtain the clusters A =
f3; 4; 5g and B = f1; 2g. Since we know the annotator backgrounds, we know
that A corresponds to the migrants with muslim background (target group) and
that B corresponds to the non-migrants (control group). This result e ectively
validates our clustering methodology based on label and feature agreement to
extract perspectives empirically.</p>
        <p>Quantitatively, we nd the following di erences between the clusters: i) The
positive label rate is much higher in A (20.5%) as compared to positive labels
in B (5.8%), indicating the annotators in A are more sensitive in this task
(all annotators 12.9%). ii) The label agreement is higher in cluster A ( A =
0:58) as compared to B = 0:44, indicating that cluster A holds more coherent
opinions as B. Both values are much higher than the average, meaning that
the groups hold polarized opinions. iii) The feature agreement is higher in both
clusters ( A = 0:86 = B) compared to the dataset feature agreement ( = 0:7),
indicating polarization of the feature agreements of the clusters as well.</p>
        <p>Qualitatively, we nd that certain words are highly correlated with the
positive label in both groups, and some words are speci c to the annotator clusters.
The shared vocabulary contains words such as \islam", \kill" and hashtags
related to US president Donald Trump (#maga, #trump2016). When inspecting
the corpus, we nd examples such as the following that exemplify the use of the
words; matched words are bold. And indeed, in this example, both annotator
groups give the positive label.</p>
        <p>RT @ davidmatheson27 : The U.K. Must ban Islam and close all mosques!
URL</p>
        <p>London should kick all Muslim Refugees out before they all kill them. #Trump2016
URL</p>
        <p>From these shared vocabulary examples, we can see that since \islam" is one
of the important words, the hate speech in this corpus appears to be at least
partially islamophobia. An inspection of the important words for the potential
target group of islamophobia, cluster A , supports this claim. We nd a speci c
and distinctive vocabulary related to muslims, invasion, terrorists.8 The following
examples illustrate the important words for cluster A. The examples got the
positive label in A and the negative label (\no, this is not hate speech") in
cluster B. Interestingly, while \islam" is a shared top word, we found it in the
combination \radical islam" typically in cluster A.</p>
        <p>FYI world, the ppl of GB supporting #Brexit know if they don't control their
own immigration/borders radical Islam will end their lives.</p>
        <p>Stealing jobs, a well-known negative prejudice towards foreigners, is also among
the examples that are important for cluster A:</p>
        <p>Bloody foreigners coming here &amp; taking our jobs though! #Brexit URL</p>
      </sec>
      <sec id="sec-5-7">
        <title>8 Words with highest relative importance for cluster A: radical, job, illegal, invasion,</title>
        <p>love, can, let, merkel, mayor, then.
Identi ed Perspectives Overall, we nd two polarized groups, both by label and
feature agreement. Cluster A - the target group - is much more likely to give the
positive label and this group of annotators consistently bases their opinion on a
speci c and distinct vocabulary which can be described as Islamophobic. Given
the background information we have on all annotators, we identify cluster A as
the Muslim perspective on the topic, highly sensitive to Islamophobic content. In
opposition, for cluster B we did not nd a characteristic vocabulary, those
annotators form more a counter position to the migrant group, therefore we describe
them as control group or non-muslim perspective. We conclude, annotators in
cluster A are very sensitive towards islamophobic and, more general, xenophobic
textual content.
5.3</p>
        <sec id="sec-5-7-1">
          <title>Perspectives in the HS Italian Dataset</title>
          <p>In this dataset, we found large di erences between the number of items
annotated by the di erent annotators. To avoid biasing our model, we only analyze
annotators with a high rating count9. When clustering using both and
agreement, we obtained the same clustering into the two clusters A; B of each 7
annotators.</p>
          <p>Quantitatively, we found an anomaly here, as the label agreement in cluster
A is almost zero ( A = 0:03), whilst in cluster B it is rather high ( B = 0:42).
This already indicates that A is a cluster of outliers. Furthermore, the feature
agreement is higher in cluster B ( B = 0:48) as compared to A ( A = 0:34).
The latter appears to be due to noise only.</p>
          <p>Qualitatively, we found that degrading talk about immigrants get positive
labels from both clusters. For cluster B, we found examples with complains about
immigrants driving up public costs by living in \hotels" as well as concerns about
\sicurezza"(security) being diminished in the country after immigration.10</p>
          <p>Identi ed Perspectives : in this dataset, we found a de ned perspective in
cluster B. The annotators tend to label a large spectrum of content - from
critical, over conservative, nationalistic, to openly hateful tweets, all as hate.
Hence, annotators in cluster B are very sensitive towards xenophobic textual
content.
5.4</p>
        </sec>
        <sec id="sec-5-7-2">
          <title>Perspectives in the Davidson Dataset</title>
          <p>Analogously to the HS Italian dataset, we only analyze annotators with a high
rating count11. We obtained di erent clusters according to and agreement.
After computing the label feature agreement, we obtained cluster A of size 45
and cluster B of size 41. Note that, in contrast with all previous datasets, we</p>
        </sec>
      </sec>
      <sec id="sec-5-8">
        <title>9 this means for this dataset at least 800 ratings per annotator</title>
        <p>10 Words with highest relative importance for cluster B: hotel, #immigrati, spesa,
se, clandestino, #gabbiaopen, giusto, tangere, succedere, #sicurezza (hotel,
#immigrants, expense, if, illegal alien, #opencage, right, touch, succeed, #safety).
11 for this dataset, at least 500 annotations per annotator
have two kinds of positive labels in this dataset, one for \o ensive language"
content and one for \hate speech" (stronger label).</p>
        <p>Quantitatively, we found cluster B to have a much higher hate speech label
rate (15.4%) over cluster A (4.1%). The base rate is 10%. While both cluster
have comparable positive label rates, this indicates that cluster B has a tendency
to give the hate speech label when the o ensive label would have been an option.
as compared to cluster A. Further, feature agreement is much lower in cluster A
( A = 0:22) as opposed to cluster B ( B = 0:46), indicating the annotators in
cluster B agree much more on their important words.</p>
        <p>Qualitatively, we found some words are understood by both clusters as
hateful. As cluster A has a much lower positive label rate, A was rarely more critical
than B. For cluster B we nd several examples with the same keywords, centered
around homophobic slurs such as \faggot".12</p>
        <p>Identi ed Perspectives: in this dataset, we nd a de ned perspective for
cluster B. Annotators in this group give harsher labels when homophobic slurs are
present in a tweet, as compared to annotators in A. We conclude that the
annotators in B are highly sensitive towards homophobic textual content.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this paper, we analyzed a number of annotated hate speech corpora, showing
how the opinions of the annotators, re ected in their annotation, are far from
uniformly distributed. In fact, the annotation of hate speech tends to be
polarized, and our methodology is able to highlight the groups of annotators sharing
similar opinions. We identi ed perspectives in the datasets, de ned as increased
sensitivity towards certain types of textual content (xenophobic, islamophobic,
homophobic). Further, we introduced an automated method to support the
manual exploration of the perspectives emerging from a polarized annotation of hate
speech, resulting in consistent patterns describing why certain groups of people
are more or less keen on judging a message as hateful.</p>
      <p>
        As future work, we plan to test our methods with deeper and more re ned
linguistic features, to abstract away from individual words and therefore provide
a more robust analysis. We also plan on investigating other NLP tasks
traditionally considered less subjective, but recently found to contain informative
disagreement [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], as well as non-linguistic tasks such as image labeling.
      </p>
      <p>
        Finally, we note how this was was only possible thanks to the availability of
non-aggregated datasets. In line with [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and the Perspectivist Data Manifesto13,
we consider this factor crucial for research like ours.
12 Words with highest relative importance for cluster B: hypocrite, til, mike, warn, fag,
spread, faggot, jealous, tat, texas.
13 https://pdai.info/
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Akhtar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
          </string-name>
          , V.:
          <article-title>A new measure of polarization in the annotation of hate speech</article-title>
          . In: Alviano,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Greco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Scarcello</surname>
          </string-name>
          ,
          <string-name>
            <surname>F</surname>
          </string-name>
          . (eds.)
          <source>AI*IA 2019 { Advances in Arti cial Intelligence</source>
          . pp.
          <volume>588</volume>
          {
          <fpage>603</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Akhtar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Modeling annotator perspective and polarized opinions to improve hate speech detection</article-title>
          .
          <source>Proceedings of the AAAI Conference on Human Computation and Crowdsourcing</source>
          <volume>8</volume>
          (
          <issue>1</issue>
          ),
          <volume>151</volume>
          {154 (Oct
          <year>2020</year>
          ), https:// ojs.aaai.org/index.php/HCOMP/article/view/7473
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Akhtar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Whose opinions matter? perspective-aware models to identify opinions of hate speech victims in abusive language detection (</article-title>
          <year>2021</year>
          ), https://arxiv.org/abs/2106.15896
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>AlDayel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magdy</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Stance detection on social media: State of the art and trends</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>58</volume>
          (
          <issue>4</issue>
          ),
          <volume>102597</volume>
          (
          <year>2021</year>
          ). https://doi.org/https://doi.org/10.1016/j.ipm.
          <year>2021</year>
          .
          <volume>102597</volume>
          , https://www. sciencedirect.com/science/article/pii/S0306457321000960
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Basile</surname>
          </string-name>
          , V.:
          <article-title>It's the end of the gold standard as we know it. on the impact of preaggregation on the evaluation of highly subjective tasks</article-title>
          .
          <source>In: 2020 AIxIA Discussion Papers Workshop</source>
          ,
          <article-title>AIxIA 2020 DP</article-title>
          . vol.
          <volume>2776</volume>
          , pp.
          <volume>31</volume>
          {
          <fpage>40</fpage>
          .
          <string-name>
            <surname>CEUR-WS</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bosco</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fersini</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nozza</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Rangel</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.M.</given-names>
            ,
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Sanguinetti</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.:</surname>
          </string-name>
          <article-title>SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in twitter</article-title>
          .
          <source>In: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          . pp.
          <volume>54</volume>
          {
          <fpage>63</fpage>
          . Association for Computational Linguistics, Minneapolis, Minnesota, USA (Jun
          <year>2019</year>
          ). https://doi.org/10.18653/v1/
          <fpage>S19</fpage>
          -2007, https://www.aclweb.org/ anthology/S19-2007
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cabitza</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campagner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Toward a perspectivist turn in ground truthing for predictive computing</article-title>
          .
          <source>In: Proceedings of the XVIII Conference of the Italian chapter of AIS - Digital Resiliance and Sustainability: People, Organizations, and Society</source>
          . pp.
          <volume>1</volume>
          {
          <fpage>16</fpage>
          .
          <article-title>Association for Intelligent Systems</article-title>
          , Trento (Oct
          <year>2021</year>
          ), http://www.itais.org/itais2021-proceedings/pdf/21.pdf
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fornaciari</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paun</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plank</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poesio</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uma</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>We need to consider disagreement in evaluation</article-title>
          .
          <source>In: Proceedings of the 1st Workshop on Benchmarking: Past, Present and Future</source>
          . pp.
          <volume>15</volume>
          {
          <fpage>21</fpage>
          . Association for Computational Linguistics,
          <source>Online (Aug</source>
          <year>2021</year>
          ). https://doi.org/10.18653/v1/
          <year>2021</year>
          .bppf-
          <volume>1</volume>
          .3, https://aclanthology.org/
          <year>2021</year>
          . bppf-
          <volume>1</volume>
          .
          <fpage>3</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lai</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanguinetti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Long-term social media data collection at the university of turin</article-title>
          . In: CLiC-it (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>Beigman</given-names>
            <surname>Klebanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Beigman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Diermeier</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>Vocabulary choice as an indicator of perspective</article-title>
          .
          <source>In: Proceedings of the ACL 2010 Conference Short Papers</source>
          . pp.
          <volume>253</volume>
          {
          <fpage>257</fpage>
          . Association for Computational Linguistics, Uppsala,
          <source>Sweden (Jul</source>
          <year>2010</year>
          ), https://aclanthology.org/P10-2047
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Campagner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciucci</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Svensson</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Figge</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cabitza</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Ground truthing from multi-rater labeling with three-way decision and possibility theory</article-title>
          .
          <source>Information Sciences</source>
          <volume>545</volume>
          ,
          <volume>771</volume>
          {
          <fpage>790</fpage>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Davidson</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warmsley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Macy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Automated hate speech detection and the problem of o ensive language</article-title>
          .
          <source>In: Proceedings of the 11th International AAAI Conference on Web and Social Media</source>
          . pp.
          <volume>512</volume>
          {
          <fpage>515</fpage>
          . ICWSM '
          <volume>17</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Florio</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lai</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Leveraging hate speech detection to investigate immigration-related phenomena in italy</article-title>
          .
          <source>In: 2019 8th International Conference on A ective Computing and Intelligent Interaction Workshops and Demos (ACIIW)</source>
          . pp.
          <volume>1</volume>
          {
          <issue>7</issue>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Izsak-Ndiaye</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Report of the special rapporteur on minority issues, rita izsak : comprehensive study of the human rights situation of roma worldwide, with a particular focus on the phenomenon of anti-gypsyism</article-title>
          .
          <source>Tech. rep., UN</source>
          ,, Geneva :.
          <year>2015</year>
          -
          <volume>05</volume>
          -11 (May
          <year>2015</year>
          ), http://digitallibrary.un.org/record/797194, submitted pursuant to Human
          <source>Rights Council resolution 26/4.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Krippendor</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Estimating the reliability, systematic error and random error of interval data</article-title>
          .
          <source>Educational and Psychological Measurement</source>
          <volume>30</volume>
          ,
          <issue>61</issue>
          {
          <fpage>70</fpage>
          (
          <year>1970</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lai</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tambuscio</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , Ru o, G.,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Stance polarity in political debates: A diachronic perspective of network homophily and conversations on twitter</article-title>
          .
          <source>Data &amp; Knowledge Engineering</source>
          <volume>124</volume>
          ,
          <volume>101738</volume>
          (09
          <year>2019</year>
          ). https://doi.org/10.1016/j.datak.
          <year>2019</year>
          .101738
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>W.H.</given-names>
          </string-name>
          , Wilson,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Wiebe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Hauptmann</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Which side are you on? identifying perspectives at the document and sentence levels</article-title>
          .
          <source>In: Proceedings of the Tenth Conference on Computational Natural Language Learning (CoNLL-X)</source>
          . pp.
          <volume>109</volume>
          {
          <fpage>116</fpage>
          . Association for Computational Linguistics, New York City (Jun
          <year>2006</year>
          ), https://www.aclweb.org/anthology/W06-2915
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Maul</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mari</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , Wilson,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Intersubjectivity of measurement across the sciences</article-title>
          .
          <source>Measurement</source>
          <volume>131</volume>
          ,
          <issue>764</issue>
          {
          <fpage>770</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arcostanzo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krasodomski-Jones</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiedlitzka</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jamali</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dale</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>From brussels to brexit: Islamophobia, xenophobia, racism and reports of hateful incidents on twitter</article-title>
          .
          <source>DEMOS</source>
          . Available at www. demos. co. uk/wpcontent/uploads/2016/07/From-Brussels-to-Brexit -
          <article-title>IslamophobiaXenophobia-Racism-and-Reports-of-Hateful-Incidents-on-Twitter-ResearchPrepared-for-</article-title>
          <string-name>
            <surname>Channel-</surname>
          </string-name>
          4
          <string-name>
            <surname>-</surname>
          </string-name>
          Dispatches-%
          <source>E2 80</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Mohammad</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiritchenko</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sobhani</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cherry</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>SemEval2016 task 6: Detecting stance in tweets</article-title>
          .
          <source>In: Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016)</source>
          . pp.
          <volume>31</volume>
          {
          <fpage>41</fpage>
          . Association for Computational Linguistics, San Diego, California (Jun
          <year>2016</year>
          ). https://doi.org/10.18653/v1/
          <fpage>S16</fpage>
          -1003, https://aclanthology.org/S16-1003
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Mossie</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          :
          <article-title>Vulnerable community identi cation using hate speech detection on social media</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>57</volume>
          ,
          <volume>102087</volume>
          (07
          <year>2019</year>
          ). https://doi.org/10.1016/j.ipm.
          <year>2019</year>
          .102087
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <given-names>O</given-names>
            <surname>'Kee e</surname>
          </string-name>
          , G.S.,
          <string-name>
            <surname>Clarke-Pearson</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>The impact of social media on children, adolescents, and families</article-title>
          .
          <source>Pediatrics</source>
          <volume>127</volume>
          ,
          <issue>800</issue>
          {
          <fpage>804</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Pamungkas</surname>
            ,
            <given-names>E.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Do you really want to hurt me? predicting abusive swearing in social media</article-title>
          .
          <source>In: The 12th Language Resources and Evaluation Conference</source>
          . pp.
          <volume>6237</volume>
          {
          <fpage>6246</fpage>
          .
          <string-name>
            <surname>European Language Resources Association</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Man</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>So</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>One event, three stories</article-title>
          .
          <source>Gazette</source>
          <volume>61</volume>
          ,
          <issue>99</issue>
          {
          <volume>112</volume>
          (04
          <year>1999</year>
          ). https://doi.org/10.1177/0016549299061002001
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Poletto</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stranisci</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanguinetti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bosco</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Hate speech annotation: Analysis of an italian twitter corpus</article-title>
          .
          <source>In: Proceedings of the Fourth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2017</year>
          ), Rome, Italy,
          <source>December 11-13</source>
          ,
          <year>2017</year>
          .
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <year>2006</year>
          .
          <article-title>CEUR-WS.org (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennacchiotti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Detecting controversial events from twitter</article-title>
          .
          <source>In: Proceedings of the 19th ACM International Conference on Information and Knowledge Management</source>
          . pp.
          <year>1873</year>
          {
          <year>1876</year>
          . CIKM '10,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2010</year>
          ). https://doi.org/10.1145/1871437.1871751, http://doi.acm.
          <source>org/ 10</source>
          .1145/1871437.1871751
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Razo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Kubler, S.:
          <article-title>Investigating sampling bias in abusive language detection</article-title>
          .
          <source>In: Proceedings of the Fourth Workshop on Online Abuse and Harms</source>
          . pp.
          <volume>70</volume>
          {
          <fpage>78</fpage>
          . Association for Computational Linguistics,
          <source>Online (Nov</source>
          <year>2020</year>
          ). https://doi.org/10.18653/v1/
          <year>2020</year>
          .alw-
          <volume>1</volume>
          .9, https://www.aclweb.org/ anthology/2020.alw-
          <volume>1</volume>
          .
          <fpage>9</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Rilo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiebe</surname>
          </string-name>
          , J.:
          <article-title>Learning extraction patterns for subjective expressions</article-title>
          .
          <source>In: Proceedings of the 2003 Conference on Empirical Methods in Natural Language Processing</source>
          . pp.
          <volume>105</volume>
          {
          <issue>112</issue>
          (
          <year>2003</year>
          ), https://www.aclweb.org/anthology/W03-1014
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Sanguinetti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poletto</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bosco</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stranisci</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>An italian twitter corpus of hate speech against immigrants</article-title>
          .
          <source>In: Proceedings of the Eleventh International Conference on Language Resources</source>
          and
          <string-name>
            <surname>Evaluation (LREC-2018). European Language Resource Association</surname>
          </string-name>
          (
          <year>2018</year>
          ), http://aclweb.org/anthology/L18-1443
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Uma</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fornaciari</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paun</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plank</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poesio</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A case for soft-loss functions</article-title>
          .
          <source>In: Proceedings of the 8th AAAI Conference on Human Computation and Crowdsourcing</source>
          . pp.
          <volume>173</volume>
          {
          <issue>177</issue>
          (
          <year>2020</year>
          ), https://ojs.aaai.org/index.php/ HCOMP/article/view/7478
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Warner</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirschberg</surname>
          </string-name>
          , J.:
          <article-title>Detecting hate speech on the world wide web</article-title>
          .
          <source>In: Proceedings of the Second Workshop on Language in Social Media</source>
          . pp.
          <volume>19</volume>
          {
          <fpage>26</fpage>
          . LSM '
          <volume>12</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, Stroudsburg, PA, USA (
          <year>2012</year>
          ), http://dl.acm.org/citation.cfm?id=
          <volume>2390374</volume>
          .
          <fpage>2390377</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Wiebe</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Wilson,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Bruce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Bell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Learning subjective language</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>30</volume>
          ,
          <issue>277</issue>
          {
          <volume>308</volume>
          (09
          <year>2004</year>
          ). https://doi.org/10.1162/0891201041850885
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Wiegand</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruppenhofer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kleinbauer</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Detection of Abusive Language: the Problem of Biased Datasets</article-title>
          .
          <source>In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers). pp.
          <volume>602</volume>
          {
          <fpage>608</fpage>
          . Association for Computational Linguistics, Minneapolis,
          <source>Minnesota (Jun</source>
          <year>2019</year>
          ). https://doi.org/10.18653/v1/
          <fpage>N19</fpage>
          -1060, https://www.aclweb. org/anthology/N19-1060
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hatzivassiloglou</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Towards answering opinion questions: Separating facts from opinions and identifying the polarity of opinion sentences</article-title>
          .
          <source>In: EMNLP</source>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Zampieri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenthal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanasova</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karadzhov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mubarak</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derczynski</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pitenis</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , Coltekin, c.: SemEval-2020
          <source>Task</source>
          <volume>12</volume>
          :
          <string-name>
            <surname>Multilingual O ensive Language</surname>
          </string-name>
          <article-title>Identi cation in Social Media (O ensEval 2020)</article-title>
          .
          <source>In: Proceedings of SemEval</source>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Hate speech detection: A solved problem? the challenging case of long tail on twitter</article-title>
          .
          <source>Semantic Web Accepted (10</source>
          <year>2018</year>
          ). https://doi.org/10.3233/SW-180338
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>