<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DisaggregHate It Corpus: A Disaggregated Italian Dataset of Hate Speech</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Madeddu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simona Frenda</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mirko Lai</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Viviana Patti</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valerio Basile</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aequa-tech srl</institution>
          ,
          <addr-line>Turin</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Università di Torino</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Recent studies in Machine Learning advocate for the exploitation of disagreement between annotators to train models in line with the diferent opinions of humans about a specific phenomenon. This means that datasets where the annotations are aggregated by majority voting are not enough. In this paper, we present an Italian disaggregated dataset concerning hate speech and encoding some information about the annotators: the DisaggregHate It Corpus. The corpus contains Italian tweets that focus on the topic of racism and has been annotated by native Italian university students. We explain how the dataset was gathered by following the recommendation of the perspectivist approach [1], encouraging the annotators to give some socio-demographic information about them. To exploit the disagreement in the learning process, we proposed two types of soft labels: softmax and standard normalization. We investigated the benefit of using disagreement by creating a baseline binary model and two regression models that were respectively trained on the 'hard' (aggregated label by majority voting) and the two types of 'soft' labels. We tested the models in an in-domain and out-of-domain setting, evaluating their performance using the cross-entropy as a metric, and showing that the models trained on the soft labels performed better.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;hate speech</kwd>
        <kwd>perspectivism</kwd>
        <kwd>disagreement</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        in the annotated data, while the latter, overcomes the
idea of “ground truth” in the construction of datasets and
The rise of the Internet and social media platforms has on the creation and evaluation of NLP models, focusing
given many users the opportunity to express their opin- more on who the annotators are.
ion online. Unfortunately, this leads to the difusion of a Our work could be considered a tentative to approach
new online phenomenon: the hate speech. To prevent the hate speech detection, exploiting the possible
disagreeviral spread of this kind of expressions on social media, ment among the annotators. Usually, models are trained
hate speech detection became a popular task in Natu- on data associated to a ‘hard’ label. In the case of binary
ral Language Processing (NLP). A lot of tools have been classification, each item is assigned a label whose value is
created to detect and counter hate speech[
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2, 3, 4</xref>
        ]. either 0 or 1. The hard label value is commonly obtained
      </p>
      <p>
        Recently, there have been studies that suggest trying through majority voting, therefore this implies that
conto shift away from the golden standard approach in Ma- troversial instances have the same label as the ones that
chine Learning, especially in tasks partly subjective and saw all annotators in agreement. This may be thought
influenced by the social and cultural context, like hate of like a loss of valuable information that can be used in
speech [
        <xref ref-type="bibr" rid="ref1 ref5">5, 1</xref>
        ]. These works advocate that diferent opin- the training phase of the models [7]. On the other hand,
ions given in the annotation process are not a noise factor ‘soft’ labels approaches try to avoid this waste of data by
but can be used to make better systems [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This shift in- assigning a real number to the label. Diferent functions
spires scholars to try diferent techniques to train models can be used in the process of determining the value of the
using datasets where the target label is not simply deter- soft label, such as standard normalization or a softmax
mined by majority voting on the annotations. In this line, function [8].
two theoretical paradigms have been established, both In this context, we created the DisaggregHate It
Corlooking for the inclusion of diferent perspectives: the pus, a new disaggregated dataset about hate speech
learning from disagreement and perspectivism. The former in the Italian language that incorporates some
sociocould be considered like a ‘soft perspectivist approach’ be- demographic information about annotators1. A corpus
cause it takes into account the presence of disagreement like this could be beneficial in exploring how diferent
segments of population are sensitive to certain social
issues like hate speech, and how this information can be
used to create better systems.
      </p>
      <p>After explaining the diferent characteristics of the
1The corpus is available here:
madeddumarco/DisaggregHateIt
https://github.com/
dataset in section 3 we will validate the corpus by using
it as the training set of diferent models in section 4. The
performed experiments show that training models on a
soft label rather than a hard label leads to better results.</p>
      <p>As suggested by [7], we used the cross entropy metric
for evaluating the models.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>One of the most famous is the Measuring of Hate Speech
corpus [10]2 available only in English, that encodes
various dimensions of hate speech (with disaggregated labels)
and also diferent information about annotators. Follow:
the HS-Brexit disaggregated dataset created by Akhtar
et al. [11], ToxCR dataset [12] and JSRPData [13], on hate
speech and toxic language. All of these datasets are in
English and contain little information about the annotators.</p>
      <p>About Italian language, to our knowledge, only
IMSyPPIT dataset [14] have been released with disaggregated
labels but without information about annotators.</p>
      <p>In this context, a dataset like DisItaggragated released
with disaggregated labels about hate speech, and that
encodes also some information about annotators,
contributes to enrich the resources for Italian community
and to encourage the modeling of perspectives and
diferent opinions in a very subjective phenomenon like hate
speech.</p>
      <p>
        The past years have seen an increase in using diferent
paradigms that try to model the diferent opinions of
human annotators, especially in cases of recognition of
subjective phenomena, like hate speech. Adopting a soft
perspectivist approach, recent challenges like Le.Wi.Di
(Learning with disagreement) shared task were proposed
at SemEval 2021 and 2023 [8, 9]. In particular, this shared
task asked participants to model various phenomena,
such as humor and hate speech detection, exploiting the
soft labels. These, contrary to the hard labels (simple
labels), are obtained computing a sort of distribution of the 3. Dataset
labels chosen by annotators. Modelling this distribution,
the systems are able to approximate the probability dis- In this section, we first introduce our dataset by
illustrattribution of the opinions about the specific phenomenon. ing the context of the annotation process and secondly
A strong perspectivist approach, instead, looks at whom the general statistics about the corpus. Further, we will
the annotators are and how to model their opinion [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. analyze the distribution of the positive and negative label
      </p>
      <p>In the experimental part of this work, we focused espe- for both the hard and soft label.
cially on the use of soft labels to model the diferent labels
without considering the information available about an- 3.1. Corpus Creation
notators, on the example of Uma et al. [7]. In this study,
the authors experimented the application of the soft la- The DisaggregHate It Corpus used for this work is
combels to detect various phenomena, employing a standard posed of 1100 tweets extracted from Contro L’Odio [15],
and a softmax normalization of the labels. They proved an Italian corpus that focuses on racist hate and in
particthat in both hard and soft evaluation settings, respec- ular on discrimination towards immigrants. The
annotatively using accuracy and cross-entropy metrics, the use tion process carried out as part of a master degree course,
of soft labels in the modelling leads to better results. so the participants are all university students aged
be</p>
      <p>
        Following their example, we evaluated the new disag- tween 21 and 30, and native of the Italian language. A
gregated dataset on hate speech, DisaggregHate It Cor- specific educational web platform has been realized on
pus, composed of Italian tweets, and enriched with some the example of the one developed by [16], for allowing
socio-demographic information about the annotators. the annotation process and the collection of some basic
Our idea was to create a dataset according perspectivist information about the annotators. For each tweet,
anrecommendations provided by Cabitza et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to en- notators have been asked to decide the presence hate
sure the transparency of the created perspective dataset. speech (yes or no), irony (yes or no) and the stance of
Among these recommendations, the authors mention the the author of the message towards immigration issues
involvement of enough and heterogeneous annotators, (positive, neutral, or negative)3. For our experiment, we
and the collection of information about them. Moreover, only considered the hate speech annotations, so from
with our work we meet also other their recommenda- this point forward when we will talk about the target
tions such as the report about the annotation process, the label we are referring to the hate speech one.
use of hard labels (computed by majority voting) and the
soft labels (to represent the distribution of the decisions
provided by annotators), and finally, we validated our 2This dataset is released on HuggingFace: https://huggingface.
models in an out-of-domain setting. co/da3Ttahseetus/suecdbgeurkideeleliyn-edslaabr/emtheaesounriensga-dhoaptete-sdpteoecahnnotate data in
      </p>
      <p>The works on hate speech that comply to some of these the HaSpeeDe context [17] (for hate speech), in the context of IronIta
recommendations and release disaggregated datasets, are [18] (for irony ), and in the context of SardiStance [19] (for stance).
few, and to our knowledge, are only in other languages. Especially the last guidelines have been adapted to the context of
immigration.</p>
      <sec id="sec-2-1">
        <title>Profile</title>
        <sec id="sec-2-1-1">
          <title>City &lt;50k</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>City &gt;50k</title>
        </sec>
        <sec id="sec-2-1-3">
          <title>TSCI</title>
        </sec>
        <sec id="sec-2-1-4">
          <title>Humanistic Men</title>
        </sec>
        <sec id="sec-2-1-5">
          <title>Women</title>
          <p>Low SM</p>
        </sec>
        <sec id="sec-2-1-6">
          <title>High SM</title>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Annotators</title>
      </sec>
      <sec id="sec-2-3">
        <title>Tweets Krippendorf’s</title>
        <p>Annotators provided basic information about their gen- an instance  and  classes , we can determine a vector
der, how many social media platforms they use, if they [1 , 2 , ..,  , ] where  is the number of votes given
live in a city with more than 50 thousands residents and by the annotators for class  to the instance . Softmax
their school background (TSCI or Humanistic). The par- normalization determines the value of the soft label  for
ticipants could choose to give one or more information each example  and class  with the following formula:
about them. 
forIneaocrhdtewretoet,csotluledcetnatss hmaavneybeaennnogtraotuiopneds ians tpeoasmsisbolef  = ∑︀(())
minimum 5 components, and each annotator was asked Meanwhile standard normalization is obtained by
applyto annotate at least 100 tweets per group. However, some ing:
sets of data have been annotated by more than 1 group of 
students and others only by few annotators. Therefore,  = ∑︀
every tweet has a number of annotation in a range from
1 to 13. This case study addresses the hate speech annotation</p>
        <p>In this context, we computed the agreement among the as a binary problem:  ∈ [, ¬]. We computed
annotators, taking into account the information that they standard and a softmax normalisations  for the sole
provided, using Krippendorf’s Alpha [ 20]. This metric, positive class.
indeed, allows evaluating agreement when the matrix Addressing the data labelling with a soft label approach
of annotations is sparse (i.e., the number of annotators prevents discarding annotations and allows for the
creis not constant for each tweet and, thus, some values ation of a more informative annotated corpus. As Uma
could miss). We did not report the value of Krippen- et al. [7] pointed out the softmax normalization, unlike
dorf’s Alpha for the ‘humanistic’ profile as the function the standard one, assigns to an instance a non zero value
requires at least two annotators (see Table 1). The value even where a class received zero votes. Therefore, the
of Krippendorf’s Alpha for the whole dataset is 0.34. softmax normalisation could be seen as a way to smooth
In table 1 we can observe that the agreement intra-group the label distribution, but it could also cause some side
is quite low as Krippendorf’s alpha values that are equal efects. Indeed, whenever  ≃ ∑︀ , i.e. there is
to 0 indicate absence of reliability meanwhile values that complete agreement among annotators but the there is
are equal to 1 show perfect agreement [20]. It means that only a very small number of annotators,  ∀ ̸=  will
annotators have diferent perception of hate speech even be however sensibly larger than 0. Therefore, the use
if they share the same socio-demographic trait. The only of standard normalization would be preferable in the
profile that shows a fair agreement is the ‘City &gt;50K’ presence of many classes and few annotators.
(living in a city with more than 50 thousands residences). In table 2 we can observe how many ties, positive and
However, the scores are low, motivating an approach negative instances are present in our dataset when we
based on learning with disagreement. apply a majority voting to obtain a hard label. We can
also observe how many tweets had an even number of
3.2. Hard and Soft Label Distribution annotators resulting in possible ties. We can see that
there is diferent percentages of positive instances in
We assigned each tweet three diferent labels: a hard some demographic division criteria like gender (Men and
label and two soft labels. The hard label matches the Women). Other category distinctions, like the one based
majority vote of annotations, while the two soft labels, on social media usage, show little diference between
respectively, employ a standard and a softmax normali- the two groups (Low SM and Hight SM). The number
sation. Using the generalization of Uma et al. [7], given of ties is very diferent between the various categories
ranging from 0% to more than 18% of the total instances. 4. Experiments
A very high number of ties indicates the presence of
controversial instances that could be very important in the The DisaggregHate It Corpus has been used to carry out
training phase of a model. The Krippedndorf’s alpha two main settings of experiments: in-domain (test set
values paired with the number of ties show that the Dis- of DisaggregHate It Corpus corpus) and out-of-domain
aggregHate It Corpus contains a not neglectable level (two test sets of two new shared tasks at EVALITA 2023).
of disagreement between the annotators. Overall, we The tested models are: a standard model trained on
agcan see that the DisaggregHate It Corpus is unbalanced gregated labels (called here Binary), and two new models
towards the negative class; therefore, in Section 4, we trained on soft labels (called here Regression) computed
proposed to train the models using weighted labels. in two diferent manners. The former trained to detect</p>
        <p>In Figure 1, we observe the label distribution using the presence or absence of hate speech in the tweets, the
standard or softmax normalization. We can observe that latter trained to give a probability about the presence of
there are more negative instances than positive ones as hate speech in the tweets in line with the distribution of
the most represented bin is the one with  &lt; 0.2. We labels provided by annotators.
can observe a mostly similar tendency comparing the
Figures 1a and 1b even if the standard normalization has 4.1. Models Description
more examples in the bins for the middle values. Overall,
we can observe that annotators usually tend to be in We built all of our models by fine-tuning an already
exagreement when there is a clear signal of hate speech, isting BERT (Bidirectional Encoder Representations from
indeed the bin with values  &gt; 0.8 has more instances Transformers) based model for Italian. BERT is the
statecompared to other ones. of-the-art family of Large Language Models based on the
transformer architecture [21]. There are a lot of BERT
models that have been trained on large amount of data,
thus they can be easily fine-tuned to perform in other
tasks by fine-tuning them with smaller data sets. The
model we chose to use is the uncased Italian BERT model task of the HaSpeeDe3 (Hate Speech Detection) shared
with the Huggingface identifier: dbmdz/bert-base-italian- task [23] annotated in regard to political and religious
uncased created by the MDZ Digital Library team [22]. hate. The used test set from HaSpeeDe3 is composed
We accessed it through the Huggingface platform and of 5600 tweets, containing 2144 positive examples. The
the Python library Transformers which ofers easy to use second dataset is the corpus from the HoDI
(Homotransfunctions to design a simple architecture for fine-tuning phobia Detection in Italian) shared task [24] containing
the pre-trained models for specific tasks like the one of 5000 tweets about homophobia. The test set of HoDI is
classification (i.e., BertForSequenceClassification ). Con- composed of 5000, containing 2008 positive examples.
sidering the characteristics of our dataset and the kind So after training our models with the in-domain training
of experiments that we wanted to perform, we designed sets we tested them on the in-domain tests, the entire
some specific techniques. HaSpeeDe 3 and HoDI training sets.</p>
        <p>The first regards the output of the network. We created
three diferent models: one trained for binary classifica- 4.3. Results
tion with the hard label of the dataset, and two regression
models respectively trained on the soft label computed In table 3 we report all the results in terms of
crosswith standard normalization and the softmax normaliza- entropy (CE) for the in-domain and out-of-domain
extion. Taking into account the need of using a soft metric periments. We decided to only report the CE scores with
(cross-entropy) to compare the performance of our mod- certain test sets to avoid an unfair comparison. Therefore,
els, as suggest by [7, 8, 9], for the binary classifier we we excluded testing the regression model trained on the
obtained soft label predictions by applying the softmax standard normalization soft labels with the softmax
norfunction to the logit outputs. The probabilities from the malized test set, and vice versa. As the binary model soft
regression models are simply obtainable thanks to the label predictions are obtained by applying the softmax
Transformers library by setting the number of labels pa- function, thus we decided it is adequate to calculate the
rameter to 1 of a classification model. As the outputs of CE with the softmax normalized test set. About the
outthe regression models are not bounded, we applied the of-domain testing, we calculated the CE between the soft
clip function to limit their value between 0 and 1. label predictions and the hard label versions of the test</p>
        <p>The second is about the diferent balance of the classes sets, as the disaggregated annotations are not available.
in our dataset. The DisaggregHate It Corpus contained, We can observe in table 3 that both regression models
indeed, more negative label examples than positive ones report better scores than the binary models in all tests
(see Table 2). To deal with this, we experimented by as- both in-domain and out-of-domain. When we compare
signing diferent weights to the positive and negative the CE score obtained with the binary model with the
label. We obtained these diferent weights through the ones obtained with regression models, we can see a very
compute_class_weight function present in the scikit learn significant diference in favor of the regression model in
Python library. These weights were used in the calcu- both scenarios. Observing in details the standard
norlation of the loss function for each model. The binary malization and softmax normalization regression models,
model was trained with a weighed cross-entropy loss we notice that the softmax normalization seems works
function and given that the training set contained hard better in general, in both experimental settings. However,
labels, we easily assigned diferent weights to each la- if in the in-domain setting, the scores report a diference
bel. The regression models were trained with a weighed of 5% in terms of ∆ , in the out-of-domain setting, the
Mean Squared Error loss function and as the label values results from both regression models are similar. These
were real number, we assigned the postive binary label results are in line with the ones obtained in the study of
weight to examples with a soft label value ≥ 0.5 and we Uma et al. [7].
assigned the negative binary label weight to the rest. The Moreover, we can observe that both regression models
models were trained for 5 epochs, each with a learning score slightly worse when compared to the in-domain
rate parameter equal to 2− 5. setting, and this could have been expected as the
crossdomain task is dificult. Another factor of this drop in
4.2. In and Out-of-Domain Testing performance could be that the target label of the
crossdomain datasets was binary and not a real number. This
encourages the releasing of datasets with disaggregated
labels.</p>
        <p>The in-domain test set has been extracted from the
DisaggregHate It Corpus, selecting 20% of the entire
dataset, while the rest was used for the training and
validation sets. As out-of-domain test sets, we used two
datasets in the Italian language that have been released
in the occasion of the 2023 edition of the EVALITA
campaign. The first one is the corpus regarding the second</p>
        <sec id="sec-2-3-1">
          <title>Binary</title>
        </sec>
        <sec id="sec-2-3-2">
          <title>Regression</title>
        </sec>
        <sec id="sec-2-3-3">
          <title>Regression</title>
        </sec>
        <sec id="sec-2-3-4">
          <title>Hard Label</title>
        </sec>
        <sec id="sec-2-3-5">
          <title>Standard Norm. Label Softmax Norm. Label</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Conclusion</title>
      <p>In this work, we presented the DisaggregHate It Corpus, a
new disaggregated dataset in the Italian language of hate
speech. To our knowledge, it is the first dataset released
with disaggregated labels and some socio-demographic
information about the annotators. Computing the
agreement among annotators with the same profile, we noticed
that the Krippendolf’ is very low. Moreover, this
information, paired with the number ties obtained by majority
voting, showed us how disagreement is a real factor in
corpora. That motivates the need to approach the hate
speech detection task with models that encode the
different opinions of humans annotators. To this purpose,
we experimented with the use of a soft label, exploring
two diferent computation of soft labels: standard and
softmax normalization.</p>
      <p>To continue our study on the usage of disagreement as
a factor in learning we carried out diferent experiments
testing the performance of our models in two specific
settings: in-domain and out-of-domain. We created a
binary model based on the hard labels and two
regression models trained on the soft labels (computed with
the two diferent normalization, regular and softmax).
Inspired by previous works [7, 8, 9], we evaluated the
models, employing the cross-entropy between the soft
labels of annotations and the model predictions.
Observing the results, we noticed that the regression models
perform better both when considering in-domain and
out-of-domain test sets. This implies that a soft label is
helpful to integrate annotators disagreement inside our
models in order to be more in line with the distribution
of the opinions of human annotators.</p>
      <p>Taking into account these results, we plan to use the
same DisaggregHate It Corpus, to explore a stronger
perspectivist approach modelling the perspectives of
different groups of annotators on the basis of their
sociodemographic traits or other commonalities.</p>
    </sec>
    <sec id="sec-4">
      <title>Ethics Statement</title>
      <p>The annotation process involved students of the
Politecnico di Torino, who performed this task in an educational
environment. The guidelines and the information about
the annotation task have been shared via the educational
platform exploited for implementing the annotation
process, and discussed during the lessons. The eforts
required to the students has been limited to their time and
oriented to complete a project work being part of the
exam of Internet e social media: tecnologie e derive della
comunicazione in rete. This annotation task has been
used, first of all, to give the students the opportunity to
discuss the disagreement, encouraging a deep reflection
on the importance of developing high quality annotated
resources, to train and evaluate machine learning models.
sider disagreement in evaluation, in: Proceed- [15] A. T. Capozzi, M. Lai, V. Basile, F. Poletto, M.
Sanings of the 1st Workshop on Benchmarking: Past, guinetti, C. Bosco, V. Patti, G. Rufo, C. Musto,
Present and Future, Association for Computational M. Polignano, et al., Computational linguistics
Linguistics, Online, 2021, pp. 15–21. URL: https: against hate: Hate speech detection and
visualiza//aclanthology.org/2021.bppf-1.3. doi:10.18653/ tion on social media in the" contro l’odio" project, in:
v1/2021.bppf-1.3. CEUR Workshop Proceedings, volume 2481,
CEUR[7] A. Uma, T. Fornaciari, D. Hovy, S. Paun, WS, 2019, pp. 1–6.</p>
      <p>B. Plank, M. Poesio, A case for soft loss func- [16] S. Frenda, A. T. Cignarella, M. A. Stranisci, M. Lai,
tions, Proceedings of the AAAI Conference C. Bosco, V. Patti, et al., Recognizing hate with
on Human Computation and Crowdsourcing 8 nlp: The teaching experience of the# deactivhate
(2020) 173–177. URL: https://ojs.aaai.org/index.php/ lab in italian high schools, in: CEUR WORKSHOP
HCOMP/article/view/7478. doi:10.1609/hcomp. PROCEEDINGS, volume 3033, CEUR-WS. org, 2021,
v8i1.7478. pp. 1–7.
[8] A. Uma, T. Fornaciari, A. Dumitrache, T. Miller, [17] C. Bosco, D. Felice, F. Poletto, M. Sanguinetti,
J. Chamberlain, B. Plank, E. Simpson, M. Poesio, T. Maurizio, et al., Overview of the evalita 2018
SemEval-2021 task 12: Learning with disagree- hate speech detection task, in: Ceur workshop
ments, in: Proceedings of the 15th International proceedings, volume 2263, CEUR, 2018, pp. 1–9.
Workshop on Semantic Evaluation (SemEval-2021), [18] A. T. Cignarella, S. Frenda, V. Basile, C. Bosco,
Association for Computational Linguistics, On- V. Patti, P. Rosso, Overview of the EVALITA 2018
line, 2021, pp. 338–347. URL: https://aclanthology. task on irony detection in Italian tweets (IronITA),
org/2021.semeval-1.41. doi:10.18653/v1/2021. in: Proceedings of the Sixth Evaluation Campaign
semeval-1.41. of Natural Language Processing and Speech Tools
[9] E. Leonardelli, A. Uma, G. Abercrombie, D. Al- for Italian (EVALITA 2018) co-located with the Fifth
manea, V. Basile, T. Fornaciari, B. Plank, V. Rieser, CLiC-it, volume 2263, 2018, pp. 1–6.
M. Poesio, Semeval-2023 task 11: Learning with dis- [19] A. T. Cignarella, M. Lai, C. Bosco, V. Patti, R. Paolo,
agreements (lewidi), 2023. arXiv:2304.14803. et al., Sardistance@ evalita2020: Overview of the
[10] C. J. Kennedy, G. Bacon, A. Sahn, C. von Vacano, task on stance detection in italian tweets, in:
ProConstructing interval variables via faceted rasch ceedings of the Seventh Evaluation Campaign of
measurement and multitask deep learning: a hate Natural Language Processing and Speech Tools for
speech application, ArXiv abs/2009.10277 (2020). Italian. Final Workshop (EVALITA 2020), Ceur, 2020,
URL: https://api.semanticscholar.org/CorpusID: pp. 1–10.</p>
      <p>221836648. [20] K. Krippendorf, Computing Krippendorf’s
[11] S. Akhtar, V. Basile, V. Patti, Modeling annota- Alpha-Reliability, Technical Report, University of
tor perspective and polarized opinions to improve PennSylvania, 2011. URL: https://www.asc.upenn.
hate speech detection, Proceedings of the AAAI edu/sites/default/files/2021-03/Computing%
Conference on Human Computation and Crowd- 20Krippendorf%27s%20Alpha-Reliability.pdf .
sourcing 8 (2020) 151–154. URL: https://ojs.aaai. [21] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova,
org/index.php/HCOMP/article/view/7473. doi:10. Bert: Pre-training of deep bidirectional
transform1609/hcomp.v8i1.7473. ers for language understanding, arXiv preprint
[12] D. Kumar, P. G. Kelley, S. Consolvo, J. Mason, arXiv:1810.04805 (2018).</p>
      <p>E. Bursztein, Z. Durumeric, K. Thomas, M. Bailey, [22] K. Clark, M.-T. Luong, Q. V. Le, C. D. Manning,
Designing toxic content classification for a diver- Electra: Pre-training text encoders as
discrimisity of perspectives, in: Seventeenth Symposium nators rather than generators, in: International
on Usable Privacy and Security (SOUPS 2021), 2021, Conference on Learning Representations, 2020,
pp. 299–318. pp. 1–14. URL: https://openreview.net/forum?id=
[13] N. Goyal, I. D. Kivlichan, R. Rosen, L. Vasserman, Is r1xMH1BtvB.</p>
      <p>your toxicity my toxicity? exploring the impact of [23] M. Lai, F. Celli, A. Ramponi, S. Tonelli, C. Bosco,
rater identity on toxicity annotation, Proceedings of V. Patti, Haspeede3 at evalita 2023: Overview of the
the ACM on Human-Computer Interaction 6 (2022) political and religious hate speech detection task, in:
1–28. Proceedings of the Eighth Evaluation Campaign of
[14] M. Cinelli, A. Pelicon, I. Mozetič, W. Quattrocioc- Natural Language Processing and Speech Tools for
chi, P. Kralj Novak, F. Zollo, Italian YouTube Hate Italian. Final Workshop (EVALITA 2023), CEUR.org,
Speech corpus, 2021. URL: http://hdl.handle.net/ Parma, Italy, 2023, pp. 1–8. URL: https://ceur-ws.
11356/1450, slovenian language resource repository org/Vol-3473/paper22.pdf .</p>
      <p>CLARIN.SI. [24] D. Nozza, A. T. Cignarella, G. Damo, T. Caselli,
V. Patti, HODI at EVALITA 2023: Overview of
the Homotransphobia Detection in Italian Task,
in: Proceedings of the Eighth Evaluation
Campaign of Natural Language Processing and Speech
Tools for Italian. Final Workshop (EVALITA 2023),
CEUR.org, Parma, Italy, 2023, pp. 1–8. URL: https:
//ceur-ws.org/Vol-3473/paper26.pdf .</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Cabitza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Campagner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <article-title>Toward a perspectivist turn in ground truthing for predictive computing</article-title>
          ,
          <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
          <volume>37</volume>
          (
          <year>2023</year>
          )
          <fpage>6860</fpage>
          -
          <lpage>6868</lpage>
          . URL: https://ojs.aaai.org/index.php/AAAI/article/ view/25840. doi:
          <volume>10</volume>
          .1609/aaai.v37i6.
          <fpage>25840</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegand</surname>
          </string-name>
          ,
          <article-title>A survey on hate speech detection using natural language processing</article-title>
          ,
          <source>in: Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media</source>
          , Association for Computational Linguistics, Valencia, Spain,
          <year>2017</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . URL: https://aclanthology. org/W17-1101.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Fortuna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nunes</surname>
          </string-name>
          ,
          <article-title>A survey on automatic detection of hate speech in text</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>51</volume>
          (
          <year>2018</year>
          )
          <volume>85</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>85</lpage>
          :
          <fpage>30</fpage>
          . URL: https://doi.org/ 10.1145/3232676.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F.</given-names>
            <surname>Poletto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <article-title>Resources and benchmark corpora for hate speech detection: A systematic review</article-title>
          ,
          <source>Language Resources and Evaluation</source>
          <volume>55</volume>
          (
          <year>2021</year>
          )
          <fpage>477</fpage>
          -
          <lpage>523</lpage>
          . URL: https://rdcu.be/cCdaB.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Plank</surname>
          </string-name>
          ,
          <article-title>The “problem” of human label variation: On ground truth in data, modeling and evaluation</article-title>
          ,
          <source>in: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Abu Dhabi, United Arab Emirates,
          <year>2022</year>
          , pp.
          <fpage>10671</fpage>
          -
          <lpage>10682</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .emnlp-main.
          <volume>731</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Fornaciari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hovy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Paun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Plank</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Poesio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Uma</surname>
          </string-name>
          , We need to con-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>