<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Integrated Gradients as Proxy of Disagreement in Hateful Content</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandro Astorino</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giulia Rizzi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elisabetta Fersini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universitat Politècnica de València</institution>
          ,
          <addr-line>Valencia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Milano-Bicocca</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Online platforms have increasingly become hotspots to spread not only opinions but also hate speech, posing substantial obstacles to developing constructive and inclusive online communities. In this paper, we propose a novel approach that leverages the integrated gradients of pre-trained language models to automatically predict both hate speech and the potential disagreement that can arise from readers. The integrated gradient attributions are used to shed light on the model's decision-making process attributing importance scores to individual tokens and enabling the identification of crucial factors contributing to disagreement and hate speech classifications. The integrated gradients' straightforwardness allows for the recognition of fundamental causes of disagreements and hate speech content. By adopting an interpretable approach, we bridge the gap between model predictions and human comprehension. Our experimental results highlight the efectiveness of our approach, outperforming traditional BERT models and state-of-the-art methods in both prediction tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Learning with Disagreement</kwd>
        <kwd>Integrated Gradients</kwd>
        <kwd>Hateful Content</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>the possibility of disagreement arises in hateful texts
shared on social media platforms (e.g. Twitter), it
beIn the modern era, human beings are constantly sub- comes critical to have a service that recognizes if that text
ject to absorbing content of various kinds generated and written in that manner causes disagreement and works
shared on the web. To ensure the sustainability of con- as a filter for these texts based on a personal perspective
tinuously produced information and promote individual Moreover, a detoxification strategy could be
impleand societal well-being in the context of online content mented to notify the authors of user-generated texts,
is important to recognize where hate content can harm cautioning them about the potential perception of their
from a personal perspective. Diferent individuals, ac- content as hateful by certain readers, and suggesting
cording to their cultural beliefs and backgrounds, may be revisions for the original message. Identifying
disagreemore or less susceptible to potentially ofensive content. ments within hateful sentences and determining the
asIt is, therefore, necessary to safeguard the perceptions of sociated disagreement-related elements can significantly
diferent individuals by defining Natural Language Pro- contribute to the creation of reliable benchmarks.
Pricessing (NLP) models that are able to capture and model marily, for contents prone to disagreements, specific
andiferent perceptions. How to deal with disagreement, in notation policies can be implemented (e.g., involving
particular related to hate speech detection problems, is more annotators, excluding samples requiring annotation
a topic that has attracted increasing interest during the from the dataset, etc.). Additionally, annotators could be
last few years [1, 2, 3, 4]. Although a good number of provided with targeted cues to focus on particular
conapproaches able to deal with disagreement in hate speech stituents that may be perceived diferently by readers
detection problems have been proposed [5, 6, 7, 8], only (e.g., underlining words, hashtags, or emojis identified as
a few of them have been focused on really modelling disagreement-related elements warranting careful
evaluperspectivism. ation).</p>
      <p>Recognizing potential disagreements within hateful In this paper, we try to connect hate speech and
discontent, especially in identifying controversial elements, agreement by determining which hateful constituents
is of paramount importance for multiple reasons. When can contribute more to predicting disagreement. In
particular, we combine pre-trained language models and
CLiC-it 2023: 9th Italian Conference on Computational Linguistics, integrated gradients providing the following main
con*NCovor3r0es—poDnedcin0g2,a2u0t2h3o,rV.enice, Italy tributions:
$ a.astorino2@campus.unimib.it (A. Astorino);
g.rizzi10@campus.unimib.it (G. Rizzi); elisabetta.fersini@unimib.it • a filtering strategy of textual constituents that
con(E. Fersini) tributes remarkably to explain hateful messages;
0000-0002-0619-0760 (G. Rizzi); 0000-0002-8987-100X (E. Fersini) • a unified model that, considering the prediction of
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ©ACt2tEr0i2bU3utCRioonpWy4r.0igoIhnrttekfornsrahtthiooisnppaalp(PCerCrboByYcite4s.0ea)u.dthionrsg.Usse( CpeErmUittRed-uWndeSr.Correagti)ve Commons License the hateful contents and the selected explanations,
1,120
943
4,050
10,753</p>
      <p>Task</p>
      <p>Hate Speech
Misogyny and sexism detection
Abusive Language detection</p>
      <p>Ofensiveness detection</p>
      <p>Annotators Pool Ann. % Full Agreement
6
3</p>
      <p>predicts if disagreement could arise when reading tween hate speech and non-hateful expressions. In recent
such contents; years, hate speech detection has extended to encompass
multimodal data analysis to keep up with the increasing</p>
      <p>The rest of the paper is organized as follows. In Section usage of images and videos in online communication.
2 an overview of the state of the art is provided. The Combining textual information with visual cues from
adopted datasets are described in Section 3. In Section 4 images and videos has shown promise for improving the
the proposed approach is detailed. The results achieved accuracy and granularity of hate speech identification
by the proposed approaches are reported in Section 5. systems [20, 14]. An increasing number of datasets are
Finally, conclusions and future research directions are collecting multimodal examples of hate content ranging
drawn in Section 6. from memes [20, 21] to advertisements [22] and videos
[23].
2. Related Work The latest datasets are addressing the problem of hate
speech under the Learning with Disagreements paradigm
The fast rise of social media and online communication reporting information both on the hard label (usually
platforms has changed the way people communicate, ex- obtained through majority voting) and on the soft
lachange information, and express their ideas, while simul- bel (with all the annotators’ labels or a confidence level
taneously increasing the spread of hate speech. Hateful attached to the labels). The inclusion of diferent
perspeccontent includes a wide range of various forms of ofen- tives allows us to address the subjectivity of the task by
sive, abusive, and discriminatory language targeted at representing the multiple perceptions of the annotators
individuals or groups based on their race, religion, ethnic- with diferent points of view and understanding [ 24]. The
ity, gender, or other protected characteristics. The prop- information that represents annotators’ disagreement is
agation of hate speech online has major implications, not only used to improve the quality of the dataset [25]
perpetuating discrimination, stoking antagonism, and but also in the training process by weighting the
saminstigating violence, necessitating the urgent need for ples according to their disagreement values [26] or by
efective anti-hate speech solutions. Over the years, sig- directly training from disagreement, without considering
nificant progress has been made in developing automatic any aggregates label [27, 28].
hate content detection systems that leverage
advancements in Natural Language Processing (NLP), machine 3. Dataset
learning, and deep learning techniques. In this section,
we highlight some of the state-of-the-art approaches and The four benchmark datasets provided by SemEval 2023
methodologies employed in hate speech detection. The task 11 related to Learning With Disagreements [29] have
dominant approach for hate speech detection is repre- been considered in order to address the problem of
presented by supervised learning [13, 14]. In particular, the dicting disagreement in hateful content. The datasets
approaches based on Language Models (LM) [15, 16, 17] have diferent characteristics for what concerns language,
have shown promising results in capturing contextual type, and goal as summarized in Table 1. All the datasets
information and semantic relationships, leading to im- have been adapted by the challenge organizer to share a
proved classification performance. common structure for what concern the textual input and</p>
      <p>One of the key challenges in hate speech detection is the hard and soft labels (additional dataset-specific
atthe ability to make sense of the context in which the ofen- tribute are present). Since in this work, the disagreement
sive language is used. Researchers have explored context- prediction is addressed as a binary task, an agreement
aware models [18, 19] that consider the surrounding text label has been derived from the soft label. This is because
or conversation to make more accurate predictions. This taking the levels of disagreement into account requires
can exploit speaker attributes, or discourse patterns to knowledge of the number of annotators, which is not
better grasp the intended meaning and diferentiate
betaken into account at this time since the objective is to
distinguish agreement and disagreement and not the
various levels of disagreement. In particular, the agreement
label is set equal to (+) when there is a 100% agreement
between the annotators, regardless of the value of the
hard label, while equal to (− ) in all the other cases.</p>
    </sec>
    <sec id="sec-2">
      <title>4. Proposed Approach</title>
      <p>The proposed approach aims at addressing the tasks
of predicting both disagreement and hate speech while
maintaining the method fully interpretable through the
adoption of integrated gradients. Integrated gradients
are used to shed light on the model’s decision-making
process attributing importance scores to individual
tokens and enabling the identification of crucial factors
contributing to the model’s decision.</p>
      <p>In particular, the proposed approach is composed of
four main steps:
contribution to explain the target label. In
particular, let  be the -th token within a text
 and  the corresponding attribution score.</p>
      <p>The token  is considered significant and
maintained for the subsequent disagreement model if
 ≥  , otherwise the token is removed from
the original input text. In our case study,  is a
specific threshold estimated according to a grid
search approach.
4. Extraction of latent representations: the
tokens considered significant according to the
previous step are used to extract the corresponding
latent representation of the filtered sentence from
the fine-tuned m-BERT model.
5. Creation of the disagreement input space:
the latent representation obtained at the previous
step is used according to the following strategies:
• Filtered Embeddings: the embedding of
the filtered sentence is obtained by
finetuned model on hate and used to train the
subsequent disagreement model.
• Predicted Label: the Boolean labels
predicted by the model fine-tuned to
distinguish hateful from non-hateful messages
are included in the input space for training
the disagreement model.
• Distribution values: the distribution
probability obtained through the sigmoid layer
of the fine-tuned models has been
alternatively considered.
1. Fine-tuning of a pre-trained LM: the
multilingual BERT (m-BERT) has been fine-tuned to
distinguish hateful content from non-hateful ones.</p>
      <p>The textual input (i.e. the tweet or the
conversation depending on the dataset) has been given as
input to the m-BERT model with a final sigmoid
layer. Additionally, to overcome the datasets’
class imbalance, in the training phase, the loss
function has been penalized accordingly to the
class distribution. The optimal decision threshold
has been determined according to the Youden’s
J statistics [30]. The statistics, which is a linear 6. Training of the disagreement model: the
combination of sensitivity and specificity, is max- derived input space (latent representation of the
imized by evaluating several cut-ofs. selected token, concatenated with the predicted
2. Estimation of the attribution score: the at- label or probability distribution) is given as input
tribution score for each textual constituent has to a trivial Neural Network with the following
been estimated using the integrated gradients pre- structure to predict disagreement labels:
sented in [31] on the fine-tuned model. This at- • Input layer: layer that reflects the shape of
tribution score assumes values from -1 to 1, 1 the input, with Relu as activation function
means that that token has a high contribution to and dropout of 0.7;
the prediction of the model and -1 the opposite. • Hidden layer: layer that halves the size of
A visual representation of the integrated gradient the input with Relu and dropout of 0.7;
on two available samples is reported in Figure • Output layer: one output neuron with a
1. On one hand, each attribution score allows sigmoid function to predict the final
agreeus to identify those tokens that contribute more ment/disagreement.
to the final prediction, and on the other hand,
those compositions of tokens characterized by The entire proposed approach is synthesized in Figure
divergent values make the content controversial 2.
potentially leading to disagreement. The
variability and the magnitude of attribution values 5. Experimental Results
within a text are subsequently exploited to detect
a potential disagreement. In this section, the results obtained by the proposed
ap3. Filtering constituents: the integrated gradi- proach are reported. We measured Precision (P), Recall
ent’s attribution scores have been used to filter (R) and F-Measure (F), distinguishing between hateful
out those tokens that do not bring a significant (+) and not hateful (− ) labels and reporting also the</p>
      <sec id="sec-2-1">
        <title>Morons...get your covid ... I mean koolaid</title>
        <p>[Hateful Tweet with Agreement]</p>
      </sec>
      <sec id="sec-2-2">
        <title>Flying in the face of science logic and common sense. People are dying and you don’t give a shit [Hateful Tweet with Disagreement]</title>
        <sec id="sec-2-2-1">
          <title>Macro F-Measure. We show in Table 2 the performance</title>
          <p>achieved by the fine-tuned model on the hate speech
detection task. The achieved results denote good prediction
capability, especially for the negative class (non-hateful).
This behaviour is mainly due to the unbalanced nature
of the datasets and in some cases to the limited number
of instances available.</p>
          <p>Now, we report in Table 3 the performance on the
disagreement prediction, distinguishing however between
agreement (+) and disagreement (− ). The results of the
proposed method are shown according to the input space
previously described. In particular, we report:
• m-BERT: a baseline m-BERT model fine-tuned
according to the disagreement label;
• NN + Filt: a neural network that takes as input
the embedding representation of the sentence
composed of the tokens selected according to the
attribution scores and trained on the
disagreement label. This configuration corresponds to the
one described in step 5(a);
• NN + Pred: a neural network that takes as input
the embedding representation of the sentence
composed of the tokens selected according to
the attribution scores with an additional Boolean
feature denoting the label predicted by the
finetuned model on the hate. This configuration
corresponds to the one described in step 5(b);
• NN + Dist: a neural network that takes as input
the embedding representation of the sentence
composed of the tokens selected according to the
attribution scores with two additional features
denoting the probability distribution associated
with the labels predicted by the fine-tuned model
on the hate. This configuration corresponds to
the one described in step 5(c);</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>In order to understand whether the proposed ap</title>
          <p>proaches obtain significant results compared with
mBERT, a McNemar Test has been performed. In
particular, the McNemar Test has been adopted to perform a
pairwise comparison between the m-BERT predictions
and each of the proposed strategies according to a
confidence level equal to 0.95. If a given model outperforms
m-BERT and its error distribution is diferent compared
to m-BERT, then the corresponding F1-Score is marked
with a wildcard symbol (* ) in Table 3.</p>
          <p>It can be easily noted that, in the majority of the
considered datasets, all of the proposed approaches significantly
outperform the considered baseline m-BERT. It is also
interesting to highlight that, considering the datasets are
even more unbalanced and with a very limited number of
samples, the proposed approach NN-Dist tends to achieve
more balanced performance between the two labels than
the other methods. The McNemar test confirms that the
NN-Dist strategy is not only the best-performing one
but also that the predictions are diferent with respect</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Dataset</title>
      </sec>
      <sec id="sec-2-4">
        <title>HS-brexit</title>
      </sec>
      <sec id="sec-2-5">
        <title>ArMIS</title>
      </sec>
      <sec id="sec-2-6">
        <title>ConvAbuse</title>
      </sec>
      <sec id="sec-2-7">
        <title>MD-Agreement</title>
        <p>F+
0.76
0.83
0.83
0.81
0.37
0.76
0.73
0.71
0.93
0.73
0.78
0.80
0.38
0.57
0.52
0.53
P−
0.51
0.62
0.62
0.57
0.32
0.50
0.46
0.47
0.33
0.23
0.24
0.27
0.58
0.67
0.66
0.66
R−
0.73
0.50
0.50
0.67
0.65
0.11
0.25
0.38
0.03
0.76
0.65
0.72
0.68
0.43
0.66
0.68</p>
      </sec>
      <sec id="sec-2-8">
        <title>Macro F</title>
        <p>to the ones given by m-BERT. This implies that the per- ported). In ConvAbuse, the misclassification of the
proformance of the proposed approach could be considered posed approach is mainly due to the reduced number of
statistically significant. tokens of the text. In fact, 40% of the original text
con</p>
        <p>An additional remark concerns the relationship that tains less than 3 tokens, making dificult the prediction of
exists between the disagreement prediction model and disagreement. Finally, in MD-Agreement the error rate
the model able to predict hateful content. The perfor- is quite higher (42.79%) compared to the other datasets.
mances of the proposed models are strictly related to In this scenario, the misclassified samples are almost
balthe recognition capabilities of the model fine-tuned to anced between the two classes, (i.e., 0.45% for the
agreedistinguish hateful content from non-hateful ones. Im- ment and 55% for the disagreement class). The main
proving the recognition capabilities of the hateful model reason behind the high classification error can be found
is expected to increase the recognition potential of the in the diferent arguments covered by the dataset. This
proposed disagreement models. suggests that disagreement is not only related to
difer</p>
        <p>For what concerns the errors of the most promising ent beliefs or backgrounds but also to specific discussed
approach, i.e., NN-Dist, we can highlight that on the topics.</p>
        <p>HS-Brexit dataset, most of the misclassifications are due
to the absence of relevant information. In particular, in
70% of the misclassified samples, there are references 6. Conclusions and Future works
to users and links that have been omitted, making the
understanding of the context even more complex. Re- The proposed paper introduces a novel approach for
degarding the ArMis dataset, most of the errors are related tecting disagreement in hateful content. The method
to the implicit language used to express hateful content leverages integrated gradients from pre-trained language
against women (no explicit insults or sexist expressions models to predict both hate speech and potential
disagreeare used, but more subtle misogynous samples are re- ment arising from diferent readers. The approach is
evaluated on four benchmark datasets related to Learning</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgments</title>
      <sec id="sec-3-1">
        <title>The work of Elisabetta Fersini has been partially funded</title>
        <p>by the European Union – NextGenerationEU under the
National Research Centre For HPC, Big Data and
Quantum Computing - Spoke 9 - Digital Society and Smart
Cities (PNRR-MUR), and by MUR under the grant
“Dipartimenti di Eccellenza 2023-2027” of the Department
of Informatics, Systems and Communication of the
University of Milano-Bicocca, Italy.</p>
        <p>With Disagreements, and the results show that the
proposed method outperforms the baseline m-BERT model
in disagreement prediction tasks. One of the proposed
strategies, namely NN + Dist, performs particularly well
and achieves statistically significant improvements
compared to a baseline model based on m-BERT. Overall, the
proposed approach demonstrates the potential to predict
disagreement in hateful content compared to bert. Future
work could focus on exploring the applicability of the
proposed approach to other languages and expanding
the scope to include multimodal data analysis,
considering the increasing use of images and videos in online
communication.
18653/v1/2021.emnlp-main.822. [25] B. Beigman Klebanov, E. Beigman, From annotator
[13] F. Poletto, V. Basile, M. Sanguinetti, C. Bosco, agreement to noise models, Computational
LinguisV. Patti, Resources and benchmark corpora for hate tics 35 (2009) 495–503.
speech detection: a systematic review, Language [26] A. Dumitrache, F. Mediagroep, L. Aroyo, C. Welty,
Resources and Evaluation 55 (2021) 477–523. A crowdsourced frame disambiguation corpus with
[14] A. Chhabra, D. K. Vishwakarma, A literature sur- ambiguity, in: Proceedings of NAACL-HLT, 2019,
vey on multimodal and multilingual automatic hate pp. 2164–2170.
speech identification, Multimedia Systems (2023) [27] A. Uma, T. Fornaciari, D. Hovy, S. Paun, B. Plank,
1–28. M. Poesio, Learning from disagreement: A survey,
[15] M. Mozafari, R. Farahbakhsh, N. Crespi, Hate Journal of Artificial Intelligence Research 72 (2021)
speech detection and racial bias mitigation in so- 1385–1470.
cial media based on bert model, PloS one 15 (2020) [28] T. Fornaciari, A. Uma, S. Paun, B. Plank, D. Hovy,
e0237861. M. Poesio, et al., Beyond black &amp; white: Leveraging
[16] H. S. Alatawi, A. M. Alhothali, K. M. Moria, Detect- annotator disagreement via soft-label multi-task
ing white supremacist hate speech using domain learning, in: Proceedings of the 2021 Conference
specific word embedding with deep learning and of the North American Chapter of the Association
bert, IEEE Access 9 (2021) 106363–106374. for Computational Linguistics: Human Language
[17] H. Saleh, A. Alhothali, K. Moria, Detection of hate Technologies, Association for Computational
Linspeech using bert and hate speech word embedding guistics, 2021.
with deep model, Applied Artificial Intelligence 37 [29] E. Leonardelli, A. Uma, G. Abercrombie, D.
Al(2023) 2166719. manea, V. Basile, T. Fornaciari, B. Plank, V. Rieser,
[18] M. Fernandez, H. Alani, Contextual semantics for M. Poesio, Semeval-2023 task 11: Learning with
disradicalisation detection on twitter (2018). agreements (lewidi), 2023. arXiv:2304.14803.
[19] M. Bilal, A. Khan, S. Jan, S. Musa, Context-aware [30] W. J. Youden, Index for rating diagnostic tests,
Candeep learning model for detection of roman urdu cer 3 (1950) 32–35.
hate speech on social media platform, IEEE Access [31] M. Sundararajan, A. Taly, Q. Yan, Axiomatic
attribu10 (2022) 121133–121151. doi:10.1109/ACCESS. tion for deep networks, in: International conference
2022.3216375. on machine learning, PMLR, 2017, pp. 3319–3328.
[20] D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh,</p>
        <p>P. Ringshia, D. Testuggine, The hateful memes
challenge: Detecting hate speech in multimodal
memes, Advances in neural information processing
systems 33 (2020) 2611–2624.
[21] E. Fersini, F. Gasparini, G. Rizzi, A. Saibene,</p>
        <p>B. Chulvi, P. Rosso, A. Lees, J. Sorensen,
SemEval2022 task 5: Multimedia automatic misogyny
identification, in: Proceedings of the 16th
International Workshop on Semantic Evaluation
(SemEval2022), Association for Computational Linguistics,
Seattle, United States, 2022, pp. 533–549. URL:
https://aclanthology.org/2022.semeval-1.74. doi:10.</p>
        <p>18653/v1/2022.semeval-1.74.
[22] F. Gasparini, I. Erba, E. Fersini, S. Corchs, et al.,</p>
        <p>Multimodal classification of sexist advertisements.,
in: ICETE (1), 2018, pp. 565–572.
[23] M. Das, R. Raj, P. Saha, B. Mathew, M. Gupta,</p>
        <p>A. Mukherjee, Hatemm: A multi-modal dataset
for hate video classification, in: Proceedings of the
International AAAI Conference on Web and Social</p>
        <p>Media, volume 17, 2023, pp. 1014–1023.
[24] A. Uma, T. Fornaciari, D. Hovy, S. Paun, B. Plank,</p>
        <p>M. Poesio, Learning from disagreement: A
survey, The Journal of Artificial Intelligence Research
72 (2021) 1385–1470. doi:https://doi.org/10.
1613/jair.1.12752.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>