<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>O-Dang at HODI and HaSpeeDe3: A Knowledge-Enhanced Approach to Homotransphobia and Hate Speech Detection in Italian</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chiara Di Bonaventura</string-name>
          <email>bonaventura@kcl.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arianna Muti</string-name>
          <email>iarianna.muti2@unibo.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Antonio Stranisci</string-name>
          <email>marcoantonio.stranisci@unito.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>King's College London</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Processing and Speech Tools for Italian</institution>
          ,
          <addr-line>Sep 7 - 8, Parma, IT</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Bologna</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Turin</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes our methods implemented during the EVALITA 2023 campaign for homotransphobia (HODI task) and hate speech detection (HaSpeeDe3 task) in Italian. We present three knowledge-enhanced approaches, namely via triple verbalisation, via prompting and via a majority vote, and we compare them to the AlBERTo baseline. These systems leverage the knowledge graph O-Dang, which contains information about named entities in Italian dangerous speech. Our knowledge-enhanced systems outperformed all the competition's baselines. Our best submissions achieved the macro-F1 score of 0.912 for HaSpeeDe3 and 0.795 for HODI, reaching the 1st and 3rd place, respectively. These results were achieved by using our baseline for HODI, and a majority voting approach for HaSpeeDe3.1 ofensive content. 1 Commons License Attribution 4.0 International (CC BY 4.0).</p>
      </abstract>
      <kwd-group>
        <kwd>Detection</kwd>
        <kwd>hate speech</kwd>
        <kwd>knowledge graph</kwd>
        <kwd>entity linking</kwd>
        <kwd>data augmentation</kwd>
        <kwd>prompting</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Technological progress and increasing online
communication have made it necessary to create automatic tools
for the detection of online abusive language to protect
users. Indeed, online abusive language not only has
increased over the past years, but also has efects that
usually expand beyond the online context2. Usually, most of
the targeted victims belong to minority groups because
of their gender identity, sexual orientation, political and
religious afiliation,</p>
      <p>
        inter alia. Therefore, their protection
is of the utmost importance. Gender- and sexual-based
violence manifests in social networks every time the
abusive language harms LGBTQIA+ individuals directly or
indirectly with homotransphobic discourse [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Politicaland religious-based hate speech instead discriminates
people based on their beliefs, afiliations, or ideologies.
      </p>
      <sec id="sec-1-1">
        <title>Within the NLP community, many recent works propose solutions to automatically classify misogynous and</title>
        <p>nEvelop-O
EVALITA 2023: 8th Evaluation Campaign of Natural Language
CEUR
Workshop
Proce dings
htp:/ceur-ws.org
ISN1613-073
© 2021 Copyright for this paper by its authors. Use permitted under Creative</p>
        <p>CEUR</p>
        <p>
          Workshop Proceedings (CEUR-WS.org)
com/dnozza/profanity-obfuscation) [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]
report-online-hate-increasing-against-minorities-says-expert
and religious hate speech [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>Evaluation campaigns have sped up the development
of innovative approaches for shared tasks. One
example is the Homotransphobia Detection in Italian (HODI)
shared task [9], which was introduced during EVALITA
2023 [10]. The task comprises two subtasks. The first
subtask, Subtask A - Homotransphobia detection, focuses
on automatically identifying whether Italian online posts
contain homotransphobic content or not. Subtask B</p>
      </sec>
      <sec id="sec-1-2">
        <title>Explainability, aims at extracting the rationales of the</title>
        <p>classification model trained for Subtask A.</p>
      </sec>
      <sec id="sec-1-3">
        <title>Another example is the shared task HaSpeeDe3 [11],</title>
        <p>presented during EVALITA 2023 to boost research on
political and religious hate speech in Italian tweets. The
task counts two subtasks. The goal of the first subtask,
Subtask A - Political Hate Speech Detection, is to
determine whether the message contains political hate speech
or not. The problem is further divided in two distinct
approaches, namely textual and contextual. In the former,
participants can only use the provided textual content
of the tweets, whereas in the latter, they can employ
additional contextual information (e.g., metadata of the
tweet and author, friends, etc.). Subtask B - Cross-domain</p>
      </sec>
      <sec id="sec-1-4">
        <title>Hate Speech Detection, instead, proposes a binary hate</title>
        <p>speech detection on test data belonging to diferent hate
domains. Precisely, two settings are evaluated:
XPolitical</p>
      </sec>
      <sec id="sec-1-5">
        <title>Hate, where the participants can use external data from any kind of other domains, and XReligiousHate, where domain.</title>
      </sec>
      <sec id="sec-1-6">
        <title>In this paper, we describe our proposed approach to ad</title>
        <p>1Profanities have been obfuscated with PrOf (https://github. submissions are tested on tweets from the religious hate
To address the two tasks, datasets were provided by the
tasks organisers. For HODI, we focus on Subtask A. 5,000
tweets were provided, manually labelled according to
two classes, homotransphobic and not. Data are slightly
skewed towards the negative class. For HaSpeeDe3, we
focus on both Subtask A and B, where A aims to classify
tweets between hateful and not, where hate is addressed
towards politicians, while in Subtask B hate is addressed
towards religious communities, in a cross-domain setting.</p>
        <p>PolicyCorpusXL (for Task A) contains 7,000 tweets about
political debates, while ReligiousHate (for Task B) is
composed by 3,000 tweets about the three main monotheistic
religions, namely Christianity, Islam and Judaism.
3. Description of our Systems
dress the HODI and HaSpeeDe3 shared tasks. We propose the Italian language trained on Twitter posts which
ina knowledge-enhanced approach on top of the AlBERTo cludes emojis, links, hashtags, and mentions. AlBERTo
baseline that leverages external knowledge (System 1), was trained on 200M tweets randomly sampled from the
internal knowledge (System 2) or both (System 3). Our re- TWITA corpus [15].
sults suggest that knowledge-enhancement can improve The following paragraphs describe the entity linking
abusive language detection depending on the hate do- step, explain the knowledge-enhancement methods we
main under study. implemented, and provide insights on its application to</p>
        <p>The rest of the paper is structured as follows. Section HODI and HaSpeeDe3 shared tasks.
2 describes the training datasets provided by the tasks’
organisers whereas Section 3 describes our proposed sys- 3.1. Entity Linking
tems. Section 4 summarises the experiments performed
and discusses the results. Section 5 includes related work Our knowledge augmentation strategy relies on an entity
in the field. Section 6 draws some conclusions and dis- linking pipeline based on a Knowledge Graph (KG)
modcusses further possible research lines. elled on the Ontology of Dangerous Speech (O-Dang)
[16]. The KG is a snapshot of a Wikidata [17] dump3
where only entities of the type ‘person’ were retained,
2. Data together with a set of 31 properties conveying relevant
information (eg: ‘member of political party’ (P102), ’place
of birth’ (P19), spouse (P26)). As a result, we obtained
9, 552, 706 entities and 69, 521, 846 triples related to them.</p>
        <p>After building the knowledge base, we implemented a
pipeline for Entity Linking organized in three steps:</p>
      </sec>
      <sec id="sec-1-7">
        <title>In this work, we adopt a knowledge-enhanced approach</title>
        <p>to address the following subtasks: Subtask A for the An example of such a pipleline is the following: Spacy
HODI shared task, and Subtask A - textual and Subtask B - identified ‘Zorzi’ as a named entity of the type person
XReligiousHate for the HaSpeeDe3 shared task. However, in the tweet: ‘@user_abcdefghij Zorzi, poveretto parla
the systems submitted for the Subtask A - textual satisfy come una ch*cca isterica @user_abcdefghij Zorzi, poor
the constraints for Subtask A - contextual and Subtask B guy, he talks like a f*ggot’. We queried Wikipedia APIs
- XPoliticalHate too. inputting the string ‘Zorzi’, obtaining the candidate
‘Tom</p>
        <p>Our intuition is to leverage knowledge about named maso Zorzi’, which was the first ranked of the list. We
entities in the training data and their association to on- computed the Ratclif/Obershelp pattern recognition to
line abusive language, which can provide auxiliary infor- obtain a similarity score between ‘Zorzi’ and ‘Tommaso
mation to solve the tasks. Abusive language detection Zorzi’, which was 0.55 and we did the same with their
systems often fail to capture diferent nuances of abusive embeddings (0.9). We averaged these scores to obtain
language because of the lack of contextual information a general score of 0.818. After a manual review of the
[12] and the target-oriented nature of hate speech [13]. scores, we decided to keep only linked entities with a
We propose knowledge-injection of relevant information score equal to or above 0.8.
into the system to help address this challenge. Firstly,
we link training instances to the named entities
associated with them. Then, we collect relevant information
about named entities from the O-Dang knowledge graph 3https://academictorrents.com/download/
and the Davinci OpenAI model, which are then injected 229c4fehbtt2p3s3:1//asdp4a3cdy4.i7o0/6efd435f6d78f40a3c438.torrent
into the AlBERTo baseline [14], a version of BERT for 5https://www.mediawiki.org/wiki/API:Search</p>
      </sec>
      <sec id="sec-1-8">
        <title>1. we identified all the entities of the type PERSON</title>
        <p>in the training sets of HaSpeeDe and HODI with
Spacy4;
2. we found all potential candidates of each detected
entity, by using the Wikipedia APIs5;
3. we generated three scores for each pair of the type
&lt;entity, candidate&gt;: string similarity based on
the Ratclif/Obershelp pattern recognition [ 18],
cosine similarity based on AlBERTo embeddings
[14], the ranking of candidates returned from
Wikipedia APIs. We retrieved 10 candidates and
we only kept the most relevant result.
quired during pretraining. We try the following prompts,
among others:
• P1: quanto è probabile che ENTITY scriva un
tweet ofensivo? Se sì, perchè? (How likely is
ENTITY to write an ofensive tweet? If so why?)
• P2: quanto è probabile che ENTITY sia associato
ad un tweet ofensivo? Se sì, perchè? (How likely
is ENTITY to be associated to an ofensive tweet? If
so why?)
• P3: quanto è probabile che ENTITY sia vittima
di un tweet ofensivo? Se sì, perchè? (How likely
is ENTITY to be victim of an ofensive tweet? If so
why?)</p>
      </sec>
      <sec id="sec-1-9">
        <title>The resulting number of entities identified in the two</title>
        <p>datasets and linked to our KG are 388 from HODI and
556 from HaSpeeDe3.</p>
        <sec id="sec-1-9-1">
          <title>3.2. Knowledge-enhancement</title>
        </sec>
      </sec>
      <sec id="sec-1-10">
        <title>We compare the AlBERTo baseline to the following three knowledge-enhanced systems:</title>
        <p>System 1: enhanced-AlBERTo with triple
verbalisation. For each named entity in the input data, we
create a verbalised description using the information
retrieved from O-Dang [16] through entity linking. Then,
this description is concatenated at the end of the input
instance, and passed to AlBERTo for the classification. Among them, only prompt P3 returned relevant results.
We generate up to three diferent verbalised templates Indeed, direct questions such as P1 and P2 triggered
per O-Dang property to account for linguistic variation. the language model into answering with a non-response
For instance, the property ‘P108’ which stands for job — i.e., ‘Come modello di intelligenza artificiale, non ho
location is translated into the following templates: (i) accesso a informazioni specifiche sulle associazioni di
‘lavora presso’ (works at), (ii) ’svolge la sua professione ENTITY con tweet ofensivi’ (As an artificial intelligence
presso’ (carries out their profession at), or (iii) ‘svolge il model, I do not have access to specific information about
suo lavoro presso’ (carries out their job at). When cre- ENTITY’s associations with ofensive tweets) . On the other
ating these templates, we ensure gender neutrality to hand, the keyword ‘victim’ does not trigger the model,
avoid any possible source of bias. Then, for each input which is thus able to return reasonable answers given its
naming an O-Dang entity, the system randomly picks pretraining knowledge. Further, we compared popular
one template per property to create the verbalisation large language models in order to choose the best one.
Taof the O-Dang triples whose heads refer to the named ble 2 and Table 3 show the diference in returned answers
entity. As a result, the verbalised descriptions for the between Davinci and Curie OpenAI models, respectively.
same named entity tend to be all linguistically diferent. Davinci performs very well, whereas Curie – which
acFurther, we fine-tune Italian BERT 6 on ItaCoLa [19] to cording to the OpenAI website8 performs almost equally
check the linguistic acceptability of the descriptions. All well – performs poorly with respect to knowledge about
descriptions passed the test since they were predicted as ofensive content.
linguistically acceptable according to the model.7 Table Similar to System 1, Davinci’s knowledge is injected
1 shows statistics of the O-Dang triples and their verbali- into AlBERTo by concatenating the returned answer to
sation: the average number of triples, the max number prompt P3 at the end of the input instance, where
ENof triples, and the average length of the verbalised triples TITY refers to the named entity in the instance. If the
of an O-Dang entity. instance names multiple entities, we concatenate all their
descriptions.
avg #triples
max #triples
avg length verbalised triples</p>
        <p>HaSpeeDe3
9.9
24
43</p>
        <p>HODI
9.7
24
40</p>
      </sec>
      <sec id="sec-1-11">
        <title>6https://huggingface.co/dbmdz/bert-base-italian-xxl-cased</title>
        <p>7The fine-tuned BERT we used for the inference on the
verbalised descriptions achieved a total MCC score of +0.57 on the
ItaCoLa dataset.
System 3: Majority vote. The third system takes
a majority vote between the AlBERTo baseline, the
enhanced-AlBERTo with triple verbalisation and the
enhanced-AlBERTo with prompting Davinci. The
intuition behind is to leverage both external knowledge
provided by O-Dang triples and internal knowledge
provided by Davinci model.</p>
        <p>OpenAI descriptions
Prompt P3 with ENTITY = Matteo Salvini
Davinci
Most capable GPT-3 model. Can do any
task the other models can do, often with
higher quality.</p>
        <p>Il leader della Lega Nord, Matteo Salvini, è
probabilmente vittima di commenti online
ofensivi a causa della sua politica
antiimmigrazione. Salvini ha preso una
posizione molto dura contro l’immigrazione
clandestina e ha fatto campagne per
limitare l’ingresso di immigrati in Italia. La
sua politica ha attirato molte critiche da
parte di coloro che sostengono che
l’immigrazione è un diritto umano fondamentale.</p>
        <p>Inoltre, Salvini è spesso criticato The leader
of Lega Nord, Matteo Salvini, is probably
the victim of ofensive online comments
because of his anti-immigration policy. Salvini
took a very tough stance against illegal
immigration and campaigned to limit the entry
of immigrants into Italy. His policy has
attracted much criticism from those who argue
that immigration is a fundamental human
right. In addition, Salvini is often criticized
Davinci
Nel corso degli ultimi anni, Matteo Renzi
è stato vittima di diversi commenti online
ofensivi. In particolare, nel 2016, è stato
preso di mira da alcuni utenti che hanno
criticato il suo modo di gestire la crisi
economica in Italia. Over the past few years,
Matteo Renzi has been the victim of several
ofensive online comments. In particular, in
2016, he was targeted by some users who
criticized his way of managing the economic
crisis in Italy.
Curie
Il commento è stato dato da una persona
molto influente, che è il presidente del
partito, che ha dato una risposta che l’ha
definita “inaccettabile”. Il partito non è
una scuola di formazione, ma una scuola di
vita, e quindi la risposta deve essere meno
formale, e più educata. Perché questa
persona è il presidente del partito? C’è una
persona che rappresenta il partito, che è
una persona che rappresenta la linea The
comment was given by a very influential
person, who is the president of the party, who
gave a response that called it “unacceptable”.</p>
        <p>The party is not a training school, but a
school of life, and therefore the answer must
be less formal, and more educated. Why is
this person the party president? There is a
person who represents the party, who is a
person who represents the political line</p>
        <sec id="sec-1-11-1">
          <title>3.4. Approach to HaSpeeDe3 Shared Task</title>
        </sec>
      </sec>
      <sec id="sec-1-12">
        <title>Subtask A - textual. For hate speech detection in the political domain, we submit our System 2 (baseline enhanced with Davinci) and the majority vote (System</title>
        <p>
          Subtask B - XReligiousHate. Similarly, we compare
System 2 and System 3. Since subtask B explores
crossdomain hate speech detection by focusing on the religious
domain, we further augment the training dataset with
CONAN [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], which is a multilingual expert-based hate
speech/counter-narrative pairs dataset on Islamophobia.
More precisely, we use only Italian hate speech content
for augmentation.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Experimental Setup and Results</title>
      <p>
        Setting For the baseline, we fine-tune AlBERTo to our
downstream tasks. We perform a minimum parameter
selection tuning on the validation set (10% of the training
set). We selected the highest performing learning rate
∈ [1e−5, 2e−5, 1e−2]; batch size ∈ [
        <xref ref-type="bibr" rid="ref4 ref8">4, 8, 16, 32</xref>
        ]; epochs in
range [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8">1 − 10</xref>
        ]. The best configuration for both models
is: lr = 1e−5, batch size = 16, epochs = 4. In order to
tune the network, we used the AdamW optimizer. As
for the pre-processing, we used the pretrained AlBERTo
tokenizer for text tokenization, and then we encoded
the data. We set the maximum length to 256 characters
for the baseline, and 512 for instanced enriched with
verbalisation or prompting.
      </p>
      <p>For System 2, we use text-davinci-002 with
temperature=0.5. To avoid too long answers, particularly for
multiple named entities, we set the max number of returned
sentences to n=5 with max_tokens=150.</p>
      <p>Task A Task B
Setting dev test test
AlBERTo 0.91 0.91 0.51
Davinci (run 2) 0.92 0.89 0.48
Verbalisation 0.91 0.91 0.51
Davinci + Verbalisation 0.90 0.89 0.47</p>
      <p>Majority Vote (run 1) - 0.91 0.52
set, leading our baseline to be our best submitted
system. In both the development and test set, enhancing
the baseline with the verbalisation cause a drop of 0.01
point with respect to the baseline. Majority voting does
not prove to be efective in this case. For HaSpeeDe3,
we obtain the best result by enhancing our baseline with
Davinci, gaining 0.01 point. Combining both the
verbalisation and Davinci leads to a drop of 0.01 point instead,
while adding only the verbalisation does not afect the
ifnal score. However, in the test set we obtain the best
result with majority vote, gaining two units over the
best-performing model, i.e., Davinci.</p>
      <sec id="sec-2-1">
        <title>Error Analysis We performed an analysis of the clas</title>
        <p>sification errors over our submitted systems. First, with
the help of NLTK, we retrieved the most frequent words
Results Tables 4 and 5 show the results on the devel- in misclassified instances. In both tasks, enhancing the
opment and test set of the HODI and HaSpeeDe3 shared models with a knowledge-base, either through Davinci
tasks. While for HODI we submitted all models, for or the verbalisation, results in less misclassified instances
HaSpeeDe3 only the models that performed the best in containing name entities, proving the eficacy of our
apthe development set were subsequently submitted. For proach, at the expenses of other instances that do not
conHaSpeeDe3 - Task B we could not produce results on tain name entities hence are not enhanced, which do not
the development set because the religious data were only get identified correctly. For what concerns HaSpeeDe3,
released during the testing phase. most false positives comes from sentences showing a
negative sentiment towards racist statements, like in the
AlBESReTttoi n(rgun 1) D0.e83v 0T.e7s9t5 following example: Se la SeaWatch fosse piena di gattini
Davinci (run 2) 0.87 0.792 sareste già partiti con i pedalò per salvarli. Mi disgustate.9
Verbalisation (run 3) 0.82 0.780 As for false negatives, we observed a pattern of
misclasMajority Vote (not submitted) - 0.80 sification when the target of the statement is implicit,
baseline - 0.669 like in the following example: Quelli che si lamentano
top 1 - 0.810 della puzza di urina per le strade di Roma sono gli stessi
che dicono #Fateliscendere o #portiaperti ! Ma secondo loro
chi c*zzo è che piscia per strada a tutte le ore, che vagano
ubriachi o senza meta? Gli alieni?.10 In this case, there is
an implicit stereotyped racist statement, i.e., those who
pee on the streets are all immigrants.</p>
        <p>In both tasks, our models struggle the most with the
identification of the positive class, which is the least
represented in the training data. For HODI, enhancing the
baseline with the text generated from Davinci helps us
gain 0.04 points with respect to the baseline in the
development set, but we observe a slight drop in the test
9If SeaWatch was full of kittens you would have already left
with paddleboats to rescue them. You disgust me.</p>
        <p>10Those who complain about the stench of urine on the streets
of Rome are the same who say #Letthemin or #opentheports ! But
according to them who the f*ck is pissing in the streets at all hours,
wandering drunk or aimlessly? The aliens?</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Related Work</title>
      <p>in Safe and Trusted Artificial Intelligence
(www.safeandtrustedai.org). CDB would like to thank her
superviRecent work has shown the eficiency of knowledge- sors, Albert Meroño-Peñuela and Barbara McGillivray,
enhancement of NLP models in many downstream tasks, for their helpful comments and mentorship.
such as sentiment classification [ 20], word sense
disambiguation [21], and semantic change detection [22].</p>
      <p>
        Sharifirad et al. [ 23] is one of the first attempts to References
leverage external world knowledge for abusive language
detection, improving performance on sexist tweet
classiifcation. They use ConceptNet [
        <xref ref-type="bibr" rid="ref10 ref9">24, 25</xref>
        ] and Wikidata [17]
to augment the original data by concatenating additional
information about concepts and their descriptions.
Similarly, Lin [
        <xref ref-type="bibr" rid="ref11">26</xref>
        ] uses an entity linking approach to link
entities mentioned in tweets to their Wikipedia descriptions
in order to leverage world knowledge for hate speech
detection. The injection of external world knowledge is a
promising avenue for explicit hate speech detection,
leading to an improvement of 10% for precision, recall and
F1-score. We build upon these works, and leverage both
external and internal knowledge about abusive language
to explore their impact on diferent hate domains. To the
best of our knowledge, we are the first to apply internal
knowledge of large language models through
promptengineering to the hate speech detection task, while most
of the works focus on natural language understanding,
question answering or text completion [
        <xref ref-type="bibr" rid="ref12">27</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>6. Conclusions</title>
      <p>We present a knowledge-enhanced classification
solution for identifying homotransphobic and hate speech
content in Italian online posts. Our first system uses
external knowledge injection via O-Dang triple
verbalisation to enhance the AlBERTo model, whereas our
second system exploits Davinci’s internal knowledge about
abusive language to enhance AlBERTo model. Lastly, a
majority vote among AlBERTo baseline, System 1 and
System 2 is adopted to improve classification particularly
on uncertain predictions. We evaluate our approach in
the Homotransphobia Detection in Italian (HODI) and
Hate Speech Detection (HaSpeeDe3) Shared Tasks of
the EVALITA 2023 campaign. Our results show that
knowledge-enhancement can improve the classification,
especially of sentences containing name entities in the
political hate domain. The resulting approach can be
expanded on multiple knowledge sources, knowledge
injection methods and tasks.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <sec id="sec-5-1">
        <title>The work of Chiara Di Bonaventura was supported</title>
        <p>by UK Research and Innovation [grant number
EP/S023356/1], in the UKRI Centre for Doctoral Training
to fight online hate speech, in: Proceedings of 7abed946e06f76b3825ae5e294ffac14.
the 57th Annual Meeting of the Association for [15] V. Basile, M. Lai, M. Sanguinetti, Long-term social
Computational Linguistics, Association for Com- media data collection at the university of turin, in:
putational Linguistics, Florence, Italy, 2019, pp. Fifth Italian Conference on Computational
Linguis2819–2829. URL: https://aclanthology.org/P19-1271. tics (CLiC-it 2018), Turin, Italy, 2018.
doi:10.18653/v1/P19- 1271. [16] M. A. Stranisci, S. Frenda, M. Lai, O. Araque, A. T.
[9] D. Nozza, A. T. Cignarella, G. Damo, T. Caselli, Cignarella, V. Basile, C. Bosco, V. Patti, O-dang! the
V. Patti, HODI at EVALITA 2023: Overview of ontology of dangerous speech messages, in:
Prothe Homotransphobia Detection in Italian Task, in: ceedings of the 2nd Workshop on Sentiment
AnalyProceedings of the Eighth Evaluation Campaign of sis and Linguistic Linked Data, European Language
Natural Language Processing and Speech Tools for Resources Association, Marseille, France, 2022, pp.
Italian. Final Workshop (EVALITA 2023), CEUR.org, 2–8. URL: https://aclanthology.org/2022.salld-1.2.</p>
        <p>Parma, Italy, 2023. [17] D. Vrandečić, M. Krötzsch, Wikidata: A free
col[10] M. Lai, S. Menini, M. Polignano, V. Russo, R. Sprug- laborative knowledgebase, Commun. ACM 57
noli, G. Venturi, Evalita 2023: Overview of the 8th (2014) 78–85. URL: https://doi.org/10.1145/2629489.
evaluation campaign of natural language process- doi:10.1145/2629489.
ing and speech tools for italian, in: Proceedings [18] J. Ratclif, D. Metzener, Ratclif-obershelp pattern
of the Eighth Evaluation Campaign of Natural Lan- recognition, Dictionary of Algorithms and Data
guage Processing and Speech Tools for Italian. Final Structures (1998).</p>
        <p>Workshop (EVALITA 2023), CEUR.org, Parma, Italy, [19] D. Trotta, R. Guarasci, E. Leonardelli, S. Tonelli,
2023. Monolingual and cross-lingual acceptability
judg[11] M. Lai, F. Celli, A. Ramponi, S. Tonelli, C. Bosco, ments with the Italian CoLA corpus, in:
FindV. Patti, Haspeede3 at evalita 2023: Overview of the ings of the Association for Computational
Linpolitical and religious hate speech detection task, in: guistics: EMNLP 2021, Association for
ComputaProceedings of the Eighth Evaluation Campaign of tional Linguistics, Punta Cana, Dominican Republic,
Natural Language Processing and Speech Tools for 2021, pp. 2929–2940. URL: https://aclanthology.org/
Italian. Final Workshop (EVALITA 2023), CEUR.org, 2021.findings-emnlp.250. doi:10.18653/v1/2021.</p>
        <p>Parma, Italy, 2023. findings- emnlp.250.
[12] B. Vidgen, D. Nguyen, H. Margetts, P. Rossini, [20] P. Ke, H. Ji, S. Liu, X. Zhu, M. Huang,
SentiR. Tromble, Introducing CAD: the contextual LARE: Sentiment-aware language representation
abuse dataset, in: Proceedings of the 2021 Con- learning with linguistic knowledge, in:
Proceedference of the North American Chapter of the As- ings of the 2020 Conference on Empirical
Methsociation for Computational Linguistics: Human ods in Natural Language Processing (EMNLP),
AsLanguage Technologies, Association for Compu- sociation for Computational Linguistics, Online,
tational Linguistics, Online, 2021, pp. 2289–2303. 2020, pp. 6975–6988. URL: https://aclanthology.
URL: https://aclanthology.org/2021.naacl-main.182. org/2020.emnlp-main.567. doi:10.18653/v1/2020.
doi:10.18653/v1/2021.naacl- main.182. emnlp- main.567.
[13] D. Nozza, Exposing the limits of zero-shot cross- [21] J. Zhou, Z. Zhang, H. Zhao, S. Zhang, LIMIT-BERT :
lingual hate speech detection, in: Proceedings Linguistics informed multi-task BERT, in: Findings
of the 59th Annual Meeting of the Association of the Association for Computational Linguistics:
for Computational Linguistics and the 11th Inter- EMNLP 2020, Association for Computational
Linnational Joint Conference on Natural Language guistics, Online, 2020, pp. 4450–4461. URL: https://
Processing (Volume 2: Short Papers), Associa- aclanthology.org/2020.findings-emnlp.399. doi:10.
tion for Computational Linguistics, Online, 2021, 18653/v1/2020.findings- emnlp.399.
pp. 907–914. URL: https://aclanthology.org/2021. [22] B. McGillivray, M. Alahapperuma, J. Cook,
acl-short.114. doi:10.18653/v1/2021.acl- short. C. Di Bonaventura, A. Meroño-Peñuela, G. Tyson,
114. S. Wilson, Leveraging time-dependent lexical
fea[14] M. Polignano, P. Basile, M. de Gemmis, G. Semer- tures for ofensive language detection, in:
Proceedaro, V. Basile, AlBERTo: Italian BERT Language ings of the The First Workshop on Ever Evolving
Understanding Model for NLP Challenging Tasks NLP (EvoNLP), Association for Computational
LinBased on Tweets, in: Proceedings of the Sixth guistics, Abu Dhabi, United Arab Emirates (Hybrid),
Italian Conference on Computational Linguistics 2022, pp. 39–54. URL: https://aclanthology.org/2022.
(CLiC-it 2019), volume 2481, CEUR, 2019. URL: evonlp-1.7.
https://www.scopus.com/inward/record.uri? [23] S. Sharifirad, B. Jafarpour, S. Matwin, Boosting
eid=2-s2.0-85074851349&amp;partnerID=40&amp;md5= text classification performance on sexist tweets</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Nozza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hovy</surname>
          </string-name>
          ,
          <article-title>The state of profanity obfuscation in natural language processing scientific publications</article-title>
          ,
          <source>in: Findings of the Association for Computational Linguistics: ACL</source>
          <year>2023</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ponnusamy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Kumaresan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sampath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Thenmozhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Thangasamy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Nallathambi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Dataset for identification of homophobia and transophobia in multilingual youtube comments</article-title>
          ,
          <source>arXiv preprint arXiv:2109.00227</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E.</given-names>
            <surname>Fersini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nozza</surname>
          </string-name>
          , G. Boifava,
          <article-title>Profiling Italian misogynist: An empirical study</article-title>
          ,
          <source>in: Proceedings of the Workshop on Resources</source>
          and
          <article-title>Techniques for User and Author Profiling in Abusive Language, European Language Resources Association (ELRA), Marseille</article-title>
          , France,
          <year>2020</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>13</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .restup-
          <volume>1</volume>
          .3.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E. W.</given-names>
            <surname>Pamungkas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <article-title>Misogyny detection in twitter: a multilingual and cross-domain study</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>57</volume>
          (
          <year>2020</year>
          )
          <fpage>102360</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H. R.</given-names>
            <surname>Kirk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Vidgen</surname>
          </string-name>
          , P. Röttger, Semeval2023 task 10:
          <article-title>Explainable detection of online sexism</article-title>
          ,
          <source>in: Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval2023)</source>
          ,
          <source>Association for Computational Linguistics</source>
          ,
          <year>2023</year>
          . URL: http://arxiv.org/abs/2303.04222. doi:
          <volume>10</volume>
          . 48550/arXiv.2303.04222.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Locatelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Damo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nozza</surname>
          </string-name>
          ,
          <article-title>A cross-lingual study of homotransphobia on Twitter</article-title>
          ,
          <source>in: Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Dubrovnik, Croatia,
          <year>2023</year>
          , pp.
          <fpage>16</fpage>
          -
          <lpage>24</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .c3nlp-
          <fpage>1</fpage>
          . 3.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Durairaj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Buitaleer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Kumaresan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ponnusamy</surname>
          </string-name>
          ,
          <article-title>Findings of the shared task on Homophobia Transphobia Detection in Social Media Comments</article-title>
          ,
          <source>in: Proceedings of the Second Workshop on Language Technology for Equality, Diversity and Inclusion</source>
          , Association for Computational Linguistics,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.-L.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kuzmenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Tekiroglu</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Guerini, CONAN - COunter NArratives through nichesourcing: a multilingual dataset of responses by text augmentation and text generation using a combination of knowledge graphs</article-title>
          ,
          <source>in: Proceedings of the 2nd Workshop on Abusive Language Online (ALW2)</source>
          ,
          <source>Association for Computational Linguistics</source>
          , Brussels, Belgium,
          <year>2018</year>
          , pp.
          <fpage>107</fpage>
          -
          <lpage>114</lpage>
          . URL: https://aclanthology.org/W18-5114. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W18</fpage>
          -5114.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <article-title>Conceptnet-a practical commonsense reasoning tool-kit</article-title>
          ,
          <source>BT technology journal 22</source>
          (
          <year>2004</year>
          )
          <fpage>211</fpage>
          -
          <lpage>226</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>R.</given-names>
            <surname>Speer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Havasi</surname>
          </string-name>
          ,
          <article-title>Conceptnet 5.5: An open multilingual graph of general knowledge</article-title>
          ,
          <source>in: Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence</source>
          , AAAI'
          <fpage>17</fpage>
          , AAAI Press,
          <year>2017</year>
          , p.
          <fpage>4444</fpage>
          -
          <lpage>4451</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Leveraging world knowledge in implicit hate speech detection</article-title>
          ,
          <source>in: Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Abu Dhabi,
          <source>United Arab Emirates (Hybrid)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>31</fpage>
          -
          <lpage>39</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .nlp4pi-
          <fpage>1</fpage>
          .4.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>R.</given-names>
            <surname>Brate</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-H. Dang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Hoppe</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>He</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>MeroñoPeñuela</surname>
          </string-name>
          , V. Sadashivaiah,
          <article-title>Improving language model predictions via prompts enriched with knowledge graphs</article-title>
          ,
          <source>in: Workshop on Deep Learning for Knowledge Graphs (DL4KG@ ISWC2022)</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>