=Paper=
{{Paper
|id=Vol-3878/118_calamita_long
|storemode=property
|title=PERSEID - Perspectivist Irony Detection: A CALAMITA Challenge
|pdfUrl=https://ceur-ws.org/Vol-3878/118_calamita_long.pdf
|volume=Vol-3878
|authors=Valerio Basile,Silvia Casola,Simona Frenda,Soda Marem Lo
|dblpUrl=https://dblp.org/rec/conf/clic-it/BasileCFL24
}}
==PERSEID - Perspectivist Irony Detection: A CALAMITA Challenge==
PERSEID - Perspectivist Irony Detection:
A CALAMITA Challenge
Valerio Basile1 , Silvia Casola2 , Simona Frenda3,4 and Soda Marem Lo1
1
University of Turin, Italy
2
MaiNLP & MCML, LMU Munich, Germany
3
Interaction Lab, Heriot-Watt University, Edinburgh, Scotland
4
aequa-tech, Turin, Italy
Abstract
Works in perspectivism and human label variation have emphasized the need to collect and leverage various voices and points
of view in the whole Natural Language Processing pipeline.
PERSEID places itself in this line of work. We consider the task of irony detection from short social media conversations in
Italian collected from Twitter (X) and Reddit. To do so, we leverage data from MultiPICO, a recent multilingual dataset with
disaggregated annotations and annotators’ metadata, containing 1000 Post, Reply pairs with five annotations each on average.
We aim to evaluate whether prompting LLMs with additional annotators’ demographic information (namely gender only, age
only, and the combination of the two) results in improved performance compared to a baseline in which only the input text is
provided.
The evaluation is zero-shot; and we evaluate the results on the disaggregated annotations using f1.
Keywords
Perspectivism, Irony Detection, Evaluation
1. Challenge: Introduction and intrinsically subjective [10], as points of view might dif-
fer depending on users’ social background, beliefs, and
Motivation demographics. Using a single aggregated label has thus
Recently, researchers have shown a growing interest in been increasingly questioned [11, 12, 13], and preserv-
human-centered technologies to make Artificial Intelli- ing disaggregated data is preferred. On the other hand,
gence (AI) models and products more attentive to the recent work has shown that design choices and biases
users’ sensitivity and needs. affect datasets and models and often result in models
In Natural Language Processing (NLP), works on per- unexpectedly aligned with a given population segment
spectivism [1] and human label variation [2] have em- more than with another [14]; in fact, aggregated data tend
phasized the intrinsic variability in human annotation to reflect a minority of perspectives, under-representing
and thus the importance of incorporating a diverse set of others [15, 4].
voices; this aspect affects all phases of the NLP pipeline, As a result, disaggregated datasets have become more
including collecting disaggregated datasets [3, 4, 5], an- popular, as listed in the Perspectivist Data Manifesto1
alyzing existing disagreement [6], learning from disag- and by Plank [2]2 .
gregated data [7, 8], and evaluating considering several Researchers are incresingly reporting annotators’ de-
voices as valid [9, 1]. mographics and other metadata when describing the
During the data collection and annotation phase, dataset, which was first advised as a good practice to
works in this area have gone beyond considering dis- avoid excluding, minimizing, and misrepresenting cer-
agreement as motivated by noise only and thus as an tain groups of users [16]. Recent work has also explored
attribute to be minimized and resolved, e.g., through whether annotators’ demographics and background — as
majority voting. In contrast, research has emphasized described by available metadata — influence their anno-
the necessity of collecting a variety of voices and con- tation [5, 17, 18, 19, 4] and can help during the modeling
sidering all such voices as valid. The reason is twofold. of the phenomenon under study [20, 8, 21].
On the one hand, researchers have argued that many Despite the increasing interest in disaggregated and
tasks that are popular in the NLP community (includ- metadata-rich datasets, few such datasets for irony de-
ing, for example, hate speech and humor detection) are tection exist. Simpson et al. [22] released a corpus for
humor detection in English, used as a benchmark in the
CLiC-it 2024: Tenth Italian Conference on Computational Linguistics, first edition of the Learning With Disagreement (LeWiDi)
Dec 04 - 06, 2024, Pisa, Italy shared task [23]. No annotators’ metadata, however, are
Envelope-Open valerio.basile@unito.it (V. Basile); s.casola@lmu.de (S. Casola);
s.frenda@hw.ac.uk (S. Frenda); sodamarem.lo@unito.it (S. M. Lo) 1
https://pdai.info/
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License 2
Attribution 4.0 International (CC BY 4.0). www.github.com/mainlp/awesome-human-label-variation
CEUR
ceur-ws.org
Workshop ISSN 1613-0073
Proceedings
included. Frenda et al. [4] proposed a dataset for irony • Gender (Task 2): the perspective is the binary
detection and investigated the influence of the annota- self-identified gender of the annotator.
tors’ demographics on their perception [6]. The dataset • Age + Gender (Task 3): in this case, both at-
contains English texts only. tributes are provided as the perspective.
For this challenge at CALAMITA [24], we propose to The post is a textual post, to which the target reply
use the Italian portion of MultiPICo (Multilingual Per- is a reply. The output of the prediction is a binary label
spectivist Irony Corpus)3 [25]. Multipico is a multilingual indicating whether the reply is ironic (or non-ironic) for
corpus of short Post-Reply conversational pairs extracted a human bearing the characteristic of the perspective
from Twitter and Reddit and annotated as ironic or not to the text. The performance of the model is evaluated
ironic by crowdsourcing workers with different demo- through a global f1 metric on the disaggregated annota-
graphics and backgrounds. MultiPICo covers 9 languages tions.
(Arabic, English, Dutch, French, German, Hindi, Italian, The challenge is zero-shot: no training, fine-tuning,
Portuguese, and Spanish) and 25 language varieties4 , or in-context learning is considered for this version of
ranging from high- to low-resourced ones. Moreover, PERSEID and the whole dataset can be used for inference.
a rich set of annotators’ sociodemographic information Note that since each annotator can be described by no
(balanced gender, age, nationality, ethnicity, student, and traits (Task 0), one single trait (Task 1 and Task 2), and
employment status) is provided. two traits (Task 3), we do not aim at optimal performance
While no perspectivist task leveraging the dataset has when considering personalized irony detection; instead,
been proposed so far, PERSEID is related to the Learn- our goal is to understand whether models improve their
ing With Disagreement task held at SemEval 2021 [11] performance when one or multiple traits is provided and
and 2023 [13]. In LeWiDi, participant systems were chal- to understand the impact of different configurations.
lenged to learn the distribution of labels, tested by cross
entropy-based metrics. In contrast, PERSEID aims at
stimulating the development of models of human per- 3. Data description
spectives, in order to explain the label distributions rather
than just quantifying them. 3.1. Origin of data
The data for the challenge are part of MultiPICo [25],
2. Challenge: Description a corpus of 18, 778 short conversations collected from
Reddit (8, 956) and Twitter (9, 822) in 9 languages, and a
The task of Perspectivist Irony Detection aims to measure total of 25 varieties.
models’ capability to detect irony in a short verbal ex- Data were collected to reproduce the structure of short
change for each annotator, conditioned on the knowledge conversations.
of demographic information about them. To this purpose, For both Reddit and Twitter, the post is typically a
we want to look at different model performances if it is message initiating a thread and the reply a direct reply
informed by one demographic trait or a combination of to that message5 .
two. In particular, we focus on the gender and age of the Reddit data were retrieved using the Pushshift reposi-
annotator, due to the balanced number of male and fe- tory6 from January 2020 to June 2021. For Italian, data
male annotators by design 3.2, and due to the fact that age were downloaded from the subreddit /r/Italy.
was shown to be one of the most polarized dimensions Pairs having at least one deleted or removed comment
in [25]. were filtered out, and the language of the messages was
The input to the task does not consist only of a text, further validated using the Python library for language
but rather of a tuple . identification LangID7 .
In this iteration of PERSEID, we considered several Twitter data were collected via Twitter Stream API,
variables for the perspective attribute: using the geolocation service and excluding quotes and
retweets. Then, the full conversation was retrieved, and
• None (Task 0): acting as a baseline, we want to tweets that directly replied to the starting ones were
investigate the models’ outputs when no infor- retained.
mation about the annotator is provided. The data collection resulted in 18, 778 instances, to-
• Age (Task 1): the perspective is one of four val- gether with their metadata, consisting of Post-Reply orig-
ues encoding the age group of the annotator. inal IDs, subreddits, and geolocation information.
3 5
MultiPICo is available at https://huggingface.co/datasets/ For Reddit, second-level replies were collected in a minority of
Multilingual-Perspectivist-NLU/MultiPICo with a CC-BY 4.0 cases; for Twitter, the post is a reply to a thread-starting message
license. in a minority of cases.
4 6
For example, texts in Austrian, German, and Swiss German are https://redditsearch.io/
7
included in the dataset. https://github.com/saffsd/langid.py
Language #Annotators #Annotations Label rate #Texts Sources Annotation mean
%not %iro #Reddit #Twitter
Arabic 68 10,609 68 32 2,181 949 1,232 4.86
Dutch 25 4,991 73 27 1,000 500 500 4.99
English 74 14,171 69 31 2,999 1,499 1,500 4.73
French 50 8,770 70 30 1,760 1,000 760 4.98
German 70 12,510 68 32 2,375 1,042 1,333 5.27
Hindi 24 4,711 65 35 786 286 500 5.99
Italian 24 4,790 69 31 1,000 500 500 4.79
Portuguese 49 9,754 62 38 1,994 997 997 4.89
Spanish 122 24,036 67 33 4,683 2,183 2,500 5.13
Total 506 94,342 68 32 18,778 8,956 9,822 5.02
Table 1
Number of annotators, annotations, texts per source, and annotation means for each language. For Italian, 1000 pairs were
collected, each annotated by 4.79 annotators. Note the label unbalance, with the negative class accounting for 69% of the total
annotations.
Figure 1: Screenshot of the annotation interface for an English instance of MultiPICo. The Italian interface was similar, with translated
question and options.
For Italian, data account for 1000 post, reply pairs, Annotators were selected based on three criteria:
equally sourced from Reddit and Twitter.
• Their completion rate had to be greater or equal
to 99%
3.2. Annotation details • They had to be native speakers of the considered
Annotators were asked to read a set of post and reply language (i.e., Italian, for the portion of data used
pairs and answer whether the text of the reply was ironic in the challenges)
or not, given the context. • The set of annotators needed to be balanced
The human annotation of the collected data was per- across genders.
formed on the crowdsourcing platform Prolific8 , through
a custom-built annotation interface designed to collect The quality of the annotation was further assured us-
a diverse and balanced set of annotators. The interface ing attention check questions in the form of “Please an-
mimicked a message conversation, having the post as swer X to this question”. Annotators had 1% probability of
context and asking whether the reply was Ironic or Not receiving these special questions. Annotators who failed
ironic. to respond correctly to at least 50% of these questions
For Italian, 24 native-speaker annotators were hired, were excluded from the final corpus.
who performed 4,790 annotations in total, resulting in a A rich set of metadata is also provided. These include
mean of 4,79 annotations per instance (see Table 1). the self-identified Gender (balanced by design), their na-
tionality, their Age Group (1 GenX, 15 GenY, 8 GenZ, for
8
Italian), Ethnicity (23 white people, 1 mixed person, for
https://www.prolific.com/
Demographics Languages
English Spanish Italian French Dutch German Hindi Arabic Portoguese
Boomer 3 2 – 2 – 5 – 1 –
GenX 22 17 1 7 4 7 3 4 1
Age group
GenY 38 66 15 23 10 36 13 36 23
GenZ 10 37 8 17 11 20 8 26 25
White 47 60 23 40 22 66 – 20 37
Mixed 1 31 1 3 2 3 – 13 10
Ethnicity Asian 18 1 – 1 1 – 22 1 –
Black 3 2 – 5 – – – 2 1
Other 3 27 – 1 – 1 8 31 1
Yes 13 39 14 16 7 14 8 29 30
Student
No 46 60 9 30 16 39 14 25 16
Full-time 25 41 9 24 10 24 10 20 15
Unemployed 11 24 7 5 4 3 1 11 8
Part-time 11 17 5 5 3 10 4 13 6
Employment
Not in paid work 4 4 1 5 4 5 – 1 –
Due to start – 3 1 1 – 2 2 – 2
Other 1 6 – 6 – 3 1 5 14
Table 2
Sociodemographic information about annotators per language.
Italian), Student status (14 yes, 9 no, for Italian), Employ- 'reply_id': 2497527360959166890,
ment status (9 in full-time jobs, 7 unemployed, 5 working 'source': 'twitter',
part-time, 1 not in paid work and 1 due to start, for Ital- 'timestamp': '2022-12-07 15:49:50'
ian), as reported in Table 2.
3.3. Data format 3.4. Example of prompts used for
The dataset is in tabular format, one row per annotation. zero-shot prediction
The data contain the text in the form of two fields (post The challenge is zero-shot, and the prompt depends on
and reply), the binary label, and a series of metadata three variables: perspective, post, and reply.
about the post, reply, and annotator. Here is an example
of instance from the Italian section of MultiPICo: Sei {perspective}.
Istruzione: Ti vengono fornite in
'Age': 29.0, input (Input) una coppia di frasi
'Country of birth': 'Italy', (Post, Reply) estratte da conversazioni
'Country of residence': 'Italy', sui social media. Il tuo compito è
'Employed': 'Yes', determinare se la Risposta (Reply) è
'Employment status': 'Part-Time', ironica nel contesto del Post (Post).
'Ethnicity simplified': 'White', Fornisci in output (Output) una singola
'Gender': 'Male', etichetta “ironia" o "non ironia".
'Generation': 'GenY', Input:
'GenerationAggregated': 'Young', Post: {post}
'Nationality': 'Italy', Reply: {reply}
'Student status': 'No', Output:
'annotator_id': 9208155880570654046,
'label': 0, Task 0 No perspective is provided, and the prompt
'language': 'it', directly starts with the instruction.
'language_variety': 'it',
'level': 1.0, Task 1 The perspective variable is a verbalization of
'post': 'Ormai il quadro è chiaro: cercare di the Generation, which is expressed as an integer
coinvolgere tutti per non farla pagare a in the dataset. It can be instantiated with the
nessuno. Se non riuscissero a corrompere i following values9 :
Pm di Torino andranno in B diretti.', 9
No workers whose age is > 42, i.e., from the baby boomer gener-
'post_id': 14071953227682835778, ations, participated in the annotation of the Italian portion of the
'reply': '@USER Magari ??', dataset
• “una persona giovane della generazione Z” In the vast majority (∼90%) of cases, the
if Generation == GenZ (Age < 26) conversation-starting messages and their direct
• “una persona giovane della generazione Y” replies were downloaded to capture the full con-
if Generation == GenY (26 ≤ Age < 42) versational context. In a few cases, the down-
• “una persona adulta della generazione X” loaded reply was not direct but rather a second-
if Generation == GenX (42 ≤ Age < 58) level reply (a reply to a direct reply); thus, some
• “una persona adulta della generazione baby conversational context might be missing.
boomer” Challenge design We describe annotators by no so-
if Generation == Boomer (Age > 58) ciodemographic traits (Task 0), one single demo-
Task 2 The perspective variable is a verbalization of graphic trait (Task 1 and Task 2), or two demo-
the Gender variable, which is expressed as a graphic traits (Task 3). We evaluate disaggregated
string in English. It can be instantiated with one annotations at inference time, having the annota-
of two values: tors represented only by those traits. Annotators’
sociodemographic information does not always
• “una donna” align with the most relevant grouping of anno-
if Gender == “Female” tators according to the language phenomenon
• “un uomo” under study [21, 28], and the limited amount of
if Gender == “Male” sociodemographic traits we provide is undoubt-
Task 3 The perspective variable is a verbalization of edly not enough to describe every single anno-
both the Age and Gender variables, e.g., “una tator. We are aware of this limitation. In fact,
giovane donna della generazione Z.” our main aim is to understand whether providing
one or more annotator traits makes the model
predictions more aligned with annotators having
4. Metrics a given characteristic.
Inspired by Mokhberian et al. [26], the Perspectivist Irony
Detection task is evaluated by means of global F1, that 6. Ethical issues
is, the F1-score computed across all the individual an-
notations in the dataset against the predictions of the This work places itself in an increasing amount of work
model. that calls to consider and include the subjectivity of
the annotators in NLP applications, encouraging reflec-
tion on the different perspectives encoded in annotated
5. Limitations datasets to minimize the amplification of biases. We hope
this challenge will be a starting point for investigating
Data The sociodemographic information about the an- and evaluating LLMs in Italian to make them suitable for
notators is partial, bound to what was available final users.
from the crowdsourcing platform, and following a The dataset used in the challenge was built by adopt-
discretization of human personal traits that could ing measures to protect the privacy of annotators, and
be perceived as forced (e.g., representing self- the data handling protocols were designed to safeguard
identified gender as a single binary label). Fur- personal information (like anonymization of users’ men-
thermore, as shown by Orlikowski et al. [21], an- tions). Although the attention during the collection of
notators’ sociodemographics do not always align data was focused on ironic content spread online, we
with the most relevant grouping of annotators acknowledge that some of the material contains racist,
according to the language phenomenon under sexist, stereotypical, violent, or generally disturbing con-
study. tent.
Annotators of the Italian portion of MultiPICO Annotators are balanced through their self-identified
tend to be young (with no annotators from the gender. However, we are aware that considering gen-
baby boomer generation and only one from der in a binary form is limited; moreover, a substantial
GenX). This aspect might influence the results. unbalance for some dimensions, like the self-identified
ethnicities, is present in the dataset. This pattern sug-
Similarly to Sachdeva et al. [5], Sap et al. [19],
gests the need to interact differently with annotators or
Forbes et al. [27], we noticed the ethnicity of an-
social communities if we want a diversity of annotators
notators is unbalanced, and all but one annotators
and perspectives in terms of social background.
are white for the considered data.
7. Data license and copyright tions in irony detection, in: ECAI 2023 Workshop
on Perspectivist Approaches to NLP, 2023.
issues [7] A. Mostafazadeh Davani, M. Díaz, V. Prabhakaran,
MultiPICo is distributed under the Creative Commons Dealing with disagreements: Looking beyond the
Attribution 4.0 (CC-BY-4.0) license. majority vote in subjective annotations, Transac-
tions of the Association for Computational Linguis-
tics 10 (2022) 92–110. URL: https://aclanthology.
Acknowledgments org/2022.tacl-1.6. doi:10.1162/tacl_a_00449 .
[8] S. Casola, S. Lo, V. Basile, S. Frenda, A. Cignarella,
This work was funded by the ‘Multilingual Perspective- V. Patti, C. Bosco, Confidence-based ensembling of
Aware NLU’ project in partnership with Amazon Alexa. perspective-aware models, in: H. Bouamor, J. Pino,
K. Bali (Eds.), Proceedings of the 2023 Conference
on Empirical Methods in Natural Language Pro-
References cessing, Association for Computational Linguis-
[1] V. Basile, M. Fell, T. Fornaciari, D. Hovy, S. Paun, tics, Singapore, 2023, pp. 3496–3507. URL: https:
B. Plank, M. Poesio, A. Uma, et al., We need to //aclanthology.org/2023.emnlp-main.212. doi:10.
consider disagreement in evaluation, in: Proceed- 18653/v1/2023.emnlp- main.212 .
ings of the 1st workshop on benchmarking: past, [9] A. N. Uma, T. Fornaciari, D. Hovy, S. Paun, B. Plank,
present and future, Association for Computational M. Poesio, Learning from disagreement: A survey,
Linguistics, 2021, pp. 15–21. Journal of Artificial Intelligence Research 72 (2021)
[2] B. Plank, The ”problem” of human label variation: 1385–1470.
On ground truth in data, modeling and evaluation, [10] L. Aroyo, C. Welty, Truth is a lie:
in: Proceedings of the 2022 Conference on Empiri- Crowd truth and the seven myths of hu-
cal Methods in Natural Language Processing, 2022, man annotation, AI Magazine 36 (2015)
pp. 10671–10682. 15–24. URL: https://ojs.aaai.org/aimagazine/
[3] F. Cabitza, A. Campagner, V. Basile, Toward a per- index.php/aimagazine/article/view/2564.
spectivist turn in ground truthing for predictive doi:10.1609/aimag.v36i1.2564 .
computing, in: Proceedings of the AAAI Con- [11] E. Leonardelli, S. Menini, A. P. Aprosio, M. Guerini,
ference on Artificial Intelligence, volume 37, 2023, S. Tonelli, Agreeing to disagree: Annotating offen-
pp. 6860–6868. URL: https://ojs.aaai.org/index.php/ sive language datasets with annotators’ disagree-
AAAI/article/view/25840. ment, in: Proceedings of the 2021 Conference on
[4] S. Frenda, A. Pedrani, V. Basile, S. M. Lo, A. T. Empirical Methods in Natural Language Processing,
Cignarella, R. Panizzon, C. Marco, B. Scarlini, 2021, p. 10528–10539.
V. Patti, C. Bosco, D. Bernardi, EPIC: Multi- [12] A. Uma, T. Fornaciari, A. Dumitrache, T. Miller,
perspective annotation of a corpus of irony, in: J. Chamberlain, B. Plank, E. Simpson, M. Poesio,
A. Rogers, J. Boyd-Graber, N. Okazaki (Eds.), Pro- Semeval-2021 task 12: Learning with disagree-
ceedings of the 61st Annual Meeting of the Associa- ments, in: Proceedings of the 15th International
tion for Computational Linguistics (Volume 1: Long Workshop on Semantic Evaluation (SemEval-2021),
Papers), Association for Computational Linguis- 2021, pp. 338–347.
tics, Toronto, Canada, 2023, pp. 13844–13857. URL: [13] E. Leonardelli, A. Uma, G. Abercrombie, D. Al-
https://aclanthology.org/2023.acl-long.774. doi:10. manea, V. Basile, T. Fornaciari, B. Plank, V. Rieser,
18653/v1/2023.acl- long.774 . M. Poesio, Semeval-2023 task 11: Learning with
[5] P. Sachdeva, R. Barreto, G. Bacon, A. Sahn, C. von disagreements (lewidi), in: Proceedings of the 17th
Vacano, C. Kennedy, The measuring hate speech International Workshop on Semantic Evaluation
corpus: Leveraging rasch measurement theory for (SemEval-2023), 2023, p. 2304–2318.
data perspectivism, in: G. Abercrombie, V. Basile, [14] S. Santy, J. Liang, R. Le Bras, K. Reinecke, M. Sap,
S. Tonelli, V. Rieser, A. Uma (Eds.), Proceedings of NLPositionality: Characterizing design biases
the 1st Workshop on Perspectivist Approaches to of datasets and models, in: Proceedings of
NLP @LREC2022, European Language Resources the 61st Annual Meeting of the Association for
Association, Marseille, France, 2022, pp. 83–94. URL: Computational Linguistics (Volume 1: Long Pa-
https://aclanthology.org/2022.nlperspectives-1.11. pers), Association for Computational Linguistics,
[6] S. Frenda, S. M. Lo, S. Casola, B. Scarlini, C. Marco, Toronto, Canada, 2023, pp. 9080–9102. URL: https://
V. Basile, D. Bernardi, Does anyone see the irony aclanthology.org/2023.acl-long.505. doi:10.18653/
here? Analysis of perspective-aware model predic- v1/2023.acl- long.505 .
[15] V. Prabhakaran, A. M. Davani, M. Diaz, On re-
leasing annotator-level labels and information in Computational Linguistics, Florence, Italy, 2019, pp.
datasets, in: Proceedings of the Joint 15th Linguis- 5716–5728. URL: https://aclanthology.org/P19-1572.
tic Annotation Workshop (LAW) and 3rd Designing doi:10.18653/v1/P19- 1572 .
Meaning Representations (DMR) Workshop, 2021, [23] A. Uma, T. Fornaciari, A. Dumitrache, T. Miller,
p. 133–138. J. Chamberlain, B. Plank, E. Simpson, M. Poesio,
[16] E. M. Bender, B. Friedman, Data statements for nat- SemEval-2021 task 12: Learning with disagree-
ural language processing: Toward mitigating sys- ments, in: A. Palmer, N. Schneider, N. Schluter,
tem bias and enabling better science, Transactions G. Emerson, A. Herbelot, X. Zhu (Eds.), Proceedings
of the Association for Computational Linguistics 6 of the 15th International Workshop on Semantic
(2018) 587–604. Evaluation (SemEval-2021), Association for Com-
[17] D. Almanea, M. Poesio, ArMIS - the Arabic Misog- putational Linguistics, Online, 2021, pp. 338–347.
yny and Sexism Corpus with Annotator Subjective URL: https://aclanthology.org/2021.semeval-1.41.
Disagreements, in: N. Calzolari, F. Béchet, P. Blache, doi:10.18653/v1/2021.semeval- 1.41 .
K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isa- [24] G. Attanasio, P. Basile, F. Borazio, D. Croce, M. Fran-
hara, B. Maegaard, J. Mariani, H. Mazo, J. Odijk, cis, J. Gili, E. Musacchio, M. Nissim, V. Patti, M. Ri-
S. Piperidis (Eds.), Proceedings of the Thirteenth naldi, D. Scalena, CALAMITA: Challenge the Abili-
Language Resources and Evaluation Conference, ties of LAnguage Models in ITAlian, in: Proceed-
European Language Resources Association, Mar- ings of the 10th Italian Conference on Computa-
seille, France, 2022, pp. 2282–2291. URL: https: tional Linguistics (CLiC-it 2024), Pisa, Italy, Decem-
//aclanthology.org/2022.lrec-1.244. ber 4 - December 6, 2024, CEUR Workshop Proceed-
[18] S. Akhtar, V. Basile, V. Patti, Whose opinions mat- ings, CEUR-WS.org, 2024.
ter? Perspective-aware models to identify opinions [25] S. Casola, S. Frenda, S. Lo, E. Sezerer, A. Uva,
of hate speech victims in abusive language detec- V. Basile, C. Bosco, A. Pedrani, C. Rubagotti, V. Patti,
tion, arXiv preprint arXiv:2106.15896 (2021). D. Bernardi, MultiPICo: Multilingual perspectivist
[19] M. Sap, S. Swayamdipta, L. Vianna, X. Zhou, Y. Choi, irony corpus, in: L.-W. Ku, A. Martins, V. Srikumar
N. A. Smith, Annotators with attitudes: How an- (Eds.), Proceedings of the 62nd Annual Meeting
notator beliefs and identities bias toxic language of the Association for Computational Linguistics
detection, in: Proceedings of the 2022 Conference (Volume 1: Long Papers), Association for Compu-
of the North American Chapter of the Association tational Linguistics, Bangkok, Thailand, 2024, pp.
for Computational Linguistics: Human Language 16008–16021. URL: https://aclanthology.org/2024.
Technologies, Association for Computational Lin- acl-long.849.
guistics, Seattle, United States, 2022, pp. 5884–5906. [26] N. Mokhberian, M. Marmarelis, F. Hopp, V. Basile,
URL: https://aclanthology.org/2022.naacl-main.431. F. Morstatter, K. Lerman, Capturing perspectives
doi:10.18653/v1/2022.naacl- main.431 . of crowdsourced annotators in subjective learning
[20] R. Wan, J. Kim, D. Kang, Everyone’s voice mat- tasks, in: K. Duh, H. Gomez, S. Bethard (Eds.),
ters: Quantifying annotation disagreement using Proceedings of the 2024 Conference of the North
demographic information, in: Proceedings of the American Chapter of the Association for Compu-
37th AAAI Conference on Anrtificial Intelligence - tational Linguistics: Human Language Technolo-
AAAI Special Track on AI for Social Impact, 2023. gies (Volume 1: Long Papers), Association for Com-
[21] M. Orlikowski, P. Röttger, P. Cimiano, D. Hovy, The putational Linguistics, Mexico City, Mexico, 2024,
ecological fallacy in annotation: Modeling human pp. 7337–7349. URL: https://aclanthology.org/2024.
label variation goes beyond sociodemographics, in: naacl-long.407.
A. Rogers, J. Boyd-Graber, N. Okazaki (Eds.), Pro- [27] M. Forbes, J. D. Hwang, V. Shwartz, M. Sap, Y. Choi,
ceedings of the 61st Annual Meeting of the Associa- Social chemistry 101: Learning to reason about so-
tion for Computational Linguistics (Volume 2: Short cial and moral norms, in: Proceedings of the 2020
Papers), Association for Computational Linguis- Conference on Empirical Methods in Natural Lan-
tics, Toronto, Canada, 2023, pp. 1017–1029. URL: guage Processing (EMNLP), Association for Com-
https://aclanthology.org/2023.acl-short.88. doi:10. putational Linguistics, Online, 2020, pp. 653–670.
18653/v1/2023.acl- short.88 . URL: https://aclanthology.org/2020.emnlp-main.48.
[22] E. Simpson, E.-L. Do Dinh, T. Miller, I. Gurevych, doi:10.18653/v1/2020.emnlp- main.48 .
Predicting humorousness and metaphor novelty [28] S. M. Lo, V. Basile, Hierarchical clustering of label-
with Gaussian process preference learning, in: based annotator representations for mining per-
A. Korhonen, D. Traum, L. Màrquez (Eds.), Proceed- spectives, in: G. Abercrombie, V. Basile, D. Bernardi,
ings of the 57th Annual Meeting of the Associa- S. Dudy, S. Frenda, L. Havens, E. Leonardelli,
tion for Computational Linguistics, Association for S. Tonelli (Eds.), Proceedings of the 2nd Workshop
on Perspectivist Approaches to NLP co-located with
26th European Conference on Artificial Intelligence
(ECAI 2023), Kraków, Poland, September 30th, 2023,
volume 3494 of CEUR Workshop Proceedings, CEUR-
WS.org, 2023. URL: https://ceur-ws.org/Vol-3494/
paper8.pdf.