=Paper= {{Paper |id=Vol-3878/118_calamita_long |storemode=property |title=PERSEID - Perspectivist Irony Detection: A CALAMITA Challenge |pdfUrl=https://ceur-ws.org/Vol-3878/118_calamita_long.pdf |volume=Vol-3878 |authors=Valerio Basile,Silvia Casola,Simona Frenda,Soda Marem Lo |dblpUrl=https://dblp.org/rec/conf/clic-it/BasileCFL24 }} ==PERSEID - Perspectivist Irony Detection: A CALAMITA Challenge== https://ceur-ws.org/Vol-3878/118_calamita_long.pdf
                                PERSEID - Perspectivist Irony Detection:
                                A CALAMITA Challenge
                                Valerio Basile1 , Silvia Casola2 , Simona Frenda3,4 and Soda Marem Lo1
                                1
                                  University of Turin, Italy
                                2
                                  MaiNLP & MCML, LMU Munich, Germany
                                3
                                  Interaction Lab, Heriot-Watt University, Edinburgh, Scotland
                                4
                                  aequa-tech, Turin, Italy


                                                 Abstract
                                                 Works in perspectivism and human label variation have emphasized the need to collect and leverage various voices and points
                                                 of view in the whole Natural Language Processing pipeline.
                                                 PERSEID places itself in this line of work. We consider the task of irony detection from short social media conversations in
                                                 Italian collected from Twitter (X) and Reddit. To do so, we leverage data from MultiPICO, a recent multilingual dataset with
                                                 disaggregated annotations and annotators’ metadata, containing 1000 Post, Reply pairs with five annotations each on average.
                                                 We aim to evaluate whether prompting LLMs with additional annotators’ demographic information (namely gender only, age
                                                 only, and the combination of the two) results in improved performance compared to a baseline in which only the input text is
                                                 provided.
                                                 The evaluation is zero-shot; and we evaluate the results on the disaggregated annotations using f1.

                                                 Keywords
                                                 Perspectivism, Irony Detection, Evaluation



                                1. Challenge: Introduction and                                                                            intrinsically subjective [10], as points of view might dif-
                                                                                                                                          fer depending on users’ social background, beliefs, and
                                   Motivation                                                                                             demographics. Using a single aggregated label has thus
                                Recently, researchers have shown a growing interest in                                                    been increasingly questioned [11, 12, 13], and preserv-
                                human-centered technologies to make Artificial Intelli-                                                   ing disaggregated data is preferred. On the other hand,
                                gence (AI) models and products more attentive to the                                                      recent work has shown that design choices and biases
                                users’ sensitivity and needs.                                                                             affect datasets and models and often result in models
                                   In Natural Language Processing (NLP), works on per-                                                    unexpectedly aligned with a given population segment
                                spectivism [1] and human label variation [2] have em-                                                     more than with another [14]; in fact, aggregated data tend
                                phasized the intrinsic variability in human annotation                                                    to reflect a minority of perspectives, under-representing
                                and thus the importance of incorporating a diverse set of                                                 others [15, 4].
                                voices; this aspect affects all phases of the NLP pipeline,                                                  As a result, disaggregated datasets have become more
                                including collecting disaggregated datasets [3, 4, 5], an-                                                popular, as listed in the Perspectivist Data Manifesto1
                                alyzing existing disagreement [6], learning from disag-                                                   and by Plank [2]2 .
                                gregated data [7, 8], and evaluating considering several                                                     Researchers are incresingly reporting annotators’ de-
                                voices as valid [9, 1].                                                                                   mographics and other metadata when describing the
                                   During the data collection and annotation phase,                                                       dataset, which was first advised as a good practice to
                                works in this area have gone beyond considering dis-                                                      avoid excluding, minimizing, and misrepresenting cer-
                                agreement as motivated by noise only and thus as an                                                       tain groups of users [16]. Recent work has also explored
                                attribute to be minimized and resolved, e.g., through                                                     whether annotators’ demographics and background — as
                                majority voting. In contrast, research has emphasized                                                     described by available metadata — influence their anno-
                                the necessity of collecting a variety of voices and con-                                                  tation [5, 17, 18, 19, 4] and can help during the modeling
                                sidering all such voices as valid. The reason is twofold.                                                 of the phenomenon under study [20, 8, 21].
                                On the one hand, researchers have argued that many                                                           Despite the increasing interest in disaggregated and
                                tasks that are popular in the NLP community (includ-                                                      metadata-rich datasets, few such datasets for irony de-
                                ing, for example, hate speech and humor detection) are                                                    tection exist. Simpson et al. [22] released a corpus for
                                                                                                                                          humor detection in English, used as a benchmark in the
                                CLiC-it 2024: Tenth Italian Conference on Computational Linguistics,                                      first edition of the Learning With Disagreement (LeWiDi)
                                Dec 04 - 06, 2024, Pisa, Italy                                                                            shared task [23]. No annotators’ metadata, however, are
                                Envelope-Open valerio.basile@unito.it (V. Basile); s.casola@lmu.de (S. Casola);
                                s.frenda@hw.ac.uk (S. Frenda); sodamarem.lo@unito.it (S. M. Lo)                                           1
                                                                                                                                              https://pdai.info/
                                           © 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License   2
                                           Attribution 4.0 International (CC BY 4.0).                                                         www.github.com/mainlp/awesome-human-label-variation




CEUR
                  ceur-ws.org
Workshop      ISSN 1613-0073
Proceedings
included. Frenda et al. [4] proposed a dataset for irony               • Gender (Task 2): the perspective is the binary
detection and investigated the influence of the annota-                  self-identified gender of the annotator.
tors’ demographics on their perception [6]. The dataset                • Age + Gender (Task 3): in this case, both at-
contains English texts only.                                             tributes are provided as the perspective.
   For this challenge at CALAMITA [24], we propose to               The post is a textual post, to which the target reply
use the Italian portion of MultiPICo (Multilingual Per-          is a reply. The output of the prediction is a binary label
spectivist Irony Corpus)3 [25]. Multipico is a multilingual      indicating whether the reply is ironic (or non-ironic) for
corpus of short Post-Reply conversational pairs extracted        a human bearing the characteristic of the perspective
from Twitter and Reddit and annotated as ironic or not           to the text. The performance of the model is evaluated
ironic by crowdsourcing workers with different demo-             through a global f1 metric on the disaggregated annota-
graphics and backgrounds. MultiPICo covers 9 languages           tions.
(Arabic, English, Dutch, French, German, Hindi, Italian,            The challenge is zero-shot: no training, fine-tuning,
Portuguese, and Spanish) and 25 language varieties4 ,            or in-context learning is considered for this version of
ranging from high- to low-resourced ones. Moreover,              PERSEID and the whole dataset can be used for inference.
a rich set of annotators’ sociodemographic information              Note that since each annotator can be described by no
(balanced gender, age, nationality, ethnicity, student, and      traits (Task 0), one single trait (Task 1 and Task 2), and
employment status) is provided.                                  two traits (Task 3), we do not aim at optimal performance
   While no perspectivist task leveraging the dataset has        when considering personalized irony detection; instead,
been proposed so far, PERSEID is related to the Learn-           our goal is to understand whether models improve their
ing With Disagreement task held at SemEval 2021 [11]             performance when one or multiple traits is provided and
and 2023 [13]. In LeWiDi, participant systems were chal-         to understand the impact of different configurations.
lenged to learn the distribution of labels, tested by cross
entropy-based metrics. In contrast, PERSEID aims at
stimulating the development of models of human per-              3. Data description
spectives, in order to explain the label distributions rather
than just quantifying them.                                      3.1. Origin of data
                                                                 The data for the challenge are part of MultiPICo [25],
2. Challenge: Description                                        a corpus of 18, 778 short conversations collected from
                                                                 Reddit (8, 956) and Twitter (9, 822) in 9 languages, and a
The task of Perspectivist Irony Detection aims to measure        total of 25 varieties.
models’ capability to detect irony in a short verbal ex-            Data were collected to reproduce the structure of short
change for each annotator, conditioned on the knowledge          conversations.
of demographic information about them. To this purpose,             For both Reddit and Twitter, the post is typically a
we want to look at different model performances if it is         message initiating a thread and the reply a direct reply
informed by one demographic trait or a combination of            to that message5 .
two. In particular, we focus on the gender and age of the           Reddit data were retrieved using the Pushshift reposi-
annotator, due to the balanced number of male and fe-            tory6 from January 2020 to June 2021. For Italian, data
male annotators by design 3.2, and due to the fact that age      were downloaded from the subreddit /r/Italy.
was shown to be one of the most polarized dimensions                Pairs having at least one deleted or removed comment
in [25].                                                         were filtered out, and the language of the messages was
   The input to the task does not consist only of a text,        further validated using the Python library for language
but rather of a tuple .                identification LangID7 .
   In this iteration of PERSEID, we considered several              Twitter data were collected via Twitter Stream API,
variables for the perspective attribute:                         using the geolocation service and excluding quotes and
                                                                 retweets. Then, the full conversation was retrieved, and
     • None (Task 0): acting as a baseline, we want to           tweets that directly replied to the starting ones were
       investigate the models’ outputs when no infor-            retained.
       mation about the annotator is provided.                      The data collection resulted in 18, 778 instances, to-
     • Age (Task 1): the perspective is one of four val-         gether with their metadata, consisting of Post-Reply orig-
       ues encoding the age group of the annotator.              inal IDs, subreddits, and geolocation information.
3                                                                5
  MultiPICo is available at https://huggingface.co/datasets/       For Reddit, second-level replies were collected in a minority of
  Multilingual-Perspectivist-NLU/MultiPICo with a CC-BY 4.0        cases; for Twitter, the post is a reply to a thread-starting message
  license.                                                         in a minority of cases.
4                                                                6
  For example, texts in Austrian, German, and Swiss German are     https://redditsearch.io/
                                                                 7
  included in the dataset.                                         https://github.com/saffsd/langid.py
       Language        #Annotators    #Annotations        Label rate        #Texts           Sources             Annotation mean
                                                         %not %iro                     #Reddit #Twitter
       Arabic               68            10,609          68      32        2,181        949       1,232                4.86
       Dutch                25            4,991           73      27        1,000        500         500                4.99
       English              74            14,171          69      31        2,999       1,499      1,500                4.73
       French               50            8,770           70      30        1,760       1,000        760                4.98
       German               70            12,510          68      32        2,375       1,042      1,333                5.27
       Hindi                24            4,711           65      35         786         286         500                5.99
       Italian              24            4,790           69      31        1,000        500        500                 4.79
       Portuguese           49            9,754           62      38        1,994        997         997                4.89
       Spanish              122           24,036          67      33        4,683       2,183      2,500                5.13
       Total                506           94,342          68      32        18,778      8,956      9,822                5.02
Table 1
Number of annotators, annotations, texts per source, and annotation means for each language. For Italian, 1000 pairs were
collected, each annotated by 4.79 annotators. Note the label unbalance, with the negative class accounting for 69% of the total
annotations.




Figure 1: Screenshot of the annotation interface for an English instance of MultiPICo. The Italian interface was similar, with translated
question and options.



  For Italian, data account for 1000 post, reply pairs,                  Annotators were selected based on three criteria:
equally sourced from Reddit and Twitter.
                                                                            • Their completion rate had to be greater or equal
                                                                              to 99%
3.2. Annotation details                                                     • They had to be native speakers of the considered
Annotators were asked to read a set of post and reply                         language (i.e., Italian, for the portion of data used
pairs and answer whether the text of the reply was ironic                     in the challenges)
or not, given the context.                                                  • The set of annotators needed to be balanced
   The human annotation of the collected data was per-                        across genders.
formed on the crowdsourcing platform Prolific8 , through
a custom-built annotation interface designed to collect                  The quality of the annotation was further assured us-
a diverse and balanced set of annotators. The interface               ing attention check questions in the form of “Please an-
mimicked a message conversation, having the post as                   swer X to this question”. Annotators had 1% probability of
context and asking whether the reply was Ironic or Not                receiving these special questions. Annotators who failed
ironic.                                                               to respond correctly to at least 50% of these questions
   For Italian, 24 native-speaker annotators were hired,              were excluded from the final corpus.
who performed 4,790 annotations in total, resulting in a                 A rich set of metadata is also provided. These include
mean of 4,79 annotations per instance (see Table 1).                  the self-identified Gender (balanced by design), their na-
                                                                      tionality, their Age Group (1 GenX, 15 GenY, 8 GenZ, for
8
                                                                      Italian), Ethnicity (23 white people, 1 mixed person, for
    https://www.prolific.com/
                   Demographics                                       Languages
                                        English Spanish Italian French Dutch German Hindi Arabic Portoguese
                           Boomer       3       2       –       2      –     5      –     1           –
                             GenX       22      17      1       7      4     7      3     4           1
             Age group
                             GenY       38      66      15      23     10    36     13    36         23
                            GenZ        10      37      8       17     11    20     8     26         25
                            White       47      60      23      40     22    66     –     20         37
                            Mixed       1       31      1       3      2     3      –     13         10
             Ethnicity       Asian      18      1       –       1      1     –      22    1           –
                             Black      3       2       –       5      –     –      –     2           1
                            Other       3       27      –       1      –     1      8     31          1
                              Yes       13      39      14      16     7     14     8     29         30
              Student
                              No        46      60      9       30     16    39     14    25         16
                          Full-time     25      41      9       24     10    24     10    20         15
                        Unemployed 11           24      7       5      4     3      1     11          8
                          Part-time     11      17      5       5      3     10     4     13          6
            Employment
                       Not in paid work 4       4       1       5      4     5      –     1           –
                         Due to start –         3       1       1      –     2      2     –           2
                            Other       1       6       –       6      –     3      1     5          14

Table 2
Sociodemographic information about annotators per language.



Italian), Student status (14 yes, 9 no, for Italian), Employ- 'reply_id': 2497527360959166890,
ment status (9 in full-time jobs, 7 unemployed, 5 working 'source': 'twitter',
part-time, 1 not in paid work and 1 due to start, for Ital- 'timestamp': '2022-12-07 15:49:50'
ian), as reported in Table 2.

3.3. Data format                                          3.4. Example of prompts used for
The dataset is in tabular format, one row per annotation.      zero-shot prediction
The data contain the text in the form of two fields (post     The challenge is zero-shot, and the prompt depends on
and reply), the binary label, and a series of metadata        three variables: perspective, post, and reply.
about the post, reply, and annotator. Here is an example
of instance from the Italian section of MultiPICo:            Sei {perspective}.
                                                              Istruzione: Ti vengono fornite in
'Age': 29.0,                                                  input (Input) una coppia di frasi
'Country of birth': 'Italy',                                  (Post, Reply) estratte da conversazioni
'Country of residence': 'Italy',                              sui social media. Il tuo compito è
'Employed': 'Yes',                                            determinare se la Risposta (Reply) è
'Employment status': 'Part-Time',                             ironica nel contesto del Post (Post).
'Ethnicity simplified': 'White',                              Fornisci in output (Output) una singola
'Gender': 'Male',                                             etichetta “ironia" o "non ironia".
'Generation': 'GenY',                                         Input:
'GenerationAggregated': 'Young',                              Post: {post}
'Nationality': 'Italy',                                       Reply: {reply}
'Student status': 'No',                                       Output:
'annotator_id': 9208155880570654046,
'label': 0,                                                   Task 0 No perspective is provided, and the prompt
'language': 'it',                                                   directly starts with the instruction.
'language_variety': 'it',
'level': 1.0,                                                 Task 1 The perspective variable is a verbalization of
'post': 'Ormai il quadro è chiaro: cercare di                       the Generation, which is expressed as an integer
   coinvolgere tutti per non farla pagare a                         in the dataset. It can be instantiated with the
   nessuno. Se non riuscissero a corrompere i                       following values9 :
   Pm di Torino andranno in B diretti.',                      9
                                                                  No workers whose age is > 42, i.e., from the baby boomer gener-
'post_id': 14071953227682835778,                                  ations, participated in the annotation of the Italian portion of the
'reply': '@USER Magari ??',                                       dataset
           • “una persona giovane della generazione Z”             In the vast majority (∼90%) of cases, the
             if Generation == GenZ (Age < 26)                      conversation-starting messages and their direct
           • “una persona giovane della generazione Y”             replies were downloaded to capture the full con-
             if Generation == GenY (26 ≤ Age < 42)                 versational context. In a few cases, the down-
           • “una persona adulta della generazione X”              loaded reply was not direct but rather a second-
             if Generation == GenX (42 ≤ Age < 58)                 level reply (a reply to a direct reply); thus, some
           • “una persona adulta della generazione baby            conversational context might be missing.
             boomer”                                       Challenge design We describe annotators by no so-
             if Generation == Boomer (Age > 58)                  ciodemographic traits (Task 0), one single demo-
Task 2 The perspective variable is a verbalization of            graphic trait (Task 1 and Task 2), or two demo-
      the Gender variable, which is expressed as a               graphic traits (Task 3). We evaluate disaggregated
      string in English. It can be instantiated with one         annotations at inference time, having the annota-
      of two values:                                             tors represented only by those traits. Annotators’
                                                                 sociodemographic information does not always
           • “una donna”                                         align with the most relevant grouping of anno-
             if Gender == “Female”                               tators according to the language phenomenon
           • “un uomo”                                           under study [21, 28], and the limited amount of
             if Gender == “Male”                                 sociodemographic traits we provide is undoubt-
Task 3 The perspective variable is a verbalization of            edly not enough to describe every single anno-
      both the Age and Gender variables, e.g., “una              tator. We are aware of this limitation. In fact,
      giovane donna della generazione Z.”                        our main aim is to understand whether providing
                                                                 one or more annotator traits makes the model
                                                                 predictions more aligned with annotators having
4. Metrics                                                       a given characteristic.

Inspired by Mokhberian et al. [26], the Perspectivist Irony
Detection task is evaluated by means of global F1, that 6. Ethical issues
is, the F1-score computed across all the individual an-
notations in the dataset against the predictions of the This work places itself in an increasing amount of work
model.                                                      that calls to consider and include the subjectivity of
                                                            the annotators in NLP applications, encouraging reflec-
                                                            tion on the different perspectives encoded in annotated
5. Limitations                                              datasets to minimize the amplification of biases. We hope
                                                            this challenge will be a starting point for investigating
Data The sociodemographic information about the an- and evaluating LLMs in Italian to make them suitable for
        notators is partial, bound to what was available final users.
        from the crowdsourcing platform, and following a       The dataset used in the challenge was built by adopt-
        discretization of human personal traits that could ing measures to protect the privacy of annotators, and
        be perceived as forced (e.g., representing self- the data handling protocols were designed to safeguard
        identified gender as a single binary label). Fur- personal information (like anonymization of users’ men-
        thermore, as shown by Orlikowski et al. [21], an- tions). Although the attention during the collection of
        notators’ sociodemographics do not always align data was focused on ironic content spread online, we
        with the most relevant grouping of annotators acknowledge that some of the material contains racist,
        according to the language phenomenon under sexist, stereotypical, violent, or generally disturbing con-
        study.                                              tent.
        Annotators of the Italian portion of MultiPICO         Annotators are balanced through their self-identified
        tend to be young (with no annotators from the gender. However, we are aware that considering gen-
        baby boomer generation and only one from der in a binary form is limited; moreover, a substantial
        GenX). This aspect might influence the results.     unbalance for some dimensions, like the self-identified
                                                            ethnicities, is present in the dataset. This pattern sug-
        Similarly to Sachdeva et al. [5], Sap et al. [19],
                                                            gests the need to interact differently with annotators or
        Forbes et al. [27], we noticed the ethnicity of an-
                                                            social communities if we want a diversity of annotators
        notators is unbalanced, and all but one annotators
                                                            and perspectives in terms of social background.
        are white for the considered data.
7. Data license and copyright                                      tions in irony detection, in: ECAI 2023 Workshop
                                                                   on Perspectivist Approaches to NLP, 2023.
   issues                                                      [7] A. Mostafazadeh Davani, M. Díaz, V. Prabhakaran,
MultiPICo is distributed under the Creative Commons                Dealing with disagreements: Looking beyond the
Attribution 4.0 (CC-BY-4.0) license.                               majority vote in subjective annotations, Transac-
                                                                   tions of the Association for Computational Linguis-
                                                                   tics 10 (2022) 92–110. URL: https://aclanthology.
Acknowledgments                                                    org/2022.tacl-1.6. doi:10.1162/tacl_a_00449 .
                                                               [8] S. Casola, S. Lo, V. Basile, S. Frenda, A. Cignarella,
This work was funded by the ‘Multilingual Perspective-             V. Patti, C. Bosco, Confidence-based ensembling of
Aware NLU’ project in partnership with Amazon Alexa.               perspective-aware models, in: H. Bouamor, J. Pino,
                                                                   K. Bali (Eds.), Proceedings of the 2023 Conference
                                                                   on Empirical Methods in Natural Language Pro-
References                                                         cessing, Association for Computational Linguis-
 [1] V. Basile, M. Fell, T. Fornaciari, D. Hovy, S. Paun,          tics, Singapore, 2023, pp. 3496–3507. URL: https:
     B. Plank, M. Poesio, A. Uma, et al., We need to               //aclanthology.org/2023.emnlp-main.212. doi:10.
     consider disagreement in evaluation, in: Proceed-             18653/v1/2023.emnlp- main.212 .
     ings of the 1st workshop on benchmarking: past,           [9] A. N. Uma, T. Fornaciari, D. Hovy, S. Paun, B. Plank,
     present and future, Association for Computational             M. Poesio, Learning from disagreement: A survey,
     Linguistics, 2021, pp. 15–21.                                 Journal of Artificial Intelligence Research 72 (2021)
 [2] B. Plank, The ”problem” of human label variation:             1385–1470.
     On ground truth in data, modeling and evaluation,        [10] L. Aroyo, C. Welty,                Truth is a lie:
     in: Proceedings of the 2022 Conference on Empiri-             Crowd truth and the seven myths of hu-
     cal Methods in Natural Language Processing, 2022,             man annotation,           AI Magazine 36 (2015)
     pp. 10671–10682.                                              15–24. URL: https://ojs.aaai.org/aimagazine/
 [3] F. Cabitza, A. Campagner, V. Basile, Toward a per-            index.php/aimagazine/article/view/2564.
     spectivist turn in ground truthing for predictive             doi:10.1609/aimag.v36i1.2564 .
     computing, in: Proceedings of the AAAI Con-              [11] E. Leonardelli, S. Menini, A. P. Aprosio, M. Guerini,
     ference on Artificial Intelligence, volume 37, 2023,          S. Tonelli, Agreeing to disagree: Annotating offen-
     pp. 6860–6868. URL: https://ojs.aaai.org/index.php/           sive language datasets with annotators’ disagree-
     AAAI/article/view/25840.                                      ment, in: Proceedings of the 2021 Conference on
 [4] S. Frenda, A. Pedrani, V. Basile, S. M. Lo, A. T.             Empirical Methods in Natural Language Processing,
     Cignarella, R. Panizzon, C. Marco, B. Scarlini,               2021, p. 10528–10539.
     V. Patti, C. Bosco, D. Bernardi, EPIC: Multi-            [12] A. Uma, T. Fornaciari, A. Dumitrache, T. Miller,
     perspective annotation of a corpus of irony, in:              J. Chamberlain, B. Plank, E. Simpson, M. Poesio,
     A. Rogers, J. Boyd-Graber, N. Okazaki (Eds.), Pro-            Semeval-2021 task 12: Learning with disagree-
     ceedings of the 61st Annual Meeting of the Associa-           ments, in: Proceedings of the 15th International
     tion for Computational Linguistics (Volume 1: Long            Workshop on Semantic Evaluation (SemEval-2021),
     Papers), Association for Computational Linguis-               2021, pp. 338–347.
     tics, Toronto, Canada, 2023, pp. 13844–13857. URL:       [13] E. Leonardelli, A. Uma, G. Abercrombie, D. Al-
     https://aclanthology.org/2023.acl-long.774. doi:10.           manea, V. Basile, T. Fornaciari, B. Plank, V. Rieser,
     18653/v1/2023.acl- long.774 .                                 M. Poesio, Semeval-2023 task 11: Learning with
 [5] P. Sachdeva, R. Barreto, G. Bacon, A. Sahn, C. von            disagreements (lewidi), in: Proceedings of the 17th
     Vacano, C. Kennedy, The measuring hate speech                 International Workshop on Semantic Evaluation
     corpus: Leveraging rasch measurement theory for               (SemEval-2023), 2023, p. 2304–2318.
     data perspectivism, in: G. Abercrombie, V. Basile,       [14] S. Santy, J. Liang, R. Le Bras, K. Reinecke, M. Sap,
     S. Tonelli, V. Rieser, A. Uma (Eds.), Proceedings of          NLPositionality: Characterizing design biases
     the 1st Workshop on Perspectivist Approaches to               of datasets and models, in: Proceedings of
     NLP @LREC2022, European Language Resources                    the 61st Annual Meeting of the Association for
     Association, Marseille, France, 2022, pp. 83–94. URL:         Computational Linguistics (Volume 1: Long Pa-
     https://aclanthology.org/2022.nlperspectives-1.11.            pers), Association for Computational Linguistics,
 [6] S. Frenda, S. M. Lo, S. Casola, B. Scarlini, C. Marco,        Toronto, Canada, 2023, pp. 9080–9102. URL: https://
     V. Basile, D. Bernardi, Does anyone see the irony             aclanthology.org/2023.acl-long.505. doi:10.18653/
     here? Analysis of perspective-aware model predic-             v1/2023.acl- long.505 .
                                                              [15] V. Prabhakaran, A. M. Davani, M. Diaz, On re-
     leasing annotator-level labels and information in             Computational Linguistics, Florence, Italy, 2019, pp.
     datasets, in: Proceedings of the Joint 15th Linguis-          5716–5728. URL: https://aclanthology.org/P19-1572.
     tic Annotation Workshop (LAW) and 3rd Designing               doi:10.18653/v1/P19- 1572 .
     Meaning Representations (DMR) Workshop, 2021,            [23] A. Uma, T. Fornaciari, A. Dumitrache, T. Miller,
     p. 133–138.                                                   J. Chamberlain, B. Plank, E. Simpson, M. Poesio,
[16] E. M. Bender, B. Friedman, Data statements for nat-           SemEval-2021 task 12: Learning with disagree-
     ural language processing: Toward mitigating sys-              ments, in: A. Palmer, N. Schneider, N. Schluter,
     tem bias and enabling better science, Transactions            G. Emerson, A. Herbelot, X. Zhu (Eds.), Proceedings
     of the Association for Computational Linguistics 6            of the 15th International Workshop on Semantic
     (2018) 587–604.                                               Evaluation (SemEval-2021), Association for Com-
[17] D. Almanea, M. Poesio, ArMIS - the Arabic Misog-              putational Linguistics, Online, 2021, pp. 338–347.
     yny and Sexism Corpus with Annotator Subjective               URL: https://aclanthology.org/2021.semeval-1.41.
     Disagreements, in: N. Calzolari, F. Béchet, P. Blache,        doi:10.18653/v1/2021.semeval- 1.41 .
     K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isa-     [24] G. Attanasio, P. Basile, F. Borazio, D. Croce, M. Fran-
     hara, B. Maegaard, J. Mariani, H. Mazo, J. Odijk,             cis, J. Gili, E. Musacchio, M. Nissim, V. Patti, M. Ri-
     S. Piperidis (Eds.), Proceedings of the Thirteenth            naldi, D. Scalena, CALAMITA: Challenge the Abili-
     Language Resources and Evaluation Conference,                 ties of LAnguage Models in ITAlian, in: Proceed-
     European Language Resources Association, Mar-                 ings of the 10th Italian Conference on Computa-
     seille, France, 2022, pp. 2282–2291. URL: https:              tional Linguistics (CLiC-it 2024), Pisa, Italy, Decem-
     //aclanthology.org/2022.lrec-1.244.                           ber 4 - December 6, 2024, CEUR Workshop Proceed-
[18] S. Akhtar, V. Basile, V. Patti, Whose opinions mat-           ings, CEUR-WS.org, 2024.
     ter? Perspective-aware models to identify opinions       [25] S. Casola, S. Frenda, S. Lo, E. Sezerer, A. Uva,
     of hate speech victims in abusive language detec-             V. Basile, C. Bosco, A. Pedrani, C. Rubagotti, V. Patti,
     tion, arXiv preprint arXiv:2106.15896 (2021).                 D. Bernardi, MultiPICo: Multilingual perspectivist
[19] M. Sap, S. Swayamdipta, L. Vianna, X. Zhou, Y. Choi,          irony corpus, in: L.-W. Ku, A. Martins, V. Srikumar
     N. A. Smith, Annotators with attitudes: How an-               (Eds.), Proceedings of the 62nd Annual Meeting
     notator beliefs and identities bias toxic language            of the Association for Computational Linguistics
     detection, in: Proceedings of the 2022 Conference             (Volume 1: Long Papers), Association for Compu-
     of the North American Chapter of the Association              tational Linguistics, Bangkok, Thailand, 2024, pp.
     for Computational Linguistics: Human Language                 16008–16021. URL: https://aclanthology.org/2024.
     Technologies, Association for Computational Lin-              acl-long.849.
     guistics, Seattle, United States, 2022, pp. 5884–5906.   [26] N. Mokhberian, M. Marmarelis, F. Hopp, V. Basile,
     URL: https://aclanthology.org/2022.naacl-main.431.            F. Morstatter, K. Lerman, Capturing perspectives
     doi:10.18653/v1/2022.naacl- main.431 .                        of crowdsourced annotators in subjective learning
[20] R. Wan, J. Kim, D. Kang, Everyone’s voice mat-                tasks, in: K. Duh, H. Gomez, S. Bethard (Eds.),
     ters: Quantifying annotation disagreement using               Proceedings of the 2024 Conference of the North
     demographic information, in: Proceedings of the               American Chapter of the Association for Compu-
     37th AAAI Conference on Anrtificial Intelligence -            tational Linguistics: Human Language Technolo-
     AAAI Special Track on AI for Social Impact, 2023.             gies (Volume 1: Long Papers), Association for Com-
[21] M. Orlikowski, P. Röttger, P. Cimiano, D. Hovy, The           putational Linguistics, Mexico City, Mexico, 2024,
     ecological fallacy in annotation: Modeling human              pp. 7337–7349. URL: https://aclanthology.org/2024.
     label variation goes beyond sociodemographics, in:            naacl-long.407.
     A. Rogers, J. Boyd-Graber, N. Okazaki (Eds.), Pro-       [27] M. Forbes, J. D. Hwang, V. Shwartz, M. Sap, Y. Choi,
     ceedings of the 61st Annual Meeting of the Associa-           Social chemistry 101: Learning to reason about so-
     tion for Computational Linguistics (Volume 2: Short           cial and moral norms, in: Proceedings of the 2020
     Papers), Association for Computational Linguis-               Conference on Empirical Methods in Natural Lan-
     tics, Toronto, Canada, 2023, pp. 1017–1029. URL:              guage Processing (EMNLP), Association for Com-
     https://aclanthology.org/2023.acl-short.88. doi:10.           putational Linguistics, Online, 2020, pp. 653–670.
     18653/v1/2023.acl- short.88 .                                 URL: https://aclanthology.org/2020.emnlp-main.48.
[22] E. Simpson, E.-L. Do Dinh, T. Miller, I. Gurevych,            doi:10.18653/v1/2020.emnlp- main.48 .
     Predicting humorousness and metaphor novelty             [28] S. M. Lo, V. Basile, Hierarchical clustering of label-
     with Gaussian process preference learning, in:                based annotator representations for mining per-
     A. Korhonen, D. Traum, L. Màrquez (Eds.), Proceed-            spectives, in: G. Abercrombie, V. Basile, D. Bernardi,
     ings of the 57th Annual Meeting of the Associa-               S. Dudy, S. Frenda, L. Havens, E. Leonardelli,
     tion for Computational Linguistics, Association for           S. Tonelli (Eds.), Proceedings of the 2nd Workshop
on Perspectivist Approaches to NLP co-located with
26th European Conference on Artificial Intelligence
(ECAI 2023), Kraków, Poland, September 30th, 2023,
volume 3494 of CEUR Workshop Proceedings, CEUR-
WS.org, 2023. URL: https://ceur-ws.org/Vol-3494/
paper8.pdf.