PejorativITy - In-Context Pejorative Language Disambiguation: A CALAMITA Challenge Arianna Muti University of Bologna - DIT Abstract Misogyny is often expressed through figurative language. Some neutral words can assume a negative connotation when functioning as pejorative epithets, and they can be used to express misogyny. Disambiguating the meaning of such terms might help the detection of misogyny. This challenge addresses a) the disambiguation of specific ambiguous words in a given context; b) the detection of misogyny in instances that contain such polysemic words. In particular, framed as a binary classification, our task is divided into two parts. In Task A, the model is asked to define if, given a tweet, the target word is used in pejorative or non-pejorative way. In Task B, the model is asked whether the whole tweet is misogynous or not. Keywords offensive language, pejorativity, misogyny Warning: This paper contains offensive words. tation in the training set and a neutral connotation in the test set results in a great number of false positives [5]. This could be overcome by decoder-based LLMs, 1. Introduction and Motivation as they could rely on their implicit knowledge to grasp the meaning of such terms. By asking models to de- This CALAMITA challenge [1] addresses the task of termine whether a term is being used in a pejorative disambiguating pejorative language to detect forms of or non-pejorative sense, we challenge the LLMs’ abil- misogyny that are masked within ambiguous and context- ity to comprehend semantic shifts in Italian. Moreover, dependent expressions. Pejorative language refers to a asking whether a sentence containing that term is misog- word or phrase that has negative connotations and is 1 ynous or not, enables us to comprehend to what extent intended to disparage or belittle. An inoffensive word LLMs understand misogyny, even when it is conveyed becoming pejorative is a form of semantic drift known as through figurative language. We expect models to strug- pejoration; thus, pejorativity is context-dependent: pe- gle with this challenge, particularly in sentences with jorative words have one primary neutral meaning, and non-standard or regional varieties of Italian, which occur another negatively connotated meaning. In this chal- in our corpus. lenge, our objective is to evaluate large language mod- els (LLM) in Italian by focusing on the disambiguation of pejorative epithets used online to express misogyny. 2. Challenge: Description In this work, misogyny is defined as a property of so- cial environments where women perceived as violating We introduce pejorative language disambiguation as a patriarchal norms are “kept down” through hostile or preliminary step to detect misogyny. Our goal is to assess benevolent reactions coming from men, other women, whether the disambiguation of potentially pejorative ep- and social structures [2, 3], in the form of sexual objec- ithets improves the detection of misogynistic language. tification, male privilege, gender discrimination, sexual Therefore, this challenge aims to address two tasks: harassment, belittling and violence [4]. An example of a pejorative epithet is balena (whale), Task A Disambiguation of in-context polysemic words whose standard meaning refers to the sea mammal, that can be used as pejorative epithets in misogy- but it is used offensively to address an overweight nistic language; woman. Encoder-based models struggle to correctly clas- Task B Misogyny detection at the sentence level. sify misogyny when sentences contain such terms: the occurrence of polysemic words with a pejorative conno- Both tasks are conceived as binary classification tasks. Fig. 1 shows the pipeline for our tasks. Assume the sen- CLiC-it 2024: Tenth Italian Conference on Computational Linguistics, tence Quella balena coi jeans non si può guardare, trans- Dec 04 — 06, 2024, Pisa, Italy lated as Can’t look at that whale with jeans. Envelope-Open arianna.muti2@unibo.it (A. Muti) Task A: First, the model is asked to identify whether Orcid 0000-0002-3387-6557 (A. Muti) © 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License the meaning of the target word (balena in our example) 1 Attribution 4.0 International (CC BY 4.0). https://www.merriam-webster.com/dictionary/pejorative is pejorative or not. The model should rely on its internal CEUR ceur-ws.org Workshop ISSN 1613-0073 Proceedings Figure 1: Visualization of our tasks. knowledge accumulated during pre-training to under- words refer to the unambiguous words used to define stand whether the term balena (whale) refers to woman polysemic words. We call these words anchors because or cetaceus. Ideally, the model should exploit the context their meaning is univocal and does not change according to perform the disambiguation, as the image of a whale to the context. For instance, the word balena (whale) is with jeans is not plausible. That is why we encourage used to refer to either a sea mammal or an overweight commonsense reasoning for this task. woman. In contrast, the anchor words cetaceo (cetacean) Task B: In the second step, the model is first informed and grassa (fat) only refer to the animal in the first case with the decision of Task A, whether the target word is and to being overweight in the second case, at least as pejorative or not, and then asked to classify the input far as their use in Twitter is concerned.4 sentence as misogynous or not. Tweets. We use Twarc5 to retrieve tweets from De- cember 2022 to February 2023 containing words in our 3. Data description lexicon. We select 50 tweets for each word in our lexicon, The compilation of our corpus involves two steps: the resulting in 1,200 tweets. We maintain a balance of pejo- creation of a lexicon of polysemic words that can function rative and neutral use of lexicon words, although an equal as pejorative epithets for women, and the retrieval of distribution for each word could not be guaranteed. We tweets containing such words. choose tweets as source of data for three reasons. First, Twitter is a prominent platform for expressing opinions, where language is varied, conversational, and often in- Lexicon. We collect our lexicon by selecting words formal, which makes it suitable to analyze misogyny con- from three distinct sources. veyed through figurative language. Second, at the time of (1) We ask ten Italian native speakers to provide a list data collection, Twitter API was public and free, which fa- of offensive words used online to address women. The cilitated our data collection process. Third, the character speakers use social media on a daily basis and their age limit on tweets encourages condensed language, limiting ranges between 27 and 39 years. the context of expression. Choosing tweets allows us to (2) We retrieve the keywords used in the two Italian challenge LLMs in disambiguating pejorative language corpora for the Automatic Misogyny Identification (AMI) for misogyny detection within the constraints of limited shared task [6, 7]. or lack of context. (3) We consult the ’List of Dirty Naughty Obscene Bad Words’.2 We only keep polysemic words whose primary mean- 3.1. Annotation Details ing is neutral and that are frequently used on Twitter We recruit six annotators with a background in linguis- with both pejorative and neutral connotations. To ensure tics, gender studies, cognitive sciences, and NLP to label the quality of our vocabulary, we qualitatively verify that our corpus for pejorative word disambiguation (word- such words are used with both connotations by manually level) and misogyny detection (sentence-level). searching them on Twitter.3 We first devise a pilot annotation study to explore the Table 1 shows our lexicon of 24 words. For each word, complexity of the task. For this purpose, we follow a we report the English translation of its literal and pe- descriptive annotation paradigm [8], which encourages jorative meaning, and their anchors in Italian. Anchor annotator subjectivity by not providing guidelines. We 2 split the annotators into two groups and assign 50 tweets https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene- and-Otherwise-Bad-Words/tree/master, consulted on January each for labeling. Each group is composed of two women 2023. and one man with ages ranging between 27 and 39 years 3 Due to their exclusive neutral or negative connotation on Twitter, the following words are discarded: barile, banco, botte, barbona, 4 In this case, the word balena has a third anchor word, from the verb facile, gatta morta, passeggiatrice, porca, principessa, privilegiata, balenare, which means ’to flash’. 5 psicopatica, scrofa, somara, travestita. https://twarc-project.readthedocs.io Word Literal Pejorative Neutral anchor Pejorative anchor acida acid/sour peevish aspra intrattabile, stronza asina female donkey stupid ciuco stupida balena whale/flash fat woman cetaceo, balenare grassa bambola doll girl (objectifying) giocattolo donna attraente cagna female dog bitch cane femmina, canide donna di facili costumi, troia cavalla female horse ugly/tall/ungainly equino brutta, alta e grossa civetta owl tease volatile rapace donna che cerca attenzioni cesso toilet ugly water, bagno, toilette brutta contadina farmer ignorant, illiterate agricoltore femmina donna ignorante cortigiana court lady prostitute dama di corte prostituta cozza mussel ugly/clingy mollusco donna brutta, appiccicosa femminista feminist feminazi femminista polemica, fastidiosa fogna sewer skanky fognatura schifosa, bocca gallina chicken stupid pennuto stupida grezza raw rude woman non lavorato rozza lesbica lesbian dyke donna a cui piacciono le donne lesbica (offensivo) lurida dirty skanky sporca promiscua, troia maiala sow whore maiale femmina promiscua, troia mucca cow bitch bovide stupida, troia oca goose stupid girl pennuto stupida, pettegola pecora sheep doormat ovino stupida strega witch hag, unpleasant maga crudele vacca cow whore bovino donna di facili costumi, troia zingara gipsy shabby gitana trasandata Table 1 Italian pejorative lexicon, their literal and pejorative translations in English, and their anchors. old. We use Krippendorff’s alpha [9] to measure the inter- 3.3. Detailed data statistics annotator agreement (IAA). The IAA of the first group is Table 3 shows the statistics of our corpus. The Pearson moderate for both pejorativity (0.48) and misogyny (0.50), correlation between misogyny and pejorativity labels is whereas the IAA of the second group is fair for pejorativ- 0.70, which is in line with our expectations. The tweets ity (0.33) and moderate for misogyny (0.50). We observe for which misogyny and pejorativity are not aligned are that, in terms of gender differences, men tend to consider mainly reported speech or men-related offensive lan- sexual objectifying compliments as non-pejorative. More guage. It is worth noting that some sentences are anno- details about the annotation process, including the dis- tated as misogynous, although they do not express any cussion of edge cases, can be found in Muti et al. [10]. form of hate towards women. However, they contain After the pilot studies, we annotate our collected cor- subtle sexist language, which we consider misogynous pus of 1,200 tweets. Only one person carries out the according to the definition provided in Section 1. For whole annotation process. We select the annotator with instance, the sentence “che bella bambola ciao tesoro” 6 the most interdisciplinary background, who is an expert does not express hate, but perpetuates the objectification in gender studies, linguistics and NLP, who has been a of women by addressing the target of the tweet as a doll, target of misogyny. This setting is considered among falling into the category of benevolent sexism [12]. the best practices for the annotation of phenomena like misogyny [11]. 3.4. Prompt Design 3.2. Data format We design two prompts to address the two task: pejorativ- ity disambiguation at word-level and misogyny detection Data are collected in an Excel file and published at at sentence-level. We adopt a zero-shot approach, al- https://github.com/arimuti/PejorativITy. Each row con- though participants are encouraged to experiment with tains the ID of the tweet, the tweet, the target word, the different prompting techniques. annotation for pejorativity at word level and the anno- tation for misogyny at sentence level. Table 2 shows examples. 6 translation: what a beautiful doll (girl), hi darling ID Tweet Pejorative Misogyny 70019 Non voglio una cagna un cane ce l’ho giaaaa 1 1 10010 Xchè avrà dato una risposta acida a lui 0 0 61209 Ma come fai a dire che sei una balena sei bellissima 1 0 Table 2 Examples of tweets with potentially pejorative words (underlined). Class Training Test Total Model Macro F1 AlBERTo 0.82 Misogynous 369 28 397 llama-3-8b 0.00 Pejorative 363 28 391 llama-3-8b-Instruct 0.64 Not pejorative 6 – 6 Table 4 Non-misogynous 735 68 803 Results on Task A. Scores are not comparable since the testing Pejorative 172 18 190 partitions are different. Not pejorative 563 50 613 Table 3 Statistics of the PejorativITy corpus. The same tweets are the Instruct version). The results are not comparable annotated for misogyny and pejorativity, for a total of 1,200 though, since llama is evaluated on the whole corpus, instances. while AlBERTo on the partition of the test set (see Ta- ble 3). Task A. We design the following zero-shot prompt that Task A. Table 4 shows the results for pejorative word asks whether the target {word} in a given {sentence} is disambiguation. The fine-tuned AlBERTo model reaches pejorative or not. We prompt the LLMs for each instance a macro F1 -measure of 0.82 ± 0.03, as reported in [10]. in our corpus. The prompt is given as follows: When it comes to decoder-based models, llama3-8b- La parola {word} è utilizzata in modo peggiorativo Instruct shows a lower score, with a difference of 0.18 nel testo in input? Rispondi solo con 1 per sì, 0 per points, showing room for improvement in the prompt no. Testo: {tweet} design. However, those scores are not comparable as the testing partitions differ. Llama-3-8b fails to complete the Task B. We design two zero-shot prompts that ask task, since it only repeats the prompt without providing whether a tweet is misogynous or not in order to evaluate an answer. For this reason, we discard llama-3-8b in the two settings: next task. It should be noted that llama has undergone A We treat Task A and Task B as independent, i.e., a safety tuning process, preventing the model from al- without having the decision for Task A to inform ways providing an answer, responding I cannot provide the decision for Task B. a response that condones hate speech. We excluded such cases from the evaluation. Of the 174 excluded instances, B We enrich the prompt for Task B with the output 123 were pejorative and 51 were not pejorative according from the previous task; to the gold standard. Although the fine-tuned version Prompt A: of AlBERTo achieves a higher performance (in a smaller Il testo in input è misogino? Rispondi solo con 1 subset of instances), llama aids in explainability by de- per sì, 0 per no. Testo: {tweet} liberately adding explanations of why it considers the target word to be pejorative or not. We will explore the Prompt B: plausibility of such explanations in future work. Il testo in input è misogino? Rispondi solo con 1 per sì, 0 per no. Considera l'informazione sulla Task B. Table 5 shows the performance regarding parola {word}: {response from Task A} Testo: {tweet} misogyny detection at sentence level. In Setting A, where the model is not informed of the 4. Evaluation of Baselines output for Task A, AlBERTo scores are much lower com- pared to Task A, achieving 0.68 ± 0.03. Llama performs In order to evaluate our tasks, we employ Macro F1 score. better in Task B compared to Task A, overcoming Al- We select two models as baselines: the fine-tuned Al- BERTo by just 0.01 point. However, the fact that all BERTo model [13] and llama-3-8b (both the standard and answers were provided in Task B (unlike in the previous misogyny detection. We have designed two tasks as binary classification problems: A) pejorative language disambiguation at word level and B) misogyny detection at sentence level. Our preliminary experiments show that a Transformer-based fine-tuned model performs better than llama-3-8b-Instruct in detecting pejorative words, while llama-3-8b-Instruct performs slightly better than the Transformer-based model in misogyny detection. In the future, we plan to explore how the unrequested ex- planations provided by llama-3-8b-Instruct about the pe- jorativity of a target word impact the classification of Figure 2: Our pipeline for injecting information about pejo- rativity for Task B (setting B) in AlBERTo. Step 1: a model misogynous sentences. identifies the connotation of possibly pejorative epithets. Step 2: the identified connotation is used to enrich (CONCAT) and substitute (SUBST) part of the textual input for misogyny de- 6. Limitations tection. Although our lexicon covers a wide variety of words that can serve as pejorative epithets for women, it is not an exhaustive list, as we have discarded all the terms task) plays a role and does not necessarily imply that that are not polysemic and that are used only with one misogyny detection is an easier task than pejorativity connotation (either positively or negatively) on Twitter. disambiguation for llama. Moreover, only 100 tweets are annotated by six an- In Setting B, the model is informed of the decision on notators, while the remaining 1,100 are labelled by only pejorativity of the target word. While for llama the infor- one annotator. Although we select an expert with an in- mation about pejorativity can be injected in the prompt, terdisciplinary background in linguistics, gender studies with AlBERTo we have adopted two approaches: i) we and NLP to carry out all the annotations, their personal concatenate the information about the pejorativity of biases, opinions, or interpretations can lead to skewed or the target word at the end of the tweet or ii) we substi- one-sided data. tute the ambiguous word with its corresponding anchor Finally, our corpus is characterized by the presence word from our lexicon. Fig. 2 shows the pipeline. We ob- of sarcasm, abbreviations, and non-standard varieties of serve a notable improvement over the baseline model for Italian, which might make the semantics of our instances concatenation (+7 absolute points) and substitution (+9 hard to be captured by current language models. absolute points) when using the predictions for Task A. Another limitation of our study concerns the substitu- On the other hand, llama does not benefit from the tion approach. First of all, some words have more than injection of knowledge about pejorative words, with a one neutral anchor words. This is the case of balena, drop of 0.09 points. This could be due to the noisy re- which has two neutral anchors: balenare (to flash) and sponse from Task A, including the refusal to answer, and cetaceo (sea mammal). In neutral examples, we substitute possible wrong explanations of why the target word is balena with both anchors. This process may alter the used pejoratively or not. semantic meaning of the tweet since only one anchor is suitable for substitution. Moreover, in some cases, we Setting Model Macro F1 replace a lexicon word with anchors that do not have A AlBERTo 0.68 the same meaning. For instance, the neutral anchor of B_concat AlBERTo 0.75 acida is aspra (sour). However, expressions like sour beer B_subst AlBERTo 0.77 A llama-3-8b-Instruct 0.69 or sour cream do not have a valid anchor replacement. B llama-3-8b-Instruct 0.60 Therefore, replacing aspra with acida is not an appropri- ate substitution. Table 5 Results on Task B. Scores are not comparable since the testing partitions are different. 7. Ethical Issues Our data collection adheres to Twitter’s terms of service and privacy policies. As this research involves the analy- 5. Conclusion sis of publicly available tweets, we do not seek explicit consent from individual users. Nevertheless, we make We have presented a new challenge for CALAMITA: pe- every effort to protect the anonymity of all individuals jorative word disambiguation as a preliminary step for mentioned. However, the exposure to misogynistic con- Language Processing and Speech tools for Italian tent still poses a mental health risk for researchers and (EVALITA 2020), CEUR.org, Online, 2020. annotators. [8] P. Röttger, B. Vidgen, D. Hovy, J. B. Pierrehumbert, Two contrasting data annotation paradigms for sub- jective NLP tasks, in: M. Carpuat, M. de Marn- Acknowledgments effe, I. V. M. Ruíz (Eds.), Proceedings of the 2022 Conference of the North American Chapter of the Thanks to Federico Ruggeri, Cagri Toraman, Alberto Association for Computational Linguistics: Hu- Barrón-Cedeño, Lorenzo Musetti, Samuel Algherini, Sil- man Language Technologies, NAACL 2022, Seattle, via Ronchi, Gianmarco Saretto, and Caterina Zapparoli WA, United States, July 10-15, 2022, Association for helping with the brainstorming for the task design. for Computational Linguistics, 2022, pp. 175–190. URL: https://doi.org/10.18653/v1/2022.naacl-main. References 13. doi:10.18653/v1/2022.naacl- main.13 . [9] K. Krippendorff, Computing krippendorff’s alpha- [1] G. Attanasio, P. Basile, F. Borazio, D. Croce, M. Fran- reliability, 2011. cis, J. Gili, E. Musacchio, M. Nissim, V. Patti, M. Ri- [10] A. Muti, F. Ruggeri, C. Toraman, A. Barrón-Cedeño, naldi, D. Scalena, CALAMITA: Challenge the Abili- S. Algherini, L. Musetti, S. Ronchi, G. Saretto, ties of LAnguage Models in ITAlian, in: Proceed- C. Zapparoli, PejorativITy: Disambiguating pe- ings of the 10th Italian Conference on Computa- jorative epithets to improve misogyny detection tional Linguistics (CLiC-it 2024), Pisa, Italy, Decem- in Italian tweets, in: N. Calzolari, M.-Y. Kan, ber 4 - December 6, 2024, CEUR Workshop Proceed- V. Hoste, A. Lenci, S. Sakti, N. Xue (Eds.), Pro- ings, CEUR-WS.org, 2024. ceedings of the 2024 Joint International Conference [2] F. M. Lopes, Perpetuating the patriarchy: Misog- on Computational Linguistics, Language Resources yny and (post-)feminist backlash, Philosophi- and Evaluation (LREC-COLING 2024), ELRA and cal Studies 176 (2019) 2517–2538. doi:10.1007/ ICCL, Torino, Italia, 2024, pp. 12700–12711. URL: s11098- 018- 1138- z . https://aclanthology.org/2024.lrec-main.1112. [3] M. Barreto, D. Doyle, Benevolent and hostile sexism [11] G. Abercrombie, A. Jiang, P. Gerrard-abbott, I. Kon- in a shifting global context, Nature reviews psy- stas, V. Rieser, Resources for automated identi- chology 2 (2023) 98–111. doi:https://doi.org/10. fication of online gender-based violence: A sys- 1038/s44159- 022- 00136- x . tematic review, in: Y.-l. Chung, P. R\”ottger, [4] K. Srivastava, S. Chaudhury, P. S. Bhat, S. Sahu, D. Nozza, Z. Talat, A. Mostafazadeh Davani (Eds.), Misogyny, feminism, and sexual harassment, The 7th Workshop on Online Abuse and Harms Industrial Psychiatry Journal 26 (2017) 111–113. (WOAH), Association for Computational Linguis- URL: https://journals.lww.com/inpj/fulltext/ tics, Toronto, Canada, 2023, pp. 170–186. URL: https: 2017/26020/misogyny,_feminism,_and_sexual_ //aclanthology.org/2023.woah-1.17. doi:10.18653/ harassment.1.aspx. doi:10.4103/ipj.ipj_32_18 . v1/2023.woah- 1.17 . [5] A. Muti, A. Barrón-Cedeño, UniBO @ AMI: A Multi- [12] C. Gothreau, K. Arceneaux, A. Friesen, Hostile, Class Approach to Misogyny and Aggressiveness Benevolent, Implicit: How Different Shades of Sex- Identification on Twitter Posts Using AlBERTo, in: ism Impact Gendered Policy Attitudes, Frontiers EVALITA Evaluation of NLP and Speech Tools for in Political Science 4 (2022). URL: https://www. Italian: Proceedings of the Final Workshop 12-13 frontiersin.org/articles/10.3389/fpos.2022.817309. December 2018, Naples, 2020. doi:10.3389/fpos.2022.817309 . [6] E. Fersini, D. Nozza, P. Rosso, Overview of [13] M. Polignano, P. Basile, M. de Gemmis, G. Semer- the evalita 2018 task on automatic misogyny aro, V. Basile, AlBERTo: Italian BERT Language identification (ami), in: EVALITA Evaluation Understanding Model for NLP Challenging Tasks of NLP and Speech Tools for Italian: Proceed- Based on Tweets, in: Proceedings of the Sixth ings of the Final Workshop 12-13 December 2018, Italian Conference on Computational Linguistics Naples, Torino: Accademia University Press, 2018, (CLiC-it 2019), volume 2481, CEUR, Bari, Italy, 2019. pp. 59–66. doi:doi:10.4000/books.aaccademia. URL: https://www.scopus.com/inward/record.uri? 4497 . eid=2-s2.0-85074851349&partnerID=40&md5= [7] E. Fersini, D. Nozza, P. Rosso, Ami @ evalita2020: 7abed946e06f76b3825ae5e294ffac14. Automatic misogyny identification, in: V. Basile, D. Croce, M. Di Maro, L. C. Passaro (Eds.), Pro- ceedings of the 7th evaluation campaign of Natural