<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of EXIST 2024 - Learning with Disagreement for Sexism Identification and Characterization in Tweets and Memes (Extended Overview)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Laura Plaza</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Carrillo-de-Albornoz</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Víctor Ruiz</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alba Maeso</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Berta Chulvi</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Rosso</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Enrique Amigó</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julio Gonzalo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roser Morante</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Damiano Spina</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>RMIT University</institution>
          ,
          <addr-line>3000 Melbourne</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidad Nacional de Educación a Distancia (UNED)</institution>
          ,
          <addr-line>28040 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universidad Politécnica de Valencia (UPV)</institution>
          ,
          <addr-line>46022 Valencia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Valencian Graduate School and Research Network Analysis of Artificial Analysis (ValgrAI)</institution>
          ,
          <addr-line>46022 Valencia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <fpage>1</fpage>
      <lpage>3</lpage>
      <abstract>
        <p>In recent years, the rapid increase in the dissemination of ofensive and discriminatory material aimed at women through social media platforms has emerged as a significant concern. This trend has had adverse efects on women's well-being and their ability to freely express themselves. The EXIST campaign has been promoting research in online sexism detection and categorization in English and Spanish since 2021. The fourth edition of EXIST, hosted at the CLEF 2024 conference, consists of three groups of tasks, continuing from EXIST 2023: sexism identification , source intention identification , and sexism categorization. However, while EXIST 2023 focused on processing tweets, the novelty of this edition is that the three tasks are also applied to memes, resulting in a total of six tasks. To address disagreements in the labeling process, the “learning with disagreement” paradigm is adopted. This approach promotes the development of equitable systems capable of learning from diferent perspectives on the sexism phenomenon. The 2024 edition of EXIST has surpassed the success of previous editions, with the participation of 57 teams submitting 412 runs. This extended lab overview describes the tasks, dataset, evaluation methodology, participant approaches, and results. Additionally, it highlights the advancements made in understanding and tackling online sexism through more diverse data sources and innovative methodologies. Finally, it briefly introduces future intended work for next editions of EXIST.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;sexism identification</kwd>
        <kwd>sexism categorization</kwd>
        <kwd>learning with disagreement</kwd>
        <kwd>memes</kwd>
        <kwd>data bias</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        EXIST (sEXism Identification in Social neTworks) is a series of scientific events and shared tasks on
sexism identification in social networks. The editions of 2021 and 2022 [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ], celebrated under the
umbrella of the IberLEF forum, were the first in proposing tasks focusing on identifying and classifying
online sexism in a broad sense, from explicit and/or hostile to other subtle or even benevolent expressions.
The 2023 edition [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] took place as a CLEF Lab and added a third task consisting in determining the
intention of the author of sexist messages with the aim of understanding the purpose behind people
posting sexist messages on social networks.. Additionally, the main novelty of the 2023 edition was
the adoption of the “Learning with Disagreements” (LwD) paradigm [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for the development of the
dataset and for the evaluation of the systems. In the LwD paradigm, models are trained to handle
and learn from conflicting or diverse annotations so that diferent annotators’ perspectives, biases, or
interpretations are taken into account. This approach fits the findings of our previous work that showed
that the perception of sexism is strongly dependent on the demographic and cultural background of the
individual. Adopting this paradigm was a distinguishing feature in comparison to the SemEval-2023
Shared Task 10: “Explainable Detection of Online Sexism” [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        EXIST 2024, organised also as a CLEF Lab, aims to continue contributing datasets and tasks that help
developing applications to combat sexism on-line, as a form of hate on-line. This edition embraces also
the LwD paradigm and, as novelty, incorporates three new tasks that center around memes. Memes
are images that are spread rapidly by social networks and Internet users. While by nature memes are
humorous, there is a growing tendency to use them for harmful purposes, as an strategy to conceal hate
speech by combining stylistic devices of humour [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], since people tolerate humorously communicated
prejudices better than explicit irrespectful remarks [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Thus, memes contribute to spreading derogatory
humour and to strengthen preexisting prejudices and maintaining hierarchies between social groups [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
As Gasparini et al. indicate [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], misogyny and sexism against women are widespread attitudes within
the social media communities, reinforcing age-old patriarchal establishments of baseless name-calling,
objectifying their appearances, and stereotyping gender roles. By including sexist memes in the EXIST
2024 dataset, we aim to encompass a broader spectrum of sexist manifestations in social networks and
to contribute to the development of automated multimodal tools capable of detecting harmful content
targeting women.
      </p>
      <p>
        Meme detection has also been the focus of other competitions. The SemEval-2022 Task 5: “Multimedia
Automatic Misogyny Identification” [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] focused on the detection of misogynous memes on the web in
English and proposed two tasks: recognising whether a meme is misogynous or not and recognising
types of misogyny in memes. The shared task on “Multitask Meme Classification - Unraveling
Misogynistic and Trolls in Online Memes” [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] consisted in classifying misogynistic content and troll memes,
focusing specifically on memes in Tamil and Malayalam languages. The originality of EXIST lies in
that the languages addressed are English and Spanish, it introduces also the task on source intention
recognition and it adopts the LwD paradigm.
      </p>
      <p>In the following sections, we provide comprehensive information about the tasks, the dataset, the
evaluation methodology, the results and the diferent approaches of the teams that participated in the
EXIST 2024 Lab. The competition features six distinct tasks: sexism identification, source intention
classification, and sexism categorization, both in tweets and in memes. A total of 148 teams from 32 diferent
countries registered to participate. Ultimately, we received 412 results from 57 teams. Interestingly,
a significant number of teams leveraged the diverse labels representing various demographic groups
and provided soft labels as the outputs of their systems. Their results showcase the efectiveness and
advantages of employing the LwD paradigm in our specific domain: sexism detection and categorization
in social networks.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Tasks</title>
      <p>The 2024 edition of EXIST feature six tasks, which are described below. The languages addressed are
English and Spanish, and the datasets are collections of tweets and memes. For the tasks on memes, all
the partitions of the dataset are new, whereas for the tasks on tweets we employ the EXIST 2023 dataset.</p>
      <sec id="sec-2-1">
        <title>2.1. Task 1: Sexism Identification in Tweets</title>
        <p>This is a binary classification task where systems must decide whether or not a given tweet expresses
ideas related to sexism in any of the three forms: it is sexist itself, it describes a sexist situation in which
discrimination towards women occurs, or criticizes a sexist behaviour. The following statements show
examples of sexist and not sexist messages, respectively.</p>
        <p>• Sexist. The tweet is sexist or describes or criticizes a sexist situation.</p>
        <p>(1)</p>
        <sec id="sec-2-1-1">
          <title>Woman driving, be careful!.</title>
          <p>(2) It’s less of #adaywithoutwomen and more of a day without feminists, which, to be quite honest,
sounds lovely.
(3) I’m sorry but women cannot drive, call me sexist or whatever but it is true.
(4)</p>
          <p>You look like a whore in those pants" - My brother of 13 when he saw me in a leather pant
• Not sexist. The tweet is not sexist, nor it describes or criticizes a sexist situation.
(5) Just saw a woman wearing a mask outside spank her very tightly leashed dog and I gotta say I
love learning absolutely everything about a stranger in a single instant.
(6)
(7)
(8)</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Where all the white women at?.</title>
          <p>The shocking video of a woman at the wheel who miraculously escapes an assassination
attempt.</p>
          <p>Don’t my arguments convince you? Let’s try to debate. Do you use "feminazi"? You stay alone.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Task 2: Source Intention in Tweets</title>
        <p>This task aims to categorize the message according to the intention of the author. This distinction will
allow us to diferentiate sexism that is actually taking place online from sexism which is being sufered
by women in other situations but that is being reported in social networks with the aim of complaining
and fighting against sexism. We propose the following ternary classification of tweets:
• Direct sexist message. The intention is to write a message that is sexist by itself or incites to be
sexist, as in:
(9)</p>
        <p>A woman needs love, to fill the fridge, if a man can give this to her in return for her services
(housework, cooking, etc), I don’t see what else she needs.</p>
        <p>(10) Women shouldn’t code...perhaps be influencer/creator instead...it’s their natural strength .
• Reported sexist message. The intention is to report and share a sexist situation sufered by a
woman or women in first or third person, as in:
(11) I doze in the subway, I open my eyes feeling something weird: the hand of the man sat next to
me on my leg #SquealOnYourPig.</p>
        <p>(12) Today, one of my year 1 class pupils could not believe he’d lost a race against a girl.
• Judgemental message. The intention is to condemn sexist situations or behaviours, as in:
(13) As usual, the woman was the one quitting her job for the family’s welfare...</p>
        <p>(14) 21st century and we are still earning 25% less than men #Idonotrenounce.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Task 3: Sexism Categorization in Tweets</title>
        <p>Many facets of a woman’s life may be the focus of sexist attitudes including domestic and parenting
roles, career opportunities, sexual image, and life expectations, to name a few. According to this, each
sexist tweet must be assigned one or more of the following categories:
• Ideological and inequality. It includes messages that discredit the feminist movement. It also
includes messages that reject inequality between men and women, or present men as victims of
gender-based oppression. Some examples of this type of sexism discourse can be found in the
following tweets:
(15) #Feminism is a war on men, but it’s also a war on women. It’s a war on female nature, a war
on femininity.
(16) Think the whole equality thing is getting out of hand. We are diferent, that’s how were made!
(17) OK. Whew, that’s good. Get a real degree and forget this poison of victimhood known as
feminism.
• Stereotyping and dominance. It includes messages that express false ideas about women that
suggest they are more suitable or inappropriate for certain tasks, and somehow inferior to men.
Some examples of this type of sexism discourse can be found in the following tweets:
(18) Most women no longer have the desire or the knowledge to develop a high quality character,
even if they wanted to.
(19) I feel like everytime I flirt with a girl they start to imagine all the ways they can utilize me .
(20) Woman driving, be careful!
• Objectification . It includes messages where women are presented as objects apart from their
dignity and personal aspects. We also include messages that assume or describe certain physical
qualities that women must have in order to fulfill traditional gender roles. Some examples of this
type of sexism discourse can be found in the following tweets:
(21) No ofense but I’ve never seen an attractive african american hooker. Not a single one.
(22) Most of the times I just want women for sex just because everytime I try to make a relationship
out of it they always break my heart.
(23) Don’t get married than blame all woman for your poor investment. You should of got a hooker
but instead you choose to go get a wedding ring.
• Sexual violence. It includes messages where sexual suggestions, requests or harassment of a
sexual nature (rape or sexual assault) are made. Some examples of this type of sexism discourse
can be found in the following tweets:
(24) I wanna touch your tits..you can’t imagine what I can do on your body.
(25) Fuck that cunt, I would with my fist.
(26) Now I understand why woman that get no action are miserable. Semen retention almost drove
me insane.
• Misogyny and non sexual violence. It includes expressions of hatred and violence towards
women. Some examples of this type of sexism discourse can be found in the following tweets:
(27) Domestic abuse is never okay. . . Unless your wife is a bitch.
(28) Some woman are so toxic they don’t even know they are draining everyone around them in
poison.</p>
        <p>(29) I hate misogyny more than I hate women.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Task 4: Sexism Identification in Memes</title>
        <p>As in Task 1, this is a binary classification consisting on deciding whether or not a meme is sexist.
Figure 1 shows examples of sexist and non sexist memes, respectively.</p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5. Task 5: Source Intention in Memes</title>
        <p>As in Task 2, this task aims to categorize the meme according to the intention of the author. Due to the
characteristics of the memes, we barely found examples of memes within the “reported” category, so
this category was not considered. As a result, in this task systems should only classify memes in two
classes: “direct” or “judgemental”, as shown in Figure 2.</p>
      </sec>
      <sec id="sec-2-6">
        <title>2.6. Task 6: Sexism Categorization in Memes</title>
        <p>This task aims to classify sexist memes according to the categorization provided for Task 3: (i) ideological
and inequality, (ii) stereotyping and dominance, (iii) objectification, (iv) sexual violence and (v) misogyny
and non-sexual violence. Figure 3 shows one meme of each category.</p>
        <p>(a) Sexist meme
(b) Non sexist meme</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset</title>
      <p>
        The EXIST 2024 dataset comprises two types of data: the tweets from the EXIST 2023 dataset and a
completely new dataset of memes. Plaza et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] provide a detailed description of the EXIST 2023 tweet
dataset. Here, we briefly describe the process followed to curate the meme dataset.
      </p>
      <p>Since we adopt the LwD paradigm, we provide all labels assigned by the diferent annotators to allow
systems to learn from conflicting and subjective information. This paradigm not only proved to improve
the systems’ accuracy, robutness and generalizability, but also helped to mitigate bias.</p>
      <sec id="sec-3-1">
        <title>3.1. Data Sampling</title>
        <p>We first curated a lexicon of terms and expressions leading to sexist memes. The set of seeds encompasses
diverse topics and contains 250 terms, with 112 in English and 138 in Spanish. The terms were used as
search queries on Google Images to obtain the top 100 images. Rigorous manual cleaning procedures
were applied, defining memes and ensuring the removal of noise such as textless images, text-only
images, ads, and duplicates. The final set consists of more than 3,000 memes per language.</p>
        <p>Since the proportion of memes per term was heterogeneous, we discarded the most unbalanced seeds
and made sure that all seeds have at least five memes. To avoid introducing selection bias, we randomly
selected memes, ensuring the appropriate distribution per seed. As a result, we have 2,000 memes per
language for the training set and 500 memes per language for the test set.</p>
        <p>(a) Ideological &amp;
inequality</p>
        <p>(b) Objectification
(c) Stereotyping &amp;
dominance
(d) Sexual
violence
(e) Misogyny &amp; non-sexual
violence</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Dealing with Label Bias</title>
        <p>We have considered some sources of “label bias”, that may be introduced by the socio-demographic
diferences of the persons that participate in the annotation process, but also when more than one
possible correct label exists or when the decision on the label is highly subjective. In particular, we
consider two sociodemographic parameters: gender (MALE/FEMALE) and age (18–22/23–45/+46 y.o.).
Each meme was annotated by six annotators selected through the Prolific crowdsourcing platform.
No personally identifiable information about the crowd workers was collected. Crowd workers were
informed that the tweets could contain ofensive information and were allowed to withdraw voluntarily
at any time. Full consent was obtained.</p>
        <p>Also, as a new feature in the datasets, both of 2023 and 2024, we have reported three additional
demographic characteristics of annotators: level of education, ethnicity, and country of residence.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Learning with Disagreement</title>
        <p>The assumption that natural language expressions have a single and clearly identifiable interpretation
in a given context is a convenient idealization, but far from reality, especially in highly subjective
task as sexism identification. The learning with disagreements paradigm aims to deal with this by
letting systems learn from datasets where no gold annotations are provided but information about the
annotations from all annotators, in an attempt to gather the diversity of views. Following methods
proposed for training directly from the data with disagreements, instead of using an aggregated label,
we will provide all annotations per instance for the six diferent strata of annotators.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Lab Setup and Participation</title>
      <p>In this section, we provide a concise overview of the approaches presented at EXIST 2024. For a
comprehensive description of the systems, please refer to Section 6 and to the participants’ papers.</p>
      <p>Although 148 teams from 32 diferent countries registered for participation, the number of participants
who finally submitted results were 57, submitting 412 runs. Teams were allowed to participate in any of
the six tasks and submit hard and/or soft outputs. Table 1 summarizes the participation in the diferent
tasks and evaluation contexts.</p>
      <p>The evaluation campaign started on March 4, 2024 with the release of the training set. The test set
was made available on April 15. The participants were provided with the oficial evaluation script. Runs
had to be submitted by May 10. Each team could submit up to three runs per task.</p>
      <p>
        A wide range of approaches and strategies were used by the participants. Nearly all participant
systems utilized large language models, both monolingual and multilingual. Most employed LLMs
include BERT, DistilBERT, MarIA, MDEBERTA, RoBERTa, DeBERTa, Llama, and GPT-4. For processing
memes, popular vision models were employed: CLIP, BEIT and VIT. Some teams employed ensembles
of multiple models to enhance the overall performance. A couple of teams made use of knowledge
integration to combine diferent language models with language features. Data augmentation techniques
were used by several teams. Prompt Engineering was also used to adapt pre-trained models to the
sexism detection task. Only two teams utilized deep learning architectures such as BiLSTM and CNN,
while four teams opted for traditional machine learning methods, including SVM, Random Forest, and
XGBoost, among others. As in EXIST 2023 [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], Twitter-specific models where employed, such as
Twitter-RoBERTa and Twitter-XML-RoBERTa.
      </p>
      <p>While 174 systems took advantage of the multiple annotations available and provided soft outputs,
238 followed the traditional approach of providing only hard labels as outputs. Textual tasks received
greater engagement, although participation is also high in the tasks on memes. The binary classification
tasks had more participants, followed by mono-label tasks, and finally, multi-label tasks, which is due
to the increasing dificulty of these tasks.</p>
      <p>For each of the six tasks, the organization also provided two diferent baseline runs:
• EXIST2024 majority, a non-informative baseline that classifies all instances as the majority
class.
• EXIST2024 minority, a non-informative baseline that classifies all instances as the minority
class.
• The evaluation metrics for the the gold standard (EXIST2024 gold) are also provided, in order to
set the upper bound for the ICM metrics.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Evaluation Methodology and Metrics</title>
      <p>As in EXIST 2023, we have carried out a “soft evaluation” and a “hard evaluation”. The soft
evaluation relates to the LwD paradigm and is intended to measure the ability of the model to capture
disagreements, by considering the probability distribution of labels in the output as a soft label and
comparing it with the probability distribution of the annotations. The hard evaluation is the standard
paradigm and assumes that a single label is provided by the systems for every instance in the dataset.</p>
      <p>From the point of view of evaluation metrics, the tasks can be described as follows:
• Tasks 1 and 4 (sexism identification): binary classification, monolabel.
• Tasks 2 and 5 (source intention): multiclass hierarchical classification, monolabel. The hierarchy
of classes has a first level with two categories, sexist/not sexist, and a second level for the sexist
category with three mutually-exclusive subcategories: direct/reported/judgemental. A suitable
evaluation metric must reflect the fact that a confusion between not sexist and a sexist category
is more severe than a confusion between two sexist subcategories.
• Tasks 3 and 6 (sexism categorization): multiclass hierarchical classification, multilabel. Again
the first level is a binary distinction between sexist/not sexist, and there is a second level for
the sexist category that includes five subcategories: ideological and inequality, stereotyping and
dominance, objectification, sexual violence, and misogyny and non-sexual violence. These classes
are not mutually exclusive: a tweet may belong to several subcategories at the same time.
The LwD paradigm can be considered in both sides of the evaluation process:
• The ground truth. In a “hard” setting, the variability in the human annotations is reduced by
selecting one and only one gold category per instance, the hard label. In a “soft” setting, the
gold standard label for one instance is the set of all the human annotations existing for that
instance. Therefore, the evaluation metric incorporates the proportion of human annotators that
have selected each category (soft labels). Note that in Tasks 1, 2, 4 and 5, which are monolabel
problems, the sum of the probabilities of each class must be one. But in Task 3, which is multilabel,
each annotator may select more than one category for a single instance. Therefore, the sum of
probabilities of each class may be larger than one.
• The system output. In a “hard”, traditional setting, the system predicts one or more categories
for each instance. In a “soft” setting, the system predicts a probability for each category, for each
instance. The evaluation score is maximized when the probabilities predicted match the actual
probabilities in a soft ground truth.</p>
      <p>
        In EXIST 2024, for each of the tasks, two types of evaluation have been performed:
1. Soft-soft evaluation. For systems that provide probabilities for each category, we perform a
soft-soft evaluation that compares the probabilities assigned by the system with the probabilities
assigned by the set of human annotators. The probabilities of the classes for each instance are
calculated according to the distribution of labels and the number of annotators for that instance.
We use a modification of the original ICM metric (Information Contrast Measure [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]), ICM-Soft
(see details below), as the oficial evaluation metric in this variant and we also provide results for
the normalized version of ICM-Soft (ICM-Soft Norm).
2. Hard-hard evaluation. For systems that provide a hard, conventional output, we perform a
hard-hard evaluation. To derive the hard labels in the ground truth from the diferent annotators’
labels, we use a probabilistic threshold computed for each task. As a result, for Tasks 1 and 4, the
class annotated by more than 3 annotators is selected; for Tasks 2 and 5, the class annotated by
more than 2 annotators is selected; and for Tasks 3 ad 6 (multilabel), the classes annotated by
more than 1 annotator are selected. The instances for which there is no majority class (i.e., no
class receives more probability than the threshold) are removed from this evaluation scheme. The
oficial metric for this task is the original ICM, as defined by [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. We also report a normalized
version of ICM (ICM Norm) and F1 (F1YES). In Tasks 1 and 4, we use F1 for the positive class. In
Tasks 2, 3, 5 and 6, we use the macro-average of F1 for all classes (Macro F1). Note, however, that
F1 is not ideal in our experimental setting: although it can handle multilabel situations, it does
not take into account the relationships between classes. In particular, a confusion between not
sexist and any of the sexist subclasses, and a confusion between two of the sexist subclasses, are
penalized equally.
      </p>
      <p>ICM is a similarity function that generalizes Pointwise Mutual Information (PMI), and can be used
to evaluate outputs in classification problems by computing their similarity to the ground truth. The
general definition of ICM is:</p>
      <p>
        ICM(, ) =  1() +  2() −  ( ∪ )
Where () is the Information Content of the instance represented by the set of features A. ICM
maps into PMI when all parameters take a value of 1. The general definition of ICM by [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] is applied
to cases where categories have a hierarchical structure and instances may belong to more than one
category. The resulting evaluation metric is proved to be analytically superior to the alternatives in the
state of the art. The definition of ICM in this context is:
      </p>
      <p>ICM((), ()) = 2(()) + 2(()) − 3(() ∪ ())
Where () stands for Information Content, () is the set of categories assigned to document  by
system , and () the set of categories assigned to document  in the gold standard. The score for
a perfect output (() = ()) is the gold standard Information Content ((()). The score for a
zero-information system (no category assignment) is − (()). We use these two boundaries for
normalisation purposes, truncating to 0 the scores lower than − (()).</p>
      <p>As there is not, to the best of our knowledge, any current metric that fits hierarchical multilabel
classification problems in a LwD scenario, we have defined an extension of ICM (ICM-soft) that accepts
both soft system outputs and soft ground truth assignments. ICM-soft works as follows: first, we define
the Information Content of a single assignment of a category  with an agreement  to a given instance
as the probability of instances in the gold standard to exceed the agrement level  for the category :
({⟨, ⟩}) = − log2( ({ ∈  : () ≥ })
In order to estimate , we compute the mean and deviation of the agreement levels for each class
across instances, and applying the cumulative probability over the inferred normal distribution. In the
case of zero variance, we must consider that the probability for values equals or below the mean is 1
(zero IC) and the probability for values above the mean must be smoothed. But this is not the case of
the EXIST datasets.</p>
      <p>
        Due to the multi-label and hierarchical nature of the classification task,for each classification instance,
the gold standard, the system output and their unions ((()) (()) and (()) ()) are
sets of category assignments. The union of the assignments (i.e. ()) ()) is calculated as fuzzy
sets, i.e. the maximum values., in order to estimate information content, we apply a recursive function
similar to the one described by Amigó and Delgado [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] for assignment sets and avoid the redundant
information of parent categories.
      </p>
      <p>︃( 
⋃︁{⟨, ⟩}
)︃
= (⟨1, 1⟩) + 
︃( 
⋃︁{⟨, ⟩}</p>
      <p>)︃
=2
=1
− 
︃( 
⋃︁{⟨lca(1, ), (1, )⟩}</p>
      <p>)︃
=2
(30)
where lca(, ) is the lowest common ancestor of categories  and .</p>
    </sec>
    <sec id="sec-6">
      <title>6. Overview of approaches EXIST 2024</title>
      <p>In this section, we provide a description of the approaches adopted by the participants. More detailed
information of each work is provided in the participants’ working notes.</p>
      <p>
        Team FraunhoferSIT [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] utilized a stacking ensemble of machine learning models. For predicting
hard labels, they used an ensemble of classification models: Multinomial Naive Bayes, Stochastic
Gradient Descent, Decision Tree, k-Nearest Neighbors, Logistic Regression and Extra Trees. For
predicting soft labels, they used an ensemble of regression models: Random Forest, Gradient Boosting,
Stochastic Gradient Descent and AdaBoost. They employed two augmentation methods at word level:
synonym replacement using WordNet for English tweets and contextual augmentation with three
transformer models: BERTIN, ALBERT Base Spanish, and RoBERTuito. They participated in the first
three tasks.
      </p>
      <p>Team frms [15] participated in Tasks 1, 2 and 3, for both hard and soft evaluations. For all three
tasks, their first run used BERT multilingual, and their second run used XLM-RoBERTa model. In the
third task, they created an ensemble of BERT and XLM-RoBERTa combining the predictions from both
models. For Task 1, RoBERTa model performed the best in the hard evaluation, but, for Task 2, BERT
performed better for both hard and soft evaluations. For Task 3, the ensemble obtained the best results
in the hard evaluation and tied with RoBERTa in the soft evaluation.</p>
      <p>Team RMIT-IR [16] proposed diferent approaches to Tasks 1-3 and Task 4. For Tasks 1–3 (on
tweets), they studied the efectiveness of zero-shot In-Context Learning (ICL) with of-the-shelf
pretrained Large Language Models (LLMs). Their approaches for meme classification (Task 4) utilize CLIP
(Contrastive Language-Image Pre-training) to experiment with multi-modal embeddings and zero-shot
sexism identification models. They participated on hard and soft evaluations, obtaining soft labels by
generating six answers and calculating the proportions. The annotator’s genders and/or study levels
were included in some runs for the first three tasks. For Task 4, they used three systems TI-CLIP
(feedforward network), TIMV-CLIP (Transformer encoder), and Prompt-CLIP (zero-shot). TIMV-CLIP
(Transformer encoder) stands out by its performance, especially in the memes in English dataset, where
it achieves the best result in soft evaluation with a ICM-Soft Norm score of 0.4998.</p>
      <p>Team dap-upv [17] proposed hard labels in Tasks 4 and 6. For Task 4, they utilized a two-stage
approach by fine-tuning the Contrastive Language-Image Pre-training (CLIP) model followed by a
classifier. A diferent classifier was tested per run, where the best one was Light Gradient Boosting
Machine (LightGBM) reaching a 0.72 in F1. For Task 6, they fine-tuned both a RoBERTa model for
text, a Google Vision Transformer (ViT) for images, and they concatenated them training a LightGBM
classifier on the concatenated embeddings. The ensemble notably improved results, obtaining 0.49 in
F1.</p>
      <p>Team Aditya [18] only participated in Task 1 submitting hard labels. They preprocessed tweets by
removing any emoji, URLs or mentions. They used XLM-RoBERTa fine-tuned on the dataset, and their
three runs difer on the number of epochs of training and whether they used development in training
too. Their best model, trained with 12 epochs, ranked 14th with F1 of 0.7691.</p>
      <p>Team BAZI [19] fine-tuned various transformer models, some monolingual in English and Spanish
and some multilingual, and they chose XLM-RoBERTa as the best performing of them. With this
method, they provided hard labels and soft labels. Hard labels were the direct outputs from the model,
and soft labels were obtained by adding a softmax function to the last layer. Their approach using
XLM-RoBERTa achieved 4th place in the soft evaluation for Task 1, and 2nd place for Task 2. Also, they
employed few-shot learning with GPT-3.5 with three examples in English and three in Spanish from
the training set, providing hard labels with this method.</p>
      <p>Team mc-mistral_2 [20] submitted hard labels for Task 1. Their approach leveraged a Mistral 7B
model along with a few-shot learning strategy and prompt engineering to address the task in the hard
labelling setup. They translated cases in Spanish to English with an online and real-time use of Google
Translator from the deep_translator library. Then they randomly selected 10 samples from the provided
labelled training set and formatted the samples of test set as: Tweet1 // NO, Tweet2 // YES. In the global
ranking, they achieved a F1 score of 0.51.</p>
      <p>
        Team CNLP-NITS-PP [21] presented two systems to cover all tasks in EXIST 2024. The model
for textual modalities is a Convolutional Neural Network - Bidirectional Long Short-Term Memory
(CNN-BiLSTM) and it is used in Tasks 1, 2 and 3, and to characterize textual data in Tasks 4, 5 and
6. Texts are tokenized into words utilizing GloVe, they are lowercased, and common stopwords are
deleted. For memes, a combination of Residual Network 50 (ResNet-50), used to analyze images, and
text-based analysis is utilized. Images are resized to 224x224 pixels, and pixel values are normalized
to the interval [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ]. To enhance model robustness, data augmentation techniques such as random
rotation, flipping, and colour jitter are applied. Hyperparameter tuning is conducted via grid search and
k-fold cross-validation. A single run was submitted for every task in both hard and soft labels. Results
stand out in Task 5, where they reached the 5th position in the ranking.
      </p>
      <p>Team Awakened [22] studied the most appropriate way of assembling transformer models by
comparing performances of diferent models and assigning diferent weights in the ensemble to the best
model. They used both English models and multilingual models, and both general language models (like
DistilBERT of xlm-RoBERTa), and domain-specific models (like twitter-xlm-roberta-base-sentiment or
roberta-hate-speech-dynabench-r4). They tested three ways of assembling models: assigning half of
the weight to the most dominant model, assigning 75% of the weight, and assigning all the weight to
the most dominant model. Results showed that, in most of the tasks, the best model was the one in
which the most dominant was assigned a 75% of the weight.</p>
      <p>The NYCU-NLP team [23] employed extensive data preprocessing techniques, which included
removing irrelevant elements, standardizing text formats, back-translation using the Google Translator
API, and implementing the AEDA method for text augmentation. Additionally, they adapted the Round
to Closest Value approach to handle non-continuous annotation values. The system relies on two
transformer-based language models: DeBERTa-v3 and xlm-RoBERTa. The team integrated annotator
information such as gender, age, and ethnicity, creating a unified vector representation for each tweet.
They further incorporated Hard Parameter Sharing to optimize shared layers across tasks, improving
generalization and computational eficiency. Notably, their model achieved outstanding performance in
the EXIST 2024 challenge, securing first place in Tasks 1, 2, and 3 in the soft evaluation setting. In the
hard setting, their system ranked first in Task 1, second in Task 2, and third in Task 3.</p>
      <p>Team shm2024 [24] participated with hard labels to Tasks 1 and 2. They implemented three diferent
models. The first system was a MultiLayer Perceptron (MLP) classifier with Language Agnostic BERT
Sentence Embeddings (LaBSE). The second one was a eXtreme Gradient Boosting (XGBoost) Classifier,
and the third approach was an ensemble of CNN models. The first model was the best performing on
the test set with a ICM-Hard Norm value of 0.6623 and F1_YES value of 0.7044 for Task 1 and, for Task
2, 0.2115 for ICM-Hard Norm and 0.1200 for Macro F1.</p>
      <p>UMUTeam [25] submitted soft predictions to Tasks 1 and 2, and hard labels to the rest of the tasks.
For textual modalities, they created an ensemble of two Spanish LLMs, BETO and MarIA; and two
multilingual LLMs, deBERTa v3 and XLMTwitter. They also extracted the linguistic features (LFs) using
the UMUTextStats tool and performed hyperparameter tuning on 10 models. They tried two ways of
ensembling: Knowledge Integration (KI), and Ensemble Learning (EL). Their three runs were KI, EL
and LFs. Their results were better at Spanish than English. Their best results were obtained in Task 2,
where KI reached the 8th position. For multi modalities, they used a CLIP model to extract images and
text. Their algorithm was formed by an Image Encoder (CLIP image encoder), a Text Encoder (CLIP
text encoder), Diagonal multiplication and a Classification head. Their results were near the baselines.</p>
      <p>The 3 Musketeers [26] team applied traditional classification machine learning algorithms such as
Logistic Regression (LR), Random Forest (RF) and Support Vector Machine (SVM), and BERT transformer
model to the EXIST Task 1. They followed a pipeline of preprocessing, including lowercasing, removing
punctuation, emoticons, links, mentions and stopwords. They lemmatized words and used TF-IDF
vectorization. They used GridSearch to optimize hyperparameters. Their best model was an SVM that
got a F1 score of 0.6299.</p>
      <p>Team TextMiner [27] submitted hard labels to Task 1. They preprocessed the text to standardize it
and text embeddings were created using TF-IDF. They performed feature selection of word n-grams and
character n-grams. They explored a diverse range of classifiers: Random Forest Classifier, Extra Trees
Classifier, LightGBM Classifier, AdaBoost Classifier, Bernoulli Naive Bayes, Support Vector Classifier.
They experimented with diferent combinations of hyperparameters, and created three ensembles of
models: one with the 10 best models, another one with the 50 best models, and the last one with the
100 best models. The best model was the one with the top 50 models, coming 39th place.</p>
      <p>The team Victor-UNED [28] participated in every task of EXIST 2024. Their system employed a
concatenation of models based on the predictions of the level of agreement of the instances. They
trained various transformer based models on the textual training dataset to generate predictions to
Tasks 1 and 4. Models with more information shown in their learning phase were used to determine soft
labels of instances with lower level of agreement. Models were trained with soft labels, and hard labels
were obtained from the resulting soft labels. For Tasks 2 &amp; 5, and Tasks 3 &amp; 6, results from previous
tasks were incorporated, taking into account the hierarchical nature of the challenge. First run was
mDeBERTa-v3 trained on this year’s dataset, second run was the concatenated models to determine low
agreement cases, and third run took into account an annotator ensemble to distinguish low agreement
cases. Their approach achieved top rankings in Tasks 4 and 5, and was one of the most consistent
models among all tasks.</p>
      <p>Team PINK [29] participated in Task 4 on sexism detection in memes. They proposed a unified,
multimodal Transformer-based architecture capable of dealing with multiple languages, namely English and
Spanish. Their architecture extracts high-level features using large-scale, pre-trained models that are
kept frozen during training. These features are then normalized and projected into the same dimensional
space. They are then conditioned based on the language of the sample and its modality before being
processed by a Transformer encoder backbone. The final classification is predicted through average
pooling and a linear projection. With this architecture, they created three types of systems: Single
Models, Majority Voting Ensembles (MVE) and Average Probability Ensembles (APE). Each of these
systems was used to obtain a set of predictions, both hard and soft labels. Their approach reached the
10th and 20th places in the final ranking for soft- and hard-label evaluations, respectively.</p>
      <p>Team RoJiNG-CL [30] investigated using large language models (LLMs), specifically GPT-4, to
extract textual descriptions from images in Task 4. They obtained these descriptions from GPT-4’s
zero-shot prompting, with textual prompts alongside memes. They integrated these descriptions with
related texts to fine-tune both monolingual and multilingual models, enhancing their ability to identify
sexist content in memes with hard labels. They experimented with several transformer-based models
and their hyperparameter’s optimization with Optuna. The first run used BERT fine-tuned in English
data, and BETO fine-tuned in Spanish data. The second run used mDeBERTa, and the third run used
GPT-4 output results. Their submissions secured the top three positions on the hard-hard evaluation
leaderboard, encompassing both English and Spanish instances. The GPT-4 based predictions emerged
as the most efective, delivering top results in a zero-shot setting.</p>
      <p>Team NICA [31] participated in tasks from 1 to 5. For textual models, they used various multilingual
transformer models to detect sexism in English and Spanish tweets. Their runs worked with
xlmRoBERTa-Large-Twitter, multilingual BERT, and multilingual DistilBERT. First, they preprocessed texts,
focusing on eliminating tags and URLs in the tweets. Their experiments showed that BERT outperformed
other models, however xlm-RoBERTa-Large-Twitter showed notable better results in the test set. For
tasks 4 and 5, they employed the CLIP model, which leverages both image and corresponding text data
to identify sexist elements. CLIP performance yielded promising results, reaching the 9th position in
the soft evaluation in Task 4, and the 4th position in both hard and soft evaluations in Task 5.</p>
      <p>Team Mind [32] presented an approach for detecting sexism in memes in Task 4. In their approach,
they used ResNet50 as image encoder and m-BERT to create the text embeddings, fine-tuned on EXIST
2024 dataset. Once they had the encodings, they applied a projection layer for dimensionality reduction
and feature transformation on input vectors. In order to combine the image and text features and get
concatenated data, they used the feature interaction matrix (FIM). Then, they trained a contrastive
learning-based model on these embeddings. They computed the cosine similarity between each test
sample and all the training samples. They used the K-Nearest Neighbors (KNN) algorithm to select
the 10 training embeddings with the highest cosine similarity to each sample. The performance of
the model achieved ICM scores of 0.2778 for English, 0.2152 for Spanish, and 0.2465 for the combined
dataset.</p>
      <p>Team DiTana-PV [33] focused on hard evaluation of Tasks 4 and 6. Their objectives were to evaluate
the efect of machine translation on model performance and explore data augmentation techniques.
They automatically translated Spanish data to English and leveraged a trained version of the BERTweet
model, BERTweet-large-sexism-detector, nfie-tuned in the dataset of SemEval-2023 Task 10 EDOS. Then
they used data augmentation techniques to increase the amount of data and reduce the imbalance. They
used BERT contextual embeddings for paraphrasing the words in the original text. One separate model
was trained for each language. For Task 4, runs included models that added a weighted-loss function
and a weighted-loss function plus data augmentation. The one that performed the best was the model
with weighted-loss function and no data augmentation. For Task 6, their models tried to predict 5 labels
or 6 labels at a time. Results showed that the model predicting 5 labels performed better.</p>
      <p>Team Atresa-I2C-UHU [34] participated in Tasks 4 and 5 submitting both hard and soft labels. They
focused on working with perspectives and Learning with Disagreement. They trained the multilingual
versions of BERT and RoBERTa with diferent hyperparameters to analyze their efect in each perspective
with enough values (gender, age, level of studies, Bachelor’s High school White). They chose the
versions of BERT and RoBERTA that worked best for every perspective, and combined them. Data was
preprocessed and translated to generate supplementary training datasets. Final runs were combinations
of BERT models or combinations of BERT and Roberta models to deal with hard and soft evaluations
in each task. For Task 4, they ranked 4th, with ICM-Hard and ICM-Soft scores of 0.5668 and 0.4476,
respectively. For Task 5, they secured 2nd and 10th places with ICM-Hard and ICM-Soft scores of 0.4119
and 0.2023, respectively.</p>
      <p>Team CAU&amp;ITU_2 [35] investigated a broad range of models, including traditional machine learning
methods, such as ensemble models and probability-based model like Random Forest and XGboost,
and deep learning architectures models with the use of multiligual BERT. In order to obtain vector
representations of texts, they implemented BOW and TF-IDF and the machine learning methods were
trained with both alternatives. Hyperparameter fine-tuning was performed with RandomizedSearchCV.
These methods were explored for binary classification in Task 1, and multi-class classification in Task 2.
Alternatively, multilingual BERT performed better than the rest of the models in Tasks 1 and 2. For
Task 3, only multilingual BERT was evaluated. Every experiment was evaluated with hard labels.</p>
      <p>Participation of team I2C-UHU_2 [36] in Tasks 1 and 2 aimed to employ Learning with Disagreement
techniques and explore annotators’ perspectives to obtain more robust models. Firstly, they preprocessed
data with cleaning and normalization techniques. Then they applied data augmentation strategies and
hyperparameter optimization with Optuna. They explored diferent transformer-based models. For
Task 1, the first run used XLM RoBERTa Base to predict all instances, and the second run separated
DeBERTa v3 Base for the English dataset and RoBERTa Base BNE for the Spanish dataset. For Task
2, they submitted three runs: the first one used XLM RoBERTa Base trained with the whole dataset,
the second one implemented an ensemble of XLM RoBERTa Base for every annotator group, and the
third one was an ensemble of XLM RoBERTa Base for every age group. In Task 1, they secured the 10th
position in the hard evaluation, and the 13th position in the soft evaluation. In Task 2, they achieved
the 11th position for the hard evaluation, and the 17th position for the soft evaluation.</p>
      <p>Team maven [37] submitted hard labels to textual Tasks 1, 2 and 3. They focused on creating a
stacking classifier composed by an ensemble of four LLMs. They built the ensemble calculating the
highest Complementary Error Correction. The stacking classifier was based on LightGBM, whose
parameters were fine-tuned with Optuna, and fed with the output scores of the four LLMs. They created
two datasets from the original EXIST dataset by translating all data from English to Spanish, and from
Spanish to English. They performed preprocessing: lowercasing and eliminating mentions, hashtags,
links, numerals, etc. From the four LLMs, two of them were trained with English data (DistilBERT +
RoBERTa), and the other two with Spanish data
(somosnlp-hackathon-2022/twitter-sexismo-finetunedrobertuito-exist2021 and annahaz/xlm-roberta-base-misogyny-sexism-indomain-mix-bal). For Tasks 2
and 3, their system was based on BERT. In Task 1, the best model obtained a F1-score of 0.7359, 0.4563
for Task 2, and 0.4491 for Task 3.</p>
      <p>Team CIMAT-CS-NLP [38] participated in Task 1, with hard and soft labels, and in Task 2 with
hard labels. The proposed methods for both tasks are based on unifying the knowledge of two
different systems: zero-shot classification with LLMs through prompting, and supervised fine-tuning of
multilingual transformer models. The zero-shot classification was performed with the Gemini API
(gemini-1.0-pro). Four kinds of results were taken into account, depending on the type of the prompt
engineering processed. On the other hand, three types of transformer models were fine-tuned on the
dataset, namely XLM-RoBERTa, mBERT and Twitter-XLM-Roberta. With seven types of results (one per
prompt and model), hard labels were obtained by three methods: creation of new input for fine-tuning,
proportion of votes and Best LLM response, or best fine-tuned model. Soft labels were obtained by
proportion of votes. The best system achieved the 3rd place in the hard evaluation for all tweets with a
F1 (positive class) of 0.7899. The highest ranked model for soft labels was in 5th place, and the two best
results obtained for Task 2 were ranked 7th and 8th.</p>
      <p>Team Medusa [39] adressed Task 3, sexism categorization in tweets, with hard and soft labels. They
aimed to study the performance of two architectural archetype: Classifier Chain and Binary Relevance.
The binary relevance architecture assumes that each label is independent of the others and can therefore
be treated separately. In the classifier chain architectures, classifiers are chained together so that
predictions from individual labels become features for other classifiers. These architectures constitute
the multi-label classifier head situated at the top of the pretrained models from the XML-RoBERTa
family. They trained various models and selected the Best BR, Best Chain, and “Best ICM-Soft”, that is,
the models with less BCE loss on the validation set. Their models achieved 4th, 5th, and 6th positions
in the soft evaluation ranking.</p>
      <p>Team Penta ML [40] participated with soft labels in Tasks 4 and 5, and with hard labels in Tasks 4, 5
and 6. They presented a multimodal architecture with five diferent components: (i) A Pretrained
VisionLanguage (ViLT) Model, which employs BERT as the text processor and ViT (Vision Transformer) as the
image processor; (ii) Semantics from Pooled Representations; (iii) Attention Enhanced Context Vector
for each Modality; (iv) Modality Fusion and (v) Classification Head, based on Cross Entropy loss. Then,
they experimented with CLIP and VILT as their baseline models. They concatenated representations
from images and texts and passed them to an MLP for classification. Since ViLT outperformed CLIP, it
was used as the model in their approach. Results showed that ViLT alone obtained the best results in
ICM metrics, however, their approach achieved better performance in Macro F1 in Task 6.</p>
      <p>Team CIMAT-GTO [41] only participated in Task 1 with hard labels. They explored the reasoning
capabilities of Llama 3 in a two-step process. In the first stage, they generated “reasoning” texts using a
LLM that aims to understand the tweets’ nature. These rationales are added to the tweets and, in the
second stage, they are processed further with a pre-trained XLM-RoBERTa model trained on multilingual
tweets. They diferentiated between various types of reasoning: positive and negative reasoning and
comparative reasoning. They included answers from the LLM (Llama 3) to questions about sexism
identification. They processed these answers with a multi-layer FFN and then concatenated with the
text. The models corresponded to a RoBERTa model with negative reasoning, an ensemble of models
with negative and comparative reasoning and an ensemble of models with negative, comparative and
answering reasoning. The latter was the best performing model, reaching the 4th position in the hard
evaluation ranking for Task 1.</p>
      <p>Team EquityExplorers [42] submitted two runs of hard labels to Task 1. These two runs corresponded
to the results of their two approaches: the Dual-Transformer Fusion Network (DTFN) and the Multimodel
Fusion Ensemble (MFE). The DTFN is based on the fusion of two Transformer models,
RoBERTaLarge and DeBERTa-V3-Large. This ensemble model leverages the distinctive characteristics of each
constituent model to enhance text classification. MFE is a more complex approach based on the ensemble
of LLMs (Mistral-7b, RoBERTa-Large and DeBERTa-V3-Large), and the DTFN using a majority voting
mechanism. Evaluation showed that these methodologies significantly outperform existing models,
with MFE and DTFN ranking 1st and 2nd, respectively, in the English segment, and 4th and 13th in the
combined English and Spanish segments of the oficial leaderboard.</p>
      <p>Umera Wajeed Pasha [43] participated in Task 4 on sexism identification in memes. The study
starts by importing and visualizing a meme dataset, then pre-processing the images using techniques
including cropping, scaling, and normalization to get them ready for model training. A pre-trained
model called CLIP is used to extract features, and the dataset is split into training and validation sets
for memes in both Spanish and English. The collected features are used to train and assess a variety of
machine learning models, such as Logistic Regression, SVM, XGBoost, Decision Trees, Random Forest,
Neural Network, AdaBoost, and SGD. The Random Forest model performed the best out of all of them.</p>
      <p>Team VerbaNex AI [44] proposed a method to deal with Task 1 in the hard setting. They implemented
a profiling approach based on demographic factors: gender, education level, and age. This allowed to
categorize the profiles into four groups based on their likelihood of labelling messages as sexist or not
sexist. Then they performed feature extraction and trained four distinct systems based on the grouped
profiles and their responses. To address class imbalance, they used K-Fold Stratified Shufle-Split. They
incorporated the Twitter-roBERTa-base model specifically fine-tuned for sentiment analysis. This
method was evaluated using the testing profiles, achieving a F1 score of 0.745. In the evaluation phase,
their approach yielded a F1 score of 0.63.</p>
      <p>Team MMICI [45] participated in every task of the EXIST challenge. For textual tasks, they used
transformer models (“cardifnlp/twitter-roberta-base-sentiment” for English and
“pysentimiento/robertuitobase-uncased” for Spanish). They created two types of ensembles. The first one used a majority vote
from the outputs of six diferent models, one for each annotator. The second ensemble used a majority
vote from the outputs of five diferent models, focusing on gender and age. For multimodal tasks, they
utilized CLIP embeddings using a Vision Transformer (ViT) model and two types of classifiers: FNNs
and Factorization Machines. In runs 1 and 2, demographic information was represented using one-hot
encoding, whereas, in run 3, a descriptive text was created for the annotator features, from which
embeddings were extracted. For runs 2 and 3, the classifier used a FNN, and in run 3 they proposed a
Factorization Machine model. Their best performances include a 10th place in Task 1, a 15th place in
Task 2, and a 13th place in Task 3 for Spanish tweets. For memes, they achieved a 3rd place in Task 4
for English.</p>
      <p>Team Penta-nlp [46] participated in Tasks 1 to 3. They explored multiple approaches: Machine
Learning (ML), Deep Learning (DL) and Transformer-based Pretrained Models. The ML models included
Support Vector Machine, Random Forest, XGBoost (Tasks 1 &amp; 2) and Logistic Regression (Task 3).
The DL models leverage both LSTM and LSTM + Attention models. Transformer models explored
XLM-RoBERTa, mBERT and BETO. They conducted experiments using various preprocessing methods:
removing usernames, URLs, punctuation and emojis. It was shown that models performed best without
URLs for Tasks 1 and 3. A single run with hard labels per task was submitted to each task, reaching the
29th position in Task 1, and the 9th position in Task 2.</p>
      <p>The ABCD [47] team participated in Tasks 1, 2 and 3 with both hard and soft labels. In their
approaches, they employed both LLMs like Llama 2 and T5 and smaller models like XLM_RoBERTa.
They divided the datasets into six subsets corresponding to each annotator group. Subsamples are
preprocessed, and prompt engineering is applied to LLMs. Then, smaller transformer models are
ifne-tuned on each subset and predictions are collected for each model. To incorporate the hierarchical
structure of Tasks 2 and 3, they only made predictions for subsamples classified as sexist in Task 1. Their
best performance model achieved 2nd in Task 1, 1st in Task 2, and 1st in Task 3 for the hard evaluation.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Results</title>
      <p>In the next subsections, we report the results of the participants and the baseline systems for each task.
Disaggregated results for each language, English and Spanish, may be found at the EXIST 2024 website
(http://nlp.uned.es/exist2024/).</p>
      <sec id="sec-7-1">
        <title>7.1. Task 1: Sexism Identification in Tweets</title>
        <p>We first report and analyze the results for Task 1, which focuses on sexism identification in tweets. This
task involves a binary classification. As discussed in Section 5, we report two sets of evaluation results
(hard and soft).
7.1.1. Soft Evaluation
Table 2 presents the results for the soft-soft evaluation for Task 1. A total of 37 runs were submitted.
Out of these, 34 runs outperformed the non-informative majority class baseline (where all instances
are labeled as “NO”), and all runs surpassed the non-informative minority class baseline (where all
instances are labeled as “YES”). We observed a significant discrepancy in performance, with ICM-Soft
Norm scores ranging from 0.6755 to 0.0374. However, if we analyze the top 5 systems, we appreciate
a diference of less than 5 percentual points. Notably, the best run achieved an ICM-Soft Norm score
of 68% for this binary classification task, surpassing the top performance of 64% recorded by the best
EXIST 2023 participant. This suggests that new models and approaches are becoming more efective at
detecting sexism in social networks. However, it also indicates that there is still room for improvement.
7.1.2. Hard Evaluation</p>
      </sec>
      <sec id="sec-7-2">
        <title>7.2. Task 2: Source Intention in Tweets</title>
        <p>In this section, we report and analyze the results for Task 2, which focuses on determing the intention
of the author when posting a sexist tweet. This task is a multi-class, mono-label classification. We
report two sets of evaluation results (hard and soft).
7.2.1. Soft Evaluation
minority class baseline (where all instances are labeled as “REPORTED”). The ICM-Soft Norm scores
range from the 0.4795 of the best system (“nycu-nlp_2”) to 0.0000 of “fmrs_2”, indicating significant
variability in the efectiveness of the submitted models. It is worth mentioning that the best system
outperforms the second-best by more than 8 percentage points. Overall, performance is considerably
lower compared to Task 1. This can be attributed to the hierarchical and multi-class nature of Task 2.</p>
        <p>It is also worth noting the correlation between the ICM-Soft and Cross-Entropy measures. The results
indicate a strong correlation between the two metrics, but some diferences can still be observed due to
the fact that cross entropy does not have into account the specificity of the diferent classes.
7.2.2. Hard Evaluation
Table 5 presents the hard-hard evaluation results for Task 2, assessing 43 systems against the hard gold
standard. Among these, 37 runs outperform the majority class baseline (where all instances are labeled
“NO”), and all systems show equal or better performance compared to the minority class baseline (where
all instances are labeled as “REPORTED”). Similar to the soft-soft evaluation, discrepancies between the
best and the worst-performing systems are more pronounced in Task 2 than in Task 1. The top-ranking
system, “ABCD Team_1,” achieved the highest ICM-Hard normalized score (0.6320). The top 5 best
systems range between 0.5937 and 0.6320. The lower end of the table includes five systems which score
0 in the ICM-Hard norm metric.</p>
        <p>The correlation between ICM-Hard and F1 is generally strong, with slight variations among the
top-ranked systems and greater variability towards the lower end of the table. This variability arises
because F1 does not account for the hierarchical nature of the task as efectively as ICM-Hard, which
more stringently penalizes misclassifications between diferent hierarchy levels.</p>
      </sec>
      <sec id="sec-7-3">
        <title>7.3. Task 3: Sexism Categorization in Tweets</title>
        <p>The third task is a hierarchical multi-class and multi-label classification problem, where systems must
determine if a tweet is sexist or not, and categorize the sexist tweets according to the five categories of
sexism defined in Section 2.
7.3.1. Soft Evaluation
Table 6 displays the results of the soft-soft evaluation for Task 3. A total of 30 runs were submitted,
with 26 runs surpassing the majority class baseline (all instances labeled as “NO”), and all systems
outperforming the minority class baseline (all instances labeled as “SEXUAL-VIOLENCE”). The
“NYCUNLP” team has the top three runs, with ”NYCU-NLP_1“ ranked first (ICM-Soft: − 1.1762, ICM-Soft Norm:
0.4379). The next two runs from the same team, ”NYCU-NLP_2“ and “NYCU-NLP_3,” follow closely,
indicating the consistency and robustness of their approach. The fourth and fifth systems, however,
show a significantly poorer performance (0.3835 and 0.3732, respectively) The range of ICM-Soft Norm
scores (from 0.4379 to 0.0000) underscores a significant variability in system performance. However,
despite the complexity of the task, it seems that systems are still able to correctly capture relevant
information concerning the diferent types of sexism.
7.3.2. Hard Evaluation
In the hard-hard evaluation context for the third task, 31 systems were submitted. As shown in Table 7,
28 systems outperformed the majority class baseline (all instances labeled as “NO”), while all systems
achieved better results than the minority class baseline (all instances labeled as “SEXUAL-VIOLENCE”).
The discrepancy between the best (“ABCD Team_1”, 0.5862 ICM-Hard norm score) and the
worstperforming system (“CAU&amp;ITU” 1, 0.000 score) is over 0.5 ICM-hard-norm, which is less than in Task 2.
Finally, comparing the performance of the three diferent textual tasks in the hard-hard evaluation, the
eficiency of the systems in this task, in terms of ICM-Hard Norm, is lower than in previous tasks. This
further highlights the complexity of categorizing sexism.
EXIST2024 gold</p>
        <p>Rank ICM-Hard ICM-Hard Norm</p>
        <p>Macro F1
0
Run
ABCD Team_1
ABCD Team_3
NYCU-NLP_3
NYCU-NLP_1
NYCU-NLP_2
Awakened_2
Awakened_3
RMIT-IR_3
RMIT-IR_2
RMIT-IR_1
Awakened_1
ABCD Team_2
NICA_2
penta-nlp_1 [46]
maven_1
UniLeon-UniBO_1
UniLeon-UniBO_2
UniLeon-UniBO_3
NICA_1
UMUTEAM_1
FraunhoferSIT_1
UMUTEAM_3
MMICI_3
UMUTEAM_2
CNLP-NITS-PP_1
CNLP-NITS-PP_2
MMICI_1
MMICI_2
fmrs_3
EXIST2024 majority
fmrs_2
fmrs_1
CAU&amp;ITU_1
EXIST2024 minority</p>
      </sec>
      <sec id="sec-7-4">
        <title>7.4. Task 4: Sexism Identification in Memes</title>
        <p>We next report and analyze the results for Task 4, which focuses on sexism identification in memes.
This task involves a binary classification. Again, we report two sets of evaluation results (hard and soft).
7.4.1. Soft Evaluation
Table 8 presents the results for the classification of memes as sexist or not sexist. The performance
results are notably low for a binary classification task: “Victor-UNED_1”, the top-ranked participant,
achieved an ICM-Soft Norm score of 0.4530 and a relatively low Cross Entropy of 1.1028. However, the
variability between the best and worst-performing systems is reduced compared to that of the tasks
described above. When comparing these results to those of Task 1 (classifying tweets as sexist or not),
we observe a significant drop in performance for image classification (0.4530 versus 0.6755 ICM-Soft
Norm). It is important to highlight that most approaches relied solely on the text within the meme for
classification, without incorporating image processing. This suggests that sexism in memes might often
be conveyed through the imagery, even when the accompanying text seems to be neutral.</p>
      </sec>
      <sec id="sec-7-5">
        <title>7.5. Task 5: Source Intention in Memes</title>
        <p>In this section, we report and analyze the results for Task 5, which focuses on determing the intention
of the author when posting a sexist meme. This task is a multi-class, mono-label classification. We
report two sets of evaluation results (hard and soft).
7.5.1. Soft Evaluation
7.5.2. Hard Evaluation
Table 11 presents the results for the hard-hard evaluation of Task 5. Out of the 19 systems submitted
for this task, only 15 ranked above the majority class baseline (all instances labeled as “NO”), while 18
systems surpassed the minority class baseline (all instances labeled as “JUDGEMENTAL”). The results
range from 0.4167 ICM-Hard Norm for the best performing system (“Victor-UNED_1”) to 0.0000 for the
worst performing systems, but are quite homogeneous among the top 5 systems.</p>
        <p>When comparing ICM-Hard Norm results with F1 scores, we again observe little correlation between
the two metrics, especially in the lower ranks of the table.</p>
        <p>Rank ICM-Hard ICM-Hard Norm
Victor-UNED_1
melialo-vcassan_3
melialo-vcassan_1
I2C-Huelva_3
I2C-Huelva_2
I2C-Huelva_1
MMICI_3
EXIST2024 majority
Penta-ML_3
Penta-ML_1
Penta-ML_2
EXIST2024 minority</p>
        <p>Rank ICM-Hard ICM-Hard Norm</p>
      </sec>
      <sec id="sec-7-6">
        <title>7.6. Task 6: Sexism Categorization in Memes</title>
        <p>The sixth task is a hierarchical multi-class and multi-label classification problem, where systems must
determine if a meme is sexist or not, and if so, categorize it according to the five categories of sexism
defined in Section 2.
7.6.1. Soft Evaluation
7.6.2. Hard Evaluation
Finally, Table 13 presents the results for classifying memes based on the aspects of women being
attacked, with outputs provided as a single class prediction. 22 runs were submitted for this task. Only
17 runs exceeded the majority class baseline (labeling all instances as “NO”), while 21 runs ranked above
the minority class (all instances labeled as “MISOGYNY-NON-SEXUAL-VIOLENCE”) The performance
for this task was low, with the top team (“DiTana-PV_1”) achieving an ICM-Soft Norm score of 0.3549.</p>
        <p>Rank ICM-Hard ICM-Hard Norm</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>8. Discussion</title>
      <p>After the study of the 34 submitted working notes, we have discovered patterns between systems and
their performance that are of the greatest interest to the task. Even though in this analysis we only
considered systems whose working notes were available, most of the best performing models were
considered. We will focus on the conclusions from the study of the top ten systems for the six tasks in
hard and soft evaluations, and the classification in English and Spanish for each one. We can divide two
types of approaches submitted: textual systems for Tasks 1, 2 and 3, and systems for Tasks 4, 5 and 6.</p>
      <sec id="sec-8-1">
        <title>8.1. Tasks 1, 2 and 3: Sexism Detection and Classification in Tweets</title>
        <p>In the textual tasks, most of the top-ranking models were encoding-based transformer, fine-tuned on
the EXIST dataset with an additional component. This additional component could be:
• A meticulous data preprocessing step. For example, that is the main contribution of the NYCU-NLP
team: they removed irrelevant elements of the text, and they increased the size of the dataset
by applying data augmentation techniques such as AEDA, and by translating from English to
Spanish and vice versa.
• Ensembles of encoding-based transformer models. A huge variety of models were employed:
most of them include BERT-like systems like BERT, RoBERTa, DeBERTa, multilingual versions of
them trained in Spanish datasets, and some were fine-tuned in tweets, hate-speech or sentiment
analysis datasets. Systems following this architecture excel particularly in the soft evaluation.
This success appears to be linked to how teams trained their models, as encoding-based models
are easily fine-tuned with soft labels. Depending on the models used to obtain results for the
ensemble, diferent types of voting methods are considered:
– If the results difer highly between models, systems seem to perform better when a higher
weight is given to the best model and the influence of the rest of the models is reduced. The
Awakened team considered diferent models in their ensembles, but their best performing
runs gave higher weight to their best model in the ensemble.
– If the diferences between the performance of models are small, a proportion of votes can
take into account aspects detected by every model. This method obtained good results for
CIMAT-CS-NLP in the first task, and Medusa in the third task.
• Sharing results with LLMs. This method proved to be successful in the hard evaluation, where
teams that submitted labels obtained by both LLMs and encoding-based transformer models
reached the best results. Due to the increasing amount of LLMs published and their constant
improvement, there is a whole range of models that the community can try. In EXIST tasks, there
were attempts testing Llama-2, Llama-3, Gemini, Mistral and GPT-4. Most of them are used to
obtained labels using zero-shot or few-shots because of their computational cost of fine-tuning.
Thus, approaches rely on the use of prompt engineering to make the model understand the task.
Because of these diferences in models and ways of formulating prompts, systems cannot be
directly compared; however, we can mention, as some of the best performing ones: the system of
CIMAT-CS-NLP, that uses an ensemble of 4 zero-shot answers of Gemini; the team ABCD, which
using prompt engineering created one Llama-2 answer per annotator; the team EquityExplorers,
who used Mistral results in their ensemble, and the team CIMAT-GTO, which studied types of
reasoning with Llama-3.</p>
        <p>In general, we can draw some general conclusions about systems for textual tasks. Models that use
encoding-based transformers performed better in the soft evaluation. Even systems which use one
model or one model trained in diferent ways, like BAZI and Victor-UNED, achieved top-10 ranking
results in the soft evaluation because they are trained with soft labels. On the other hand, LLMs
performance stands out in the hard evaluation, but drops in the soft evaluation. This can be clearly
seen with ABCD, where runs obtained by Llama-2 got better results than encoding models in the hard
evaluation for the three tasks, whereas encoding models outperformed them in the soft evaluation. In
general, demographic information was not included in models, but those teams who have included it
have obtained improvements in their systems in Task 3.</p>
      </sec>
      <sec id="sec-8-2">
        <title>8.2. Tasks 4, 5 and 6: Sexism Detection and Classification in Memes</title>
        <p>Tasks 4, 5 and 6 are multimodal, that is, teams could use both images and text to detect and categorize
sexism in the instances. Although most of the teams have faced these tasks as multimodal, the best
performances correspond to models which only use text to analyze memes. Top ranking positions
for most of the tasks, especially Tasks 4 and 5, were obtained by textual models. However, images
remain important. RoJiNG-CL used GPT-4 to create a textual description of the image and analyzed it,
outperforming the other approaches. Models that were used to analyze the text were mostly
encodingbased transformers: combinations of BERT, DeBERTa and RoBERTa with diferent weights or diferent
ifne-tunings. Systems that only considered text where similar to those presented in the previous section.</p>
        <p>Next in the ranking are multimodal approaches. The influence of textual analysis implies that even
in multimodal approaches the ones that focus on textual treatment with a ViT module using BERT or
similar outperform models than do not include it. This phenomenon is shown in RMIT-IR runs, whose
best run is the one that used mBERT and obtained several positions ahead of their next run which did
not include it. Most of the systems utilized CLIP to analyze images. However, the use of CLIP alone
led to poor results. The concatenation of text and images stands as the necessary way of dealing with
these tasks, but new ways of obtaining representations of images are needed to outperform text-only
systems.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>9. Conclusions</title>
      <p>The objective of the EXIST challenge is to encourage research on the automated detection and modeling
of sexism in online environments, with a specific focus on social networks. The EXIST 2024 Lab held as
part of CLEF attracted nearly 60 participant teams, and received more than 400 runs. Participants adopted
a wide range of approaches, including vision transformer models, data augmentation through automatic
translation, data duplication, utilization of data from past EXIST editions, multilingual language models,
Twitter-specific language models, and transfer learning techniques from domains like hate speech,
toxicity, and sentiment analysis. While many systems opted for the traditional approach of providing
only hard labels as outputs, a significant number of systems leveraged the multiple annotations available
in the dataset, and provided soft outputs, proving that there is an increasing interest by the research
community in developing systems able to deal with disagreements and with diferent perspectives.</p>
      <p>Concerning the results, in the textual tasks (Tasks 1, 2 and 3), top-performing models were typically
encoding-based Transformers fine-tuned on the EXIST dataset with an additional component, such
as meticulous data preprocessing, data augmentation or the use of model ensembles. In multimodal
tasks (Tasks 4, 5, and 6), where both images and text could be used to detect and categorize sexism, top
performances were achieved by models focusing solely on text.</p>
      <p>For future editions of EXIST, we plan to expand our study in order to include additional communication
channels and media formats, such as TikTok videos. By doing so, we aim to address the nuances and
unique challenges presented by diferent formats, enhancing the robustness and applicability of research
on automated sexism detection. Additionally, this expansion will allow us to capture a broader spectrum
of online interactions and cultural contexts.</p>
    </sec>
    <sec id="sec-10">
      <title>Acknowledgments</title>
      <p>This work has been financed by the European Union (NextGenerationEU funds) through the “Plan
de Recuperación, Transformación y Resiliencia”, by the Ministry of Economic Afairs and Digital
Transformation and by the UNED University. It has also been financed by the Spanish Ministry of
Science and Innovation (project FairTransNLP (PID2021-124361OB-C31 and PID2021-124361OB-C32))
funded by MCIN/AEI/10.13039/501100011033 and by ERDF, EU A way of making Europe, and by the
Australian Research Council (DE200100064 and CE200100005).
Learning for Sexism Detection, in: Working Notes of CLEF 2024 – Conference and Labs of the
Evaluation Forum, 2024.
[15] M. Usmani, R. Siddiqui, S. Rizwan, F. Khan, F. Alvi, A. Samad, Sexism Identification in Tweets
using BERT and XLM – Roberta, in: Working Notes of CLEF 2024 – Conference and Labs of the
Evaluation Forum, 2024.
[16] T. Smith, R. Nie, J. Trippas, D. Spina, RMIT-IR at EXIST Lab at CLEF 2024, in: Working Notes of</p>
      <p>CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[17] M. Obrador Reina, A. García Cucó, LightGMB for Sexism Identification in Memes, in: Working</p>
      <p>Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[18] A. Shah, A. Gokhale, Team Aditya at EXIST 2024 — Detecting Sexism In Multilingual Tweets
Using Contrastive Learning Approach, in: Working Notes of CLEF 2024 – Conference and Labs of
the Evaluation Forum, 2024.
[19] A. Azadi, B. Ansari, S. Zamani, Bilingual Sexism Classification: Fine-Tuned XLM-RoBERTa and
GPT-3.5 Few-Shot Learning, in: Working Notes of CLEF 2024 – Conference and Labs of the
Evaluation Forum, 2024.
[20] M. Siino, I. Tinnirello, Prompt Engineering for Identifying Sexism using GPT Mistral 7B, in:</p>
      <p>Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[21] A. Vetagiri, P. Mogha, P. Pakray, Cracking Down on Digital Misogyny with MULTILATE a
MULTImodaL hATE Detection System, in: Working Notes of CLEF 2024 – Conference and Labs of
the Evaluation Forum, 2024.
[22] A. Petrescu, C.-O. Truică, E.-S. Apostol, Language-based Mixture of Transformers for EXIST2024,
in: Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[23] Y.-Z. Fang, L.-H. Lee, J.-D. Huang, NYCU-NLP at EXIST 2024 – Leveraging Transformers with
Diverse Annotations for Sexism Identification in Social Networks, in: Working Notes of CLEF
2024 – Conference and Labs of the Evaluation Forum, 2024.
[24] G. Shimi, J. Mahibha, D. Thenmozhi, Automatic Classification of Gender Stereotypes in Social
Media Post, in: Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum,
2024.
[25] R. Pan, J. A. García Díaz, T. Bernal Beltrán, R. Valencia-Garcia, UMUTeam at EXIST 2024:
Multimodal Identification and Categorization of Sexism by Feature Integration, in: Working Notes of
CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[26] M. Sreekumar, S. K, T. Durairaj, S. Gopalakrishnan, K. Swaminathan, Sexism Identification in
Tweets using Traditional Machine Learning Approaches, in: Working Notes of CLEF 2024 –
Conference and Labs of the Evaluation Forum, 2024.
[27] R. Keinan, Sexism Identicfiation in Social Networks using TF-IDF Embeddings, PreProccessing,
Feature Selection, Word/Char N-Grams and Various Machine Learning Models In Spanish and
English, in: Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[28] V. Ruiz, J. Carrillo-de-Albornoz, L. Plaza, Concatenated Transformer Models based on Levels of
AgreementsfFor Sexism Detection, in: Working Notes of CLEF 2024 – Conference and Labs of the
Evaluation Forum, 2024.
[29] G. Rizzi, D. Gimeno-Gómez, E. Fersini, C.-D. Martínez-Hinarejos, PINK at EXIST2024: A
CrossLingual and Multi-Modal Transformer Approach for Sexism Detection in Memes, in: Working
Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[30] J. Ma, R. Li, RoJiNG-CL at EXIST 2024: Sexism Identification in Memes by Integrating Prompting
and Fine-Tuning, in: Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum,
2024.
[31] A. Naebzadeh, M. Nobakhtian, S. Eetemadi, NICA at EXIST CLEF Tasks 2024, in: Working Notes
of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[32] F. Maqbool, E. Fersini, A Contrastive Learning based Approach to Detect Sexism in Memes, in:</p>
      <p>Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[33] A. Menárguez Box, D. Torres Bertomeu, DiTana-PV at sEXism Identification in Social neTworks
(EXIST) Tasks 4 and 6: The Efect of Translation in Sexism Identification, in: Working Notes of</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rodríguez-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          , L. Plaza,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Comet</surname>
          </string-name>
          , T. Donoso, Overview of EXIST 2021:
          <article-title>Sexism identification in social networks</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <year>2021</year>
          )
          <fpage>195</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rodríguez-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mendieta-Aragón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Marco-Remón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Makeienko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2022:
          <article-title>Sexism identification in social networks</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>69</volume>
          (
          <year>2022</year>
          )
          <fpage>229</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          , J. C. de Albornoz,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2023 -
          <article-title>Learning with Disagreement for Sexism Identification and Characterization (Extended Overview)</article-title>
          , in: M.
          <string-name>
            <surname>Aliannejadi</surname>
            , G. Faggioli,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
          </string-name>
          , M. Vlachos (Eds.),
          <source>Working Notes of the Conference and Labs of the Evaluation Forum (CLEF</source>
          <year>2023</year>
          ), volume
          <volume>497</volume>
          , CEUR Working Notes,
          <year>2023</year>
          , pp.
          <fpage>813</fpage>
          -
          <lpage>854</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Uma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Fornaciari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dumitrache</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chamberlain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Plank</surname>
          </string-name>
          , E. Simpson, M. Poesio, SemEval
          <article-title>-2021 task 12: Learning with disagreements</article-title>
          ,
          <source>in: Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>338</fpage>
          -
          <lpage>347</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H. R.</given-names>
            <surname>Kirk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Vidgen</surname>
          </string-name>
          , P. Röttger, SemEval-2023
          <source>Task 10: Explainable Detection of Online Sexism, in: Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval)</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Billig</surname>
          </string-name>
          ,
          <article-title>Humour and hatred: the racist jokes of the Ku Klux Klan</article-title>
          ,
          <source>Discourse &amp; Society</source>
          <volume>12</volume>
          (
          <year>2014</year>
          )
          <fpage>267</fpage>
          -
          <lpage>289</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mendiburo-Seguel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. E.</given-names>
            <surname>Ford</surname>
          </string-name>
          ,
          <article-title>The Efect of Disparagement Humor on the Acceptability of Prejudice., Current Psychology: A Journal for Diverse Perspectives on Diverse Psychological Issues (</article-title>
          <year>2019</year>
          )
          <article-title>No Pagination Specified-No Pagination Specified</article-title>
          .
          <source>doi: 10.1007/s12144-019-00354-2.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>G.</given-names>
            <surname>Hodson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rush</surname>
          </string-name>
          , C. C.
          <article-title>MacInnis, A Joke Is Just a Joke (except When It Isn't): Cavalier Humor Beliefs Facilitate the Expression of Group Dominance Motives</article-title>
          .,
          <source>Journal of Personality and Social Psychology</source>
          <volume>99</volume>
          (
          <year>2010</year>
          )
          <fpage>660</fpage>
          -
          <lpage>682</lpage>
          . doi:
          <volume>10</volume>
          .1037/a0019627.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>F.</given-names>
            <surname>Gasparini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Saibene</surname>
          </string-name>
          , E. Fersini,
          <article-title>Benchmark Dataset of Memes with Text Transcriptions for Automatic Detection of Multi-modal Misogynistic Content, Data in Brief 44 (</article-title>
          <year>2022</year>
          )
          <fpage>108526</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Fersini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gasparini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Saibene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chulvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lees</surname>
          </string-name>
          , J. Sorensen, SemEval
          <article-title>-2022 Task 5: Multimedia Automatic Misogyny Identification</article-title>
          ,
          <source>in: Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>533</fpage>
          -
          <lpage>549</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rajiakodi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ponnusamy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Pannerselvam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Madasamy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rajalakshmi</surname>
          </string-name>
          , H. LekshmiAmmal,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kizhakkeparambil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sivagnanam</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Rajkumar, Overview of Shared Task on Multitask Meme Classification - Unraveling Misogynistic and Trolls in Online Memes</article-title>
          ,
          <source>in: Proceedings of the Fourth Workshop on Language Technology for Equality</source>
          , Diversity, Inclusion,
          <year>2024</year>
          , pp.
          <fpage>139</fpage>
          -
          <lpage>144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2023 -
          <article-title>Learning with Disagreement for Sexism Identification and Characterization (Extended Overview)</article-title>
          ,
          <source>in: Working Notes of CLEF 2023 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Delgado</surname>
          </string-name>
          ,
          <article-title>Evaluating Extreme Hierarchical Multi-label Classification, in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics</article-title>
          , volume Volume
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers</given-names>
          </string-name>
          ,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          , Dublin, Ireland,
          <year>2022</year>
          , p.
          <fpage>5809</fpage>
          -
          <lpage>5819</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Frick</surname>
          </string-name>
          , M. Steinebach, FraunhoferSIT@EXIST2024: Leveraging Stacking Ensemble
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>