<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NYCU-NLP at EXIST 2025: An Empirical Study of Annotator-Aware Two-Stage Pipeline for Sexism Detection in Tweets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Joy Chrissetyo Prajogo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lung-Hao Lee</string-name>
          <email>lhlee@nycu.edu.tw</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hsien-I Lin</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Electrical Engineering and Computer Science, National Yang Ming Chiao Tung University</institution>
          ,
          <country country="TW">Taiwan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Artificial Intelligence Innovation, National Yang Ming Chiao Tung University</institution>
          ,
          <country country="TW">Taiwan</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Electrical and Control Engineering, National Yang Ming Chiao Tung University</institution>
          ,
          <country country="TW">Taiwan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>This paper presents a comprehensive approach for automated sexism detection in bilingual tweets, evaluated within the context of the EXIST 2025 shared task. Our proposed framework explicitly integrates annotator demographics and leverages bilingual fusion, combining original and cross-translated tweets. We implement a novel two-stage hierarchical pipeline consisting of sexism identification (binary classification) followed by misogynistic intent classification (multi-class) and sexism type categorization (multi-label). Our experiments systematically compare three distinct modeling strategies within this pipeline: a fine-tuned transformer-based dual-encoder architecture with early and late fusion, a zero-shot auto-regressive (AR) large language model (LLM), and a zero-shot difusion-based LLM. Evaluation results demonstrate that our transformer-based approach consistently achieves the highest performance across most metrics, emphasizing the efectiveness of explicitly modeling annotator disagreement and demographic context. Notably, the difusion-based LLM demonstrates performance on par with, and in some cases superior to, the AR LLM, highlighting the potential of difusion models as an alternative paradigm for complex text classification tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Transformers</kwd>
        <kwd>Auto-Regressive LLM</kwd>
        <kwd>Difusion LLM</kwd>
        <kwd>Sexism Identification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        • EXIST 2021 initiated tasks with binary sexism identification and categorization [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
• EXIST 2022 further emphasized the automation and scalability of sexism detection [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
• EXIST 2023 introduced the detection of sexist intent (direct, reported, judgmental) and adopted
the Learning with Disagreements (LeWiDi) paradigm, recognizing that diferent annotators might
disagree on labels due to inconsistent subjective perceptions of sexism [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ].
• EXIST 2024 expanded from text-based tweets to visual memes, still using the LeWiDi paradigm
[
        <xref ref-type="bibr" rid="ref5">5, 6</xref>
        ].
• EXIST 2025 further broadens the task scope to multimedia data, including TikTok videos, thus
covering text, images, and videos under the LeWiDi framework [7, 8].
      </p>
      <p>In this edition, our team specifically addresses sexism detection in tweets, tackling the three subtasks
designed for textual data:
• Task 1: Binary sexism detection (sexist vs. non-sexist tweets).
• Task 2: Misogynistic intent classification (direct, reported, judgmental).
• Task 3: Fine-grained sexism categorization, including ideological inequality, stereotyping and
dominance, objectification, sexual violence, and misogyny (non-sexual violence).</p>
      <p>Due to the continued use of Learning with Disagreements (LeWiDi) in EXIST 2025, each tweet is
annotated by six annotators whose judgments may vary. Capturing and modeling this disagreement is
crucial for developing robust detection systems.</p>
      <p>To address these challenges, we propose a novel annotator-aware, two-stage bilingual sexism detection
pipeline. Our approach integrates annotator demographics directly into the modeling process and
exploits bilingual fusion through original tweets and their translations. We test this pipeline using
three modeling strategies: a transformer-based dual-encoder architecture, an auto-regressive (AR)
large language model (LLM), and a difusion-based LLM. Empirical results indicate that difusion-based
LLM achieves performance comparable to that of the AR LLM, highlighting promising directions in
leveraging difusion models for text classification tasks.</p>
      <p>The remainder of this paper is structured as follows: Section 2 reviews related work, Section 3 presents
the methodologies and system architecture, Section 4 describes our experimental setup, including the
dataset, parameters, and results, and Section 5 concludes with our findings and future directions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Sexism involves unjustified distinctions or discriminatory actions based on gender, primarily targeted
against women, varying across cultural contexts. Various studies highlight the pervasive nature and
diverse manifestations of online sexism. Early prominent sexism detection datasets include AMI [9]
and HatEval [10], initially established for English-language detection. EXIST, launched at CLEF in 2021,
uniquely combines binary detection, intent classification, and fine-grained categorization tasks [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>Learning with Disagreement (LeWiDi) addresses scenarios where multiple annotators provide
conlficting labels or interpretations due to subjective or ambiguous data. Instead of forcing consensus into
a single "correct" label, LeWiDi explicitly incorporates annotator disagreement into model training and
evaluation as informative signals. This approach enhances robustness and generalization by
systematically leveraging diverse annotator perspectives, particularly useful in sentiment analysis, hate speech
detection, and other subjective classification tasks [11].</p>
      <p>
        EXIST adopted the LeWiDi paradigm starting from EXIST-2023 edition [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ], recognizing that
annotator disagreement is inherent due to difering subjective perceptions of sexism [ 12]. Soft labeling, which
quantifies annotator disagreements, has been shown to improve robustness compared to traditional
single-label training [13, 11]. Prior solutions [14, 15] explicitly incorporated soft labeling, significantly
enhancing system performance by modeling the inherent annotator disagreements.
      </p>
      <p>Annotators’ demographic backgrounds strongly influence their perception of what constitutes
sexism, leading to inconsistent results. Tredici et al. [16] demonstrated that incorporating annotator
demographics via demographic embeddings can improve the accuracy of hate speech detection.
Multiannotator attention mechanisms have also been explored, notably using personalized toxicity classifiers
[17]. Similarly, EXIST-2024 participants, such as Fang et al. [14] and Quan and Thin [15] efectively
used annotator demographics and metadata, showcasing that explicitly modeling annotator features
significantly enhances sexism detection performance.</p>
      <p>Transformer-based architectures, particularly multilingual models like XLM-R [18] have proven
highly efective for multilingual hate speech detection [ 19]. Data augmentation and multilingual fusion
strategies, including cross-translation and back-translation, have been shown to enhance generalization
in low-resource scenarios[20, 21]. Our pipeline uniquely uses bilingual fusion by simultaneously
modeling original and cross-translated tweets, thereby enhancing robustness against linguistic variability
and annotator disagreement.</p>
      <p>Recently, zero-shot and few-shot large language models (LLMs) have shown impressive performance
in toxicity detection tasks without task-specific fine-tuning. GPT-3 demonstrated that prompt-based
LLMs can efectively identify toxic content through careful prompt engineering [ 22]. The LLaMA
family[23], particularly the LLaMA 3.1 series, have been optimized for multilingual dialogue use cases
and outperform many existing open-source and closed-source chat models on common industry
benchmarks. Two teams [24, 10] conducted comprehensive evaluations demonstrating LLM competitiveness
against fine-tuned models for detecting hate and toxic speech.</p>
      <p>Difusion-based language models represent a novel and promising approach to language generation
tasks. Introduced by [25, 26], difusion models generate text through iterative denoising steps, yielding
more controlled and coherent outputs compared to autoregressive models. The recent Dream-v0 Instruct
model, developed by the HKU NLP Group [27], demonstrated encouraging results for
instructionfollowing tasks, showcasing the potential of difusion models in text classification . Similarly, LLaDA, a
difusion model trained from scratch under the pre-training and supervised fine-tuning paradigm[ 28],
has shown strong scalability and performance, rivaling autoregressive models like LLaMA 3 in various
benchmarks . These developments highlight the viability of difusion models as an alternative to
traditional autoregressive approaches in language modeling.</p>
      <p>Our proposed bilingual fusion approach, combining original tweets with their translated counterparts,
is inspired by robust back-translation methods. Multi-view ensemble strategies further enhance
robustness and generalization, as demonstrated by previous EXIST participants. Additionally, our two-stage
pipeline (binary gate followed by multi-label classification) efectively reduces the complexity of sexism
categorization, allowing models to specialize sequentially and improve overall detection accuracy.</p>
      <p>In summary, while EXIST has significantly advanced the state-of-the-art in sexism detection through
the introduction of multimedia tasks and the LeWiDi paradigm, no prior EXIST entry has systematically
compared supervised dual-encoder transformers, zero-shot auto-regressive LLMs, or difusion-based
LLMs within a unified annotator-aware bilingual pipeline. Our study fills this important research gap.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodologies and Architecture</title>
      <p>In this section, we detail our proposed framework for bilingual sexism detection, comprising multiple
interconnected stages that form a comprehensive pipeline. The pipeline starts with an extensive data
pre-processing stage that transforms raw tweet data into structured, annotated, bilingual data suitable
for modeling. We then introduce our annotator-aware modeling pipeline, implemented as a two-stage
hierarchical system. Within this pipeline, we explore and compare three distinct modeling strategies:
a transformer-based dual-encoder, an auto-regressive large language model (LLM), and a
difusionbased LLM. The pipeline concludes with a data post-processing step that formats the final predictions
appropriately for evaluation.</p>
      <p>Our approach explicitly leverages annotator-provided metadata and directly addresses annotator
disagreements. Furthermore, it exploits bilingual data by integrating original and cross-translated
tweets to enhance the prediction accuracy.</p>
      <sec id="sec-3-1">
        <title>3.1. Data Pre-Processing</title>
        <p>Figure 1 illustrates our data pre-processing pipeline, which converts raw JSON annotation files into a
structured, machine-readable CSV format. Below we describe each step in detail.</p>
        <p>First, raw JSON annotation files are transformed into a structured CSV table, facilitating easier
manipulation and subsequent processing. Following this, we generate bilingual tweet pairs by using</p>
        <p>Convert JSON
to CSV
Final CSV</p>
        <p>Original Tweet</p>
        <p>[ES]
@BestKabest Esta gringa sigue llorando por
el gamergate, que "coincidencia" que tenga
pronombres en su perfil</p>
        <p>[EN]
@BarbieReports @Londonist
@TimeOutLondon @visitlondon the curse of
phallogocentrism ! god forbid a woman
should be upright :)</p>
        <p>New Cross-Translated Tweet</p>
        <p>[EN]
@Bestkabest This gringa continues to cry for
the gamergate, that "coincidence" that has
pronouns in its profile</p>
        <p>[ES]
@Barbiereports @londonist @timeoutlondon</p>
        <p>@visitlondon ¡La maldición del
fallogocentrismo! Dios no permita que una
mujer debería estar en posición vertical :)
Original Text</p>
        <p>[ES]
No queremos hijos valientes. Queremos hijos
felices. Protejan a los menores con diversidad
funcional para que dejen de ser víctimas del
acoso escolar. #detidepende</p>
        <p>#laEspañaqueQueremos
#derechoshumanos@sanchezcastejon
@IreneMontero https://t.co/cZtiZfX5O6</p>
        <p>[EN]
We do not want brave children. We want
happy children. They protect minors with
functional diversity to cease to be victims of
bullying. #Detidepende #laspaña we will see
#Haries Humanos @SanchezCastejon
@irenemontero https://t.co/cztizfx5O6
Cleaned Text</p>
        <p>[ES]
no queremos hijos valientes. queremos hijos
felices. protejan a los menores con diversidad
funcional para que dejen de ser víctimas del
acoso escolar.</p>
        <p>[EN]
we do not want brave children. we want
happy children. they protect minors with
functional diversity to cease to be victims of
bullying. we will see humanos
Google Translate via Python’s deep-translator library, translating Spanish tweets into English and
vice versa. Table 1 shows samples of the original and cross-translated tweets.</p>
        <p>Annotator metadata, originally aggregated in list form, is then expanded into separate, structured
columns, allowing for eficient integration into downstream modeling. Next, rigorous data cleaning is
applied to both the original and cross-translated tweets. This process involves removing usernames,
URLs, emails, percentages, timestamps, phone numbers, hashtags, emojis, and other symbols.
Additionally, tweets are converted to lowercase, and multiple spaces between words are removed. Table 2
illustrates tweet examples before and after cleaning.</p>
        <p>Finally, annotator metadata and labels are encoded. Metadata and labels are transformed into
numerical representations through one-hot encoding, ordinal encoding, and symbolic letter combinations,
thereby reducing complexity and optimizing resource utilization during modeling. This culminates in a
JSON Output</p>
        <p>Prediction
fully structured CSV dataset ready for use by the modeling pipeline.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Annotator-Aware Two-Stage Pipeline</title>
        <p>Figure 2 illustrates our annotator-aware two-stage modeling pipeline, explicitly designed to address
annotator disagreement. The pipeline operates hierarchically with two clearly defined stages:</p>
        <p>Stage 1 addresses Task 1_1: Sexism Identification in Tweets, a binary classification task determining
whether a tweet contains sexist content (YES or NO).</p>
        <p>Stage 2 encompasses Task 1_2: Source Intention Classification and Task 1_3: Sexism Categorization.
Task 1_2 is a multi-class classification of the author’s intention behind sexist tweets, categorizing them
as DIRECT, REPORTED, or JUDGEMENTAL. Task 1_3 performs a multi-label classification to determine
the specific type of sexism, assigning labels such as IDEOLOGICAL-INEQUALITY,
STEREOTYPINGDOMINANCE, OBJECTIFICATION, SEXUAL-VIOLENCE, and MISOGYNY-NON-SEXUAL-VIOLENCE.</p>
        <p>To comprehensively evaluate framework performance, we implement three distinct modeling
approaches: (1) a fine-tuned transformer-based dual-encoder using both early and late fusion, (2) a
zero-shot auto-regressive large language model (LLM) with late fusion, and (3) a zero-shot
difusionbased LLM with late fusion. Each method explicitly incorporates bilingual information and annotator
metadata, using an identical pipeline architecture (as depicted in Fig. 2) but difering in terms of
modeling paradigms.
Annotator</p>
        <p>Metadata</p>
        <sec id="sec-3-2-1">
          <title>3.2.1. First System: Transformer-Based Approach</title>
          <p>Our first system, depicted in Fig. 3, uses a transformer-based dual-encoder architecture with
XLMRoBERTa-Base (XLM-R), explicitly integrating bilingual tweet data and annotator metadata. The detailed
steps are as follows:</p>
          <p>Tweets (original and cross-translated) undergo independent tokenization using the XLM-R tokenizer,
producing token IDs and attention masks. Simultaneously, annotator metadata is mapped into numerical
indices and then aggregated, forming structured tensors. The resulting tokenized tweets and metadata
tensors are combined within a dual encoder that comprises the XLM-R backbone and an annotator-aware
sub-encoder using Multi-Head Attention (MHA).</p>
          <p>During training (Stages 1 and 2), embeddings from the original and cross-translated tweets are fused
early using a weighted sum (alpha = 0.8). Subsequently, during inference, predictions from each encoder
branch (original and cross-translated) are fused using the same weighted averaging (alpha = 0.8). This
produces both hard (binary or categorical) and soft (probability) predictions. Importantly, the same
transformer-based framework is reused for both stages, difering only in terms of the classification
heads and labels specific to each task.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>3.2.2. Second System: Auto-regressive LLM-Based Approach</title>
          <p>The second system, shown in Fig. 4, adopts a zero-shot auto-regressive LLM approach using the
LLaMA-3.1-8B-Instruct model without additional fine-tuning. The approach difers notably from the
transformer-based method by relying on carefully engineered textual prompts instead of encoder
embeddings.</p>
          <p>In this system, each tweet and its associated annotator metadata are concatenated into structured
textual prompts, provided separately for the original and cross-translated tweets. The auto-regressive
LLM directly computes log probabilities for targeted tokens such as "Yes," "No," or category codes ("1,"
"2," "3"), depending on the specific task and classification type (binary or multi-class). Original and
translated log probabilities are subsequently fused using late fusion (alpha = 0.55), enhancing prediction
robustness without task-specific training.</p>
          <p>Similar to the first system, an identical process is repeated for Stages 1 and Stage 2, difering only in
terms of prompt content corresponding to task objectives.</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>3.2.3. Third System: Difusion LLM-Based Approach</title>
          <p>Our third modeling approach introduces a difusion-based LLM, specifically the Dream-v0-Instruct-7B
model. The pipeline architecture for this system (illustrated in Fig. 4) matches the second system exactly,
difering only in the underlying model. This approach demonstrates a novel text classification paradigm
that leverages difusion models’ generative capabilities for classification tasks.</p>
          <p>The difusion-based model uses the same prompt-engineering strategy, combining bilingual tweet
pairs and annotator metadata into structured prompts. It similarly extracts log probabilities for target
tokens ("Yes," "No," or category labels), which are fused using late fusion (alpha = 0.65). Comparing
the difusion LLM against traditional transformer and auto-regressive LLM methods, we evaluate and
highlight its efectiveness for complex tasks such as sexism detection.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Data Post-Processing</title>
        <p>Following prediction generation, the results are consolidated into unified CSV files containing all fused
predictions. We then use a lightweight post-processing script to convert these predictions into the
six separate JSON files required by the evaluation toolkit. For each dataset (dev and test) and each
system, the script iterates over each tweet prediction and extracts the following: a single probability
and binary label for Task 1_1, a four-element probability vector with a corresponding integer label for
Task 1_2, and a six-element probability vector paired with six binary labels for Task 1_3. The extracted
predictions are formatted and written into distinct JSON files, thus conforming precisely to the format
expected by the oficial evaluation scorers.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation</title>
      <sec id="sec-4-1">
        <title>4.1. Dataset</title>
        <p>In this section, we present a detailed evaluation of our proposed systems. The evaluation is structured
into subsections covering the dataset, parameters and settings, results and ablation studies, and the
ifnal outcomes in terms of performance metrics.</p>
        <p>We use the EXIST 2025 dataset [7, 8], which includes tweets specifically curated for Tasks 1_1, 1_2,
and 1_3. The dataset contains 6,920 tweets for training, 1,038 tweets for development, and 2,076
tweets for testing. The tweets are evenly balanced between English and Spanish. Each tweet in the
dataset is provided as a separate JSON file with the following attributes: 1) id_EXIST: a unique
identifier for the tweet; 2) lang: the language of the text (“en” or “es”); 3) tweet: the text content; 4)
number_annotators: the number of annotators; 5) annotators: unique identifiers for each
annotator; 6) gender_annotators: the gender of each annotator (“F” or “M”); 7) age_annotators: the age
group of each annotator (“18–22”, “23–45”, or “46+” ); 8) ethnicity_annotators: the self-reported
ethnicity (“Black or African American”, “Hispanic or Latino”, “White or Caucasian”, “Multiracial”, “Asian”,
“Asian Indian”, or “Middle Eastern”); 9) study_level_annotators: self-reported education level (“No
high school diploma”, “High school degree or equivalent”, “Bachelor’s degree”, “Master’s degree”, or
“Doctorate”); 10) country_annotators: the self-reported country of residence; 11) labels_task1_1:
labels indicating whether the tweet contains sexist expressions or refers to sexist behaviors (“YES” or
“NO”); 12) labels_task1_2: labels for the author’s intention (“DIRECT”, “REPORTED”,
“JUDGMENTAL”, “-”, or “UNKNOWN”); 13) labels_task1_3: arrays of labels indicating the type(s) of sexism in
the tweet (“IDEOLOGICAL_INEQUALITY”, “STEREOTYPING-DOMINANCE”, “OBJECTIFICATION”,
“SEXUAL-VIOLENCE”, “MISOGYNY-NON-SEXUAL-VIOLENCE”, “-”, or “UNKNOWN”); and 14) split:
the subset (“TRAIN”, “DEV”, or “TEST”) with language code (“EN”/“ES”).</p>
        <p>Each tweet is annotated by multiple annotators, who may difer in demographic characteristics,
thus promoting diverse viewpoints aligned with the organizers’ concept of learning with disagreement
(LeWiDi).</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Parameters and Settings</title>
        <p>We configure distinct parameters for each of our three systems. The first system fine-tunes an
XLM-RoBERTa-base backbone with a maximum sequence length of 256 tokens, augmented by a small
annotator metadata encoder (AnnEncoder). Within this encoder, annotator demographics (gender,
country, ethnicity, and study level) are respectively embedded into vectors with 4, 8, 8, and 4 dimensions,
while annotator age is projected to a 4-dimensional vector. These embeddings are aggregated via a
4-head multi-head attention layer to produce a unified 32-dimensional metadata summary. During
Stage 1, we produce a 768-dimensional CLS embedding from both the original and cross-translated
tweets, fuse them using an alpha weight ( enc = 0.8), concatenate the result with the metadata
embedding (yielding an 800-dimensional vector), and train a binary classifier head ( Linear(800→1))
using BCEWithLogitsLoss for 2 epochs (batch size = 16, learning rate = 2 × 10− 5, dropout = 0.1).
Encoder weights are then saved and reloaded for Stage 2, which adds two additional classification
heads: a 4-class head for Task 1_2 and a 6-class head for Task 1_3. The Stage 1 predicted probability is
embedded (1→4 dimensions) and concatenated (resulting in 804 dimensions). Stage 2 training occurs
over 3 epochs (batch size = 16, learning rate = 2 × 10− 5), using CrossEntropyLoss (Task 1_2) and
BCEWithLogitsLoss (Task 1_3). Inference for both stages performs late fusion of logits (original
vs. cross-translated) with  inf = 0.8. All models are run on GPUs with random seeds fixed to 42 for
reproducibility.</p>
        <p>Our second system uses a quantized, autoregressive LLaMA-3.1-8B-Instruct model with 4-bit NF4
quantization and double quantization (compute dtype = torch.float16). This model, along with
its corresponding tokenizer (maximum prompt length = 256), is used in a zero-shot setting without
ifne-tuning. Tweets and annotator metadata are structured into textual prompts, and predictions are
generated directly from model-derived token log probabilities. Original and cross-translated predictions
are combined using late fusion ( = 0.55), producing both soft (probabilities) and hard
(binary/categorical) outputs. Inference occurs on a GPU with a fixed random seed of 42. Predictions are saved as CSV
ifles and later converted to JSON format for evaluation.</p>
        <p>The third system mirrors the second but replaces the autoregressive LLaMA model with the
difusionbased Dream-v0-Instruct-7B model, using identical quantization settings (4-bit NF4, compute dtype =
torch.float16) and tokenizer parameters. The fusion weight for combining original and translated
log probabilities is set slightly higher ( = 0.65). As in the second system, inference runs without
ifne-tuning, and results are similarly post-processed from CSV to JSON format.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Results</title>
        <p>Tables 3,4, and5 summarize the evaluation results on the development set across the three tasks (Task
1_1, Task 1_2, and Task 1_3). The evaluation metrics include ICM, ICM-Norm, F1 Score, ICM-Soft,
ICM-Soft-Norm, and Cross Entropy [7, 8], with higher scores indicating better performance except for
Cross Entropy, where lower values are preferable.</p>
        <p>For Task 1_1 (Table 3), the Transformer-Based system achieves the best results in the hard label
evaluation metrics (ICM, ICM-Norm, and F1 Score). Interestingly, in the soft probability metrics
(ICMSoft, ICM-Soft-Norm, and Cross Entropy), both LLM-based systems outperform the Transformer-Based
model. Specifically, the Difusion LLM slightly surpasses the AR LLM in ICM-Soft and ICM-Soft-Norm,
whereas the AR LLM achieves the lowest Cross Entropy.</p>
        <p>In Task 1_2 (Table 4), the Transformer-Based system consistently outperforms both LLM-based
systems across all metrics for both hard and soft evaluations, while the two LLM systems provide similar
performance. The Difusion LLM performs slightly better on ICM and ICM-Norm, while the AR LLM
achieves a higher F1 Score in the hard metrics. In the soft metrics, the AR LLM system significantly
outperforms the Difusion-based model.</p>
        <p>For Task 1_3 (Table 5), the Transformer-Based system achieves superior performance across all
metrics. The AR LLM system ranks second, followed by the Difusion LLM system, indicating a clear
performance advantage for the Transformer-based architecture on this complex multi-label task.</p>
        <p>These results reflect the best performance achieved following extensive ablation studies. These
studies explored multiple configurations, including data pre-processing variants (original vs. fully
translated vs. cross-translated tweets), diferent fusion strategies (equal vs. adjustable weighting), epoch
counts, and various early and late fusion  values.</p>
        <p>Overall, our evaluation demonstrates that the Transformer-Based dual-encoder model ofers robust
and superior performance on most metrics and tasks. Meanwhile, the novel Difusion LLM-based
system shows performance comparable to or slightly better than the more established AR LLM-based
approach on selected metrics, indicating promising potential for difusion-based LLMs in challenging
classification scenarios.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Final Rankings</title>
        <p>This section presents the final rankings and oficial results of our systems evaluated on the test set.
Tables 6, 7, and 8 show the published rankings from the organizers across all three tasks. Our three
submitted systems correspond to: System 1 (run 1) — the Transformer-based model; System 2 (run 2)
— the AR LLM-based model; and System 3 (run 3) — the Difusion LLM-based model. The number of
participating submissions per task varies: for hard parameter evaluation, there were respectively 160,
140, and 132 submissions for Task1_1, Task1_2, and Task1_3, as opposed to 67, 56, and 53 for the soft
parameter evaluation.</p>
        <p>From the oficial test results, it is evident that System 1 (transformer-based) consistently outperformed
the other two across all tasks, securing the highest rankings and results metrics among our submissions.
When comparing the two LLM-based systems, System 3 (Difusion LLM-based) sometimes outperformed
System 2 (AR LLM-based), depending on the task and language-specific setting.</p>
        <p>Overall, these final test rankings align with our development set findings. The transformer-based
system demonstrated superior performance across the board, while the difusion-based LLM showed
competitive results, occasionally matching or even surpassing the AR LLM-based system. This highlights
the potential of difusion models as viable alternatives for multilingual, multi-task text classification.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>We introduce an annotator-aware, bilingual sexism detection pipeline evaluated on the EXIST-2025
shared task. Our methodology explicitly addresses annotator disagreement by incorporating annotator
demographics directly into the modeling process, enhancing prediction robustness through the bilingual
fusion of original and cross-translated tweets. We implemented and evaluated three distinct modeling
strategies: a transformer-based dual-encoder architecture with fine-tuning and fusion, a zero-shot
auto-regressive LLM, and a zero-shot difusion-based LLM.</p>
      <p>Experimental results clearly indicate that our transformer-based approach delivers superior
performance across all three EXIST subtasks —sexism detection, misogynistic intent classification, and sexism
type categorization— and this advantage is especially evident in metrics such as ICM, ICM-Norm, and F1
scores. While the AR LLM performed efectively in zero-shot scenarios, the difusion-based LLM notably
demonstrated competitive performance, occasionally surpassing the AR LLM on selected metrics. This
ifnding suggests that difusion-based LLMs hold considerable potential for text classification tasks,
warranting further exploration and optimization. Our findings were also confirmed by the final rankings
published by the organizers.</p>
      <p>Future research directions can include deeper integration of annotator disagreement into the
difusionbased models and hybrid architectures combining transformer encoders with difusion mechanisms.
Extending the proposed pipeline to multimedia contexts, such as image-based memes and video content,
could help broaden the applicability and generalization capabilities of automated sexism detection
systems.</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used Open AI ChatGPT model GPT-4.5 for: Abstract
drafting, drafting content, generating the literature review, grammar and spelling checking, paraphrasing
and rewording, and plagiarism detection. After using these tools, the authors reviewed and edited the
content as needed and take full responsibility for the content.
characterization in tweets and memes, in: L. Goeuriot, et al. (Eds.), Experimental IR Meets
Multilinguality, Multimodality, and Interaction, Lecture Notes in Computer Science, Springer
Nature Switzerland, Cham, 2024, pp. 93–117. To appear.
[6] L. Plaza, J. C. de Albornoz, R. Morante, E. Amigó, J. Gonzalo, D. Spina, P. Rosso, Overview of
exist 2024 – learning with disagreement for sexism identification and characterization in tweets
and memes (extended overview), in: Working Notes of CLEF 2024 – Conference and Labs of the
Evaluation Forum, volume 3740 of CEUR-WS.org, CEUR Workshop Proceedings, Grenoble, France,
2024. URL: https://ceur-ws.org/Vol-3740/, presented at CLEF 2024, September 9–12, 2024.
[7] L. Plaza, J. C. de Albornoz, I. Arcos, P. Rosso, D. Spina, E. Amigó, J. Gonzalo, R. Morante, Overview
of exist 2025: Learning with disagreement for sexism identification and characterization in tweets,
memes, and tiktok videos, in: J. C. de Albornoz, J. Gonzalo, L. Plaza, A. G. S. de Herrera, J. Mothe,
F. Piroi, P. Rosso, D. Spina, G. Faggioli, N. Ferro (Eds.), Experimental IR Meets Multilinguality,
Multimodality, and Interaction, Proceedings of the Sixteenth International Conference of the CLEF
Association (CLEF 2025), 2025.
[8] L. Plaza, J. C. de Albornoz, I. Arcos, P. Rosso, D. Spina, E. Amigó, J. Gonzalo, R. Morante, Overview
of exist 2025: Learning with disagreement for sexism identification and characterization in tweets,
memes, and tiktok videos (extended overview), in: G. Faggioli, N. Ferro, P. Rosso, D. Spina (Eds.),
CLEF 2025 Working Notes, 2025.
[9] E. Fersini, D. Nozza, P. Rosso, Overview of the evalita 2018 task on automatic misogyny
identification (ami), in: Proceedings of the Sixth Evaluation Campaign of Natural Language Processing
and Speech Tools for Italian (EVALITA 2018), volume 2263 of CEUR Workshop Proceedings,
CEURWS.org, Turin, Italy, 2018. URL: http://ceur-ws.org/Vol-2263/.
[10] V. Basile, C. Bosco, E. Fersini, D. Nozza, V. Patti, F. M. Rangel Pardo, P. Rosso, M. Sanguinetti,
SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in
Twitter, in: J. May, E. Shutova, A. Herbelot, X. Zhu, M. Apidianaki, S. M. Mohammad (Eds.), Proceedings
of the 13th International Workshop on Semantic Evaluation, Association for Computational
Linguistics, Minneapolis, Minnesota, USA, 2019, pp. 54–63. URL: https://aclanthology.org/S19-2007/.
doi:10.18653/v1/S19-2007.
[11] A. Uma, T. Fornaciari, A. Dumitrache, T. Miller, J. Chamberlain, B. Plank, E. Simpson, M. Poesio,
SemEval-2021 task 12: Learning with disagreements, in: A. Palmer, N. Schneider, N. Schluter,
G. Emerson, A. Herbelot, X. Zhu (Eds.), Proceedings of the 15th International Workshop on
Semantic Evaluation (SemEval-2021), Association for Computational Linguistics, Online, 2021, pp. 338–
347. URL: https://aclanthology.org/2021.semeval-1.41/. doi:10.18653/v1/2021.semeval-1.
41.
[12] E. Leonardelli, G. Abercrombie, D. Almanea, V. Basile, T. Fornaciari, B. Plank, V. Rieser, A. Uma,
M. Poesio, SemEval-2023 task 11: Learning with disagreements (LeWiDi), in: A. K. Ojha, A. S.
Doğruöz, G. Da San Martino, H. Tayyar Madabushi, R. Kumar, E. Sartori (Eds.), Proceedings
of the 17th International Workshop on Semantic Evaluation (SemEval-2023), Association for
Computational Linguistics, Toronto, Canada, 2023, pp. 2304–2318. URL: https://aclanthology.org/
2023.semeval-1.314/. doi:10.18653/v1/2023.semeval-1.314.
[13] A. Uma, T. Fornaciari, D. Hovy, S. Paun, B. Plank, M. Poesio, Learning from disagreement: A
survey, International Journal of Language, Translation and Intercultural Communication 10 (2021)
1392–1470. URL: https://doi.org/10.12681/ijltic.29238.
[14] Y.-Z. Fang, L.-H. Lee, J.-D. Huang, Nycu-nlp at exist 2024: Leveraging transformers with diverse
annotations for sexism identification in social networks, in: Working Notes of CLEF 2024 –
Conference and Labs of the Evaluation Forum, CEUR Workshop Proceedings, CEUR-WS.org,
Grenoble, France, 2024. URL: http://ceur-ws.org/, notebook for the EXIST Lab at CLEF 2024.
[15] L. M. Quan, D. V. Thin, Sexism identification in social networks with generation-based language
models: Notebook for the exist lab at clef 2024, in: Working Notes of CLEF 2024 – Conference and
Labs of the Evaluation Forum, CEUR Workshop Proceedings, CEUR-WS.org, Grenoble, France,
2024. URL: http://ceur-ws.org/, to appear.
[16] M. Del Tredici, D. Marcheggiani, S. Schulte im Walde, R. Fernández, You shall know a user by the
company it keeps: Dynamic representations for social media users in NLP, in: K. Inui, J. Jiang, V. Ng,
X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language
Processing and the 9th International Joint Conference on Natural Language Processing
(EMNLPIJCNLP), Association for Computational Linguistics, Hong Kong, China, 2019, pp. 4707–4717. URL:
https://aclanthology.org/D19-1477/. doi:10.18653/v1/D19-1477.
[17] B. Vidgen, A. Harris, D. Nguyen, R. Tromble, S. Hale, H. Margetts, Challenges and frontiers
in abusive content detection, in: S. T. Roberts, J. Tetreault, V. Prabhakaran, Z. Waseem (Eds.),
Proceedings of the Third Workshop on Abusive Language Online, Association for Computational
Linguistics, Florence, Italy, 2019, pp. 80–93. URL: https://aclanthology.org/W19-3509/. doi:10.
18653/v1/W19-3509.
[18] A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott,
L. Zettlemoyer, V. Stoyanov, Unsupervised cross-lingual representation learning at scale, in:
D. Jurafsky, J. Chai, N. Schluter, J. Tetreault (Eds.), Proceedings of the 58th Annual Meeting of the
Association for Computational Linguistics, Association for Computational Linguistics, Online,
2020, pp. 8440–8451. URL: https://aclanthology.org/2020.acl-main.747/. doi:10.18653/v1/2020.
acl-main.747.
[19] D. Nozza, F. Bianchi, D. Hovy, Pipelines for social bias testing of large language models, in:
A. Fan, S. Ilic, T. Wolf, M. Gallé (Eds.), Proceedings of BigScience Episode #5 – Workshop on
Challenges &amp; Perspectives in Creating Large Language Models, Association for Computational
Linguistics, virtual+Dublin, 2022, pp. 68–74. URL: https://aclanthology.org/2022.bigscience-1.6/.
doi:10.18653/v1/2022.bigscience-1.6.
[20] R. Sennrich, B. Haddow, A. Birch, Improving neural machine translation models with monolingual
data, in: K. Erk, N. A. Smith (Eds.), Proceedings of the 54th Annual Meeting of the Association for
Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics,
Berlin, Germany, 2016, pp. 86–96. URL: https://aclanthology.org/P16-1009/. doi:10.18653/v1/
P16-1009.
[21] T. O. Ajayi, M. Arcan, P. Buitelaar, Cross-lingual transfer and multilingual learning for detecting
harmful behaviour in African under-resourced language dialogue, in: T. Kawahara, V. Demberg,
S. Ultes, K. Inoue, S. Mehri, D. Howcroft, K. Komatani (Eds.), Proceedings of the 25th Annual
Meeting of the Special Interest Group on Discourse and Dialogue, Association for Computational
Linguistics, Kyoto, Japan, 2024, pp. 579–589. URL: https://aclanthology.org/2024.sigdial-1.49/.
doi:10.18653/v1/2024.sigdial-1.49.
[22] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan,
P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child,
A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray,
B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, D. Amodei,
Language models are few-shot learners, in: H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan,
H. Lin (Eds.), Advances in Neural Information Processing Systems, volume 33, Curran
Associates, Inc., 2020, pp. 1877–1901. URL: https://proceedings.neurips.cc/paper_files/paper/2020/file/
1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
[23] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal,
E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, G. Lample, Llama: Open and eficient
foundation language models, 2023. URL: https://arxiv.org/abs/2302.13971. arXiv:2302.13971.
[24] T. Hartvigsen, S. Gabriel, H. Palangi, M. Sap, D. Ray, E. Kamar, ToxiGen: A large-scale
machine-generated dataset for adversarial and implicit hate speech detection, in: S. Muresan,
P. Nakov, A. Villavicencio (Eds.), Proceedings of the 60th Annual Meeting of the Association
for Computational Linguistics (Volume 1: Long Papers), Association for Computational
Linguistics, Dublin, Ireland, 2022, pp. 3309–3326. URL: https://aclanthology.org/2022.acl-long.234/.
doi:10.18653/v1/2022.acl-long.234.
[25] S. Gong, M. Li, J. Feng, Z. Wu, L. Kong, DifuSeq: Sequence to sequence text generation with
difusion models, in: International Conference on Learning Representations, ICLR, 2023.
[26] J. Austin, D. D. Johnson, J. Ho, D. Tarlow, R. van den Berg, Structured denoising difusion
models in discrete state-spaces, in: M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, J. W.
Vaughan (Eds.), Advances in Neural Information Processing Systems, volume 34, Curran
Associates, Inc., 2021, pp. 17981–17993. URL: https://proceedings.neurips.cc/paper_files/paper/2021/
ifle/958c530554f78bcd8e97125b70e6973d-Paper.pdf.
[27] J. Ye, Z. Xie, L. Zheng, J. Gao, Z. Wu, X. Jiang, Z. Li, L. Kong, Dream 7b, 2025. URL: https:
//hkunlp.github.io/blog/2025/dream.
[28] S. Nie, F. Zhu, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y. Lin, J.-R. Wen, C. Li, Large language
difusion models, arXiv preprint arXiv:2502.09992 (2025).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rodríguez-Sánchez</surname>
          </string-name>
          , J. C. de Albornoz, L. Plaza,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Comet</surname>
          </string-name>
          , T. Donoso, Overview of exist 2021:
          <article-title>sexism identification in social networks</article-title>
          ,
          <source>in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021) at SEPLN</source>
          <year>2021</year>
          , volume
          <volume>2943</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Málaga, Spain,
          <year>2021</year>
          , pp.
          <fpage>195</fpage>
          -
          <lpage>207</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2943</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rodríguez-Sánchez</surname>
          </string-name>
          , J. C. de Albornoz,
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mendieta-Aragón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Marco-Remón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Makeienko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of exist 2022:
          <article-title>sexism identiifcation in social networks</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>69</volume>
          (
          <year>2022</year>
          )
          <fpage>229</fpage>
          -
          <lpage>240</lpage>
          . URL: http: //journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6443. doi:
          <volume>10</volume>
          .26342/2022-69-20.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          , J. C. de Albornoz,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <article-title>Overview of exist 2023 - learning with disagreement for sexism identification and characterization</article-title>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
          </string-name>
          , et al. (Eds.),
          <article-title>Experimental IR Meets Multilinguality, Multimodality, and</article-title>
          <string-name>
            <surname>Interaction. CLEF</surname>
          </string-name>
          <year>2023</year>
          , volume
          <volume>14163</volume>
          of Lecture Notes in Computer Science, Springer, Cham,
          <year>2023</year>
          , pp.
          <fpage>398</fpage>
          -
          <lpage>418</lpage>
          . URL: https://doi.org/10.1007/978-3-
          <fpage>031</fpage>
          -42448-9_
          <fpage>23</fpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -42448-9_
          <fpage>23</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          , J. C. de Albornoz,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <article-title>Overview of exist 2023 - learning with disagreement for sexism identification and characterization (extended overview)</article-title>
          ,
          <source>in: Working Notes of CLEF 2023 - Conference and Labs of the Evaluation Forum</source>
          , volume
          <volume>3497</volume>
          <source>of CEUR-WS.org, CEUR Workshop Proceedings</source>
          , Thessaloniki, Greece,
          <year>2023</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3497</volume>
          /, presented at CLEF 2023, September 18-
          <issue>21</issue>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. C. de Albornoz</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Ruiz</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Maeso</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chulvi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Amigó</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Morante</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Spina</surname>
          </string-name>
          ,
          <article-title>Overview of exist 2024 - learning with disagreement for sexism identification and</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>