<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Modalities for Hate Speech Detection in Memes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fadi Hassan</string-name>
          <email>fadi.hassan@huawei.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Evgenii Migaev</string-name>
          <email>evgenii.migaev2@huawei.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Md Saroar Jahan</string-name>
          <email>saroar.jahan@huawei.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Frank Mtumbuka</string-name>
          <email>frank.mtumbuka@h-partners.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Hate Speech Detection, Multimodal Classification, Memes, Few-Shot Learning</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Huawei Finland Research Center</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <abstract>
        <p>This paper presents our system developed by our research team at FiRC-NLP for the HASOC-meme 2025 shared task, which focuses on detecting hate speech and ofensive content in multilingual memes. We explore three distinct approaches: prompt-based inference using Gemini, cross-modality encoders combining image and text, and a text-only modality leveraging Optical Character Recognition (OCR), OCR English translation and image descriptions. Our final ensemble system fuses these modalities, achieving macro  1-scores up to 67.50 on test sets and top-1 ranking in 3 out of 4 tracks, demonstrating the efectiveness of multimodal fusion and robust training strategies.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Memes pose unique challenges for hate speech detection due to their multimodal nature [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Visual
elements often carry implicit ofensive cues that are not captured by text alone [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. To address this, we
designed a system that explores multiple modalities and fusion strategies, aiming to maximize semantic
coverage and classification accuracy. The proliferation of memes on social media has established them
as a dominant mode of online communication, capable of conveying ideas through the interplay of
text and imagery. While memes are often used for entertainment and cultural commentary, they are
increasingly exploited to disseminate ofensive and hateful content [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. Detecting such harmful
content presents unique challenges compared to traditional text-based hate speech detection. This
dificulty arises from the multimodal nature of memes: textual components may appear benign in
isolation but gain ofensive meaning when combined with visual cues, and conversely, images may
implicitly reinforce or alter the semantics of overlaid text. As a result, efective detection requires
approaches that can account for the rich cross-modal interactions inherent in memes [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        Recent advances in multimodal machine learning have motivated research into systems capable of
integrating textual and visual signals for classification tasks. Prior work has explored multimodal fusion
for hate detection [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], including CLIP-based approaches [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and vision-language pretraining [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
However, challenges like linguistic diversity in datasets such as MUTE [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], and language-specific
resources such as BanglaAbuseMeme [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] persist. Designing robust systems for hate speech detection
in memes remains an open problem due to issues such as noisy or stylized embedded text, linguistic
diversity, implicit visual symbolism, and the limited availability of annotated datasets. Related eforts
on hate mitigation in Indic languages, such as SafeSpeech [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], further highlight the importance of
addressing linguistic and cultural diversity. To address these challenges, the HASOC-meme 2025 shared
task provides a benchmark for evaluating multimodal approaches to hate speech and ofensive content
detection in memes.
      </p>
      <p>In this paper, we present our system developed for the shared task, which combines multiple
modalities and fusion strategies to maximize semantic coverage. Specifically, we explore three complementary
approaches: (i) prompt-based inference with large multimodal models, (ii) cross-modality encoders</p>
      <p>CEUR
Workshop</p>
      <p>ISSN1613-0073
jointly processing images and text, and (iii) text-only pipelines enhanced with optical character
recognition (OCR), translation into English, and image description generation. We further investigate ensemble
strategies that integrate these modalities, demonstrating that their combination yields significant
performance improvements over single-modality baselines.</p>
      <p>Our contributions are threefold:
• We design and evaluate distinct modality-specific pipelines for meme classification.
• We propose a fusion strategy that leverages the complementary strengths of prompt-based,
cross-modal, and text-enriched approaches.
• We provide an empirical analysis of system performance on the HASOC-meme dataset, achieving
top-1 ranking in three out of four tracks of the competition.</p>
      <p>These findings underline the importance of multimodal integration for robust hate speech detection
in memes and suggest promising directions for future research, such as adaptive fusion strategies and
meme-specific captioning models.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task Description</title>
      <p>
        The HASOC shared task series (Hate Speech and Ofensive Content Identification) has reached its seventh
edition in 2025. The HASOC-meme 2025 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] task extends previous editions by focusing specifically
on memes, which combine textual and visual modalities, thereby presenting unique challenges for
automatic classification. Unlike purely textual corpora, memes often encode ofensive cues through
both explicit and implicit multimodal signals such as sarcastic phrasing, symbolic imagery, or subtle
cultural references.
      </p>
      <p>The 2025 edition introduces a five-part classification challenge, requiring participants to analyze
memes across multiple dimensions of ofensive communication:
• Positive: The meme conveys support, humor, or appreciation.
• Neutral: The meme is neither overtly positive nor negative.</p>
      <p>• Negative: The meme expresses hostility, mockery, or criticism.</p>
      <sec id="sec-2-1">
        <title>Sentiment Detection</title>
      </sec>
      <sec id="sec-2-2">
        <title>Sarcasm Detection</title>
      </sec>
      <sec id="sec-2-3">
        <title>Vulgarity Detection</title>
        <p>Abuse Detection
• Abusive: The meme includes harmful, derogatory, or ofensive elements targeting individuals or
groups.
• Non-Abusive: The meme lacks abusive or derogatory content.
• Sarcastic: The meme implies the opposite of its literal meaning, often mocking or ridiculing.
• Non-Sarcastic: The meme conveys its message directly, without irony.
• Vulgar: The meme contains explicit or ofensive words, gestures, or depictions.
• Not Vulgar: The meme does not include such content.</p>
        <p>Target Community Identification
• Gender: Mentions gender identities (male, female, non-binary, transgender).
• Religion: References to religious beliefs, practices, or symbols.
• Individual: Mentions or depicts a specific person.
• Political: Refers to political parties, ideologies, politicians, or policies.
• National Origin: Targets groups by country or ethnicity.
• Social Sub-groups: Refers to socio-economic, occupational, or cultural groups.
• Others: Any community not covered above.</p>
        <p>• None: No explicit target community.</p>
        <p>Through this multi-dimensional design, HASOC-meme 2025 emphasizes the complexity of abusive
content detection, encouraging participants to go beyond binary classification and address fine-grained
aspects such as sarcasm, sentiment, and vulgarity.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset and Preprocessing</title>
      <p>We used the oficial HASOC-meme dataset. To evaluate and compare diferent strategies, we employed
5-fold cross-validation, partitioning the training data into five subsets and iteratively using four folds
for training and one for validation.</p>
      <p>
        While the organizers provided baseline OCR text, we hypothesized that a more detailed extraction of
visual and textual elements would improve performance [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]. We developed an advanced feature
extraction pipeline using the cost-efective and fast Gemini 2.5 Flash-Lite [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] model, which involved
the following steps:
1. Image Decomposition: Each meme image was programmatically divided into several distinct
sub-images to isolate diferent visual components.
2. Detailed Extraction: For each sub-image, we prompted the model to perform OCR and generate
a concise description of the visual content.
3. Aggregation and Translation: The extracted texts and descriptions from all sub-images were
aggregated. The combined OCR text was also translated into English to create a unified linguistic
representation for subsequent retrieval tasks.
      </p>
      <p>This process yielded a rich set of features for each meme: the original image, a comprehensive OCR text
(in its native language), an English translation of the OCR, and a structured description of the visual
elements.</p>
    </sec>
    <sec id="sec-4">
      <title>4. System description</title>
      <p>One of the key challenges in this work was building a strong classifier from limited training data in
low-resource languages. To address this, we unified all datasets into a single multilingual corpus and
explored two complementary strategies: (1) prompt-based inference using large multimodal foundation
models, and (2) fine-tuning compact multi- and single-modal encoders. Our system integrates multiple
components, detailed in the following sections:</p>
      <sec id="sec-4-1">
        <title>4.1. Prompt-Based Inference</title>
        <p>Our core methodology revolved around prompt-based inference using a powerful multimodal foundation
model. We explored both zero-shot and few-shot paradigms to leverage the model’s capabilities.</p>
        <sec id="sec-4-1-1">
          <title>4.1.1. Foundation Model Selection</title>
          <p>Our initial step was to select the most suitable foundation model. We conducted a preliminary evaluation
on a small subset of 20 training samples, testing OpenAI’s GPT-4, Anthropic’s Claude 3 Opus, and
Google’s Gemini 2.5 Flash. Each model was prompted to classify the meme images across the four
target labels. In this comparison, Gemini 2.5 Flash demonstrated consistently superior performance,
leading us to select it for all subsequent experiments.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.1.2. Zero-Shot Classification</title>
          <p>Our first strategy was a zero-shot approach, designed to test the model’s intrinsic understanding of the
task without any prior examples. The model was provided with the raw meme image and a carefully
crafted prompt instructing it to classify the image according to the four labels (sentiment, abuse, vulgar,
sarcasm). The prompt 1 also required the model to provide a brief explanation for its decision, which
ofered valuable qualitative insights into its reasoning process. This approach used no samples from
our prepared training set.</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>4.1.3. Few-Shot In-Context Learning via Retrieval</title>
          <p>Our second, more advanced strategy was a few-shot approach that provided the model with relevant
examples from the training set to guide its prediction. This method, often referred to as
RetrievalAugmented Generation (RAG), relied on our extracted features rather than the raw images. The process
was as follows:
1. Vector Database Creation: We used a Gemini text embedding model to generate vector
embeddings for the English-translated OCR and the image descriptions of all samples in the
training set. These embeddings were stored in a Chroma vector database.
2. Dynamic Example Retrieval: For each new sample, we performed two similarity searches
against the vector database to retrieve the top 5 training examples with the most similar translated
OCR and the top 5 examples with the most similar image description.
3. Prompt Engineering: The retrieved 10 examples, along with their ground-truth labels, were
concatenated and formatted into the prompt. The model was then tasked with classifying the
target sample based on this contextual information. Prompt 2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Cross-Modality Encoders</title>
        <p>
          We fused the vision encoder from google/siglip-base-patch16-224 (SigLIP) [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] with
XLM-RoBERTalarge (XLM-R) [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] in a similar way to Figure 4.3, leveraging the strengths of both modalities: SigLIP
provides robust visual representations, while XLM-R excels in handling multilingual text, especially in
low-resource languages. This hybrid design allowed us to build a more efective cross-modal encoder
capable of capturing nuanced relationships between meme images and their associated text across
diverse linguistic contexts.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Text Modality via OCR and Description</title>
        <p>
          We used the OCR of the images and applied Gemini (using Prompt 3) to generate detailed image
descriptions and translate the content into English. The outputs from Gemini and the OCR process
were then fed into two language models: XLM-R and l3cube-pune/bengali-bert (l3cube-bangla) [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], for
classification, as illustrated in Figure 4.3.
        </p>
        <sec id="sec-4-3-1">
          <title>Abuse</title>
        </sec>
        <sec id="sec-4-3-2">
          <title>Vulgar</title>
        </sec>
        <sec id="sec-4-3-3">
          <title>Sarcasm</title>
        </sec>
        <sec id="sec-4-3-4">
          <title>Sentiment</title>
          <p>XLM-R
l3Cube-Bangla</p>
          <p>AbuseJoint ClassifiVcuatligoanr Head</p>
        </sec>
        <sec id="sec-4-3-5">
          <title>Sarcasm</title>
        </sec>
        <sec id="sec-4-3-6">
          <title>Sentiment</title>
          <p>JCoiLnSt PColoalsinsigfication Head</p>
        </sec>
        <sec id="sec-4-3-7">
          <title>CLS Pooling</title>
        </sec>
        <sec id="sec-4-3-8">
          <title>Image OCR</title>
        </sec>
        <sec id="sec-4-3-9">
          <title>Image Description</title>
        </sec>
        <sec id="sec-4-3-10">
          <title>OCR EN Translation</title>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Fusion Strategy for Final Submission</title>
        <p>For the final submission, we adopted a hybrid selection strategy to maximize the overall macro 
1score. For each of the 16 tasks (4 languages × 4 labels), we compared the performance of all our methods,
including zero-shot LLM, few-shot LLM, and our encoder-based models in the internal validation set.</p>
        <p>The approach that yields the highest validation  1 score for a specific task was selected to generate
the final prediction for the test set. For encoder-based models, we further optimized the  1 score by
setting a prediction threshold for each label based on its positive class distribution in the training
data. This ensured that the predicted label ratios were aligned with the observed data. This strategic
combination allowed us to take advantage of the distinct strengths of each modeling technique.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experimental Setup</title>
      <p>For the fine-tuned models, we employed 5-fold cross-validation, where four splits were used for training
and one split for validation (Valid) in each fold. At the end of training, predictions from the five
folds were combined into an ensemble to score the test set. All models were implemented using the
HuggingFace Transformers framework and optimized with early stopping to prevent overfitting. A
constant learning rate schedule with warmup was applied to enhance training stability and overall
performance.</p>
      <p>We trained the models on four out of the five classes (sentiment, sarcasm, vulgarity, and abusive),
excluding the target community class since it was not requested in the final submission and, in preliminary
experiments, its inclusion degraded performance due to potentially low-quality annotations.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Results and Discussion</title>
      <p>We evaluate the performance of multiple approaches across Bangla, Bodo, Gujarati, and Hindi, reporting
macro  1-scores on both validation and test sets. Table 2 presents the performance comparison of
diferent approaches.</p>
      <p>Language →
Approach ↓ Backbone
Prompt-Based (Zero-Shot) Gemini
Prompt-Based (Few-Shot) Gemini
Cross-Modality Encoder XLM-R + SigLIP
Text Modality (OCR only) XLM-R + l3cube-Bangla
Text Modality (OCR + Image
Description + OCR Translation) XLM-R + l3cube-Bangla
Combined Ensemble</p>
      <p>Bangla
Training set Valid Test
- 58.10 59.45
all 52.34 54.04
all 55.95 56.48
all 58.56 59.48
all 61.02 59.91
Bodo -
all - 62.75</p>
      <sec id="sec-6-1">
        <title>6.1. Prompt-Based Methods</title>
        <p>Zero-shot prompt-based approaches using Gemini perform surprisingly well for the test set of the
Bangla language, outperforming the fine-tuned models. However, when applied to other languages, its
performance tends to be more modest. On the other hand, few-shot prompting occasionally outperforms
zero-shot prompting but not consistently. This variability makes it challenging to draw a clear conclusion
about its overall efectiveness.</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Text-Modality Encoders</title>
        <p>
          Approaches leveraging text-only encoders (XLM-R + l3cube-Bangla) generally outperform prompt-based
strategies, particularly for Gujarati and Hindi. However, their performance varies across languages,
with weaker results observed for Bodo. The performance on Bodo was notably low when the model
was trained jointly with other languages. We hypothesize that this may be due to inconsistencies
in the annotation of Bodo data, warranting further investigation of the annotation process. Similar
low-resource challenges were also discussed in prior shared task analyses [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. To address this, we
trained a monolingual classifier for Bodo, which resulted in improved performance.
        </p>
      </sec>
      <sec id="sec-6-3">
        <title>6.3. Cross-Modality Fusion</title>
        <p>Models combining multimodal cues (OCR, image descriptions, and OCR translations) consistently
outperform text-only methods. This fusion strategy leads to the strongest results, indicating that
meme-based hate speech detection benefits substantially from integrating multimodal signals.</p>
      </sec>
      <sec id="sec-6-4">
        <title>6.4. Combined Ensemble</title>
        <p>The ensemble approach achieves the best overall macro  1-score. These results confirm that
ensemblebased strategies, which leverage diverse models and representations, provide robust improvements
across all languages.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>Our system demonstrates that combining prompt-based reasoning, cross-modality encoding, and
enriched text-only pipelines yields strong performance in meme-based hate speech detection.</p>
      <p>Future work will focus on exploring dynamic fusion strategies and fine-tuning captioning models
for meme-specific semantics. In addition, we plan to investigate encoder fusion using cross-attention
mechanisms, as the current implementation relies solely on late aggregation (where inputs are
independently processed through encoders and fused only at the end). Cross attention has the potential to
capture richer cross-encoder interactions and improve overall system efectiveness.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Acknowledgments</title>
      <p>We would like to thank the HASOC team and all participants for their eforts in organizing and
contributing to the shared task.</p>
    </sec>
    <sec id="sec-9">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the author(s) used GPT-4-based tools in order to: Grammar and
spelling check. After using this tool, the author(s) reviewed and edited the content as needed and take(s)
full responsibility for the publication’s content.</p>
    </sec>
    <sec id="sec-10">
      <title>A. Prompts in YAML Format</title>
      <p>Listing 1: Prompt used for the Zero-Shot approach
prompt: &gt;</p>
      <p>You are a helpful content moderator. You need to analyze meme images in Bangla, Hindi,</p>
      <p>Gujarati, or Bodo.</p>
      <p>You should detect:
1) sentiment (one from 3 values):
- positive - The meme conveys a supportive, humorous, or appreciative tone.
- neutral - The meme is neither overtly positive nor negative in tone.</p>
      <p>- negative - The meme expresses hostility, mockery, or criticism.
2) sarcasm (True or False):</p>
      <p>- True - The meme presents statements or visuals that imply the opposite of their literal
meaning, often to mock or ridicule.</p>
      <p>- False - The meme directly conveys its message without sarcasm or irony.
3) vulgar (True or False):
- True - The meme contains explicit or offensive words, gestures, or depictions.
- False - The meme does not include any such content.
4) abuse (True or False):</p>
      <p>- True - The meme includes offensive, harmful, or derogatory language, imagery, or
implications targeting an individual or a group.</p>
      <p>- False - The meme does not contain any offensive, harmful, or derogatory content.
5) target - (one from 8 values):
- "gender" - Any reference to male, female, non-binary, or transgender identities.
- "religion" - Mentions or imagery related to any religious belief, deity, or practice.
- "individual" - Specifically mentions or portrays a particular person.
- "political" - Targets political ideologies, parties, politicians, or policies.
- "national" - Targets people based on their country or ethnicity.</p>
      <p>- "social subgroups" - Groups based on socio-economic status, occupation, cultural
identity, or other affiliations.</p>
      <p>- "other" - Any target that does not fall into the above categories.</p>
      <p>- "non-targeted" - If the meme does not target any specific community, no target label is
assigned.
6) description - finally provide a short explanation of why you make these predictions.</p>
      <p>Listing 2: Prompt used for the Few-Shot approach
prompt_few_shots: &gt;</p>
      <p>You are a helpful content moderator. You need to analyze meme images in Bangla, Hindi,</p>
      <p>Gujarati or Bodo.</p>
      <p>I will provide you description of every part of the meme (images and texts)
You should detect:
1) sentiment (one from 3 values):
- positive - The meme conveys a supportive, humorous, or appreciative tone.
- neutral - The meme is neither overtly positive nor negative in tone.</p>
      <p>- negative - The meme expresses hostility, mockery, or criticism.
2) sarcasm (True or False):</p>
      <p>- True - The meme presents statements or visuals that imply the opposite of their literal
meaning, often to mock or ridicule.</p>
      <p>- False - The meme directly conveys its message without sarcasm or irony.
3) vulgar (True or False):
- True - The meme contains explicit or offensive words, gestures, or depictions.
- False - The meme does not include any such content.
4) abuse (True or False):</p>
      <p>- True - The meme includes offensive, harmful, or derogatory language, imagery, or
implications targeting an individual or a group.</p>
      <p>- False - The meme does not contain any offensive, harmful, or derogatory content.
Use the examples below as a guide:
{examples}
Meme description to analyze:
{meme_description}</p>
      <p>Text description: {text_description}
description_ocr_template: &gt;</p>
      <p>OCR from a meme: {ocr}
example_template: &gt;
-- {description}
Ground Truth: {ground_truth}</p>
      <p>Listing 3: Prompt used for Image OCR + Image Description + OCR Translation extraction
ocr_prompt: &gt;</p>
      <p>You are an expert at analyzing and describing internet meme images in Bangla, Hindi, Gujarati,
or Bodo. Your task is to extract the text and provide a concise description of the image
content for each distinct part of a meme.</p>
      <p>You need to return a list of descriptions of every part of the meme. For every part (subimage)
, return the location of the part, image_description and text in the original language it
exists on this part and translated_text of this text to the English language (if the
original text is English, then return the same text). Your example output:
[{"location": "Top", "image_description": None, "text": "work in IT", "translated_text": "
work in IT"}, {"location": "right bottom image", "image_description": "picture of software
developer working at night", "text": "I will work only when I want", "translated_text": "I
will work only when I want"}, {"location": "left bottom image", "image_description": "
picture of tired and angry software developer", "text": "And also when my boss will ask me",
"translated_text": "And also when my boss will ask me"}]</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kiela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Firooz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Goswami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ringshia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Testuggine</surname>
          </string-name>
          ,
          <article-title>The hateful memes challenge: Detecting hate speech in multimodal memes</article-title>
          , CoRR abs/
          <year>2005</year>
          .04790 (
          <year>2020</year>
          ). URL: https://arxiv.org/abs/
          <year>2005</year>
          .04790. arXiv:
          <year>2005</year>
          .04790.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>M.-H. Van</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>Detecting and mitigating hateful content in multimodal memes with visionlanguage models</article-title>
          ,
          <year>2025</year>
          . URL: https://arxiv.org/abs/2505.00150. arXiv:
          <volume>2505</volume>
          .
          <fpage>00150</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kapil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ekbal</surname>
          </string-name>
          ,
          <article-title>A transformer based multi task learning approach to multimodal hate speech detection</article-title>
          ,
          <source>Natural Language Processing Journal</source>
          <volume>11</volume>
          (
          <year>2025</year>
          )
          <article-title>100133</article-title>
          . URL: https://www. sciencedirect.com/science/article/pii/S2949719125000093. doi:https://doi.org/10.1016/j.nlp.
          <year>2025</year>
          .
          <volume>100133</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          , J. Han,
          <string-name>
            <surname>S</surname>
          </string-name>
          . Hu,
          <article-title>Uncertainty-aware cross-modal alignment for hate speech detection</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            , M.-
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Hoste</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Lenci</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sakti</surname>
          </string-name>
          , N. Xue (Eds.),
          <source>Proceedings of the 2024 Joint International Conference on Computational Linguistics</source>
          ,
          <article-title>Language Resources and Evaluation (LREC-COLING 2024), ELRA</article-title>
          and
          <string-name>
            <given-names>ICCL</given-names>
            ,
            <surname>Torino</surname>
          </string-name>
          , Italia,
          <year>2024</year>
          , pp.
          <fpage>16973</fpage>
          -
          <lpage>16983</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .lrec-main.
          <volume>1475</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E. L. T.</given-names>
            <surname>Tchokote</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. F.</given-names>
            <surname>Tagne</surname>
          </string-name>
          ,
          <article-title>Efective multimodal hate speech detection on facebook hate memes dataset using incremental pca, smote, and adversarial learning</article-title>
          ,
          <source>Machine Learning with Applications</source>
          <volume>20</volume>
          (
          <year>2025</year>
          )
          <article-title>100647</article-title>
          . URL: https://www.sciencedirect.com/science/article/pii/S2666827025000301. doi:https://doi.org/10.1016/j.mlwa.
          <year>2025</year>
          .
          <volume>100647</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Koutlis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schinas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Papadopoulos</surname>
          </string-name>
          , Memefier:
          <article-title>Dual-stage modality fusion for image meme classification</article-title>
          ,
          <source>in: Proceedings of the 2023 ACM International Conference on Multimedia Retrieval</source>
          , ICMR '23,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2023</year>
          , p.
          <fpage>586</fpage>
          -
          <lpage>591</lpage>
          . URL: https://doi.org/10.1145/3591106.3592254. doi:
          <volume>10</volume>
          .1145/3591106.3592254.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Das</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Narzary</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Saha</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Barman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mukherjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Modha</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Ganguly</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          <string-name>
            <surname>Garain</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Jaki</surname>
          </string-name>
          , T. Mandl,
          <article-title>Overview of the hasoc track at fire 2025: Abusive meme identification - shadows behind the laughter</article-title>
          , in: K. Ghosh,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Majumdar</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Chakraborty (Eds.),
          <source>Forum for Information Retrieval Evaluation (Working Notes) (FIRE</source>
          <year>2025</year>
          ), December 17-20, Varanasi, India, CEUR-WS.org,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Prabhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Seethalakshmi</surname>
          </string-name>
          ,
          <article-title>A comprehensive framework for multi-modal hate speech detection in social media using deep learning</article-title>
          ,
          <source>Scientific Reports</source>
          <volume>15</volume>
          (
          <year>2025</year>
          )
          <article-title>13020</article-title>
          . URL: https://doi.org/10. 1038/s41598-025-94069-z. doi:
          <volume>10</volume>
          .1038/s41598- 025- 94069- z.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>G.</given-names>
            <surname>Arya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. K.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bagwari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Safie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Islam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. R. A.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          , A. De,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Ghazal</surname>
          </string-name>
          ,
          <article-title>Multimodal hate speech detection in memes using contrastive language-image pre-training</article-title>
          ,
          <source>IEEE Access 12</source>
          (
          <year>2024</year>
          )
          <fpage>22359</fpage>
          -
          <lpage>22375</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2024</year>
          .
          <volume>3361322</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <article-title>Multimodal detection of hateful memes by applying a vision-language pretraining model</article-title>
          ,
          <source>PLOS ONE 17</source>
          (
          <year>2022</year>
          )
          <article-title>e0274300</article-title>
          . URL: https://doi.org/10.1371/journal.pone.0274300. doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0274300</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hossain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Sharif</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>M. Hoque, MUTE: A multimodal dataset for detecting hateful memes</article-title>
          , in: Y. Hanqi,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zonghan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruder</surname>
          </string-name>
          , W. Xiaojun (Eds.),
          <source>Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing: Student Research Workshop</source>
          , Association for Computational Linguistics, Online,
          <year>2022</year>
          , pp.
          <fpage>32</fpage>
          -
          <lpage>39</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .aacl-srw.5/. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          .aacl- srw.5.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>M. Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <article-title>Banglaabusememe: A dataset for bengali abusive meme classification</article-title>
          ,
          <source>in: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>15498</fpage>
          -
          <lpage>15512</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. K.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mahapatra</surname>
          </string-name>
          , et al.,
          <article-title>Safespeech: a three-module pipeline for hate intensity mitigation of social media texts in indic languages</article-title>
          ,
          <source>Social Network Analysis and Mining</source>
          <volume>14</volume>
          (
          <year>2024</year>
          ). URL: https://doi.org/10.1007/s13278-024-01393-9. doi:
          <volume>10</volume>
          .1007/s13278- 024- 01393- 9.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Multimodal learning for hateful memes detection</article-title>
          ,
          <source>in: 2021 IEEE International Conference on Multimedia Expo Workshops (ICMEW)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . doi:
          <volume>10</volume>
          .1109/ ICMEW53276.
          <year>2021</year>
          .
          <volume>9455994</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Anaissi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Akram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chaturvedi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Braytee</surname>
          </string-name>
          ,
          <article-title>Detecting and understanding hateful contents in memes through captioning and visual question-answering</article-title>
          ,
          <year>2025</year>
          . URL: https://arxiv.org/abs/2504. 16723. arXiv:
          <volume>2504</volume>
          .
          <fpage>16723</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>R.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <surname>R. K.-W. Lee</surname>
            ,
            <given-names>W.-H.</given-names>
          </string-name>
          <string-name>
            <surname>Chong</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Jiang</surname>
          </string-name>
          ,
          <article-title>Prompting for multimodal hateful meme classification</article-title>
          , in: Y.
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Kozareva</surname>
          </string-name>
          , Y. Zhang (Eds.),
          <source>Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Abu Dhabi, United Arab Emirates,
          <year>2022</year>
          , pp.
          <fpage>321</fpage>
          -
          <lpage>332</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .emnlp-main.
          <volume>22</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          .emnlp- main.22.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mustafa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          , L. Beyer,
          <article-title>Sigmoid loss for language image pre-training</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2303</volume>
          .
          <fpage>15343</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chaudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wenzek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Guzmán</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Unsupervised cross-lingual representation learning at scale</article-title>
          , CoRR abs/
          <year>1911</year>
          .02116 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1911</year>
          .02116. arXiv:
          <year>1911</year>
          .02116.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>R.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <article-title>L3cube-hindbert and devbert: Pre-trained bert transformer models for devanagari based hindi and marathi languages</article-title>
          ,
          <source>arXiv preprint arXiv:2211.11418</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Saha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Modha</surname>
          </string-name>
          ,
          <article-title>Findings from shared tasks on hate speech detection: Performance patterns for low-resource languages, Pattern Recognition Letters (</article-title>
          <year>2025</year>
          ). URL: https://www.sciencedirect.com/science/article/pii/S0167865525003150. doi:
          <volume>10</volume>
          .1016/j.patrec.
          <year>2025</year>
          .
          <volume>09</volume>
          .004.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>