<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>RoJiNG-CL at EXIST 2024: Leveraging Large Language Models for Multimodal Sexism Detection in Memes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jing Ma</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rong Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Zurich</institution>
          ,
          <addr-line>Zurich</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>This paper addressed Task 4 of the challenge posed by Sexism Identification in Social Networks (EXIST) at Conference and Labs of the Evaluation Forum (CLEF) 2024, which involves binary classification to determine the presence of sexism in memes. The task dataset contains memes in both English and Spanish. We explored the application of Large Language Models (LLMs), specifically GPT-4, for extracting textual descriptions from memes. Our methodology integrated these descriptions with associated texts to fine-tune various models, both monolingual and multilingual, to enhance the classifiers' ability to identify sexist content in memes using hard labels. By experimenting with diverse models and hyperparameters, we tailored our approach to optimize performance. Our submissions achieved the top three positions on the hard-hard evaluation leaderboard, which includes both English and Spanish instances.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Memes sexism identification</kwd>
        <kwd>Classification</kwd>
        <kwd>Large Language Models (LLMs)</kwd>
        <kwd>Prompt engineering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Various social networking platforms provide a virtual space where internet users can freely express
themselves. However, this freedom is tainted by the presence of sexist or misogynistic content,
potentially leading to physical and psychological harm to women [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Thus, developing efective mechanisms
to detect and identify such content is crucial. The rapid expansion of Natural Language Processing
(NLP) in the social sciences has prompted researchers to explore its capabilities for identifying sexist
content in textual data. Previous studies have applied Long-Short-Term Memory networks (LSTMs)
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Convolutional Neural Networks (CNNs) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to classify such content [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Additionally, more
advanced language models such as Electra [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], RoBERTa [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and GPT-2 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] have demonstrated
significant eficacy in the classification of sexist text.
      </p>
      <p>
        The challenge extends beyond textual analysis as sexism in online content often includes visual
elements, particularly in memes. Memes, often considered jokes, gain attention through their rapid
digital dissemination within online communities [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. While frequently humorous, memes can also
subtly propagate hate messages, including sexism and misogyny, causing harm at both individual and
societal levels. The Sexism Identification in EXIST [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], part of CLEF 2024 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], reflects this complexity
by addressing sexism in both tweets and memes. Our research specifically focuses on Task 4: identifying
sexism within memes, aiming to efectively classify these multimodal expressions. The integration of
text and image in detecting sexism necessitates sophisticated vision-language models. While models like
Residual Network (ResNet) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and Vision Transformer (ViT) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] are essential for processing images,
recent advancements have introduced more integrated models capable of handling the complexities
of memes. For instance, CLIP [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and multimodal models like mPLUG-Owl [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and OpenFlamingo
[16] have shown substantial proficiency in image classification tasks. Despite their efectiveness, these
architectures require significant computational resources and extensive processing time.
      </p>
      <p>To address these challenges and enhance cost-eficiency, our work incorporates zero-shot prompting
techniques with GPT-4 [17] to extract textual descriptions from memes, considering both text and
image information. This approach allows for a nuanced understanding of the meme’s context, crucial
for detecting underlying sexist themes. We refine this approach by fine-tuning models on a dataset
comprising both provided texts and GPT-4 generated descriptions. Given the multilingual nature of
social media content, we process memes in both English and Spanish, employing language-specific
models to process memes in the respective language, and using multilingual models for the entire
dataset.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <sec id="sec-2-1">
        <title>2.1. Text-Based Sexism Detection</title>
        <p>
          Research on text-based sexism detection has mainly centered around analyzing social media texts.
The pioneering work by Waseem and Hovy [18] on detecting hate speech on Twitter, including sexist
content, highlighted the importance of linguistic and extra-linguistic features and expert annotations in
training classifiers. The research indicated the potential of character n-grams to outperform other textual
features like word n-grams and user demographic metadata. However, they also noted the challenge of
scalability due to the labor-intensive nature of manual tagging. More recent researchers have explored
automated feature extraction using transformer-based models like BERT and its variants [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ][
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], which
excel in contextual understanding and have shown remarkable improvements in detecting complex and
subtle sexist expressions. This advancement is evident in [19], which created the first Spanish corpus
for sexism on Twitter. They involved a combination of traditional classifiers like Logistic Regression
and Random Forest, neural network approaches including Bi-LSTM networks, and BERT. Their results
showed that BERT outperformed other methods. EXIST 2023 [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] tackled the challenge of detecting
sexism in tweets, focusing on identifying sexism, determining the source’s intention, and categorizing
types of sexism. The approaches for these tasks primarily involved fine-tuning models such as mBERT,
XLM-Roberta, GPT-NeoX, BERTIN-GPT-J-6B, and Bernice [20] [21] [22], employing techniques like
ensembling and contrastive learning. These methods demonstrated excellent performance on the tasks.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Multimodal Approaches to Sexism Identification</title>
        <p>Given the complex nature of internet expressions, particularly within the realm of social media, sexism
detection has necessarily expanded beyond textual data to include visual content, where images often
carry implicit messages not evident in text alone. Fusion of multimodal information has become a
popular method in this domain. For instance,[23] employed a methodology for detecting sexism in
advertisements by combining outputs from visual and textual classifiers. The visual classifier analyzed
features like Local Binary Pattern (LBP) and deep learning features extracted using a pre-trained CNN
AlexNet[24], while the textual classifier utilized n-grams, syntactic tags, metadata about word usage,
and word embeddings. A notable study by [25] pioneered the challenge of identifying sexist content
in memes, proposing a framework that uses both unimodal and multimodal classifiers. This research
developed unimodal classifiers that analyzed either textual or visual meme features independently using
models like Support Vector Machines (SVM), Naive Bayes (NB), Decision Tree (DT), and 1-Nearest
Neighbors (1NN). The study explored both early and late fusion techniques for integrating these
modalities. Early fusion combined features at the input level before classification, whereas late fusion
aggregated outputs from the unimodal classifiers post-analysis to determine the presence of sexism.
Their findings indicated that textual classifiers typically outperformed visual classifiers, suggesting that
textual cues are stronger indicators of sexism in memes. Additionally, late fusion was found to be more
efective than early fusion, demonstrating that preserving the integrity of modality-specific features
by combining classifier outputs after individual analyses can enhance overall accuracy. The study
concluded that while unimodal approaches hold value, particularly in textual analysis, their integration
with multimodal strategies significantly improves the efectiveness and reliability of sexism detection
in memes.</p>
        <p>
          More recent works on detecting sexist or misogynous memes used state-of-the-art (SOTA)
pretrained models [26] [27]. Visual features are extracted using CLIP [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], and multimodal models such
as mPLUG-Owl [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and OpenFlamingo [16] have also been employed. These pipelines achieve high
performance but are computationally intensive, necessitating substantial computational power and
memory. Considering these limitations, our work uses GPT-4, which can simultaneously process image
and text inputs, generating detailed descriptions of image content and integrating textual information
for comprehensive analysis. This capability is crucial for understanding the humor, context, and cultural
symbols in memes, addressing challenges highlighted by [25], which emphasize the dificulties of relying
solely on visual features, as they can be ambiguous and less directly indicative of sexism compared to
text.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Method</title>
      <p>Our work focused on Task 4: Sexism Identification in Memes, which is a binary classification task aimed
at determining whether a given meme is sexist or not. Our approach pipeline is outlined in Figure 1: First,
inputs consisting of English and Spanish prompts, along with memes in the corresponding languages,
were processed using zero-shot Chain-of-Thought (CoT) prompting by the GPT-4 model. This resulted
in a one-sentence description of the meme and a hard-label output. Subsequently, this descriptive
text was concatenated with the text extracted from the memes. This combined text, which included
both visual and textual elements of the memes, was employed to fine-tune various language models.
Depending on the language-specific requirements, we either used the entire dataset or the Spanish and
English dataset respectively. Ultimately, the model made binary decisions, producing a definitive ’YES’
or ’NO’ hard-label.</p>
      <sec id="sec-3-1">
        <title>3.1. Dataset</title>
        <p>The dataset used in this study is provided by EXIST 2024 and encompasses both training and test
datasets. The training dataset comprises a total of 4,044 memes, distributed between 2,034 Spanish
memes and 2,010 English memes. The test dataset consists of 1,053 memes, with 540 classified as
Spanish and 513 as English. The associated metadata for these memes is archived in a JSON file, which
includes extracted texts from the memes and task-specific annotations: six labels annotated by six
diferent annotators. Additionally, it captures detailed demographic information about each annotator,
including their gender, age, ethnicity, educational attainment, and country of residence. The gold labels
are provided in a separate JSON file; for hard-labels, when the human-annotated labels for a meme are
evenly split between "YES" and "NO," such memes are excluded from the training dataset.</p>
        <p>To facilitate a more efective evaluation of the results, we subdivide the training sets into separate
training and validation datasets, using a random allocation of 20% for validation. The validation
dataset is designed to monitor model performance on unseen data and plays a crucial role in adjusting
hyperparameters to mitigate overfitting. For the multilingual mDeBERTa-v3-base model[ 28], we split the
entire dataset into training and validation subsets as a whole. In contrast, when fine-tuning monolingual
models such as BETO-uncased[29] for Spanish and bert-base-uncased[30] for English, we performed
splits for the Spanish and English memes first, followed by targeted fine-tuning for each language
group.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Image Information Extraction</title>
        <p>Our task focuses on memes that typically contain both visual and textual elements. We employed a
two-step process to analyze these memes. First, we extracted the text from the memes provided in the
JSON dataset. Following this, we analyzed the visual content using the GPT-4 Turbo model API, chosen
for its robust multimodal understanding capabilities.
Prompt Engineering. Prompt engineering is crucial to optimize the performance of the model.
We aim to keep the input and output not only informative, but also with a relatively short length.
Recent studies have highlighted the potential of LLMs in role-playing scenarios. Assigning specific
roles to an LLM can enhance the naturalness and interactivity of its responses [31, 32], and improve its
performance in complex tasks [33]. We therefore configured the system’s role as an expert in sexism
Model Configuration.
eficiency:
detection in memes. We experimented with prompts directing the model to analyze solely the image or
the combination of image and text. Results indicated that the latter approach yielded more informative
insights. This is due to the complementary nature of the image and text in memes, which, when analyzed
together, provide a fuller and more accurate understanding of the meme’s thematic message. In contrast,
the analysis based on images alone often leads to neutral descriptions that are less relevant to the
memes’ themes. This aligns with observations by [25], which highlight the interpretative challenges
posed by memes: first, identical images can be perceived as sexist or not based on the accompanying
text, which can alter the conveyed message. Second, sexism may be manifested through the image
alone, the text alone, or a combination of both.</p>
        <p>We conducted prompt engineering primarily in English, and subsequently translated it into Spanish
to accommodate the Spanish memes in the dataset using GPT-4. We also involved three bilingual
(Spanish-native) speakers to evaluate the quality of the translations, ensuring the translations maintain
the eficacy and accuracy of the original prompts. The final prompts, along with the input meme and
the model output, are shown in Figure 2 (Spanish) and Figure 3 (English).</p>
        <p>The model was configured with specific settings to enhance performance and
• Model Setting: gpt-4-turbo
• Temperature: 0.75 (to modulate the randomness of the outputs, ensuring their coherence and
relevance)
• Seed: 1234 (to promote consistency in model responses across various runs, though absolute
consistency cannot be guaranteed)
• Detail: Low (to process images in a resource-eficient manner, as fine details are not critical for
our task)</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Model Fine-tuning</title>
        <p>At this stage, we experimented with various models, including multilingual models as well as
monolingual models for English and Spanish. For the Spanish dataset, we fine-tuned the BETO-uncased
model[29], and for the English dataset, the BERT-uncased model[30]. Additionally, we employed
multilingual models including mBERT [30] (both cased and uncased), mDeBERTa[28], XLM-R[34], and
XLM-Twitter[35]. The hyperparameter optimization was facilitated using Optuna[36], a framework that
automates the search for optimal hyperparameters through systematic exploration, considering factors
like learning rate, number of training epochs, batch size, warmup steps, and weight decay. The learning
rate was varied between 1 × 10− 5 and 5 × 10− 5, with the number of training epochs ranging from 3 to
5. Batch sizes are set at 8 and 16, warmup steps ranged from 0 to 500, and weight decay from 0.0 to 0.3
to add a regularization term to the loss function to minimize overfitting. Early stopping mechanisms
were also incorporated to curtail training upon stabilization of validation losses. The objective function
for optimization was defined based on accuracy, with a total of 15 trials conducted to strike a balance
between obtaining the best hyperparameters and managing computational resources.</p>
        <p>Table 1 presents the optimal hyperparameter settings for the three NLP models: mDeBERTa-v3-base,
BETO-uncased, and bert-base-uncased. The table includes values for learning rate, epoch, train batch
size, warmup steps, weight decay, and dropout rate. Each model’s settings are specifically configured
to enhance its training eficacy and overall performance in tasks, reflecting a strategic approach to
ifne-tuning.</p>
        <p>Following the identification of the best hyperparameters by Optuna, the models were fine-tuned again
on the full training dataset. The performance of the fine-tuned models was then evaluated using the
validation dataset, with particular focus on accuracy and the F1 score for the positive class (pos_label=1).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>With an emphasis on evaluation loss, accuracy, and F1 Score Positive, Table 2 compares the
performance of several NLP models on a number of metrics. The mDeBERTa-v3-base model comes in second
at 0.614, while the bert-base-uncased model has a greater evaluation loss of 0.722. The BETO-uncased
model has the lowest evaluation loss at 0.540, suggesting a better capacity to decrease mistakes during
the evaluation phase. The bert-base-uncased model has the highest accuracy at 78.07%, indicating a
noteworthy improvement of almost 9.94% above the baseline’s observed 68.13%. The BETO-uncased
model has the greatest F1 Score Positive 81.38%, which is 6.44% higher than the baseline. This indicates
that BETO-uncased, which efectively balances recall and precision, is especially good at properly
recognizing the positive class even with a reduced evaluation loss. Overall, these findings highlight
the distinct advantages of each model, with bert-base-uncased achieving the highest accuracy and
BETO-uncased performs well at evaluation loss minimization and F1 score optimization.</p>
      <p>Table 3 presents the oficial rankings for Task 4 on the Leaderboard. Our first run combined BERT
predictions fine-tuned on English data and BETO fine-tuned on Spanish data. Our second run employed
mDeBERTa, while the third run was the GPT-4 output results. We achieved top rankings out of more
than 50 results, although surprisingly, the GPT-4 based predictions emerged as the most efective,
delivering top results in a zero-shot setting, showcasing its exceptional capacity to comprehend and
analyze complex sexist memes.</p>
      <p>Further analysis, as shown in Table 4, indicates a consistent trend where all models achieved higher
scores on the English dataset compared to the Spanish one. Despite the size similarities between
BERT and BETO, BETO’s lower performance relative to BERT highlights the challenge of achieving
efectiveness gap between English and Spanish.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Ablation Study</title>
      <p>This ablation study was initiated to address the performance discrepancies observed between the
English and Spanish datasets. To explore whether translating the English datasets into Spanish could
serve as a method of data augmentation, we translated the entire English dataset, including both the
original meme texts and descriptions generated by GPT-4, into Spanish using the DeepL API. This
translated data was then combined with the existing Spanish training dataset. Following the integration,
the combined dataset was divided into training (80%) and validation sets (20%). For model evaluation
purposes, the validation portion of the original Spanish dataset was repurposed as our test set due to
the unavailability of gold labels for the original Spanish test dataset.</p>
      <p>As illustrated in Table 5, the ablation study revealed no performance improvement with the augmented
Spanish dataset. Several factors might have influenced this result. Primarily, the translations provided
by the DeepL API could have introduced semantic inaccuracies or noise, complicating the training
process. Although these translation tools ofer a quick method for converting large datasets from
one language to another, they may not always capture the nuanced cultural contexts and idiomatic
expressions necessary for accurate sentiment and thematic analysis. These translation errors likely
reduced the model’s ability to learn efectively, resulting in worse performance. Additionally, the
BETO-uncased model may not have been optimally fine-tuned for the nuances of the Spanish language,
potentially limiting its processing and comprehension abilities on the Spanish dataset. To mitigate
these issues, future research should concentrate on enhancing the quality of translations, employing
advanced data augmentation strategies, and ensuring thorough fine-tuning of the models for specific
language contexts.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>In this working notes, we have demonstrated that LLMs, particularly GPT-4, can serve as competitive
tools for extracting textual information from memes. Our methodology, with the strategic use of prompt
engineering, has sidestepped the complexities typically associated with multimodal approaches and
focused on generating descriptive texts directly from meme content. This approach not only simplifies
the computing resources needed but also enhances our system’s ability to detect subtlety that cannot
be fully understood by mere texts. Our results are promising, showing that the application of LLMs,
when finely tuned with tailored prompts, can efectively interpret and describe meme content. This is
crucial for tasks requiring not just textual extraction but also an understanding of underlying societal
and cultural contexts conveyed through humor and satire in memes.</p>
      <p>The scope and generalizability of this study are constrained by several factors. First, the outputs
of LLMs may exhibit intrinsic biases originating from their training data. These biases, particularly
gender biases, could potentially lead to descriptions that are not accurate or appropriate, thereby
misrepresenting the intent or sentiment of the memes. Another notable limitation is our reliance on
binary (’YES’ or ’NO’) hard labels for sexism classification. However, in reality, sexism often exists
on a continuous spectrum rather than a simple classification question. This complexity is particularly
pronounced in the context of memes, which are inherently open to interpretation. Variability in
perceptions among diferent individuals is common, as reflected in our dataset, where annotators
frequently disagree. In future work, we aim to explore this diversity of human perspectives more
thoroughly by incorporating soft labels that better capture the spectrum of responses.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We would like to express our gratitude to Simon Clematide and Andrianos Michail for their invaluable
support and insightful suggestions. We also extend our thanks to the Department of Computational
Linguistics at the University of Zurich for their financial support in making this project possible.
with multimodality, ArXiv abs/2304.14178 (2023). URL: https://api.semanticscholar.org/CorpusID:
258352455.
[16] A. Awadalla, I. Gao, J. Gardner, J. Hessel, Y. Hanafy, W. Zhu, K. Marathe, Y. Bitton, S. Y. Gadre,
S. Sagawa, J. Jitsev, S. Kornblith, P. W. Koh, G. Ilharco, M. Wortsman, L. Schmidt, Openflamingo:
An open-source framework for training large autoregressive vision-language models, ArXiv
abs/2308.01390 (2023). URL: https://api.semanticscholar.org/CorpusID:261043320.
[17] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt,</p>
      <p>S. Altman, S. Anadkat, et al., Gpt-4 technical report, arXiv preprint arXiv:2303.08774 (2023).
[18] Z. Waseem, D. Hovy, Hateful symbols or hateful people? predictive features for hate speech
detection on Twitter, in: J. Andreas, E. Choi, A. Lazaridou (Eds.), Proceedings of the NAACL
Student Research Workshop, Association for Computational Linguistics, San Diego, California,
2016, pp. 88–93. URL: https://aclanthology.org/N16-2013. doi:10.18653/v1/N16-2013.
[19] F. Rodríguez-Sánchez, J. Carrillo-de Albornoz, L. Plaza, Automatic classification of sexism in social
networks: An empirical study on twitter data, IEEE Access 8 (2020) 219563–219576. doi:10.1109/
ACCESS.2020.3042604.
[20] A. F. M. de Paula, G. Rizzi, E. Fersini, D. Spina, Ai-upv at exist 2023 - sexism characterization using
large language models under the learning with disagreement regime, ArXiv abs/2307.03385 (2023).</p>
      <p>URL: https://api.semanticscholar.org/CorpusID:259376983.
[21] L. Tian, N. Huang, X. Zhang, Eficient multilingual sexism detection via large language model
cascades., 2023.
[22] J. Angel, S. T. Aroyehun, A. F. Gelbukh, Multilingual sexism identification using contrastive
learning., in: CLEF (Working Notes), 2023, pp. 855–861.
[23] F. Gasparini, I. Erba, E. Fersini, S. Corchs, Multimodal classification of sexist advertisements,
in: International Conference on E-Business and Telecommunication Networks, 2018. URL: https:
//api.semanticscholar.org/CorpusID:52121016.
[24] A. Krizhevsky, I. Sutskever, G. E. Hinton, Imagenet classification with deep convolutional neural
networks, in: F. Pereira, C. Burges, L. Bottou, K. Weinberger (Eds.), Advances in Neural Information
Processing Systems, volume 25, Curran Associates, Inc., 2012. URL: https://proceedings.neurips.
cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf.
[25] E. Fersini, F. Gasparini, S. Corchs, Detecting sexist meme on the web: A study on textual and visual
cues, in: 2019 8th International Conference on Afective Computing and Intelligent Interaction
Workshops and Demos (ACIIW), 2019, pp. 226–231. doi:10.1109/ACIIW.2019.8925199.
[26] H. B. Zia, I. Castro, G. Tyson, Racist or sexist meme? classifying memes beyond hateful, in:
A. Mostafazadeh Davani, D. Kiela, M. Lambert, B. Vidgen, V. Prabhakaran, Z. Waseem (Eds.),
Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021), Association for
Computational Linguistics, Online, 2021, pp. 215–219. URL: https://aclanthology.org/2021.woah-1.
23. doi:10.18653/v1/2021.woah-1.23.
[27] S. Chen, U. Naseem, I. Razzak, F. Salim, Unveiling misogyny memes: A multimodal analysis of
modality efects on identicfiation, in: Companion Proceedings of the ACM on Web Conference
2024, WWW ’24, Association for Computing Machinery, New York, NY, USA, 2024, p. 1864–1871.</p>
      <p>URL: https://doi.org/10.1145/3589335.3651974. doi:10.1145/3589335.3651974.
[28] M. Laurer, W. v. Atteveldt, A. S. Casas, K. Welbers, Less Annotating, More Classifying – Addressing
the Data Scarcity Issue of Supervised Machine Learning with Deep Transfer Learning and BERT
NLI, Preprint (2022). URL: https://osf.io/74b8k, publisher: Open Science Framework.
[29] J. Cañete, G. Chaperon, R. Fuentes, J.-H. Ho, H. Kang, J. Pérez, Spanish pre-trained bert model and
evaluation data, in: PML4DC at ICLR 2020, 2020.
[30] J. Devlin, M. Chang, K. Lee, K. Toutanova, BERT: pre-training of deep bidirectional transformers
for language understanding, CoRR abs/1810.04805 (2018). URL: http://arxiv.org/abs/1810.04805.
arXiv:1810.04805.
[31] Y. Tao, A. Agrawal, J. Dombi, T. Sydorenko, J. I. Lee, Chatgpt role-play dataset: Analysis of
user motives and model naturalness, in: International Conference on Language Resources and
Evaluation, 2024. URL: https://api.semanticscholar.org/CorpusID:268723733.
[32] Z. Wang, Z. Peng, H. Que, J. Liu, W. Zhou, Y. Wu, H. Guo, R. Gan, Z. Ni, M. Zhang, Z. Zhang,
W. Ouyang, K. Xu, W. Chen, J. Fu, J. Peng, Rolellm: Benchmarking, eliciting, and enhancing
role-playing abilities of large language models, ArXiv abs/2310.00746 (2023). URL: https://api.
semanticscholar.org/CorpusID:263334495.
[33] G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, B. Ghanem, Camel: Communicative agents
for" mind" exploration of large scale language model society (2023).
[34] A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott,
L. Zettlemoyer, V. Stoyanov, Unsupervised cross-lingual representation learning at scale, in:
D. Jurafsky, J. Chai, N. Schluter, J. Tetreault (Eds.), Proceedings of the 58th Annual Meeting of the
Association for Computational Linguistics, Association for Computational Linguistics, Online,
2020, pp. 8440–8451. URL: https://aclanthology.org/2020.acl-main.747. doi:10.18653/v1/2020.
acl-main.747.
[35] F. Barbieri, L. Espinosa Anke, J. Camacho-Collados, XLM-T: Multilingual language models in
Twitter for sentiment analysis and beyond, in: Proceedings of the Thirteenth Language Resources
and Evaluation Conference, European Language Resources Association, Marseille, France, 2022,
pp. 258–266. URL: https://aclanthology.org/2022.lrec-1.27.
[36] T. Akiba, S. Sano, T. Yanase, T. Ohta, M. Koyama, Optuna: A next-generation hyperparameter
optimization framework, in: Proceedings of the 25th ACM SIGKDD International Conference on
Knowledge Discovery and Data Mining, 2019.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Karthikeyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sundarraj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sampathkumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Mouthami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Yuvaraj</surname>
          </string-name>
          ,
          <article-title>Sexism classification in social media using machine learning algorithms</article-title>
          , in: A.
          <string-name>
            <surname>Abraham</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Hanne</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Gandhi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Manghirmalani Mishra</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Bajaj</surname>
          </string-name>
          , P. Siarry (Eds.),
          <source>Proceedings of the 14th International Conference on Soft Computing and Pattern Recognition (SoCPaR</source>
          <year>2022</year>
          ), Springer Nature Switzerland, Cham,
          <year>2023</year>
          , pp.
          <fpage>14</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <article-title>Long short-term memory</article-title>
          ,
          <source>Neural Computation</source>
          <volume>9</volume>
          (
          <year>1997</year>
          )
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>K. O'Shea</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Nash</surname>
          </string-name>
          ,
          <article-title>An introduction to convolutional neural networks</article-title>
          ,
          <source>ArXiv abs/1511</source>
          .08458 (
          <year>2015</year>
          ). URL: https://api.semanticscholar.org/CorpusID:9398408.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kalra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zubiaga</surname>
          </string-name>
          ,
          <article-title>Sexism identification in tweets and gabs using deep neural networks</article-title>
          ,
          <source>ArXiv abs/2111</source>
          .03612 (
          <year>2021</year>
          ). URL: https://api.semanticscholar.org/CorpusID:243832598.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Clark</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>T.</given-names>
            <surname>Luong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          , Electra:
          <article-title>Pre-training text encoders as discriminators rather than generators</article-title>
          , arXiv preprint arXiv:
          <year>2003</year>
          .
          <volume>10555</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: North American Chapter of the Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          . URL: https://api.semanticscholar.org/CorpusID:52967399.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          , ArXiv abs/
          <year>1907</year>
          .11692 (
          <year>2019</year>
          ). URL: https://api.semanticscholar.org/CorpusID:198953378.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Language models are unsupervised multitask learners</article-title>
          ,
          <year>2019</year>
          . URL: https://api.semanticscholar.org/CorpusID:160025533.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Iloh</surname>
          </string-name>
          ,
          <article-title>Do it for the culture: The case for memes in qualitative research</article-title>
          ,
          <source>International Journal of Qualitative Methods</source>
          <volume>20</volume>
          (
          <year>2021</year>
          )
          <article-title>16094069211025896</article-title>
          . URL: https://doi.org/10.1177/16094069211025896. doi:
          <volume>10</volume>
          .1177/16094069211025896. arXiv:https://doi.org/10.1177/16094069211025896.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of exist 2023:
          <article-title>sexism identification in social networks</article-title>
          ,
          <source>in: Proceedings of ECIR'23</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>593</fpage>
          -
          <lpage>599</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -28241-6_
          <fpage>68</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maeso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chulvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <article-title>Overview of exist 2024 - learning with disagreement for sexism identification and characterization in social networks and memes (extended overview)</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . García Seco de Herrera (Eds.),
          <source>Working Notes of CLEF 2024 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <source>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          (
          <year>2015</year>
          )
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          . URL: https: //api.semanticscholar.org/CorpusID:206594692.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. M.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>in: Neural Information Processing Systems</source>
          ,
          <year>2017</year>
          . URL: https://api. semanticscholar.org/CorpusID:13756489.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          , et al.,
          <article-title>Learning transferable visual models from natural language supervision</article-title>
          ,
          <source>in: International conference on machine learning, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>8748</fpage>
          -
          <lpage>8763</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , F. Huang, mplug-owl:
          <article-title>Modularization empowers large language models</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>