<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LightGBM for Sexism Identification in Memes⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arnau Garcia i Cucó</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miquel Obrador Reina</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universitat Politècnica de València (UPV)</institution>
          ,
          <addr-line>València</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universitat Politècnica de València (UPV)</institution>
          ,
          <addr-line>València</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>In this paper, we present our participation in the EXIST2024 competition, focusing on the detection and categorization of sexism in memes through binary and multiclass classification tasks. For Task 4, which involves binary classification to identify whether a meme is sexist, we utilized a two-stage approach by fine-tuning the Contrastive Language-Image Pre-training (CLIP) model followed by training a Light Gradient Boosting Machine (LightGBM) classifier on the obtained embeddings. For Task 6, which categorizes sexist memes into various types, we fine-tuned both a RoBERTa model for text and a Google Vision Transformer (ViT) for images, subsequently training a LightGBM classifier on the concatenated embeddings. Our results demonstrate that the combination of CLIP and LightGBM outperforms other models in binary classification, while an ensemble of text and image models enhances performance in multiclass categorization. These findings highlight the potential of leveraging advanced machine learning models and multi-modal embeddings for efective sexism detection and classification in memes, addressing a critical issue in combating online misogyny.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Sexism detection</kwd>
        <kwd>meme classification</kwd>
        <kwd>binary classification</kwd>
        <kwd>multiclass classification</kwd>
        <kwd>ViT</kwd>
        <kwd>Transformers</kwd>
        <kwd>Gradient Boosting Methods</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        According to Glick and Fiske’s concept of "sexism" [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], sexism is a multidimensional construct that
encompasses two sets of sexist attitudes: hostile and benevolent. While hostile sexism communicates
a clear antipathy toward women, benevolent sexism takes the form of seemingly positive but in fact
patronizing beliefs about women. These beliefs often involve traditional stereotyping and masculine
dominance, ultimately leading to a restriction of women’s roles and damaging consequences for gender
equality.
      </p>
      <p>
        EXIST is a series of scientific events and shared tasks on sexism identification in social
networks.Building upon previous EXIST challenges (EXIST 2021, EXIST 2022, and EXIST 2023), this
edition, EXIST2024 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] aims to capture sexism in a broad sense, from explicit misogyny to subtle
expressions involving implicit sexist behaviors. While the three previous editions focused solely on
detecting and classifying sexist textual messages, this new edition incorporates new tasks that center
around images, particularly memes. Memes are images, typically humorous in nature, that are spread
rapidly by social networks and Internet users. The challenge consists of five tasks in two languages,
English and Spanish.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Tasks performed</title>
      <p>In the EXIST2024 competition, our team chose to participate in Task 4, which focuses on identifying
sexism in memes through binary classification, determining whether a given meme is sexist or not.
Additionally, we took part in Task 6, which involves categorizing sexist memes into various types
of sexism, such as ideological and inequality, stereotyping and dominance, objectification, sexual
violence, and misogyny and non-sexual violence. This categorization is based on the classification
system provided for Task 3.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Main Objectives of Experiments</title>
      <p>In the EXIST2024 competition, our primary objective is to investigate the efectiveness of machine
learning algorithms in identifying and categorizing sexist memes through binary classification and
multiclass classification. Specifically, we aim to develop models that can accurately determine whether
a given meme is sexist or not, and further classify sexist memes into various types of sexism.</p>
      <p>Our research questions include:
• How efective are machine learning algorithms in identifying sexist memes through binary
classification?
• Can we develop a model that accurately categorizes sexist memes into various types of sexism,
such as ideological and inequality, stereotyping and dominance, objectification, sexual violence,
and misogyny and non-sexual violence?
• What are the challenges and limitations of using machine learning algorithms for sexism detection
in memes, and how can we address these issues?
To address these questions, our experiments will focus on the following tasks:
• Binary Classification: We will develop and evaluate machine learning models that can accurately
classify memes as sexist or not sexist.
• Multiclass Classification: We will develop and evaluate machine learning models that can
accurately categorize sexist memes into various types of sexism, as defined by the EXIST2024
competition’s classification system.
• Model Generalization: We will evaluate the performance of our models on a diverse range of
memes, including those from diferent cultural and linguistic backgrounds. This will help us
assess the generalizability of our models and identify potential issues related to cross-cultural
diferences and language barriers.</p>
      <p>By addressing these objectives, we aim to contribute to the development of more accurate, fair, and
robust machine learning models for sexism detection in memes, ultimately helping to combat the spread
of hateful content against women on social media platforms.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Approaches and methodology</title>
      <p>4.1. Task 4
In the context of the EXIST2024 Task 4 for binary classification of memes, our methodology involved a
two-stage approach to address the classification problem. The dataset was splitted into train/val test in
order to evaluate the model.</p>
      <p>Initially, we fine-tuned the pre-trained Contrastive Language-Image Pre-training (CLIP) model [6], a
deep learning model developed by OpenAI that has demonstrated exceptional performance in vision
and language tasks. CLIP is trained on a vast dataset of 400 million (image, text) pairs, enabling it to
learn rich visual and textual representations.</p>
      <p>During the first stage of our approach, we incorporated a fully connected network specifically
designed for binary classification within the CLIP model. This allowed the model to learn task-specific
features and improve its ability to diferentiate between the two classes of memes. Once the fine-tuning
process was complete, we discarded the fully connected layer, as our primary focus was on obtaining
high-quality embeddings that could be used to train a more specialized classifier.</p>
      <p>In the second stage, we employed a Light Gradient Boosting Machine (LightGBM) [4] classifier, a
highly eficient and efective gradient boosting framework that uses tree-based learning algorithms. By
training the LightGBM classifier with the CLIP embeddings obtained from the first stage, we aimed to
further enhance the classification performance. This two-stage training strategy allowed us to leverage
the strengths of both CLIP and LightGBM, resulting in improved binary classification of memes in the
EXIST2024 Task 4. We tried a logistic regression classifer instead of LightGBM for the final classification
(See table 1).
4.2. Task 6
For the EXIST2024 Task 6, after segmenting the dataset into train/val and test, ensuring a robust
evaluation process, we meticulously crafted a methodology involving three stages: two fine tunings
and a LightGBM training.</p>
      <p>The classification problem we tackled was multifaceted, encompassing six categories. Initially, we
had to take into account binary classification to distinguish between sexist and non-sexist memes. Then,
if the meme was deemed sexist, we further dissected the binary category into five classes that could
overlap: stereotyping, objectification, ideological, sexual violence, and misogyny. It’s important to
note that there was no overlap when the binary category was ’no ’. Thus, we defined only the five
sexism categories as labels, and if a meme didn’t fall into any of these categories, we assumed it to be
non-sexist.</p>
      <p>Furthermore, this task was labelled by a group of experts who voted on the classification they
considered optimal. Thus, we had not only the hard label but also the soft label given by those experts.
We aimed to use all this data to configure an accurate model.</p>
      <p>As stated previously, we had images and text for the meme data. This information could not be used
as such, so we tried to obtain embeddings that would represent task-specific features. To do so, we
ifne-tuned a RoBERTa model [2] in the first stage, incorporating a final fully connected layer. Then, we
retrained the model to learn the soft labels, using the Kullback-Leibler divergence as a loss function. In
the second phase, we did the same for the image, using, in this case, a Google ViT [1] as the base model.
The idea in those steps was to represent the memes in arrays that sum the features the experts use to
discriminate them.</p>
      <p>Finally, in the last stage of the process, we concatenated both embeddings (thus enhancing the
information achieved in image and text), and we applied a LightGBM to finally learn to classify into
hard labels.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Resources employed</title>
      <p>To develop and evaluate our machine learning models for sexism detection in memes, we utilized
various resources, including open-source datasets, pre-trained models, and computational platforms.
Our primary focus was on utilizing Google Colab’s free GPU resources to train and test our models
eficiently. Here is a detailed breakdown of the resources we used for model development:
• Datasets: We relied on publicly available datasets, such as the EXIST2024 competition’s training
dataset, which provided a diverse collection of memes labeled as sexist or not sexist. Additionally,
we used external datasets to augment our training data and improve model performance.
• Pre-trained Models: We leveraged pre-trained models, such as RoBERTa, Google ViT and CLIP,
as feature extractors for our meme classification tasks. These models were fine-tuned on our
training data to improve their performance on the sexism detection and categorization tasks.
• Google Colab: We used Google Colab’s free GPU resources to train and test our models. This
allowed us to eficiently experiment with various algorithms, hyperparameters, and feature
engineering techniques without incurring significant computational costs.</p>
      <p>By leveraging these resources, we were able to develop and evaluate machine learning models for
sexism detection and categorization in memes. Our experiments focused on addressing the research
questions outlined in the main objectives section, ultimately contributing to the development of more
accurate, fair, and robust models for combating hateful content against women on social media platforms.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Results</title>
      <p>6.1. Task 4</p>
      <p>The results presented in Table1 showcase the performance of diferent models on the binary
classiifcation task for memes. The models evaluated include CLIP with a logistic regression (LR) classifier,
CLIP with Light Gradient Boosting Machine (LGBM) classifier, and CLIP with a simple binary classifier
(). The performance metrics used for evaluation are F1-score and Matthews Correlation
Coeficient (MCC).</p>
      <p>The results indicate that the CLIP model with the LightGBM classifier (  ) outperforms
the other models in both F1-score and MCC. The  model achieves an F1-score of 0.72 and
an MCC of 0.34, which is a slight improvement over the  model (F1-score of 0.71 and MCC of
0.32) and the  model (F1-score of 0.71 and MCC of 0.33).
6.2. Task 6
As for task 6, we observed the weighted F1-score when training with text had more potential than
when training with image, as seen in table 3. The Google ViT did not learn to classify misogyny and
sexual violence, thus causing a massive drop in its performance. However, the ensemble was the best
performing so far, hence showing that the image data had relevant information for the classification.</p>
      <p>In Figure 1, we observe that even though the model predicts ideological inequality and objectification
pretty well, it fails to predict accurately misogyny (non-sexual violence).</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions and Future work</title>
      <p>The results showcase the efectiveness of the proposed two-stage training strategy, which combines
the strengths of CLIP and LightGBM, resulting in improved binary classification performance. Future
research could explore the potential of integrating other advanced classifiers or fine-tuning techniques
to further enhance the classification capabilities of the CLIP model. Furthermore, we would like to
experiment by applying data augmentation since we did not use the English data at all for task 6, and
we could have included it in training by translating it, for example. We would also like to explore other
ways to cross both embeddings rather than concatenating them, such as using attention mechanisms.</p>
      <p>Building on the findings and methodologies presented in this paper, future research can explore
several promising avenues to further enhance sexism detection and categorization in memes:
• Data Augmentation and Multilingual Training: Integrating additional data sources and employing
data augmentation techniques, such as translating non-English memes, can improve model
robustness and performance. This approach will help address the challenges of cross-cultural
diferences and language barriers.
• Advanced Embedding Techniques: Investigating alternative methods for combining text and image
embeddings, such as attention mechanisms, can potentially yield richer and more contextually
relevant features. This could enhance the model’s ability to capture nuanced information from
both modalities.
• User Feedback and Iterative Improvement: Incorporating user feedback into the model
development cycle can help refine the models and make them more aligned with real-world expectations
and nuances.
• Explainability and Interpretability: Enhance the interpretability of the models by incorporating
explainable AI techniques. This would help in understanding the decision-making process of the
models and identifying potential biases or errors.</p>
      <p>By exploring these directions, future work can contribute to the development of more sophisticated,
accurate, and ethical machine learning models for detecting and combating sexism in memes, ultimately
fostering a safer and more inclusive online environment.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Alexey</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          et al.
          <article-title>An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale</article-title>
          .
          <year>2021</year>
          . arXiv:
          <year>2010</year>
          .
          <article-title>11929 [cs</article-title>
          .CV].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Asier</given-names>
            <surname>Gutiérrez</surname>
          </string-name>
          Fandiño et al. “
          <article-title>MarIA: Spanish Language Models”</article-title>
          .
          <source>In: Procesamiento del Lenguaje Natural</source>
          <volume>68</volume>
          (
          <year>2022</year>
          ). issn:
          <fpage>1135</fpage>
          -
          <lpage>5948</lpage>
          . doi:
          <volume>10</volume>
          .26342/2022-68-3. url: https://upcommons.upc.edu/ handle/2117/367156#.YyMTB4X9A-0.mendeley.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Glick</surname>
          </string-name>
          and Susan T. Fiske. “
          <article-title>Ambivalent sexism”</article-title>
          . In: vol.
          <volume>33</volume>
          . Advances in Experimental Social Psychology. Academic Press,
          <year>2001</year>
          , pp.
          <fpage>115</fpage>
          -
          <lpage>188</lpage>
          . doi: https://doi.org/10.1016/S0065-
          <volume>2601</volume>
          (
          <issue>01</issue>
          )
          <fpage>80005</fpage>
          -
          <lpage>8</lpage>
          . url: https://www.sciencedirect.com/science/article/pii/S0065260101800058.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Guolin</given-names>
            <surname>Ke</surname>
          </string-name>
          et al. “
          <article-title>LightGBM: A Highly Eficient Gradient Boosting Decision Tree”</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems. Ed. by I. Guyon et al</source>
          . Vol.
          <volume>30</volume>
          . Curran Associates, Inc.,
          <year>2017</year>
          . url: https : / / proceedings . neurips . cc / paper _ files / paper / 2017 / file / 6449f44a102fde848669bdd9eb6b76fa-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Laura</given-names>
            <surname>Plaza</surname>
          </string-name>
          et al. “
          <article-title>Overview of EXIST 2024 - Learning with Disagreement for Sexism Identification and Characterization in Social Networks and Memes”</article-title>
          . In:
          <article-title>Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the Fifteenth International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ).
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Alec</given-names>
            <surname>Radford</surname>
          </string-name>
          et al. “
          <article-title>Learning Transferable Visual Models From Natural Language Supervision”</article-title>
          .
          <source>In: CoRR abs/2103</source>
          .00020 (
          <year>2021</year>
          ). arXiv:
          <volume>2103</volume>
          .00020. url: https://arxiv.org/abs/2103.00020.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>