<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>BertModel) ast_hidden_state=(N
one</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Approach with BERT, XLM-RoBERTa, and DistilBERT for EXIST 2023 Tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hadi Mohammadi</string-name>
          <email>email1h.mohammadi@uu.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anastasia Giachanou</string-name>
          <email>a.giachanou@uu.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ayoub Bagheri</string-name>
          <email>a.bagheri@uu.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Methodology and Statistics, Utrecht University</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <volume>256</volume>
      <issue>768</issue>
      <abstract>
        <p>This research investigates the application of pre-trained transformer-based models, including BERT, XLM- RoBERTa, and DistilBERT, in the context of the EXIST 2023 shared task, which focuses on identifying and categorizing online sexism. The study emphasizes the crucial role of Natural Language Processing (NLP) in detecting harmful content, and it draws on previous competitions that have incorporated tasks to detect hate speech and abusive language. The methodology combines various advanced techniques from the text classification domain, including the use of additional datasets, data preprocessing, and model building. The research also explores data augmentation techniques and label encoding as preprocessing steps. The study's findings indicate that the developed model performs optimally in English, and it suggests that the use of a voting system and the combination of outputs from multiple models contribute to the overall performance. The research concludes with a call for sustained initiatives to curb the prevalence of harmful content on digital platforms, and it outlines future work directions, including incorporating additional information about annotators, the assessment of annotator reliability, and exploring more sophisticated techniques for handling imbalances.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Online Sexism</kwd>
        <kwd>Natural Language Processing (NLP)</kwd>
        <kwd>Transformer-based Models</kwd>
        <kwd>BERT</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The advent of social media has revolutionized the way we communicate, interact, and disseminate
information. However, this digital revolution that allows everyone to publish posts quickly and easily
has resulted in a concerning amount of harmful content, such as hate speech, discrimination, and sexism.
Despite efforts to mitigate these issues, the identification and categorization of sexism in online
platforms remain challenging due to such expressions' nuanced and context-dependent nature [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
      </p>
      <p>
        Natural Language Processing (NLP) has witnessed remarkable progress in the past decade, primarily
due to the emergence of deep learning algorithms for text classification [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. These advancements have
resulted in the development of models that can better understand and interpret human language, thereby
paving the way for many applications, including identifying and classifying sexist content. Advanced
Natural Language Processing (NLP) methods and Language Models can comprehend the subtleties and
context-dependent language often employed in sexist expressions and, therefore, can be used to detect
harmful content [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Several shared tasks and competitions, such as SemEval, have incorporated tasks to detect hate
speech and abusive language. In 2023, the EXIST (sEXism Identification in Social neTworks) lab was
launched as part of the CLEF (Conference and Labs of the Evaluation Forum) conference focusing on
addressing the problem of sexism detection [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. To address those challenges, we participated in the
EXIST 2023 competition on sexism detection in social media platforms [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. EXIST 2023 provides a
dataset of over 10,000 labeled tweets in English and Spanish which comprises three tasks: Identification
of sexism, source intention, and categorization of sexism. Sexism identification is a binary classification
task that requires systems to decide whether a tweet contains or describes sexist expressions or
behaviors. Source intention is a categorization task that classifies the intention behind the sexist
messages into one of three pre-defined categories: direct, reported, or judgmental. The final task, sexism
categorization, is a multiclass and multilabel task that categorizes the sexist messages according to the
type(s) of sexism they contain. Those categories, such as ideological and inequality, stereotyping and
dominance, objectification, and sexual violence, after considering the undermined aspects of women
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>In the remainder of this paper, we present our methodology for tackling the EXIST 2023 tasks in
Section 2, we discuss our experimental results in Section 3, and finally, in Section 4, we summarize our
work and present future work directions.</p>
      <p>Methodology</p>
      <p>We have developed a methodology combining various advanced techniques from the text
classification domain to address the sexism detection task. Our methodology consists of several steps,
which are shown in Figure 1:
1.1.</p>
    </sec>
    <sec id="sec-2">
      <title>Additional Dataset Description</title>
      <p>
        In addition to the EXIST 2023 dataset, we also utilized another valuable resource for our research:
the dataset provided by the SemEval-2023 Task 10: Explainable Detection of Online Sexism (EDOS);
this dataset was introduced to address the issue of binary detection of sexism, which often overlooks
the diversity of sexist content and fails to provide clear explanations for why something is considered
sexist [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This dataset contains 20,000 social media comments with fine-grained labels. By integrating
the EDOS dataset into our methodology, we enhanced the robustness and explainability of our models
for online sexism detection.
      </p>
      <p>1.2.</p>
    </sec>
    <sec id="sec-3">
      <title>Model Building</title>
      <p>The model-building phase of this research is a distinctive element of the methodology, with a unique
architecture that incorporates three different transformer-based models (BERT, XLM- RoBERTa,
DistilBERT). Transformer-based models are a category of models that use self-attention mechanisms
and have achieved state-of-the-art results in various natural language processing tasks.</p>
      <p>BERT (Bidirectional Encoder Representations from Transformers) is a transformer-based model
pre-trained on a large corpus of text data. BERT considers the full context of a word by looking at the
words that come before and after it, which is why it is described as bidirectional. Understanding the
context from both directions allows BERT better to understand the semantic meaning of words and
sentences.</p>
      <p>As our second model, we used XLM-Roberta, a variant of the RoBERTa model (a robustly
optimized version of BERT), which is trained on a large amount of multilingual data making it highly
effective for multilingual tasks.</p>
      <p>DistilBERT is a lighter version of BERT, created through a process called distillation, where the
knowledge of the larger BERT model is transferred to a smaller, faster model. Despite being smaller
and faster, DistilBERT maintains most of the performance of BERT, making it a good choice for tasks
where computational resources are limited.</p>
      <p>Each transformer model processes the input text independently, generating a unique data
representation. These representations are then concatenated and passed through a Conv1D layer.
Following the Conv1D layer, a MaxPooling1D layer is used. After the pooling layer, the data is
flattened. Flattening is a simple operation that transforms a multi-dimensional array into a
onedimensional array. This step is necessary because the output of the MaxPooling1D layer is a 2D tensor,
while the subsequent layers require a 1D tensor. The flattened data is then subjected to a Dropout layer.
Finally, two Dense layers are used. Dense layers are the regular deeply connected neural network layers.
In this case, one Dense layer is used for binary classification (task1), and the other for multiclass
classification (task2 and 3). These layers output the final predictions of the model, which can then be
evaluated and used for making decisions.</p>
    </sec>
    <sec id="sec-4">
      <title>2. Experimental Setup</title>
      <p>In this section, we describe the experimental setup we followed to assess the performance of our
models. This setup encompassed a range of processes, from system configuration and hyperparameter
optimization to model training and validation, all aimed at ensuring the most accurate and reliable
results.</p>
      <p>2.1.</p>
    </sec>
    <sec id="sec-5">
      <title>Data Preprocessing</title>
      <p>Data loading and preprocessing are critical initial phases in machine learning or data analysis tasks.
This process prepares the raw data for subsequent steps, such as model training and testing. After we
loaded the data, we continued with the preprocessing phase. The first step in preprocessing was to
convert all textual content to lowercase. This step is essential for normalizing the data and reducing
complexity, as machine learning algorithms treat uppercase and lowercase versions of the same
character as different entities. By converting all text to lowercase, we can ensure that the same word in
different cases is recognized as the same word by the algorithm. Next, we removed elements that
contained no information for this task, including URLs, memorable characters, and punctuation.
Removing them helps reduce the data’s dimensionality and makes extracting meaningful information
easier for the algorithm. Following this, the text was tokenized and lemmatized. Tokenization is the
process of dividing text into individual words or tokens. This step is necessary because machine learning
algorithms work with individual units (like words) rather than entire texts. Lemmatization is a further
refinement of this process. It reduces words to their base or root form, grouping different grammatical
forms of the same word. This helps reduce the data’s complexity and makes it easier for the algorithm
to identify patterns. Finally, we incorporated the EDOS dataset to enrich the data pool. We followed
the same preprocessing steps. Incorporating additional data can help improve the model’s robustness
and increase its generalizability to unseen data.</p>
      <p>Data augmentation is a strategy used to increase the diversity and amount of training data without
collecting new data. This technique is primarily used to prevent overfitting, a common problem where
a model performs well on training data but poorly on unseen data. By creating modified versions of the
existing data, the model can learn more robust features and generalize better to new data.</p>
      <p>In this research, we used the following data augmentation techniques from the” nlpaug” library in
Python:
• Synonym Replacement: This technique involves replacing words in the text with their
synonyms. This increases the linguistic diversity of the dataset and helps the model to
understand that different words can have the same meaning.
• Random Word Swapping: This technique involves randomly swapping pairs of words in the
text. This introduces a controlled noise level into the data and helps the model learn that the
order of words can vary while preserving the overall meaning.
• Random Character Insertion: This technique involves randomly inserting characters into
words. This also introduces a controlled noise level and helps the model learn to handle typos
or other minor errors in the input data.</p>
      <p>Label encoding is a preprocessing step that converts categorical labels into a format that machine
learning algorithms can use. Most algorithms expect numerical input and output, so it is necessary to
transform categorical labels into numerical values. In this research, we used a majority voting system
to decide the labels for the sexism detection task (task 1). This means that the label assigned to each
data point was the one that was chosen by the majority of annotators. These labels were binary
(indicating the presence or absence of a particular feature) and were encoded into a binary format for
compatibility with the machine learning model. For tasks 2 and 3, a multiclass label encoding process
was used. This process involved encoding each class as a unique integer. This is a standard method for
handling multiclass problems, where each data point can belong to one of several classes. A variety of
machine-learning algorithms can easily use the resulting numerical labels.</p>
    </sec>
    <sec id="sec-6">
      <title>2.2.Model Training</title>
      <p>In the training phase, the model learns to map inputs (features) to outputs (labels) based on the
training data. We used the Adam optimizer for training. Adam, short for Adaptive Moment Estimation,
is a popular choice for deep learning applications due to its adaptive learning rates, meaning it adjusts
the learning rate for each weight in the model individually. A learning rate scheduler from callbacks in
TensorFlow with a learning rate of 3e-05 and warmup steps of 200 was incorporated to adjust the
learning rate during training dynamically. This helps to fine-tune the learning process, often leading to
better model performance. An early stopping mechanism was also utilized to prevent needless training
once the model’s performance ceased to improve significantly. This not only saves computational
resources but also helps prevent overfitting. Mixed precision training was employed to expedite the
training process. This method involves using a mix of single-precision (float32) and
halfprecision(float16) data types during training, which can significantly reduce the use of computational
resources without compromising the model’s performance.</p>
    </sec>
    <sec id="sec-7">
      <title>2.3.Different Runs</title>
      <p>In We conducted three different runs for Task 1, which involved Sexism Identification. Each run
represented a unique combination of methods and data, allowing us to explore various strategies for
improving the performance of our model.</p>
      <p>• Task 1 - Run 1 (Binary label- Original dataset): In the first run, we trained our model Without
using additional data. We used a voting system that combined the predictions of several models
(BERT, XLM-RoBERTa, and DistilBERT) to make a final decision. This approach, often called
ensemble learning, can help improve the robustness and accuracy of predictions by leveraging
the strengths of multiple models. The task was treated as a binary classification problem, with
the model predicting whether each instance was sexist.
• Task 1 - Run 2 (Binary label- Original and additional dataset): In the second run, we
included additional features derived from the text, metadata, and external resources (EDOS).
Like the first run, we used a voting system with several models and treated the task as a binary
classification problem. The additional information enriched the feature space and gave the
model more context for making predictions.
• Task 1 - Run 3 (Multi-label- Original and additional dataset): In the third run, we again
incorporated additional information and used a voting system with several models. However,
we treated the task as a multilabel classification problem in this run. This means that each
instance could belong to more than one class. We had two labels for the original data (indicating
whether the instance was sexist) and two based on the additional data. This approach allowed
us to capture more nuanced information about the instances and potentially improve the model’s
ability to detect different forms or levels of sexism.</p>
      <p>Our methodology for tasks 2 and 3 is like task 1- run 3. Through different runs, we explored a range
of strategies for sexism identification and gained insights into their relative strengths and weaknesses.
These findings can inform future work in this area, helping to guide the development of more effective
tools for detecting and combating online sexism.</p>
      <p>2.4.</p>
    </sec>
    <sec id="sec-8">
      <title>Evaluation and Validation</title>
      <p>Model performance is generally defined by its ability to predict unseen data accurately.
One of the most popular of these techniques is cross-validation. Cross-validation aims to
estimate the model’s accuracy on the test set during the model training stage. In this regard,
we divide the original data into k partitions of equal size. One partition is designated a
validation dataset, and the remaining k-1 partitions are used as training data. These stages
enable iterative enhancements to the model and its training process, ultimately creating a
reliable machine-learning model that can effectively generalize to new, unseen data.</p>
    </sec>
    <sec id="sec-9">
      <title>2.4.1. System Configuration</title>
      <p>Our system leverages a suite of transformers for text processing sourced from the
Hugging Face Transformers library, a state-of-the-art platform for natural language
processing. After</p>
      <p>preprocessing, we set a max length of 256 for tokenization, reflecting the average length
of the text data. The num-labels parameter was set to 1, indicating binary classification.
Hyper-parameters were optimized via Random Search within a defined range of possible
values. We examined learning rate values between 1e-5 and 1e-4 and assessed batch sizes 16,
32, and 64 to determine the optimal balance between computational efficiency and model
performance. We also implemented a learning rate scheduler that dynamically adjusts the
learning rate according to a cosine decay schedule. The warmup period was set to 200 steps,
with incremental steps computed based on the number of epochs and the dataset size. We
utilized early stopping based on validation loss to prevent overfitting, with the patience set to
3 epochs.</p>
    </sec>
    <sec id="sec-10">
      <title>2.4.2. Model Training and Validation Process</title>
      <p>The models were trained using binary cross-entropy loss and the Adam optimizer, a
popular choice for deep learning applications due to its adaptive learning rates. We
utilized the mixed-precision training policy ’mixed float16’ to expedite training without
compromising model performance. We employed a custom function to construct models,
with each model utilizing a different transformer:  −  −  −
,  −  − , and  −  −  − .
Outputs from these transformers were concatenated and passed through a 1
convolutional layer, a max-pooling layer, a dropout layer (rate of 0.5), and a dense layer
for binary classification with 2 regularization.</p>
      <p>Validation was carried out via stratified dataset splitting, with the test size constituting
20 percent of the entire data. This data was tokenized and passed through the trained
models. The model training process involved searching for the optimal set of
hyperparameters for 3 trials using the validation set. A final model was trained and
evaluated after identifying the best hyperparameters (learning rate and batch size).
2.5.</p>
    </sec>
    <sec id="sec-11">
      <title>Evaluation and Validation</title>
      <p>To gain a comprehensive understanding of our results, we visualized the data distribution
and the proportion of the labels in our data using plots like bar charts, pie charts, or histograms.
This provided valuable insights into the imbalance of the dataset. The distribution of task one
is given in the diagram below; due to the large volume of task two and three labels, we avoided
bringing its diagram.</p>
      <p>We also conducted a text length analysis to examine the lengths of the text data before and
after preprocessing. This helped us determine an appropriate max-length value for the
tokenization and padding step.</p>
      <p>We plotted the loss and accuracy progress over the epochs during the training phase.
This provided insight into whether the model was learning effectively or
overfitting/underfitting. After training, we evaluated our model on the test data and
generated a classification report and a confusion matrix.
dataset was evaluated based on several key metrics: accuracy, precision, recall, and F1
score. In the below table, performance metrics for each task in the test dataset have been
shown.</p>
      <p>According to table 2, in task one, it is clear that the best accuracy was obtained in run
two, where the total data of the original and additional data were used. The classification
was considered in binary form, indicating that the model has been more accurate on the
total data; however, the F index has decreased due to the additional data set in English.
Also, the best F1 score is obtained in the third run, related to using the total data and
multi-labeling (two labels for the original data and two for the additional data).</p>
      <p>In our model’s evaluation, we have utilized Accuracy, Precision, Recall, and F1-score.
How- ever, the evaluation competition adopted metrics such as ICM-Hard, ICM-Soft,
their normalized versions, and Cross-Entropy, designed to evaluate the model’s
performance under various scenarios, including hard and soft classifications. The ’hard’
scenarios deal with discrete, categorical classifications, similar to Accuracy in our current
framework, while the ’soft’ scenarios handle continuous, probabilistic classifications.
Majority and minority class classifiers were also used as baselines to compare the models’
performance over a naïve approach. Our team’s best run was ranked 3rd among the
participants which was related to  3  −  . The details of the official
result are shown in the below table.</p>
      <p>EXIST_2023_Leaderboard_Task3</p>
      <sec id="sec-11-1">
        <title>Task 3 Soft-Soft ALL</title>
      </sec>
      <sec id="sec-11-2">
        <title>Task 3 Hard-Hard ALL</title>
      </sec>
      <sec id="sec-11-3">
        <title>Task 3 Hard-Soft ALL</title>
      </sec>
      <sec id="sec-11-4">
        <title>Task 3 Soft-Soft ES</title>
      </sec>
      <sec id="sec-11-5">
        <title>Task 3 Hard-Hard ES</title>
      </sec>
      <sec id="sec-11-6">
        <title>Task 3 Hard-Soft ES</title>
      </sec>
      <sec id="sec-11-7">
        <title>Task 3 Soft-Soft EN</title>
      </sec>
      <sec id="sec-11-8">
        <title>Task 3 Hard-Hard EN</title>
        <p>M&amp;S_NLP_1
M&amp;S_NLP_1
M&amp;S_NLP_1
M&amp;S_NLP_1
M&amp;S_NLP_1
M&amp;S_NLP_1
M&amp;S_NLP_1
Run
Run
Run
Run
Run
Run
Run</p>
        <sec id="sec-11-8-1">
          <title>Rank</title>
          <p>14</p>
        </sec>
        <sec id="sec-11-8-2">
          <title>Rank 31</title>
        </sec>
        <sec id="sec-11-8-3">
          <title>Rank</title>
          <p>4</p>
        </sec>
        <sec id="sec-11-8-4">
          <title>Rank 11</title>
        </sec>
        <sec id="sec-11-8-5">
          <title>Rank 30</title>
        </sec>
        <sec id="sec-11-8-6">
          <title>Rank</title>
          <p>6</p>
        </sec>
        <sec id="sec-11-8-7">
          <title>Rank 11 Rank</title>
        </sec>
        <sec id="sec-11-8-8">
          <title>ICM-Soft</title>
          <p>-8,3574</p>
        </sec>
        <sec id="sec-11-8-9">
          <title>ICM-Hard</title>
          <p>-2,1587</p>
        </sec>
        <sec id="sec-11-8-10">
          <title>ICM-Soft</title>
          <p>-9,504</p>
        </sec>
        <sec id="sec-11-8-11">
          <title>ICM-Soft</title>
          <p>-8,5493</p>
        </sec>
        <sec id="sec-11-8-12">
          <title>ICM-Hard</title>
          <p>-2,2525</p>
        </sec>
        <sec id="sec-11-8-13">
          <title>ICM-Soft</title>
          <p>-9,6746</p>
        </sec>
        <sec id="sec-11-8-14">
          <title>ICM-Soft</title>
          <p>-7,939</p>
        </sec>
        <sec id="sec-11-8-15">
          <title>ICM-Hard</title>
        </sec>
        <sec id="sec-11-8-16">
          <title>ICM-Soft Norm 0,6793</title>
        </sec>
        <sec id="sec-11-8-17">
          <title>ICM-Hard Norm</title>
        </sec>
        <sec id="sec-11-8-18">
          <title>ICM-Soft Norm</title>
        </sec>
        <sec id="sec-11-8-19">
          <title>ICM-Soft Norm</title>
          <p>0,1838
0,6586
0,6701
0,192</p>
        </sec>
        <sec id="sec-11-8-20">
          <title>ICM-Hard Norm</title>
        </sec>
        <sec id="sec-11-8-21">
          <title>ICM-Soft Norm 0,6496</title>
        </sec>
        <sec id="sec-11-8-22">
          <title>ICM-Soft Norm 0,6957</title>
        </sec>
        <sec id="sec-11-8-23">
          <title>ICM-Hard Norm</title>
          <p>F1
0,0017</p>
          <p>F1
0,0026</p>
        </sec>
      </sec>
      <sec id="sec-11-9">
        <title>Task 3 Hard-Soft EN</title>
        <p>Run
M&amp;S_NLP_1</p>
        <sec id="sec-11-9-1">
          <title>Rank</title>
          <p>3</p>
        </sec>
        <sec id="sec-11-9-2">
          <title>ICM-Soft Norm 0,6745 0,0007</title>
          <p>The official results, along with a thorough soft and hard analysis, reveal that the
developed model performs optimally in English. This outcome may be attributed to the use
of a multilingual model as opposed to a Spanish-focused one, as well as the relative
abundance of English language information in these models. The use of additional datasets,
which were exclusively in English, could also have influenced this result. Future endeavors
include experimentation with multiple models catered to different languages. It’s also
noteworthy to highlight that the developed model outperformed competing teams in the
multi-class classification problem (the third task). This success is likely due to the
implementation of a voting system and the combination of outputs from multiple models,
which collectively contributed to the overall performance.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-12">
      <title>3. Conclusions and Future Work</title>
      <p>Our research presents a comprehensive approach to online sexism detection, leveraging
advanced machine learning techniques and natural language understanding methodologies. We
utilized a multi-model approach incorporating BERT, XLM-RoBERTa, and DistilBERT to
tackle this complex issue. Our methodology was evaluated through a robust experimental setup,
and the results demonstrated the effectiveness of our approach in both data sets, but it needs to
be improved in dealing with unbalanced datasets by considering the lack of sexist text in the
original and additional datasets.</p>
      <p>
        Looking ahead, we plan to extend our research by incorporating additional information about
the annotators, such as gender, age, and other demographic details. This information could
provide valuable context that may influence the interpretation and categorization of sexist
content. For instance, research has shown that perceptions of sexism can vary significantly based
on an individual’s personal experiences and perspectives (Burn, 2000; Swim et al., 2005)
Moreover, we aim to assess the reliability of the annotators. Annotator reliability is a crucial
aspect of any study involving human annotation, as it can significantly impact the quality and
validity of the data [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. By evaluating the reliability of our annotators, we can ensure that our
dataset is robust and reliable, thereby enhancing the validity of our findings. In addition, we
plan to explore more sophisticated techniques for handling imbalanced data and methods for
dealing with subtlety and context dependency in sexist expressions. We also aim to test our
model on different datasets and contexts to assess its generalizability.
      </p>
    </sec>
    <sec id="sec-13">
      <title>3.1.Model Training</title>
      <p>In the training phase, the model learns to map inputs (features) to outputs (labels) based on
the training data. We used the Adam optimizer for training. Adam, short for Adaptive Moment
Estimation, is a popular choice for deep learning applications due to its adaptive learning rates,
meaning it adjusts the learning rate for each weight in the model individually. A learning rate
scheduler from callbacks in TensorFlow with a learning rate of 3e-05 and warmup steps of 200
was incorporated to adjust the learning rate during training dynamically. This helps to fine-tune
the learning process, often leading to better model performance. An early stopping mechanism
was also utilized to prevent needless training once the model’s performance ceased to improve
significantly. This not only saves computational resources but also helps prevent overfitting.
Mixed precision training was employed to expedite the training process. This method involves
using a mix of single-precision (float32) and half-precision(float16) data types during training,
which can significantly reduce the use of computational resources without compromising the
model’s performance.</p>
      <p>We plotted the loss and accuracy progress over the epochs during the training phase. This
provided insight into whether the model was learning effectively or overfitting/underfitting.
After training, we evaluated our model on the test data and generated a classification report
and a confusion matrix.</p>
    </sec>
    <sec id="sec-14">
      <title>4. Data Availability</title>
      <p>In this article, the dataset utilized is specifically associated with EXIST 2023
competition. Furthermore, the code utilized in the study was made accessible through a
dedicated Zenodo [10].</p>
    </sec>
    <sec id="sec-15">
      <title>5. References</title>
      <p>To gain a deeper insight into the model utilized, refer to the diagram above which illustrates its
structure. The outputs from three pre-trained language models are compiled by a single CNN model,
which then categorizes the information based on the task at hand. Detailed information about the
number of parameters can be found in the table provided below.</p>
      <p>Model: "model"
_________________________________________________________________________________________
_________</p>
      <p>Layer (type) Output Shape Param # Connected to
=========================================================================================
=========
text_input (InputLayer) [(None, 256)] 0 []
concatenate (Concatenate)
(None, 256, 2304)</p>
      <p>['tf_bert_model[0][0]',
['concatenate[0][0]']
max_pooling1d (MaxPooling1D)</p>
      <p>(None, 128, 128)
flatten (Flatten)
dropout_93 (Dropout)
dense (Dense)
=====================================================================================
Total params: 581,035,393
Trainable params: 581,035,393
Non-trainable params: 0
_________________________________________________________________________________
16385</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Waseem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hovy</surname>
          </string-name>
          ,
          <article-title>Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter</article-title>
          ,
          <source>in: Proceedings of the NAACL Student Research Workshop</source>
          , Association for Computational Linguistics, San Diego, California,
          <year>2016</year>
          , pp.
          <fpage>88</fpage>
          -
          <lpage>93</lpage>
          . URL: https://aclanthology.org/N16-2013. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N16</fpage>
          -2013.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Young</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hazarika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Poria</surname>
          </string-name>
          , E. Cambria,
          <source>Recent Trends in Deep Learning Based Natural Language Processing</source>
          ,
          <year>2018</year>
          . URL: http://arxiv.org/abs/1708.02709. doi:
          <volume>10</volume>
          .48550/arXiv.1708. 02709, arXiv:
          <fpage>1708</fpage>
          .02709 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bolukbasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.-W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Saligrama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Kalai</surname>
          </string-name>
          ,
          <article-title>Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          , volume
          <volume>29</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2016</year>
          . URL: https://proceedings.neurips.cc/paper/2016/hash/ a486cd07e4ac3d270571622f4f316ec5-Abstract.html.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pavlopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Malakasiotis</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Androutsopoulos</surname>
          </string-name>
          ,
          <article-title>Deeper Attention to Abusive User Content Moderation</article-title>
          ,
          <source>in: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Copenhagen, Denmark,
          <year>2017</year>
          , pp.
          <fpage>1125</fpage>
          -
          <lpage>1135</lpage>
          . URL: https://aclanthology.org/D17-1117. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D17</fpage>
          -1117.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <source>Overview of EXIST</source>
          <year>2023</year>
          :
          <article-title>sEXism Identification in Social NeTworks</article-title>
          ,
          <source>in: Advances in Information Retrieval: 45th European Conference on Information Retrieval</source>
          ,
          <string-name>
            <surname>ECIR</surname>
          </string-name>
          <year>2023</year>
          , Dublin, Ireland, April 2-
          <issue>6</issue>
          ,
          <year>2023</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>III</given-names>
          </string-name>
          , Springer-Verlag, Berlin, Heidelberg,
          <year>2023</year>
          , pp.
          <fpage>593</fpage>
          -
          <lpage>599</lpage>
          . URL: https://doi.org/10.1007/978-3-
          <fpage>031</fpage>
          -28241-6_
          <fpage>68</fpage>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>031</fpage>
          -28241-6_
          <fpage>68</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of exist 2023:
          <article-title>sexism identification in social networks</article-title>
          ,
          <source>in: Proceedings of ECIR'23</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>593</fpage>
          -
          <lpage>599</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -28241-6_
          <fpage>68</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H. R.</given-names>
            <surname>Kirk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Vidgen</surname>
          </string-name>
          , P. Röttger, SemEval-2023
          <source>Task 10: Explainable Detection of Online Sexism</source>
          ,
          <year>2023</year>
          . URL: http://arxiv.org/abs/2303.04222. doi:
          <volume>10</volume>
          .48550/arXiv.2303.04222, arXiv:
          <fpage>2303</fpage>
          .04222 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Artstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Poesio</surname>
          </string-name>
          , Survey Article:
          <article-title>Inter-Coder Agreement for Computational Linguis-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9] tics,
          <source>Computational Linguistics</source>
          <volume>34</volume>
          (
          <year>2008</year>
          )
          <fpage>555</fpage>
          -
          <lpage>596</lpage>
          . URL: https://aclanthology.org/J08- 4004. doi:
          <volume>10</volume>
          .1162/coli.07-034
          <string-name>
            <surname>-R2. H. Mohammadi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Giachanou</surname>
            , &amp;
            <given-names>A. Bagheri.</given-names>
          </string-name>
          (
          <year>2023</year>
          ).
          <article-title>Code for "Towards Robust Online Sexism Detection: A Multi-Model Approach with BERT, XLM-RoBERTa, and DistilBERT for EXIST 2023 Tasks"</article-title>
          . URL: https://doi.org/10.5281/zenodo.8144300.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>