<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>IUEXIST: Multilingual Pre-trained Language Models for Sexism Detection on Twitter in EXIST2023</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yash A. Hatekar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Muhammad S. Abdo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Snigdha Khanna</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandra Kübler</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indiana University</institution>
          ,
          <addr-line>Bloomingtonm, IN</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>8</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>We describe an approach towards sexism detection in tweets, for the EXIST 2023-Task 1, a shared task on sexism identification. The dataset for this task consists of English and Spanish tweets. Task 1 is a binary classification task, where our system needs to decide whether a given tweet contains sexist expressions or behaviors. We describe our experiments with diferent machine learning algorithms and vector lengths, algorithms including Multinomial Naive Bayes, SVM, XGBoost, transformers, and Distilbert. The best model performance was achieved by an ensemble of transformers including XLM-Roberta small and large and TwHIN-BERT base and large, combined using XGBoost. The ensemble was trained on the original tweets dataset plus additional training data from the 2021 shared task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The past two decades witnessed an unprecedented surge in the amount of online content
produced by social network users. Unfortunately, the rapid growth and ubiquity of this content
made them a fertile ground for darker human emotions, including sexism. The definition of
sexism often varies, but it generally refers to discriminatory practices or beliefs on the basis of
sex or gender. It can take on various forms, which may range from subtle and indirect to overt
and hidden expressions. Most often, these forms of discrimination are expressed against women
with the aim of humiliating or objectifying them, destroying their reputation, undervaluing
their skills and opinions, or making them feel fearful and vulnerable [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4">1, 2, 3, 4</xref>
        ]. Hence, hatred,
threats, harassment, intimidation, and disparagement may all be the results of such sexist
content. Fox et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] argue that sexist behavior is promoted due to the ‘online dis-inhibition
efect’, i.e., online users who remain anonymous may exhibit behaviors that they would not
typically display in face-to-face situations or when their identity is known. They also argue
that engaging with sexist content online may lead to sexist attitudes ofline. As a result, the
      </p>
      <p>Sexist Call me sexist but it just feels wrong that women are refing the NBA
like go ref the WNBA.</p>
      <p>Esta gringa sigue llorando por el gamergate, que c¨oincidenciaq¨ue tenga
pronombres en su perfil
Non-Sexist Even if you get embarrassed and blush, you can still confront hard
things. #KeepMoving
Los políticos acostumbran a hablarle al pueblo como si fueran una
manada de estúpidos pero lamanada no hacemos nada por contradecirlos.
automatic detection and classification of this content into distinct categories have become a
critical task to promote gender equality and create a safe online environment for everyone.</p>
      <p>
        Machine learning techniques have proven to be efective in detecting and classifying sexist
content. By training machine learning models on large datasets labeled for sexism, algorithms
can learn the patterns and features that characterize such content. Both binary classification
(i.e., sexist and non-sexist) or a more fine-grained classification, such as implicit and direct
sexism exist [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6, 7, 8</xref>
        ]. Nevertheless, detecting sexist content on online platforms is challenging,
especially on Twitter. Tweets are typically short, making it dificult for the models to extract
unique patterns and features which discriminate sexist from non-sexist content. Also, because
Twitter users have to limit their tweets to a small number of words, they resort to using
nonstandard language, emojis, and abbreviations, among other ways, to send their messages in the
shortest form. Additionally, sarcasm, irony, and vague language make it dificult for the models
to perform well [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        In this paper, we present our IUEXIST team submissions for the EXIST 2023 [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ] Shared
Task 1. For this task, the system needs to decide whether a given tweet in English or Spanish
is sexist or not. As per the task, sexist is chosen if the tweet i) is sexist, ii) describes a sexist
situation, iii) criticizes a sexist behavior. Table 1 shows sample sexist and non-sexist tweets in
both English and Spanish.
      </p>
      <p>The remainder of the paper is organized as follows. In section 2, we present a review of related
work on detecting sexist content on various online platforms. In section 3, we describe our
methodology, including the dataset, data pre-processing, and the machine learning algorithms
used in our experiments. We then present our experimental results and discuss the implications
of our findings in section 4. Finally, in section 5, we conclude the paper by summarizing our
contributions, discussing the limitations of our study, and outlining avenues for future research.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        The SemEval-2023 Shared Task 10 [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] aimed to improve the automatic detection of online
sexism. Unlike previous studies that focused on the binary classification of sexist content, this
task introduces a new hierarchical taxonomy of sexist content that contains granular vectors
of sexism. The study uses a dataset of 20k social media comments and aims to create more
accurate and explainable models for sexism detection. The taxonomy included four categories
(Threats, Derogation, Animosity, and Prejudiced discussion) and 11 subcategories (e.g., threats
of harm, aggressive attacks, gender stereotypes, and supporting mistreatment of women). The
data used in the study was compiled from both Reddit and Gab, and sexist content was later
annotated by highly-trained female annotators. The shared task involved three main tasks:
a) a binary classification (sexist vs non-sexist), b) a four-category classification, and c) an
11ifne-grained-vector classification. The leading system in Task A employed a multi-task DNN
structure and performed additional pretraining of DeBERTa-v3 and TwHIN-BERT on the starter
kit unlabelled data, as well as an extra dataset. In Task B, the top-performing system utilized an
instruction-tuned Pathways Language Model (PaLM) with the model, with a prompt that was
parameter-eficient and tuned specifically for the task data. The system used majority voting
over six iterations. Lastly, for Task C, the best-performing system conducted further training of
DeBERTa-v3 using the starter kit unlabelled data and incorporated a second loss term known
as normalized temperature-scaled cross entropy.
      </p>
      <p>
        Almanea and Poesio [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] created a corpus of Arabic misogynistic tweets, annotated by three
annotators, and used it to train a model for classifying tweets for misogyny using AraBERT.
They trained a binary classifier for each coder with soft loss functions and a majority vote hard
training. The results showed that the model trained using CE soft loss had the highest accuracy
(77.79%) and F1-score (77.38), but had a relatively higher cross-entropy (0.586) and JSD (0.244)
compared to the other models. The overall agreement between the three annotators was low,
and the results suggest that annotator subjectivity has a significant impact on the accuracy of
machine learning models for classifying sexist language.
      </p>
      <p>
        Parikh et ak. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] developed a semi-supervised multi-task learning neural framework for the
multi-label fine-grained sexism classification of accounts of sexism, using sentence
representations from word embeddings and pre-trained models. The study used a dataset of 13 023 accounts
of sexism that was tagged with 23 diferent categories of sexism, created by trained annotators
who had formal experience with studying gender and/or sexuality. The study explored diferent
baselines and approaches for classifying sexism. With regard to baselines, random labeling
and traditional machine learning methods such as SVM, random forest, and logistic regression
were explored. The features chosen in these methods include TF-IDF on character n-grams,
word unigrams, and bigrams, ELMo embeddings, and a composite set of features. Additionally,
they explored various deep learning architectures, including LSTM-based architectures such as
biLSTM, biLSTM-Attention, and hierarchical-biLSTM-Attention. They also included sentence
embeddings with biLSTM-attention, CNN-based architectures such as CNN-Kim and C-biLSTM,
and CNN-biLSTM-Attention. Logistic regression with averaged ELMo embeddings as features
was found to perform best among the traditional ML methods with an F1 score of 0.595, and a
macro-F of 0.479. Among the deep learning baselines, biLSTM-Attention (F1: 0.728, macro-F:
0.650) and Hierarchical-biLSTM-Attention (F1: 0.725, macro-F: 0.650) were the best models.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>
        3.1. Data
For the development of the EXIST dataset, over 400 popular expressions and terms that are
commonly used to undermine women’s roles in society, in English and Spanish, were used
as search terms. Overall, the original training data consists of 3,660 Spanish tweets and 3,260
English tweets. We also used additional data from the EXIST task1 datasets in 2021 and 2022
[
        <xref ref-type="bibr" rid="ref14 ref4">4, 14</xref>
        ]. The final size of the training dataset is 8 960 tweets in Spanish and English out of which
5 593 were sexist and 3 367 non-Sexist.
      </p>
      <p>Since the tweets were provided with six annotator votes, we use a majority voting scheme
to label the tweets as either sexist or not-sexist. For tweets for which there was a tie in the
annotations, we consider them sexist to partly address the class imbalance.</p>
      <sec id="sec-3-1">
        <title>3.2. Data Pre-Processing</title>
        <p>The data pre-processing step involved five steps. 1) Any URLs were replaced by ’URL’. 2)
Retweet ’RT’, which is not relevant to the task, was removed. 3) Usernames were replaced with
the word ’USER’. 4) Emojis were converted to their corresponding text equivalents using the
Python library ’emoji’ (https://pypi.org/project/emoji/). 5) All non-alphanumeric characters,
except apostrophes and spaces, were removed.</p>
        <p>A first attempt to eliminate hashtags showed that they are helpful and should not be deleted.
For instance, the tweet "#Catcalling is #Harassment. It’s Not a Compliment. It’s never okay.
#feminist #feminism #stopstreetharassment https:// t.co/ g5nJy12sIl’", is labeled as sexist by the
majority of annotators in the training dataset. Removing the hashtags results in the removal of
all relevant content.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Classifiers</title>
        <p>
          We used a range of classifiers: multinomial Naive Bayes and Support Vector Machines using the
scikit-learn implementation [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], XGBoost [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], DistilBERT [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], RoBERTa [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], XLM-RoBERTa
[
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], and TwHIN [20]. We used HuggingFace to fine-tune these transformers for our Twitter
dataset.
        </p>
        <p>Since the transformers can only accept input of a predetermined maximum length, we
experiment with diferent vector lengths and found the following lengths optimal: 95 for
XML-RoBERTa base and 128 for the remaining transformers.</p>
        <p>For the multinomial Naive Bayes and SVM, we used the default parameters. To
hyperparameterize XGBoost, we used GridSearchCV along with a five-fold cross-validation.The best
hyperparameters were identified as a maximum depth of 128, a learning rate of 0.1, the number
of estimators set to 200, a seed of 47, and the internal evaluation metric set to logloss. We used
HuggingFace’s Autotrain to fine-tune the transformer models.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Evaluation</title>
        <p>The oficial score in task 1 is ICM (Information Contrast Measure) [ 21], thus we report our
results using this metric. We also report macro-averaged F1 in the HARD-HARD setting.</p>
        <p>After considering former studies and the gold labels provided by the guidelines, we decided
to focus on the Hard-Hard evaluation.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <sec id="sec-4-1">
        <title>4.1. Oficial Results</title>
        <p>We submitted two systems for evaluation. IUEXIST_1 uses XLM-RoBERTa Large trained on
the oficial training set provided by the shared task. IUEXIST_2 uses an ensemble of four
transformers, XLM-RoBERTa base and large, and TwHIN base and large. We then train XGBoost
on the output of the transformers. All the ensemble models are trained on the combination of
the oficial training set and the additional data (see Section 3.1).</p>
        <p>Table 2 provides a summary of the oficial result or our team’s submission. These results
show that IUEXIST_2 performs slightly better than IUEXIST_1 in the hard evaluation while
IUEXIST_1 performs significantly better in the soft evaluation. The gains of IUEXIST_2 in
the hard evaluation are due to gains in Spanish, where the ICM is about 0.02 higher than
for IUEXIST_1 (0.5460 for IUEXIST_2 and 0.5294 for IUEXIST_1). These gains are ofset by a
smaller loss for English. The good performance of IUEXIST_1 in the soft evaluation is due to its
performance in English (ICM: 0.7115 vs. 0.6141 for IUEXIST_2).</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Results on the Development Set</title>
        <p>In addition to the oficially submitted systems, we performed a more extensive evaluation on
the development set.</p>
        <p>We trained and evaluated the 10 diferent classifiers and the ensemble described in Section 3.3,
the individual classifiers trained on the original training set, and the ensemble on the extended
training set.</p>
        <p>The results of these experiments are shown in Table 3. Since ICM and the F-scores show
the same trends, we will focus on ICM scores in this analysis. The best model performance
was achieved by the ensemble without pre-processing with an ICM of 0.5873. The non-neural
classifiers generally show lower performance than the transformers. Among the latter, the
XLM-RoBERTa base shows the best performance, with an ICM of 0.5716.</p>
        <p>When we look at the question of whether pre-processing is useful, we see that some classifiers,
Pre-processing</p>
        <p>Classifier
original
pre-processing
multinomial Naive Bayes
Support Vector Machines
XGBoost
DistilBERT
ROBERTA base
XLM-RoBERTa base
XLM-RoBERTa large
TwHIN base
TwHIN large
ensemble
multinomial Naive Bayes
Support Vector Machines
XGBoost
DistilBERT
RoBERTa base
XLM-RoBERTa base
XLM-RoBERTa large
TwHIN base
TwHIN large
such as the multinomial Naive Bayes and RoBERTa base profit from pre-processing, but for
most models, pre-processing is detrimental.</p>
        <p>Since our oficial results leave us with the question of whether the gains of the IUEXIST_2
model over IUEXIST_1 are due to the ensemble approach or to the extended training set, we
perform a comparison including an experiment where we use the ensemble with only the
current training set. The results are shown in Table 4. They show that the gains are mostly due
to the additional training data, the ICM for the ensemble with this year’s training data only
reaches 0.5634, in comparison to 0.5534 for the XLM-RoBERTa model. Adding the 2021 training
data adds a larger gain to 0.5873.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>We have presented our submissions to the EXIST 2023 shared task 1. We found that an ensemble
of four diferent transformers with XGBoost for voting provides the best results in the
HARDHARD evaluation (rank 15). However, for the SOFT-SOFT evaluation, we found XML-RoBERTa
to reach significantly higher results. This system was ranked 9th out of 70 submissions.</p>
      <p>As described above, this system was developed as a project in a course on machine learning.
For most of us, this was our first time collaborating on a shared NLP task, drawing on experiences
and discussions with individuals from various academic backgrounds and cultural perspectives.
One of the key challenges we faced was being able to confine ourselves to the definition of
Sexism, as multiple examples in the data set seemed to be open to interpretation depending on
the context.</p>
      <p>Yet another challenge was knowing how to sift through the extensive amount of technical
information and resources available to us, and to direct our attention to problem-solving through
continuous learning.</p>
      <p>For the future, we plan on investigating the efect of using additional training data for the
diferent systems since the results showed that adding more training data was more successful
than going from a single transformer to the ensemble. On the one hand, adding training data
can help combat data sparsity, but it also adds the risk of distorting the class distribution. We are
also interested in a long-term evaluation to see to what degree the temporal distance between
the training and test has a negative efect on performance.
[20] X. Zhang, Y. Malkov, O. Florez, S. Park, B. McWilliams, J. Han, , A. El-Kishky, TwHIN-BERT:
A Socially-Enriched Pre-trained Language Model for Multilingual Tweet Representations,
Technical Report arXiv:2209.07562, arXiv, 2022.
[21] E. Amigó, A. Delgado, Evaluating extreme hierarchical multi-label classification, in:
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics
(ACL), 2022, pp. 5809–5819.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Abburi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chhaya</surname>
          </string-name>
          , Niyati ad Varma,
          <article-title>Fine-grained multi-label sexism classification using a semi-supervised multi-level neural approach</article-title>
          ,
          <source>Data Science and Engineering</source>
          <volume>6</volume>
          (
          <year>2021</year>
          )
          <fpage>359</fpage>
          -
          <lpage>379</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Chiril</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Moriceau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Benamara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mari</surname>
          </string-name>
          , G. Origgi,
          <string-name>
            <given-names>M.</given-names>
            <surname>Coulomb-Gully</surname>
          </string-name>
          ,
          <article-title>An annotated corpus for sexism detection in French tweets</article-title>
          ,
          <source>in: Proceedings of the Twelfth Language Resources and Evaluation Conference</source>
          , Marseille, France,
          <year>2020</year>
          , pp.
          <fpage>1397</fpage>
          -
          <lpage>1403</lpage>
          . URL: https: //aclanthology.org/
          <year>2020</year>
          .lrec-
          <volume>1</volume>
          .
          <fpage>175</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. K.</given-names>
            <surname>Swim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mallett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Stangor</surname>
          </string-name>
          ,
          <article-title>Understanding subtle sexism: Detection and use of sexist language</article-title>
          ,
          <source>Sex Roles</source>
          <volume>51</volume>
          (
          <year>2004</year>
          )
          <fpage>117</fpage>
          -
          <lpage>128</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rodríguez-Sánchez</surname>
          </string-name>
          , J. C. de Albornoz, L. Plaza,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Comet</surname>
          </string-name>
          , T. Donoso,
          <source>Overview of EXIST</source>
          <year>2021</year>
          :
          <article-title>sEXism Identification in Social neTworks</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <year>2021</year>
          )
          <fpage>195</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Fox</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Cruz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Perpetuating online sexism ofline: Anonymity, interactivity, and the efects of sexist hashtags on social media</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>52</volume>
          (
          <year>2015</year>
          )
          <fpage>436</fpage>
          -
          <lpage>442</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mamidi</surname>
          </string-name>
          ,
          <article-title>When does a compliment become sexist? Analysis and classification of ambivalent sexism using Twitter data</article-title>
          ,
          <source>in: Proceedings of the Second Workshop on NLP and Computational Social Science</source>
          , Vancouver, Canada,
          <year>2017</year>
          , pp.
          <fpage>7</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sharifirad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jacovi</surname>
          </string-name>
          ,
          <article-title>Learning and understanding diferent categories of sexism using convolutional neural network's filters</article-title>
          ,
          <source>in: Proceedings of the 2019 Workshop on Widening NLP</source>
          , Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Abburi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Badjatiya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Krishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Chhaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Varma</surname>
          </string-name>
          ,
          <article-title>Multilabel categorization of accounts of sexism using a neural framework</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          ,
          <source>Hong Kong, China</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1642</fpage>
          -
          <lpage>165</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Butt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ashraf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          ,
          <article-title>Sexism identification using BERT and data augmentation - EXIST2021</article-title>
          , in: IberLEF@ SEPLN,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de-Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2023 -
          <article-title>Learning with Disagreement for Sexism Identification and Characterization</article-title>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Giachanou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Aliannejadi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Vlachos</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality</source>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          , Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de-Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2023 -
          <article-title>Learning with Disagreement for Sexism Identification and Characterization (Extended Overview)</article-title>
          , in: M.
          <string-name>
            <surname>Aliannejadi</surname>
            , G. Faggioli,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
          </string-name>
          , M. Vlachos (Eds.),
          <source>Working Notes of CLEF 2023 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H. R.</given-names>
            <surname>Kirk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Vidgen</surname>
          </string-name>
          , P. Röttger, SemEval-2023
          <source>Task 10: Explainable Detection of Online Sexism</source>
          ,
          <source>Technical Report arXiv:2303.04222</source>
          , arXiv,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D.</given-names>
            <surname>Almanea</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Poesio, ArMIS - the Arabic misogyny and sexism corpus with annotator subjective disagreements</article-title>
          ,
          <source>in: Proceedings of the Thirteenth Language Resources and Evaluation Conference</source>
          , Marseille, France,
          <year>2022</year>
          , pp.
          <fpage>2282</fpage>
          -
          <lpage>2291</lpage>
          . URL: https://aclanthology. org/
          <year>2022</year>
          .lrec-
          <volume>1</volume>
          .
          <fpage>244</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rodríguez-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mendieta-Aragón</surname>
          </string-name>
          , G. MarcoRemón, M. Makeienko,
          <string-name>
            <given-names>M.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2022:
          <article-title>Sexism identification in social networks</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>69</volume>
          (
          <year>2022</year>
          )
          <fpage>229</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          , et al.,
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <year>2011</year>
          )
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Guestrin</surname>
          </string-name>
          ,
          <article-title>Xgboost: A scalable tree boosting system</article-title>
          ,
          <source>in: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining(KDD)</source>
          , New York, NY,
          <year>2016</year>
          , pp.
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Debut</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. C.</surname>
          </string-name>
          andThomas Wolf,
          <article-title>DistilBERT, a distilled version of BERT: Smaller, Faster, Cheaper and Lighter</article-title>
          ,
          <source>Technical Report abs/1910</source>
          .01108,
          <issue>ArXiv</issue>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , V. Stoyanov,
          <string-name>
            <surname>RoBERTa: A Robustly Optimized BERT Pretraining Approach</surname>
          </string-name>
          ,
          <source>Technical Report abs/1907</source>
          .11692, arXiv,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chaudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wenzek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Guzmán</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Unsupervised cross-lingual representation learning at scale</article-title>
          ,
          <source>in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>8440</fpage>
          -
          <lpage>8451</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .acl-main.
          <volume>747</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .acl-main.
          <volume>747</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>