<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IU-NLP-JeDi: Investigating Sexism Detection in English and Spanish</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Matthew Buzzell</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jeremy Dickinson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Natasha Singh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandra Kübler</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indiana University</institution>
          ,
          <addr-line>Bloomington, IN</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <abstract>
        <p>In this paper we present the results from three diferent classification algorithms our team (IU-NLP-JeDi) developed for Task 1 of the EXIST 2023 shared task on Sexism Identification on Social Networks. The task consists of identifying sexism within English and Spanish tweets. We separated the English and Spanish tweets and then developed two diferent neural model approaches and an SVM model for each language. We achieved our highest ICM score on the test set from the RNN model.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;TweetTokenizer</kwd>
        <kwd>SVM</kwd>
        <kwd>RNN</kwd>
        <kwd>CNN</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>goal of this shared task is to develop systems that can decide whether or not a given tweet
contains sexist expressions or behaviors.</p>
      <p>Our work focuses on investigating diferent machine learning models with regard to their
suitability for the task. In this study, we conducted an in-depth analysis of the impact of several
factors on the performance of various models. Specifically, we examined :
1. The efects of data pre-processing techniques on model performance.
2. The implications of utilizing word-level versus character-level models.</p>
      <p>3. The use of soft labels versus hard labels during the training process.</p>
      <p>These approaches were developed during a course on machine learning in NLP.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        There is a growing body of research on the topic of sexism detection. In line with our task,
Vaca-Serrano [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] built a system for sexism identification and for sexism categorization for both
English and Spanish. For the English tweets, DeBERTa-v3-large, RoBERTa-large and
BERTweetlarge were trained, while for Spanish tweets, BERTIN, MarIA-base, BETO, and RoBERTuito
were trained. After training, Vaca-Serrano implemented ensemble learning using weighted
majority voting over all the models to decide on the final classification for a tweet. For sexism
identification, BERTweet-large reached the highest F1-score of 0.903 for English tweets while
MarIA-base reached the highest F1-score of 0.883 for Spanish tweets. For sexism
identification, DeBERTa-v3-large had the highest F1-score of 0.729 for English tweets while BETO had
the highest F1-score of 0.820 for Spanish tweets. Another approach taken by Chiril et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
examined the efectiveness of data augmentation methods based on sentence similarity and
the use of gender stereotype detection for sexism classification on a multilingual data set. The
best results were obtained by a SentenceBERT model trained to detect both sexism and
gender stereotypes (multiclass classification), which achieved precision and recall scores of 0.816
and 0.827 respectively, outperforming a BERT model trained on word embeddings, linguistic
features and generalization strategies. Difering from the previous two approaches, Jha and
Mamidi [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] used an SVM and a sequence-to-sequence model to detect sexism according to
ambivalent sexism theory [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], according to which sexism comes in hostile and benevolent forms.
By training an SVM model and a sequence-to-sequence model on a dataset labeled for hostile
and benevolent sexism they found that the SVM model outperformed the sequence-to-sequence
model for benevolent tweets, while the sequence-to-sequence model outperformed the SVM
model for the detection of hostile tweets.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>3.1. Data
We used the EXIST 2023 shared task training, validation, and test set provided by the shared task
organizers. This dataset was constructed by compiling a list of over 400 commonly used sexist
expressions in English and Spanish, and tweets containing these expressions were extracted for
both languages. Each tweet was then classified by six crowd-sourced annotators as “sexist” or
“nonsexist”. Additionally, the annotators’ gender (male or female) and age group (18-22 years,
23-45 years, or 46 or more years) was recorded. Since these annotations were crowdsourced,
the annotators were provided with guidelines created by two experts in gender issues. It is
important to note that this dataset follows the learning with disagreements paradigm, i.e., there
were no gold annotations as such provided. Given that six annotations were provided for
each tweet, there were cases in which there was a tie for the majority class. To generate the
hard labels for the dataset, any tweet labeled as sexist by three or more of the annotators was
considered “sexist”.</p>
      <sec id="sec-3-1">
        <title>3.2. Pre-Processing</title>
        <p>HTML characters were converted to Unicode (e.g., the HTML “&amp;gt;” was converted to “&gt;”, the
text font was normalized to Roman characters, removing any non-Roman characters while
ensuring Spanish characters with diacritics were not removed. Any emoticons were converted to their
emoji equivalents and URL links were removed. Spaces were added between consecutive
usernames, and any symbols were converted to their literals (e.g., the ° symbol to the word “degrees”).
All special characters were removed before passing each tweet through NLTK’s TweetTokenizer.
Additionally, any numbers or words with numbers were removed. Subsequently, all hashtags
were passed through a parser that removed the “#” symbol and tokenized the contents of the
hashtag using Wordninja (https://github.com/keredson/wordninja), a probabilistic parser of
concatenated words that was trained on the Spanish Billion Words Corpus [8] and the Kaggle English
Word Frequency dataset (https://www.kaggle.com/datasets/rtatman/english-word-frequency).
Pairs of upside-down question/exclamation marks and rightside-up question/exclamation marks
were condensed down to a single question mark and exclamation point respectively. Finally,
duplicate usernames, exclamation points, and question marks were counted before being removed.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Models</title>
        <p>The following classification algorithms were chosen for the task of sexism detection in
Spanish and English: Support Vector Machines with term frequency-inverse document frequency
(TF-IDF) using bigrams and trigrams. Recurrent Neural Network (RNN) with one embedding
layer, two bidirectional long short-term memory layers, followed by a dense layer with softmax
activation. A Convolutional Neural Network (CNN) model with an embedding layer, a
convolutional layer, a max pooling layer, a flatten layer, and two dense layers with ReLU and sigmoid
activations respectively.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Data Transformation &amp; Feature Extraction</title>
        <sec id="sec-3-3-1">
          <title>3.4.1. Support Vector Machines</title>
          <p>Support Vector Machines seek to find a line or hyperplane that maximizes the margin between
two classes that are projected into vector space [9]. The SVM models trained in the present
study utilize a mixture of features for sexism detection. In this study, five SVM models were
trained for Spanish, and five SVM models were trained for English, resulting in ten SVM models
total.</p>
          <p>All models were trained using TF-IDF of word / character bigrams and trigrams. The additional
features used to train the SVM models, except for the TF-IDF bigrams and trigrams, were obtained
using the tokenizer described in section 3.2. This tokenizer counts and returns the number of
usernames in each tweet, the number of exclamation points, questions marks, usernames used
in the possessive, and the number of hashtags present in the tweet.</p>
          <p>Furthermore, upsampling was conducted on two clean and two original datasets for each
language. The upsampling step consists of duplicating the sexist tweets present in the training
set. The purpose of conducting upsampling was to increase the number of sexist tweets to
improve recall on the minority class (sexism).</p>
          <p>To summarize, the following models were trained for each language, all with TF-IDF character
bigrams and trigrams: one with clean data, one with the original data, one with upsampling on
the clean data, and one with upsampling on the original data, and one with upsampling and the
additional features described (username counts, hashtag counts, etc.) on the clean data.</p>
          <p>For parameter optimization, we performed grid search to obtain the best regularization(C)
and gamma parameter for each SVM model. The optimal parameters for each SVM model are
shown in Table 1.</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>3.4.2. Neural Models</title>
          <p>First, unique words from the tokenized input data are collected to form a vocabulary. Then two
special tokens ’[UNK]’ and ’[PAD]’ are added to vocabulary. [UNK] is used to mask any word in
test data which is not present in the vocabulary. [PAD] is used to make all the sentences of equal
length. The vocabulary thus obtained is then used to encode the words in an input sentence to
numbers based on the index at which they are present in the vocabulary. All unknown words
are mapped to the index of [UNK] token. Finally, all the sentences are extended to a sentence of
length ‘max_len’ (100 words or 300 chars) by adding [PAD] tokens at the end. Similarly, the
output labels are one hot-encoded and are used to train the RNN and CNN model for 10 epochs.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>3.5. Evaluation Metrics</title>
        <p>To allow for the comparison of the classification algorithms implemented in this study, the
key metrics calculated were Accuracy, Precision, Recall and F1-score. The performance of the
models were assessed primarily using the F1-score during the training and validation phase.</p>
        <p>The final evaluation metric used to evaluate the performance of models on the test set was
ICM metric [10]. ICM calculates the below scores for each model:
1. HARD-HARD: hard system output and hard ground truth.
2. HARD-SOFT: hard system output and soft ground truth.</p>
        <p>3. SOFT-SOFT: soft system output and soft ground truth.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <sec id="sec-4-1">
        <title>4.1. Performance in the Oficial Evaluation</title>
        <p>For the oficial ICM evaluation on the test set, we submitted the predictions obtained by an
RNN model (IU-NLP-JeDi_3) that was trained on the word-level processed sequences with soft
labels, a CNN model (IU-NLP-JeDi_2) that was trained on word-level raw sequence with soft
labels, and an SVM model (IU-NLP-JeDi_1) that was trained using TF-IDF of character bigrams
and trigrams of a processed sequence.</p>
        <p>Table 2 shows the results of the oficial evaluation on the text set for these models. Among
these models, the RNN model exhibits the highest performance across all three evaluation types,
followed by the SVM model and then the CNN model. Notably, both the RNN model and the
SVM model attained the highest ICM scores of 0.2753 and 0.2676, respectively, in the Hard-hard
evaluation.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Performance on the Validation Set</title>
        <p>We performed more extensive experiments on the validation set. Table 3 shows the results of
the ICM evaluations for the models with the best performance. Table 4 shows the results of
a range of CNN, RNN, and SVM experiments respectively. The best macro F1 scores for each
model are marked in bold for each language.</p>
        <p>Based on the macro F1-scores, the highest performing SVM model for English was trained
on TF-IDF character bigrams and trigrams constructed from English tweets that had not been
preprocessed. This English model included upsampling of the minority class and achieved
a macro F1-score of 0.746. The best performing CNN and RNN models were trained using
word-level processed input sequences and hard-label outputs. These models achieved a macro
F1-score of 0.721 and 0.7546 respectively. As for Spanish, the best performing SVM model was
trained on TF-IDF character bigrams and trigrams taken from preprocessed Spanish tweets
without any additional features. This model did not include upsampling and it reached a macro
F1-score of 0.740. The highest performing CNN and RNN models were trained on word level
raw sequences with hard labels. They attained an F1-score of 0.702 and 0.701 respectively.</p>
        <p>For the ICM evaluation on the validation set, the prediction result of each model for both
languages was consolidated and evaluated based on Hard-hard and Hard-soft scores. Among
the various models, the SVM model trained using TF-IDF of processed input and the RNN
model trained on word-level processed sequences and soft labels emerged as the top-performing
models, exhibiting Hard-hard scores of 0.2992 and 0.2923, respectively. Only the RNN trained
on word-level sequences and soft labels achieved a positive Hard-soft score of 0.2831. All the
remaining models performed poorly on this metric and obtained negative results.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Investigating the Efects of Preprocessing</title>
        <p>Table 4 also shows the experiments using pre-processing for all models and for upsampling
with the SVM. These results show that the cleaning and preprocessing step results in a higher
performance for both languages. For English, the macro F1-score increases from 0.727 to 0.739,
and for Spanish, it increases from 0.692 to 0.740. Upsampling sexist tweets also improves the
SVM models. The results also show that preprocessing boosts the performance of the RNN and
CNN models for English, but afects the performance negatively for Spanish. In the case of
English, the RNN model’s F1-score improves from 0.722 to 0.7546 and the CNN model’s F1-score
improves from 0.713 to 0.721. However, for Spanish, the RNN model’s F1-score declines from
0.701 to 0.6849 and the CNN model’s F1-score drops from 0.705 to 0.6947.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Limitation and Challenges</title>
      <p>The main limitations in our study comes from two sources: (i) the inability of neural models, such
as RNNs and CNNs, to handle long sequences and (ii) the information loss due to preprocessing.
Some of the tweets in the EXIST 2023 dataset were around 100 words long and in such cases,
the neural models are not able to efectively model long dependencies. Neural models tend
to forget the initial words of the sequence, leading to a decline in model performance. This
problem is more prominent when using a character level model where the sequence length
was roughly 300 characters long. This was evident from the experiments as the character-level
models consistently under-performed against the word-level models. The other major limitation
of information loss due to preprocessing was a combined efect of non-standard spelling in the
tweets, the presence of both English and Spanish words in the same tweet, and the presence of
words and characters from other languages. All words which were not seen during the training
phase were replaced by [UNK] token, thus preventing the models from accessing information
which may be crucial in determining if a tweet was sexist or not. In our training setup, we
created a separate model for English and Spanish tweets. Due to this, we categorized the tweets
with both English and Spanish words as Spanish and treated English words as out of vocabulary
words, thus replacing them with the token [UNK]. Lastly, some tweets had characters/words
from other languages which were completely removed in our analysis. All these factors lead to
an information loss for our models, hence decreasing the model’s performance.</p>
      <p>Additionally, there were many challenges we faced throughout our study, specifically focused
on the creation of an accurate tokenizer. While analyzing the tweets, we encountered tweets
using non-Roman characters, leading to issues when trying to tokenize these tweets. These
included characters from other languages, such as Arabic, Gregorian, CJK, Hangul, and Hiranaga
characters, and non-traditional styles of Roman characters, such as Gothic letters. To handle
this issue, we used the unicode values of the characters we wanted and removed any characters
outside of that range of unicode values. Another challenge with the tokenizer came from
punctuation marks and the diferences in punctuation between Spanish and English. In order
to simplify processing, we decided to delete the Spanish inverted punctuation. However,
punctuation tends to be irregular or missing in tweets. For this reason, we developed a heuristic
that deleted an inverted punctuation mark only if a regular one could be found in the tweet.
We faced additional challenges when handling emoji, removing usernames and inconsistent
diacritics. For all of those cases, we developed diacritics.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Future Work</title>
      <p>In this study we demonstrated that when optimized on F1-score, the neural models trained on
word level with hard labels and an SVM trained on TF-IDF of character bigrams and trigrams
are our best performing models. However, when optimized on ICM metric, the neural models
show better performance with soft labels and the behavior of the SVM remains unchanged.
These findings illustrate the dependency of model behavior on the choice of evaluation metric.
Additionally, we showed that applying upsampling techniques on the minority class can enhance
the performance of SVM models when dealing with an imbalanced datasets.</p>
      <p>For future work, we will investigate how to utilize the gender and age information of the
annotators, either by modeling this latent variable in the models, or by choosing reliable training
data or labels. Additionally, we will investigate whether the labels are influenced by the gender
bias of the annotators.
[8] C. Cardellino, Spanish Billion Words Corpus and embeddings, 2019. https://crscardellino.</p>
      <p>github.io/SBWCE/.
[9] N. Christianini, J. Shawe-Taylor, An Introduction to Support Vector Machines and Other</p>
      <p>Kernel-Based Learning Methods, Cambridge University Press, Cambridge, 2000.
[10] E. Amigó, A. Delgado, Evaluating extreme hierarchical multi-label classification, in:
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics
(ACL), 2022, pp. 5809–5819.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. K.</given-names>
            <surname>Swim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. L.</given-names>
            <surname>Hyers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. L.</given-names>
            and
            <surname>Cohen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Ferguson</surname>
          </string-name>
          ,
          <article-title>Everyday sexism: Evidence for its incidence, nature, and psychological impact from three daily diary studies</article-title>
          ,
          <source>Journal of Social Issues</source>
          <volume>57</volume>
          (
          <year>2001</year>
          )
          <fpage>31</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de-Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2023 -
          <article-title>Learning with Disagreement for Sexism Identification and Characterization</article-title>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Giachanou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Aliannejadi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Vlachos</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality</source>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          , Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de-Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigó</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of EXIST 2023 -
          <article-title>Learning with Disagreement for Sexism Identification and Characterization (Extended Overview)</article-title>
          , in: M.
          <string-name>
            <surname>Aliannejadi</surname>
            , G. Faggioli,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
          </string-name>
          , M. Vlachos (Eds.),
          <source>Working Notes of CLEF 2023 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaca-Serrano</surname>
          </string-name>
          ,
          <article-title>Detecting and classifying sexism by ensembling transformers models</article-title>
          , in: IberLEF@SEPLN,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Chiril</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Benamara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Moriceau</surname>
          </string-name>
          , “
          <article-title>be nice to your wife! the restaurants are closed”: Can gender stereotype detection improve sexism classification?, in: Findings of the Association for Computational Linguistics</article-title>
          : EMNLP,
          <year>2021</year>
          , pp.
          <fpage>2833</fpage>
          -
          <lpage>2844</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mamidi</surname>
          </string-name>
          ,
          <article-title>When does a compliment become sexist? Analysis and classification of ambivalent sexism using Twitter data</article-title>
          ,
          <source>in: Proceedings of the Second Workshop on NLP and Computational Social Science</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>7</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Glick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Fiske</surname>
          </string-name>
          ,
          <article-title>The ambivalent sexism inventory: Diferentiating hostile and benevolent sexism</article-title>
          ,
          <source>Journal of Personality and Social Psychology</source>
          (
          <year>1996</year>
          )
          <fpage>491</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>