<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Sexism identification in Social Networks EXIST Task proposal 1</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Álvaro Faubel Sanchis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Clara Martí Torregrosa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Polytechnic University of Valencia</institution>
          ,
          <addr-line>Valencia 46022</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Twitter is a well-known microblogging social site where users express their views and opinions. Consequently, tweets tend to be a clear view of society. With the use of different techniques like pretrained models, this paper describes the process to try some approaches such as from feature extraction to deep learning models (BERT), in order to achieve a system that allows classifying tweets as sexist or non-sexist and disjoin them into levels. Finally, we will conclude with the best system, which is BERT pre-trained model.</p>
      </abstract>
      <kwd-group>
        <kwd>Sexism</kwd>
        <kwd>BERT</kwd>
        <kwd>Ensemble Methods</kwd>
        <kwd>IDF and Word Embeddings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Nowadays, the society is still sexist, and it can be seen on social media, that is a clear
reflection of opinions, knowledge, and culture population. Among them, Twitter stands
out as social platform is a bidirectional communication service, which is perfectly
structured to share from personal experiences to opinions on current news as fast as possible.
Nevertheless, these opinions could perturb other users, for example the case of sexist,
aggressive or toxic tweets. These comments could have serious consequences in real
life, and legal issues towards social platforms. As a result, a need for language models
specific to social media domain arises.</p>
      <p>
        To address these types of problems, it has been created a bunch of international
competitions. In this case, we participated in EXISTS 2021 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], the first shared task on
Sexism Identification in Social networks at IberLEF 2021 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        In this work, based on the main idea of previous studies, we understand that the best
techniques for this type of problem use: deep learning models [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], feature vectors and
word embeddings [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. We address three techniques, selecting and sending those that
achieve the best results.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Experimental Setting</title>
      <sec id="sec-2-1">
        <title>Data Collection</title>
        <p>Once we enrolled in the task, we received two sets of tweets and gabs messages, both
with text in Spanish and English. The first one refers to the train set with 6977 tweets
and others 3386 for testing in the second set.</p>
        <p>There are two tasks, the first one on sexism identification in a binary classification
(sexist or non-sexists). The following task aims to categorize the sexism into five
different types: ideological and inequality, stereotyping and dominance, objectification,
sexual violence, misogyny, and non-sexual violence.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Preprocessing</title>
        <p>The first step for any text analysis is the preprocessing of the text. Text in tweets is
characterized by its informal or colloquial language. In order to address it, initially we
started by the Python package tweet-preprocessor.</p>
        <p>Also, it was necessary to make an exhaustive cleaning, which it goes from deleting all
non-alphanumeric characters, convert all text to lowercase to delete stop-words, with
the help of the Regex library and Nltk package.</p>
        <p>Following, we performed a tokenization process and additionally for Spanish a text
stemming and for English tweets a lemmatization.</p>
        <p>This process was done for our approaches based on Words Embeddings and IDF
Matrix.</p>
        <p>In case of pre-trained BERT model, we only replace URLs, Mentions, Reserve words,
Emojis and Smileys with special tokens.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Text Representation &amp; Models</title>
        <p>We proposed different strategies in order to represent the tweets and then we used
different models in each representation.</p>
        <p>The first one uses Word Embeddings to represent text, with the use of the pre-trained
model GloVe Twitter 200 from the Python library Gensim. GloVe is an unsupervised
learning algorithm for obtaining weight vector for words. In this case, we used a weight
vector of size 200 for each term of our tweets. For each tweet we add the vectors of
their terms, dividing the sum by the number of words, and finally obtaining an average
weight vector. Here there was the possibility that terms in our tweets do not have their
weight vector in the pre-trained model, so first we fitted another model only with our
data. It is obvious that weights of this new model will not be as robust as the previous
ones. However, it is a better way than filling a zero vector.</p>
        <p>For the second strategy instead of using Word Embedding, we focus on IDF matrix
using the Python sklearn package for tweet representation. As a rule, there are two
metrics that evaluate the term's weight based on its occurrences, these are Term Frequency
(TF) and Inverse Document Frequency (IDF). While TF considers all term equally
important, IDF reduces weight based on the number of occurrences.</p>
        <p>( ) =  (

Deep neural network models have a significant value in this context, one example of
these models is BERT, which we use in our last approach.</p>
        <p>Bidirectional Encoder Representations from Transformers (BERT) is a technique
created in 2018 for researchers at Google. Since the release of this model, many users have
been using this type of neural networks in several tasks. The pre-trained models are
useful to solve one biggest challenge in NLP, that is the lack of enough training data.
Because starting with one of these pre-trained models, with a little fine-tuning, it could
help us in our task, we decided to use it.</p>
        <p>
          We choose BETO for Spanish tweets which is a BERT model trained on a big Spanish
corpus [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and Twitter-roBERTa-base for English text, which was trained on more than
58M tweets. It is necessary to mention that, based on the results presented in the
TweetEval benchmark paper [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], other models for English (i.e. BERTweet [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]).
To implement BERT in this project, it was necessary to use Pytorch and HuggingFace
Tranformers.
2.4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Results</title>
        <p>Task 1 has been organized with a typical binary classification, while task 2 is a
multilabel problem. In order to get better results in our second task we decided to use only
the tweets classified as sexist, leaving aside those not. It was done because non-sexist
tweets are irrelevant to determine the level of sexism.</p>
        <p>In both tasks, we have tried several models from sklearn and have validated them with
the test set by Cross Validation 5 K-fold. Initially we started with baseline models, then
with grid parameters and finally we combined these models with hybrid approximations
and ensemble methods, the obtained results shown in Table 1 and Table 2:
To sum up, for the feature extraction proposals, the best model when Word Embedding
is used Stacking Classifier with SVM and Logistic Regression, and with IDF matrix,
Bagging with Logistic Regression.</p>
        <p>In our last proposal, we used BERT’s pretrained tokenizers for each language. These
tokenizers basically add special tokens to the text and transform it into numerical
vectors that serve as input to the neural network (see Fig. 1).</p>
        <p>It is relevant to say that the max length selected for the Spanish language was 155 for
task 1 and 120 for task 2, and for English texts, 170 in both cases. Thus, the columns
of those tweets that are less than the maximum length will be filled to 0 in the creation
of the tokenized matrix. In addition, those texts of greater length will be truncated to
that size. This is called padding and truncation.</p>
        <p>After having all the tweets tokenized, we pass them to the pretrained neural network,
that will give us outputs on which we can extract the final predictions. Fig. 2 shows
how it works:</p>
        <p>In this approach we tried two classifiers. The first only considers the first vector for
each layer which is called CLS Vector (see Fig. 3).</p>
        <p>At first, we fit a linear layer with only one neuron per class with the first output vector
CLS, on which is applied a Softmax activation function.</p>
        <p>Nevertheless, another classifier returns us better results. In lieu of only selecting the
first array, this methodology considers all the outputs, creating an interconnected dense
layer to which the Relu activation function is applied. Then, another dense layer is
created with less input vectors and with as many neurons as classes, in which the activation
function applied is the LogSoftmax.</p>
        <p>Focusing on technical details, we have used the Cross Entropy as loss function, a
learning rate of 1 −5 and an epsilon of 1 −8. Finally, the dropout to avoid the overfitting
over forward step was 0.1.</p>
        <p>In addition, we choose 10 epochs and a batch size of 16 for the training loop. So, in
order to train and validate our model, we split the data in three sets, 90% for training,
5% for validation over the epoch’s iterations and 5% for testing the performance once
finished the training.</p>
        <p>Next tables show the BERT’s results in each task and language over the test subset.
BERT Spanish
BERT English
Since with BERT’s modelling we obtained better results, this was our main proposal in
the competition. However, we also sent the other two.</p>
        <sec id="sec-2-4-1">
          <title>Finally, we got 13th position in task 1 and 7th in task 2.</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions and Future work</title>
      <p>We proposed three different approaches for the Exits task on sexism detection in
Twitter. The first two models use feature extraction techniques: IDF and pre-trained word
embeddings. And the last based on BERT pre-trained model.</p>
      <p>The best results were obtained with the last model, since such a model has been trained
with a corpus large enough to handle one of the biggest issues in this field, which is the
lack of input data.</p>
      <p>However, the amount of data is never enough, so an improvement would be, to train
with more data and to improve model parameters. Furthermore, we could test the
combination of neural network models, as well as LSTM techniques.</p>
      <sec id="sec-3-1">
        <title>IberLEF proceedings [9].</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>F.</given-names>
            <surname>Rodriguez Sanchez</surname>
          </string-name>
          , J. Carillo de Albornoz, L. Plaza,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Comet</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Donoso</surname>
          </string-name>
          , “
          <article-title>Overview of EXIST 2021: sEXism Identification in Social neTworks</article-title>
          ,
          <source>” Procesamiento del Lenguaje Natural</source>
          , vol.
          <volume>67</volume>
          , (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. “IberLEF 2021:
          <article-title>Iberian Languages Evaluation Forum</article-title>
          .”, https://sites.google.com/view/iberlef2021/, (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>J. Carrillo de Albornoz</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Plaza</surname>
            and
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Rodrígez</surname>
            <given-names>Sanchéz</given-names>
          </string-name>
          , “
          <article-title>Automatic Classification of Sexism in Social,” NLP</article-title>
          &amp; IR Group, UNED, December
          <volume>17</volume>
          , (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>S.</given-names>
            <surname>Frenda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ghanem</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Montes y Gómez and</article-title>
          P. Rosso, “Special Section:
          <article-title>Intelligent and Fuzzy Systems applied to Language &amp; Knowledge Engineering</article-title>
          .”, vol.
          <volume>34</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>2959</fpage>
          -
          <lpage>2969</lpage>
          , (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>P.</given-names>
            <surname>Chiril</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Moriceau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Benamara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Origgi and M.</surname>
          </string-name>
          Coulomb-Gully, “
          <article-title>He said “who's gonna take care of your children when you are at ACL?”, Reported Sexist Acts are Not Sexist</article-title>
          .
          <source>Association for Computational Linguistic</source>
          , pp.
          <fpage>4055</fpage>
          -
          <lpage>4066</lpage>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>6. D. UChile, “BETO: Spanish BERT.”, https://github.com/dccuchile/beto.</mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>J.Camacho</given-names>
            <surname>Collados</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Neves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.Espinosa</given-names>
            <surname>Anke</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Barbereri</surname>
          </string-name>
          , “
          <article-title>TweetEval: Unified Benchmarck and Comparative Evaluation for Tweet Classification</article-title>
          .”, School of Computer Science and Informatics, Cardiff University, United Kingdom (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. V. Research, “
          <article-title>BERTweet: A pre-trained language model for English Tweets (</article-title>
          <year>EMNLP2020</year>
          ).”, https://github.com/VinAIResearch/BERTweet, last accessed 04/19/
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. IberLEF 2021 proceedings: Manuel Montes, Paolo Rosso, Julio Gonzalo, Ezra Aragón, Rodrigo Agerri,
          <string-name>
            <surname>Miguel Ángel</surname>
          </string-name>
          Álvarez-Carmona, Elena Álvarez Mellado, Jorge Carrillo-deAlbornoz, Luis Chiruzzo, Larissa Freitas, Helena Gómez Adorno, Yoan Gutiérrez, Salud María Jiménez Zafra, Salvador Lima,
          <string-name>
            <surname>Flor Miriam</surname>
          </string-name>
          Plaza-de-Arco and Mariona Taulé (eds.):
          <source>Proceedings of the Iberian Languages Evaluation Forum (IberLEF</source>
          <year>2021</year>
          ),
          <source>CEUR Workshop Proceedings</source>
          , 2021
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>