<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>No Place For Hate Speech @ AMI: Convolutional Neural Network and Word Embedding for the Identification of Misogyny in Italian</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Adriano dos S.R. da Silva</string-name>
          <email>adriano.santos.silva@usp.br</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Norton T. Roman</string-name>
          <email>norton@usp.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Schoool of Arts, Sciences and Humanities, University of Sao Paulo</institution>
          ,
          <addr-line>Sao Paulo -</addr-line>
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Schoool of Arts, Sciences and, Humanities - University of Sao Paulo</institution>
          ,
          <addr-line>Sao Paulo -</addr-line>
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. In this article, we describe two classification models (a Convolutional Neural Network and a Logistic Regression classifier), arranged according to three different strategies, submitted to subtask A of Automatic Misogyny Identification at EVALITA 2020. Results were very encouraging for detecting misogyny, even though aggressiveness was less accurate. Our second strategy, consisting of a Convolutional Neural Network and logistic regression to identify misogyny and aggressiveness, respectively, won the sixth place in the competition.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. In questo articolo,
descriviamo due modelli di classificazione (i.e.,
Convolutional Neural Network e
Regressione Logistica), organizzati secondo tre
diverse strategie, per il subtask A dello
shared task Automatic Misogyny
Identification a EVALITA 2020. I risultati sono
stati molto incoraggianti nel rilevamento
della misoginia, anche se l’aggressivita`
viene riconosciuta con una precisione piu`
basse. La nostra seconda strategia
(Convolutional Neural Network per misoginia
e Regressione Logistica per aggressivita`)
ci ha permesso di ottenere il sesto posto
nella competizione.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>Hate speech is a problem that has been gaining
space both in the media and in academic research.
Political organizations have been working to
combat this type of discourse. As is the case with the</p>
      <p>Copyright © 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
code of conduct1 created by the European Union
Commission, and signed by some of the main
social networks, such as Facebook, YouTube,
Twitter, which aims to monitor and remove this type of
content within 24 hours of its disclosure.</p>
      <p>The subject has even become a marketing
problem, to the extent that recently several
companies stopped advertising on Facebook2, only to put
some pressure at the network to have it remove
this type of publication from the posts within it.
Advertisers point, in this case, is that they do not
want their brand to be linked to this type of speech.</p>
      <p>
        Defined as “language which attacks or demeans
a group based on race, ethnic origin, religion,
gender, age, disability, or sexual orientation/gender
identity“
        <xref ref-type="bibr" rid="ref12">(Nobata et al., 2016)</xref>
        , hate speech
represents a problem that cannot be allowed to grow,
under the risk of having it lead to more concrete
actions, by some people, with truly undesired
results.
      </p>
      <p>
        When this hate speech is targeted specifically
at women, it is called misogyny
        <xref ref-type="bibr" rid="ref11">(Manne, 2017)</xref>
        .
The problem with misogyny is such an issue that
it has already been related to real crime cases and
cybercrimes
        <xref ref-type="bibr" rid="ref9">(Fulper et al., 2014)</xref>
        . In this case,
correlations were found between rape cases and
the amount of misogynous tweets per state in the
United States.
      </p>
      <p>
        Some academic work and several competitions
have proposed some tasks to promote studies and
advances in the area. Much of this work and
data sets focus on English
        <xref ref-type="bibr" rid="ref1 ref15 ref6 ref7 ref8">(Fortuna and Nunes,
2018)</xref>
        only, even though this is a widespread
phenomenon that happens in any language.
      </p>
      <p>It is extremely important, therefore, to
encourage the development of this kind of study
1https://ec.europa.eu/info/policies/justice-andfundamental-rights/combatting-discrimination/racismand-xenophobia/eu-code-conduct-countering-illegal-hatespeech-online en</p>
      <p>
        2https://www.nytimes.com/2020/08/01/
business/media/facebook-boycott.html
in different languages and competitions, such as
IberEval
        <xref ref-type="bibr" rid="ref1 ref6 ref7">(Fersini et al., 2018b)</xref>
        , SemEval
        <xref ref-type="bibr" rid="ref2">(Basile
et al., 2019)</xref>
        and EVALITA
        <xref ref-type="bibr" rid="ref1 ref6 ref7">(Fersini et al., 2018a)</xref>
        ,
which have already proposed activities to identify
misogynous discourse in Spanish, English, and
Italian.
      </p>
      <p>In this work, we help address this problem
by testing two classification models as part of
EVALITA 2020’s subtask A on Automatic
Misogyny Identification (AMI). Tested models were a
Convolutional Neural Network (CNN) and a
Logistic Regression (LR) classifier. Three different
strategies were designed and tested, with one of
them scoring 6th in the competition.</p>
      <p>The rest of this article is organized as follows.
Section 2 presents some related work in the
identification of misogyny or hate speech. Section 3,
in turn, gives an overview of EVALITA’s AMI.
Next, in section 4, we describe our experimental
set-up, giving details of the implemented methods
and tested strategies. Finally, in Section 5 we
discuss our results, whereas in Section 6 we present
our final remarks on this task.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>
        IberEval
        <xref ref-type="bibr" rid="ref1 ref6 ref7">(Fersini et al., 2018b)</xref>
        proposed a task
to identify misogynous discourse in tweets in
English and Spanish. Several teams participated in
this competition and the best team reached an
accuracy of 0.91 and 0.81 for Spanish and English,
respectively, with the use of an SVM as a
classifier and with the addition of some lexical features
to characterize the tweets.
      </p>
      <p>
        SVMs were also proposed to identify racism
in Twitter messages in English, achieving an F1
score of 0.76
        <xref ref-type="bibr" rid="ref10">(Hasanuzzaman et al., 2017)</xref>
        . In
SemEval 2019, a Convolutional Neural Network
(CNN) performed competitively in the task of
identifying hate speech against immigrants and
women in English
        <xref ref-type="bibr" rid="ref2">(Basile et al., 2019)</xref>
        . The team
that presented this architecture ranked fourth with
an F1 score of 0.535.
      </p>
      <p>
        During the Automatic Misogyny Identification
shared task at EVALITA 2018, it was proposed a
subtask A, which consisted of identifying
misogyny
        <xref ref-type="bibr" rid="ref1 ref1 ref6 ref7 ref7">(Fersini et al., 2018a; Anzovino et al., 2018)</xref>
        .
For this subtask, Logistic Regression was the
model to deliver the best performance with an
accuracy of 0.704
        <xref ref-type="bibr" rid="ref15">(Saha et al., 2018)</xref>
        .
      </p>
    </sec>
    <sec id="sec-4">
      <title>Subtask</title>
      <p>
        The second edition of misogyny identification
at EVALITA 2020 consists of two subtasks: A
and B. The purpose of subtask A is to
identify the presence or absence of misogyny and
aggressiveness in tweets
        <xref ref-type="bibr" rid="ref5">(Elisabetta Fersini, 2020)</xref>
        ,
whereas subtask B checks whether the model is
capable of distinguishing misogynous from
nonmisogynous content, also ensuring fairness
(unintended bias)
        <xref ref-type="bibr" rid="ref13 ref2">(Nozza et al., 2019)</xref>
        .
      </p>
      <p>
        The ”No Place For Hate Speech” team
participated only in subtask A, and all discussions that
will be followed are related to this subtask. Within
EVALITA 2020, the subtask consisted of
identifying the presence or absence of misogynous speech
and aggressiveness in tweets in Italian
        <xref ref-type="bibr" rid="ref3 ref5">(Basile et
al., 2020; Elisabetta Fersini, 2020)</xref>
        .
      </p>
      <p>
        The training dataset consisted of 5,000 tweets.
The class that determines the presence or absence
of misogyny is nearly balanced. However,
aggressiveness is not balanced at all, with approximately
35% of tweets containing aggressiveness. Table 1
shows the distribution of each class in the training
set.
In subtask A, we tested two different
classifiers within different configurations. These were
a Convolutional Neural Network (CNN), using
BERT
        <xref ref-type="bibr" rid="ref4">(Devlin et al., 2018)</xref>
        as its language model;
and a Logistic Regression (LR) classifier, with L2
regularisation.
      </p>
      <p>
        The LR classifier used a 4-gram language
model, with tf-idf
        <xref ref-type="bibr" rid="ref14">(Rajaraman and Ullman, 2011)</xref>
        normalization. Both models were developed in
Python, with the aid of the TensorFlow3 and
Sklearn4 libraries.
      </p>
      <p>Since the subtask A at EVALITA allows each
team to submit up to three classifiers, we decided
to approach the problem according to three
different strategies, involving different combinations
of these classifiers, along with different subsets of
data on which they should be trained.</p>
      <p>3https://www.tensorflow.org/
4https://scikit-learn.org/stable/
In all cases, the training set was divided in a
90% subset, used for training purposes, with the
remaining 10% used for out-of-sample testing. All
classifiers used this same proportion both to
identify misogyny and aggressiveness. Tweets were
used in their raw form and no preprocessing was
used.</p>
      <p>All CNNs used in the experiments had the same
configuration, being trained for 15 epochs. They
also have three convolution layers, relu activation
functions, and dropout rate of 0.10, with adam
optimisation. Finally, cross-entropy was used as their
loss function. In what follows, we will describe,
with more details, each of the strategies followed
during our tests.</p>
      <sec id="sec-4-1">
        <title>4.1 Strategy 1</title>
        <p>The first strategy consisted of training two CNNs,
one for each specific sub-problem separately, i.e.
one for misogyny and another for aggressiveness
classification. In both cases, the entire data set was
used for training.</p>
        <p>At the testing stage, the CNNs were arranged as
a pipeline, in which the first CNN was responsible
for identifying whether a tweet had some
misogynous content, whereas the second CNN was
responsible for identifying the presence or absence
of aggressiveness only in those tweets marked as
misogynous by the first CNN.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2 Strategy 2</title>
        <p>Similar to Strategy 1, the second strategy also
consisted of training a CNN to detect misogynous
content in tweets. This time, however, the
classification of aggressiveness was left to a Linear
Regression classifier. As in the first strategy, both
models were trained in the entire data set.</p>
        <p>During testing, once again models were
arranged in a pipeline, with the CNN coming first,
to detect misogyny in tweets. In the sequence, all
tweets classified as misogynous by the CNN were
then fed to the LR classifier, so it could determine
the presence or absence of aggressiveness.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3 Strategy 3</title>
        <p>Our third strategy is similar to Strategy 1, in that it
also consists of two CNNs trained separately over
the data set. The only difference, however, lies
during the training stage. In this case, whereas the
first CNN (i.e. the one responsible for misogyny
identification) was trained using the entire data set,
the second CNN (the one responsible for detecting
aggressiveness) was trained only on those
examples labeled as misogynous.</p>
        <p>During testing the same set-up as in Strategy 1
was followed. As such, both CNNs were arranged
in a pipeline, with the first one responsible for
detecting misogynous tweets, and the second one
responsible for identifying aggressiveness, amongst
those tweets held misogynous by the first CNN.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results and Discussion</title>
      <p>Results for aggressiveness detection, on the
other hand, varied substantially, with the
Logistic Regression classifier (Strategy 2) performing
worst, when compared to the CNNs used for the
same task in the other strategies (7% against
Strategy 1, and 18% against Strategy 3).</p>
      <p>Interestingly, the CNN trained only on
examples labeled as misogynous (Strategy 3) performed
better (around 13%) than its counterpart trained
over the entire data set (Strategy 1). It is important
to recall that this was the only difference between
both strategies.</p>
      <p>Final results at the competition’s private test set
can be seen in Table 3. As it turns out, Strategy
2 was the best ranked of our models, reaching the
sixth place at the competition (being only F =
0:03 worse than the winning model).
Puzzling enough, this was the model that scored
worse in our test set. One possible
explanation for this fact might be that our CNN was not
capable of generalising over different data sets.
Differences in the balance between misogynous
and non-misogynous, and between aggressive and
non-aggressive examples, in both data sets, might
also explain this behaviour. Whatever the reason,
we leave this investigation for future work.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this work, we described two models
submitted to EVALITA 2020’s subtask A on Automatic
Misogyny Identification. To this task, a CNN and
an LR classifier were trained, and arranged as a
pipeline following three different strategies, with
one of them coming at sixth place in the
competition.</p>
      <p>Even though our classifier turned out to be
competitive, we believe improvements could be made
to achieve better results, such as the addition of
lexical features, for example. Also, it might be
that following some preprocessing strategy, such
as removing stop words, for example, might result
in a better performance.</p>
      <p>As for future work, besides testing the above
cited changes, it would be interesting investigating
why the worst model at the test set (as distributed
to all participants) turned out to be the best model
at the competition’s private data set. The reasons
for this behaviour are something to be determined.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Maria</given-names>
            <surname>Anzovino</surname>
          </string-name>
          , Elisabetta Fersini, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Automatic identification and classification of misogynistic language on twitter</article-title>
          .
          <source>In International Conference on Applications of Natural Language to Information Systems</source>
          , pages
          <fpage>57</fpage>
          -
          <lpage>64</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Cristina Bosco, Elisabetta Fersini, Nozza Debora, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso,
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , et al.
          <year>2019</year>
          .
          <article-title>Semeval-2019 task 5: Multilingual detection of hate speech against immigrants and women in twitter</article-title>
          .
          <source>In 13th International Workshop on Semantic Evaluation</source>
          , pages
          <fpage>54</fpage>
          -
          <lpage>63</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Evalita 2020: Overview of the 7th evaluation campaign of natural language processing and speech tools for italian</article-title>
          .
          <source>In Valerio Basile</source>
          , Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso Elisabetta Fersini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Debora</given-names>
            <surname>Nozza</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Ami @ evalita2020: Automatic misogyny identification</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA 2020), Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Elisabetta</given-names>
            <surname>Fersini</surname>
          </string-name>
          , Debora Nozza, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          . 2018a.
          <article-title>Overview of the evalita 2018 task on automatic misogyny identification (ami)</article-title>
          .
          <source>EVALITA Evaluation of NLP and Speech Tools for Italian</source>
          ,
          <volume>12</volume>
          :
          <fpage>59</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Elisabetta</given-names>
            <surname>Fersini</surname>
          </string-name>
          , Paolo Rosso, and
          <string-name>
            <given-names>Maria</given-names>
            <surname>Anzovino</surname>
          </string-name>
          . 2018b.
          <article-title>Overview of the task on automatic misogyny identification at ibereval 2018</article-title>
          . IberEval@ SEPLN,
          <volume>2150</volume>
          :
          <fpage>214</fpage>
          -
          <lpage>228</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Paula</given-names>
            <surname>Fortuna</surname>
          </string-name>
          and Se´rgio Nunes.
          <year>2018</year>
          .
          <article-title>A survey on automatic detection of hate speech in text</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>51</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Rachael</given-names>
            <surname>Fulper</surname>
          </string-name>
          , Giovanni Luca Ciampaglia, Emilio Ferrara,
          <string-name>
            <surname>Y Ahn</surname>
          </string-name>
          , Alessandro Flammini, Filippo Menczer,
          <string-name>
            <given-names>Bryce</given-names>
            <surname>Lewis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Kehontas</given-names>
            <surname>Rowe</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Misogynistic language on twitter and sexual violence</article-title>
          .
          <source>In Proc. ACM Web Science Workshop on Computational Approaches</source>
          to Social Modeling (ChASM).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Mohammed</given-names>
            <surname>Hasanuzzaman</surname>
          </string-name>
          , Gae¨l Dias, and
          <string-name>
            <given-names>Andy</given-names>
            <surname>Way</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Demographic word embeddings for racism detection on twitter</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Kate</given-names>
            <surname>Manne</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Down girl: The logic of misogyny</article-title>
          . Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Chikashi</given-names>
            <surname>Nobata</surname>
          </string-name>
          , Joel Tetreault, Achint Thomas,
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Mehdad</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Yi</given-names>
            <surname>Chang</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Abusive language detection in online user content</article-title>
          .
          <source>In Proceedings of the 25th international conference on world wide web.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Debora</given-names>
            <surname>Nozza</surname>
          </string-name>
          , Claudia Volpetti, and
          <string-name>
            <given-names>Elisabetta</given-names>
            <surname>Fersini</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Unintended bias in misogyny detection</article-title>
          .
          <source>In IEEE/WIC/ACM International Conference on Web Intelligence</source>
          , pages
          <fpage>149</fpage>
          -
          <lpage>155</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Anand</given-names>
            <surname>Rajaraman</surname>
          </string-name>
          and Jeffrey David Ullman.
          <year>2011</year>
          .
          <article-title>Mining of massive datasets</article-title>
          . Cambridge.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Punyajoy</given-names>
            <surname>Saha</surname>
          </string-name>
          , Binny Mathew, Pawan Goyal, and
          <string-name>
            <given-names>Animesh</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hateminers : Detecting hate speech against women</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1812</year>
          .06700.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>