<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the Evalita 2018 Task on Automatic Misogyny Identification (AMI)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elisabetta Fersini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Debora Nozza</string-name>
          <email>debora.nozzag@disco.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Rosso</string-name>
          <email>prosso@dsic.upv.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DISCo, Universita ́ degli Studi di Milano-Bicocca</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>PRHLT Research Center, Universitat Polite`cnica de Vale`ncia</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. Automatic Misogyny Identification (AMI) is a new shared task proposed for the first time at the Evalita 2018 evaluation campaign. The AMI challenge, based on both Italian and English tweets, is distinguished into two subtasks, i.e. Subtask A on misogyny identification and Subtask B about misogynistic behaviour categorization and target classification. Regarding the Italian language, we have received a total of 13 runs for Subtask A and 11 runs for Subtask B. Concerning the English language, we received 26 submissions for Subtask A and 23 runs for Subtask B. The participating systems have been distinguished according to the language, counting 6 teams for Italian and 10 teams for English. We present here an overview of the AMI shared task, the datasets, the evaluation methodology, the results obtained by the participants and a discussion of the methodology adopted by the teams. Finally, we draw some conclusions and discuss future work.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. Automatic Misogyny
Identification (AMI) e` un nuovo shared task
proposto per la prima volta nella campagna
di valutazione Evalita 2018. La sfida AMI,
basata su tweet italiani e inglesi, si
distingue in due sottotask ossia Subtask A
relativo al riconoscimento della misoginia e
Subtask B relativo alla categorizzazione di
espressioni misogine e alla classificazione
del soggetto target. Per quanto riguarda la
lingua italiana, sono stati ricevuti un
totale di 13 run per il Subtask A e 11 run
per il Subtask B. Per quanto riguarda la
lingua inglese, sono stati ricevuti 26 run
per il Subtask A e 23 per Subtask B. I
sistemi partecipanti sono stati distinti in
base alla lingua, raccogliendo un totale
di 6 team partecipanti per l’italiano e 10
team per l’inglese. Presentiamo di
seguito una sintesi dello shared task AMI,
i dataset, la metodologia di valutazione,
i risultati ottenuti dai partecipanti e una
discussione sulle metodologie adottate dai
diversi team. Infine, vengono discusse
conclusioni e delineati gli sviluppi futuri.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        During the last years, the phenomenon of hate
against women increased exponentially especially
in online environment such as microblogs
        <xref ref-type="bibr" rid="ref17 ref20">(Hewitt et al., 2016; Poland, 2016)</xref>
        . According to
the Pew Research Center Online Harassment
report (2017)
        <xref ref-type="bibr" rid="ref10">(Duggan, 2017)</xref>
        , we can highlight that
41% of people were personally targeted, whose
18% were subjected to serious kinds of
harassment because of the gender (8%) and that women
are more likely to be targeted than men (11% vs
5%). Misogyny, defined as the hate or prejudice
against women, can be linguistically manifested in
numerous ways, ranging from less aggressive
behaviours like social exclusion and discrimination
to more dangerous expressions related to threats
of violence and sexual objectification
        <xref ref-type="bibr" rid="ref12 ref14 ref2">(Anzovino
et al., 2018)</xref>
        . Given this relevant social problem,
the Automatic Misogyny Identification (AMI) task
has been proposed first at IberEval 2018
(Spanish and English)
        <xref ref-type="bibr" rid="ref12 ref14 ref2">(Fersini et al., 2018)</xref>
        and later at
Evalita 2018 (Italian and English)
        <xref ref-type="bibr" rid="ref14 ref9">(Caselli et al.,
2018)</xref>
        . The main goal of AMI is to distinguish
misogynous contents from non-misogynous ones,
to categorize misogynistic behaviours and finally
to classify the target of a tweet.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Task Description</title>
      <p>The AMI shared task is organized according to
two main subtasks:</p>
      <sec id="sec-3-1">
        <title>Subtask A - Misogyny Identification: a sys</title>
        <p>tem must discriminate misogynistic contents
from the non-misogynistic ones. Examples
of misogynous and non-misogynous tweets
are reported in Table 1.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Subtask B - Misogynistic Behaviour and</title>
      </sec>
      <sec id="sec-3-3">
        <title>Target Classification: a system must rec</title>
        <p>ognize the targets that can be either specific
users or groups of women together with the
identification of the type of misogyny against
women.</p>
        <p>Regarding the misogynistic behaviour, a tweet
must be classified as belonging to one of the
following categories:</p>
        <p>Stereotype &amp; Objectification: a widely held
but fixed and oversimplified image or idea of
a woman; description of women’s physical
appeal and/or comparisons to narrow
standards.</p>
        <p>Dominance: to assert the superiority of men
over women to highlight gender inequality.
Derailing: to justify woman abuse,
rejecting male responsibility; an attempt to disrupt
the conversation in order to redirect women’s
conversations on something more
comfortable for men.</p>
        <p>Sexual Harassment &amp; Threats of Violence: to
describe actions as sexual advances, requests
for sexual favours, harassment of a sexual
nature; intent to physically assert power over
women through threats of violence.</p>
        <p>Discredit: slurring over women with no other
larger intention.</p>
        <p>Examples of Misogynistic Behaviours are
reported in Table 2.</p>
        <p>Concerning the target classification, the main
goal is to classify each misogynous tweet as
belonging to one of the following two target
categories:</p>
        <p>Active (individual): the text includes
offensive messages purposely sent to a specific
target;
Passive (generic): it refers to messages
posted to many potential receivers (e.g.
groups of women).</p>
        <p>Examples of targets of misogynous tweets are
reported in Table 3.
3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Training and Testing Data</title>
      <p>In order to provide training and testing data both
for Italian and English, three approaches were
employed to collect misogynous text on Twitter:
Streaming download using a set of manually
defined representative keywords, e.g. bi**h,
w**re, c*nt for English and pu****a, tr**a,
f**a di legno for Italian;
Monitoring of potential victims’ accounts,
e.g. gamergate victims and public feminist
women;
Downloading the history of identified
misogynist, i.e. explicitly declared hate against
women on their Twitter profiles.</p>
      <p>Among all the collected tweets we selected a
subset of text querying the database with the
copresence of keywords, originating two corpora
initially composed of 10000 tweets for each
language. In order to label both the Italian and
English datasets, we involved a group of 6 experts
exploiting the CrowdFlower1 platform for internal
use. At the end of the labelling phase, we provided
one corpus for Italian and one corpus for English
to all the participants. The inter-rater annotator
agreement on the English dataset for the fields of
“misogynous”, “misogyny category” and “target”
1Now Figure Eight: https://figure-eight.com/
is 0.81, 0.45 and 0.49 respectively, while for the
Italian dataset is 0.96, 0.68 and 0.76. Each corpus
is distinguished in Training and Test datasets.
Regarding the training data, both the Italian and
English corpora are composed of 4000 tweets.
Concerning the test data, we provided 1000 tweets for
each language. The training data has been
provided as tab-separated, according to the following
fields:
id denotes a unique identifier of the tweet.
text represents the tweet text.
misogynous defines if the tweet is
misogynous or not misogynous; it takes values as 1
if the tweet is misogynous, 0 if the tweet is
not misogynous.
misogyny category denotes the type of
misogynistic behaviour; it takes value as:
– stereotype: denotes the category
“Stereotype &amp; Objectification”;
– dominance: denotes the category
“Dominance”;
– derailing: denotes the category
“Derailing”;
– sexual harassment: denotes the
category “Sexual Harassment &amp; Threats of
Violence”;
– discredit: denotes the category
“Discredit”;
– 0 if the tweet is not misogynous.
target denotes the subject of the misogynous
tweet; it takes value as:
– active: denotes a specific target
(individual);
– passive: denotes potential receivers
(generic);
– 0 if the tweet is not misogynous.</p>
      <p>Concerning the test data, only “id” and “text”
have been provided to the participants.
Examples of all possible allowed combinations are
reported below. Additionally to the field “id”, we
report all the combinations of labels to be predicted,
i.e. “misogynous”, “misogyny category” and
“target”:
0 0 0
1 stereotype active
1 stereotype passive
1 dominance active
1 dominance passive
1 derailing active
1 derailing passive
1 sexual harassment active
1 sexual harassment passive
1 discredit active
1 discredit passive</p>
      <p>The label distribution related to the Training and
Test datasets is reported in Table 4. While the
distribution of labels related to the field
“misogynous” is almost balanced (for both languages), the
classes related to the other fields are quite
unbalanced. Regarding the “misogyny category”, we
can distinguish between the two considered
languages. In particular, for the Italian language,
the most frequent label is related to the category
Stereotype &amp; Objectification, while for English
the most predominant one is Discredit.
Concerning the “target”, the most predominant victims are
specific users (active) with a strong imbalanced
distribution on the Italian corpus, while it is
almost balanced for the English training dataset and
strongly imbalanced on the (active) targets for the
corresponding test dataset.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Evaluation Measures and Baseline</title>
      <p>Considering the distribution of labels of the
dataset, we have chosen different evaluation
metrics. In particular, we distinguished as follows:
Subtask A. Systems have been evaluated on
the field “misogynous” using the standard
accuracy measure, and ranked accordingly.</p>
      <p>Subtask B. Each field to be predicted has
been evaluated independently on the other using
a Macro F1-score. In particular, the Macro
F1-score for the “misogyny category” field has
been computed as average of F1-scores obtained
for each category (stereotype, dominance,
derailing, sexual harassment, discredit), estimating
F1(misogyny category). Analogously, the
Macro F1-score for the “target” field has been
computed as average of F1-scores obtained for
each category (active, passive), F1(target).
The final ranking of the systems participating
to Subtask B was based on the Average Macro
F1-score (F1), computed as follows:</p>
      <p>F1 = F1(misogyny category)+F1(target)
2
(1)</p>
      <p>In order to compare the submitted runs with a
baseline model, we provided a benchmark
(AMIBASELINE) based on Support Vector Machine
trained on a unigram representation of tweets. In
particular, we created one training set for each
field to be predicted, i.e. “misogynous”,
“misogyny category” and “target”, where each tweet has
been represented as a bag-of-words (composed of
1000 terms) coupled with the corresponding label.
Once the representations have been obtained,
Support Vector Machines with linear kernel have been
trained, and provided as AMI-BASELINE.</p>
    </sec>
    <sec id="sec-6">
      <title>Participants and Results</title>
      <p>A total of 6 teams for Italian and 10 teams for
English from 10 different countries participated in at
least one of the two subtasks of AMI. Each team
had the chance to submit up to three runs for
English and three runs for Italian. Runs could be
constrained, where only the provided training data and
lexicons were admitted, and unconstrained, where
additional data for training were allowed. Table 5
shows an overview of the teams2 reporting their
affiliation, their country, the number of
submissions for each language and the subtasks they
addressed.</p>
      <sec id="sec-6-1">
        <title>5.1 Subtask A: Misogyny Identification</title>
        <p>Table 6 reports the results for the Misogyny
Identification task, which received 13 submissions for
Italian and 26 runs for English submitted
respectively from 6 and 10 teams. The highest Accuracy
has been achieved by bakarov at 0.844 for Italian
and by hateminers at 0.704 for English, both in
a constrained setting. Most of the systems have
shown an improvement with respect to the
AMIBASELINE. While the bakarov team submitted
only one run based on TF-IDF coupled with
Singular Value Decomposition and Boosting
classifier, hateminers achieved the highest performance
with a run based on vector representation that
concatenates sentence embedding, TF-IDF and
average word embeddings coupled with a Logistic
Regression model.</p>
      </sec>
      <sec id="sec-6-2">
        <title>5.2 Subtask B: Misogynistic Behaviour and</title>
      </sec>
      <sec id="sec-6-3">
        <title>Target Classification</title>
        <p>
          Table 7 reports the results for the Misogynistic
Behaviour and Target Classification task, which
received 11 submissions by 5 teams for Italian and
23 submissions by 9 teams for English. The
highest Average Macro F1-score has been achieved by
bakarov at 0.501 for Italian (even if the amended
run of CrotoneMilano achieved the highest
effective performance) and by himani at 0.406 for
English, both in a constrained setting. On the
contrary of the previous task, most of the systems have
shown lower performance compared to the
AMIBASELINE. It can be easily noted by looking at
the Average Macro F1-score of all the approaches,
that the problem of recognizing the misogyny
category and the target is more difficult than the
2The teams himani and resham described their systems in
the same report
          <xref ref-type="bibr" rid="ref1 ref14">(Ahluwalia et al., 2018)</xref>
          .
misogyny identification task.
        </p>
        <p>This is due to the fact that there can be a high
overlapping between textual expressions of
different misogyny categories, therefore it is highly
subjective for an annotator (and consequently for a
system) to select a category rather than another
one. Regarding the target classification, systems
can be easily misled by the presence of mentions
that are not the target of the misogynous content.</p>
        <p>While for the bakarov team the system for
Subtask B is the same one of Subtask A, himani
achieved the highest performance on the English
language with a run based on a Bag of N-Gram
representation coupled with an Ensemble of 5
models for classifying the Misogynistic Behaviour
and 2 models for Target Classification.
6</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Discussion</title>
      <p>The submitted systems can be compared by taking
into consideration the kind of input features that
they have considered for representing tweets and
the machine learning model that has been used as
classification model.</p>
      <p>Textual Feature Representation. The systems
submitted by the challenge participants’ consider
various techniques for representing the tweet
contents. Some teams have concentrated the effort on
considering a single type of representation, i.e. the
team ITT adopted the traditional TF-IDF
representation, while bakarov and RCLN proposed
systems considering only weighted n-grams at
character level for better dealing with misspellings and
capturing few stylistic aspects.</p>
      <p>
        Additionally to the traditional textual
feature representation techniques (i.e. bag of
words/characters, n-grams of words/characters
eventually weighted with TF-IDF) several teams
proposed specific lexical features for improving
the input space and consequently the classification
performances. The team of CrotoneMilano
experimented feature abstraction following the
bleaching approach proposed by Goot et al.
        <xref ref-type="bibr" rid="ref14 ref16">(Goot et al.,
2018)</xref>
        for modelling gender through the language.
Specific lexicons for dealing with hate speech
language have been included as features in the
systems of SB, resham and 14-exlab. In particular,
resham and 14-exlab made also use of
environmentspecific features, such as links, hashtags and
emojis, and task-specific features, such as swear word,
sexist slurs and women-related words.
      </p>
      <p>Differently from these approaches,
StopPropagHate and hateminers teams proposed
systems that consider the popular Embeddings
techniques both at word and sentence level.</p>
      <sec id="sec-7-1">
        <title>Machine Learning Models. Concerning the</title>
        <p>machine learning models, we can distinguish
between approaches that work with traditional
Support Vector Machines and Logistic
Regression, Ensemble Models and finally Deep
Learning methods. Following, we report the models
adopted by the systems that participated in the
AMI shared task, according to the type of the
machine learning model that has been adopted:
Support Vector Machines have been
exploited by 14-exlab by using both linear and
RBF kernel, by SB investigating only a radial
basis function kernel, and by CrotoneMilano
by adopting again a simple linear kernel;
Logistic Regression has been used by
bakarov and hateminers;
Ensemble Models have been adopted by three
teams according to different settings, i.e. ITT
and himani used a Simple Voting of different
classifiers, resham induced a Simple Voting
over different input features and RCLN used
an Ensemble based on Random Forest;
A Deep Learning classifier has been adopted
by only one team, i.e StopPropagHate that
trained a simple dense neural network.</p>
        <p>External Resources Several participants
exploited external resources for providing
taskspecific lexical features.</p>
        <p>
          The lexicons for addressing AMI for Italian
have been mostly obtained from lists available
online. The team SB used an available specific Italian
lexicon called “Le parole per ferire” built by Tullio
De Mauro3. Starting from this lexicon provided
by De Mauro, the HurtLex multilingual lexicon
has been created
          <xref ref-type="bibr" rid="ref14">(Bassignana et al., 2018)</xref>
          .
Beyond HurtLex, the team 14-exlab gathered a swear
word list from several sources4 including a
translated version of the noswearing dictionary5 and a
list of swear words from
          <xref ref-type="bibr" rid="ref8">(Capuano, 2007)</xref>
          .
        </p>
        <p>
          Regarding the English language, both resham
and 14-exlab used the list of swear words from
noswearing dictionary and the sexist slur list
provided by
          <xref ref-type="bibr" rid="ref11">(Fasoli et al., 2015)</xref>
          . The team
resham further investigated the sentiment
polarity retrieved from SentiWordNet
          <xref ref-type="bibr" rid="ref3">(Baccianella et
al., 2010)</xref>
          . Differently, the team SB exploited a
manually modeled lexicon for the misogyny
detection task proposed in
          <xref ref-type="bibr" rid="ref14 ref15">(Frenda et al., 2018a)</xref>
          .
The HurtLex lexicon has been used by the team
14-exlab also for the English task.
        </p>
        <p>Finally, pre-trained Word Embeddings have
3https://www.internazionale.it/
opinione/tullio-de-mauro/2016/09/27/
razzismo-parole-ferire</p>
        <p>4https://www.parolacce.org/2016/12/
20/dati-frequenza-turpiloquio/ and https:
//it.wikipedia.org/wiki/Turpiloquio_
nella_lingua_italiana
5https://www.noswearing.com/dictionary
Rank
We presented here a new shared task about
Automatic Misogyny Identification on Twitter for
Italian and English. By analysing the runs submitted
by the participants we can conclude that the
problem of misogyny identification has been
satisfactorily addressed by all the teams, while the
misogynistic behaviour and target classification still
remains a challenging problem. Concerning the
future work, several issues should be considered to
improve the quality of the collected data,
especially for capturing those less frequent
misogynistic behaviours such as Dominance and Derailing.
The problem of hate speech against women will
be further addressed in the HatEval shared task at
SemEval in English and Spanish tweets6.</p>
        <p>6SemEval 2019 Task 5: HatEval: Multilingual
Detection of Hate Speech Against Immigrants and Women
in Twitter https://competitions.codalab.org/
competitions/19935</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>The work of the third author was partially funded
by the SomEMBED TIN2015-71147-C2-1-P
research project (MINECO/FEDER). We thank
Maria Anzovino for her initial help in
collecting the tweets subsequently used for the labelling
phase and the final creation of the Italian and
English corpora used for the AMI shared task.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Resham</given-names>
            <surname>Ahluwalia</surname>
          </string-name>
          , Himani Soni, Edward Callow,
          <string-name>
            <surname>Anderson Nascimento</surname>
          </string-name>
          , and Martine De Cock.
          <year>2018</year>
          .
          <article-title>Detecting Hate Speech Against Women in English Tweets</article-title>
          .
          <source>In Proceedings of Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ), Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Maria</given-names>
            <surname>Anzovino</surname>
          </string-name>
          , Elisabetta Fersini, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Automatic Identification and Classification of Misogynistic Language on Twitter</article-title>
          .
          <source>In International Conference on Applications of Natural Language to Information Systems</source>
          , pages
          <fpage>57</fpage>
          -
          <lpage>64</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Baccianella</surname>
          </string-name>
          , Andrea Esuli, and
          <string-name>
            <given-names>Fabrizio</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Sentiwordnet 3.0: An enhanced lexical resource for sentiment analysis and opinion mining</article-title>
          .
          <source>In Proceedings of the International Conference on Language Resources and Evaluation.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Amir</given-names>
            <surname>Bakarov</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Vector Space Models for Automatic Misogyny Identification</article-title>
          .
          <source>In Proceedings of Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ), Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Pierpaolo</given-names>
            <surname>Basile</surname>
          </string-name>
          and
          <string-name>
            <given-names>Nicole</given-names>
            <surname>Novielli</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>UNIBA at EVALITA 2014-SENTIPOLC Task: Predicting tweet sentiment polarity combining micro-blogging, lexicon and semantic features</article-title>
          .
          <source>In Proceedings of Fourth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2014</year>
          ). CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Angelo</given-names>
            <surname>Basile</surname>
          </string-name>
          and
          <string-name>
            <given-names>Chiara</given-names>
            <surname>Rubagotti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Automatic Identification of Misogyny in English and Italian Tweets at EVALITA 2018 with a Multilingual Hate Lexicon</article-title>
          .
          <source>In Proceedings of Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ), Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Davide</given-names>
            <surname>Buscaldi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Tweetaneuse AMI EVALITA2018: Character-based Models for the Automatic Misogyny Identification Task</article-title>
          .
          <source>In Proceedings of Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ), Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>R.G.</given-names>
            <surname>Capuano</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Turpia: sociologia del turpiloquio e della bestemmia</article-title>
          .
          <source>Riscontri (Milan</source>
          , Italy).
          <source>Costa &amp; Nolan.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Tommaso</given-names>
            <surname>Caselli</surname>
          </string-name>
          , Nicole Novielli, Viviana Patti, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>EVALITA 2018: Overview of the 6th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          . In Tommaso Caselli, Nicole Novielli, Viviana Patti, and Paolo Rosso, editors,
          <source>Proceedings of Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ), Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Maeve</given-names>
            <surname>Duggan</surname>
          </string-name>
          .
          <year>2017</year>
          . Online Harassment. http://www.pewinternet.org/
          <year>2017</year>
          / 07/11/online-harassment-2017/. Last accessed 2018-
          <volume>10</volume>
          -28.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Fasoli</surname>
          </string-name>
          , Andrea Carnaghi, and Maria Paola Paladino.
          <year>2015</year>
          .
          <article-title>Social acceptability of sexist derogatory and sexist objectifying slurs across contexts</article-title>
          .
          <source>Language Sciences</source>
          ,
          <volume>52</volume>
          :
          <fpage>98</fpage>
          -
          <lpage>107</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Elisabetta</given-names>
            <surname>Fersini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Anzovino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and P</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of the task on automatic misogyny identification at ibereval</article-title>
          .
          <source>In Proceedings of the Third Workshop on Evaluation of Human Language Technologies for Iberian Languages (IberEval</source>
          <year>2018</year>
          ).
          <source>CEUR Workshop Proceedings. CEUR-WS. org, Seville</source>
          , Spain.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Paula</given-names>
            <surname>Fortuna</surname>
          </string-name>
          , Ilaria Bonavita, and Se´rgio Nunes.
          <year>2018</year>
          .
          <article-title>INESC TEC, Eurecat</article-title>
          and Porto University.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Simona</given-names>
            <surname>Frenda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ghanem</given-names>
            <surname>Bilal</surname>
          </string-name>
          , et al. 2018a.
          <article-title>Exploration of Misogyny in Spanish and English tweets</article-title>
          .
          <source>In Third Workshop on Evaluation of Human Language Technologies for Iberian Languages (IberEval</source>
          <year>2018</year>
          ), volume
          <volume>2150</volume>
          , pages
          <fpage>260</fpage>
          -
          <lpage>267</lpage>
          . Ceur Workshop Proceedings.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Simona</given-names>
            <surname>Frenda</surname>
          </string-name>
          , Bilal Ghanem,
          <article-title>Estefan´ıa Guzma´nFalco´n, Manuel Montes-y-Go´mez, and Luis Villasen˜or-Pineda. 2018b. Automatic Lexicons Expansion for Multilingual Misogyny Detection</article-title>
          .
          <source>In Proceedings of Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ), Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Rob</given-names>
            <surname>Goot</surname>
          </string-name>
          , Nikola Ljubesˇic´,
          <string-name>
            <surname>Ian</surname>
            <given-names>Matroos</given-names>
          </string-name>
          , Malvina Nissim, and
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Plank</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bleaching Text: Abstract Features for Cross-lingual Gender Prediction</article-title>
          .
          <source>In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          , volume
          <volume>2</volume>
          , pages
          <fpage>383</fpage>
          -
          <lpage>389</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Sarah</given-names>
            <surname>Hewitt</surname>
          </string-name>
          , Thanassis Tiropanis, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bokhove</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>The Problem of identifying Misogynist Language on Twitter (and other online social spaces)</article-title>
          .
          <source>In Proceedings of the 8th ACM Conference on Web Science</source>
          , pages
          <fpage>333</fpage>
          -
          <lpage>335</lpage>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Endang</given-names>
            <surname>Wahyu</surname>
          </string-name>
          <string-name>
            <surname>Pamungkas</surname>
          </string-name>
          , Alessandra Teresa Cignarella, Valerio Basile, and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Automatic Identification of Misogyny in English and Italian Tweets at EVALITA 2018 with a Multilingual Hate Lexicon</article-title>
          .
          <source>In Proceedings of Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ), Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey</surname>
            <given-names>Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Glove: Global vectors for word representation</article-title>
          .
          <source>In Proceedings of the 2014 conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Bailey</given-names>
            <surname>Poland</surname>
          </string-name>
          .
          <year>2016</year>
          . Haters: Harassment, Abuse, and
          <string-name>
            <given-names>Violence</given-names>
            <surname>Online</surname>
          </string-name>
          . Potomac Books, Incorporated.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Punyajoy</given-names>
            <surname>Saha</surname>
          </string-name>
          , Binny Mathew, Pawan Goyal, and
          <string-name>
            <given-names>Animesh</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Indian Institute of Engineering Science and Technology (Shibpur), Indian Institute of Technology (Kharagpur).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Elena</given-names>
            <surname>Shushkevich and John Cardiff</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Misogyny detection and classification in English tweets</article-title>
          .
          <source>In Proceedings of Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ), Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>