<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic Expansion of Lexicons for Multilingual Misogyny Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simona Frenda</string-name>
          <email>simona.frenda@unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bilal Ghanem Universitat Polite`cnica de Vale`ncia</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universita` degli Studi di Torino, Italy Universitat Polite`cnica de Vale`ncia</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. The automatic misogyny identification (AMI) task proposed at IberEval and EVALITA 2018 is an example of the active involvement of scientific Research to face up the online spread of hate contents against women. Considering the encouraging results obtained for Spanish and English in the precedent edition of AMI, in the EVALITA framework we tested the robustness of a similar approach based on topic and stylistic information on a new collection of Italian and English tweets. Moreover, to deal with the dynamism of the language on social platforms, we also propose an approach based on automatically-enriched lexica. Despite resources like the lexica prove to be useful for a specific domain like misogyny, the analysis of the results reveals the limitations of the proposed approaches.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. Il task AMI circa
l’identificatione automatica della
misoginia proposto a IberEval e a EVALITA
2018 e` un chiaro esempio dell’attivo
coinvolgimento della Ricerca per
fronteggiare la diffusione online di contenuti
di odio contro le donne. Considerando i
promettenti risultati ottenuti per spagnolo
e inglese nella precedente edizione di
AMI, nel contesto di EVALITA abbiamo
testato la robustezza di un approccio
simile, basato su informationi stilistiche e di
dominio, su una nuova collezione di tweet
in inglese e in italiano. Tenendo conto
dei repentini cambiamenti del linguaggio
nei social network, proponiamo anche un
approccio basato su lessici
automaticamente estesi. Nonostante risorse come i
lessici risultano utili per domini specifici
come quello della misoginia, analizzando
i risultati emergono i limiti degli approcci
proposti.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        The anonymity and the interactivity, typical of
computer-mediated communication, facilitate the
spread of hate messages and the perpetuated
presence of hate contents online. As investigated by
Fox et al. (2015), these factors increase and
influence social misbehaviors also offline. In order
to foster scientific research to find optimal
solutions that could help to monitor the spread of hate
speech contents, different tasks have been
proposed in various campaigns of evaluation. An
example is the AMI shared task proposed at IberEval
20181 and later at EVALITA 20182. This task
focuses on the automatic identification of misogyny
in different languages. In particular, the first
edition focuses on Spanish and English languages,
and the second one on a new English corpus and
Italian language. The multilingual context
allows to observe the analogies and differences
between different languages. The AMI’s organizers
        <xref ref-type="bibr" rid="ref1 ref1 ref10 ref10 ref9 ref9">(Fersini et al., 2018a; Fersini et al., 2018b)</xref>
        asked
participants to detect firstly misogynistic tweets
and then classify the misogynistic categories and
the kind of target (individuals or groups). In the
first edition, we proposed an approach based on
stylistic and topic information captured
respectively by means of character n-grams and a set of
modeled lexica
        <xref ref-type="bibr" rid="ref13">(Frenda et al., 2018)</xref>
        . Considering
the encouraging results obtained with the
lexiconbased approach in Spanish and English languages,
we re-proposed a similar approach for Italian
language and a new collection of English tweets in
1http://amiibereval2018.wordpress.com/
2http://amievalita2018.wordpress.com/
order to test the performance and robustness of
this approach. Actually, in this paper we
propose two approaches. The first one, similar to
previous work
        <xref ref-type="bibr" rid="ref13">(Frenda et al., 2018)</xref>
        , involves topic,
linguistic and stylistic information. The second
one focuses mainly on the automatic extension of
the original lexica. Indeed, to deal with the
continuous variation of the language on social
platforms, the modeled lexica are enriched
considering the contextual similarity of lexica by the use
of pre-trained word embeddings. This technique
helps the system to consider also new terms
relative to the topic information of the original
lexica. It could be considered as a good methodology
to upgrade automatically the existing list of words
used to block offensive contents in real
applications of Internet companies. Indeed, a
comparison between the two approaches reveals that the
automatic enrichment of the lexica improves the
results especially for English language. However,
comparing the results obtained in both
competitions and observing the error analyses, we notice
that lexica represent a good resource for a specific
domain like misogyny, but they are not sufficient
to detect misogyny online.
      </p>
      <p>Following, Section 2 describes the studies that
inspired our work. Section 3 explains the
approaches employed in both languages. Section 4
discusses the obtained results and delineates some
conclusions.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>
        A first work about misogyny detection is
proposed in Anzovino et al. (2018). In this study, the
authors compared the performance of different
supervised approaches using word embeddings,
stylistic and syntactic features. In particular,
their results reveal that the best machine learning
approach for identification of misogyny is the
linear Support Vector Machine (SVM) classifier.
In general machine learning techniques are the
most used in hate speech detection
        <xref ref-type="bibr" rid="ref17 ref8">(Escalante
et al., 2017; Nobata et al., 2016)</xref>
        , because they
allow researchers for exploring closely the issue
exploiting different features, such as textual
        <xref ref-type="bibr" rid="ref6">(Chen
et al., 2012)</xref>
        and syntactical aspects
        <xref ref-type="bibr" rid="ref18 ref3 ref5">(Burnap and
Williams, 2014)</xref>
        or semantic and sentiment
information
        <xref ref-type="bibr" rid="ref14 ref17 ref20">(Samghabadi et al., 2017; Nobata et
al., 2016; Gitari et al., 2015)</xref>
        . Finally, some recent
works have investigated also the potential of
deep learning techniques
        <xref ref-type="bibr" rid="ref16 ref17 ref7">(Mehdad and Tetreault,
2016; Del Vigna et al., 2017)</xref>
        . Considering
the specific domain concerning the hate against
women, this work exploits stylistic, linguistic and
topic information about the misogynistic speech.
In particular, differently from previous studies,
we use specific lexica relative to offensiveness
and discredit of women for English and Italian
languages, and we extend them with new words
relative to the issues of the considered lexica.
Considering the fact that commercial methods
rely currently on the use of blacklists to
monitor or block offensive contents, the proposed
approach could help to upgrade their blacklists
automatizing the process of the lexicon building.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Proposed Approaches</title>
      <p>The AMI shared task proposed at EVALITA 2018
aims to detect misogyny in English and Italian
collections of tweets. The organizers asked
participants to detect misogynistic texts (Task A),
and then, if the tweet is predicted as
misogynistic, to distinguish the nature of target (individuals
or groups labeled respectively “active” and
“passive”), and identify the type of misogyny (Task
B), according to the following classes proposed
by Poland (2016): (a) stereotype and
objectification, (b) dominance, (c) derailing, (d) sexual
harassment and threats of violence, and (e) discredit.</p>
      <p>Actually, these classes represent the different
manifestations and the various aspects of this social
misbehavior. Table 1 shows the composition of
the datasets.</p>
      <p>Considering the promising results obtained at
the IberEval campaign, in this work we use two
approaches mainly based on lexica. The first one
(Section 3.1) is similar to the approach used in
Frenda et al. (2018), based on topic, linguistic and
stylistic information captured by means of
modeled lexica and n-grams of characters and words.</p>
      <p>
        The second one (Section 3.2) principally involves
the automatically extended versions of the
original lexica
        <xref ref-type="bibr" rid="ref15">(Guzma´n Falco´n, 2018)</xref>
        . In particular,
we aim: 1) to test the robustness of lexicon based
approaches in the new collections of tweets and in
a new language, and 2) to understand the impact of
automatically enriched lexica to face up the
variation of the language in the multilingual
computermediated communication.
(a)
(b) (c)
      </p>
      <p>Misogynistic</p>
      <p>(d)
The first proposed approach aims to capture topic,
linguistic and stylistic information by means of
manually-modeled lexica and n-grams of words
and characters. Below the features description for
each language.</p>
      <p>English Features. For the detection of
misogyny in English tweets, we employed the
manuallymodeled lexica proposed in Frenda et al. (2018).</p>
      <p>These lexica concerns sexuality, profanity,
femininity and human body as described in Table 2.</p>
      <p>These lexica contain also slang expressions.</p>
      <p>Moreover, we take into account hashtags and
abbreviations collected in Frenda et al. (2018): 40
misogynistic hashtags, such as: #ihatef emales
or #bitchesstink; and a list of 50 negative
abbreviations, such as wtf or stf u. Considering
the most relevant n-grams of words, we employ
the bigrams for the first task and the
combination of unigrams, bigrams and trigrams (hence
defined as UBT) for the second task. Moreover,
the bag of characters (BoC) in a range from 1
to 7 grams is employed to manage misspellings
and to capture stylistic aspects of digital
writing. In order to perform the experiments, each
tweet is represented as a vector. The presence
of words in each lexicon is pondered with
Information Gain, and character and word n-grams
are weighted with Term Frequency-Inverse
Document Frequency (TF-IDF) measure. In
addition, considering the fact that in Frenda et al.
(2018) several misclassified misogynistic tweets
were ironic or sarcastic, we try to analyze the
impact of irony in misogyny detection in English.</p>
      <p>Indeed, Ford and Boxer (2011) reveal that
sexist jokes that in general are considered innocent,
truthfully they are experienced by women as
sexual harassment. In particular, inspired by Barbieri
and Saggion (2014), we calculate the imbalance of
the sentiment polarities (positive and negative) in
each tweet using SentiWordNet provided by
Baccianella et al. (2010). For each degree of
imbalance, we associate a weight used in the vectorial
representation of the tweets. Despite our
hypothesis is well funded, we obtained lower results for
the runs that contain sentiment imbalance among
the features (see Table 4).</p>
      <p>Italian Features. For the Italian language, we
selected some specific issue groups, described in
Bassignana et al. (2018), from the Italian
lexicon “Le parole per ferire” provided by Tullio De
Mauro3. In particular, we consider the lists of
words described in Table 3. Differently from
English, the experiments reveal that: the UBT is
useful for both tasks and the best range for BoC is
from 3 to 5 grams4. Indeed, in a morphological
complex language like Italian the desinences of
the words (such as the extracted n-grams “tona” or
“ana ”) contain relevant linguistic information.
Diversely, in English, longer sequences of characters
could help to capture multi-word expressions
containing also pronouns, adjectives or prepositions,
such as “ing at” or “ss bitc”.</p>
      <p>To extract the features correctly, in order to
train our models, we pre-process the data
deleting emoticons, emojis and URLs. Indeed, from
our experiments, the emoticons and emojis do not
prove to be relevant for these tasks. In order to
perform a correct match between the dictionaries of
the corpora and the single lexicon, we use the
lemmatizer provided by the Natural Language Toolkit
(NLTK5) for English, and the Snowball Stemmer
for Italian. Differently from English, the use of
lemmatizer for Italian tweets hinders the match.</p>
      <p>3http://www.internazionale.it/
opinione/tullio-de-mauro/2016/09/27/
razzismo-parole-ferire
4The experiments are carried out using the Grid Search.
5http://www.nltk.org/
Profanity
Femininity
Human body</p>
      <p>Words
290</p>
      <p>Definition
contains words relative to sexual subject (orgasm, orgy, pussy) and especially male domination on
women (rape, pimp, slave)
is a collection of vulgar words such as motherf ucker, slut and scum
is a list of terms used to identify the women as target. It contains personal pronouns or possessive
adjectives (such as she, her, herself ), common words used to refer to women (girl, mother) and
also offensive words towards women (such as barbie, hooker or non male)
is a lexicon strongly connected with sexuality collecting words referred especially to feminine body
also with negative connotations (such as holes, throat or boobs)</p>
      <p>Definition
collects words relative to animals, such as sanguisuga or pecora
contains terms referred to female genitalia, such as f essa
contains terms referred to male genitalia, such as verga
is a list of derogatory words, such as bastardo or spazzatura
contains words derived from plants but that are used as offensive words, such as f inocchio or rapa
is a list of professions or jobs that have also a negative connotations, such as portinaia or impiegato
contains terms about prostitution, such as bagascia or zoccolona
is a list of words relative to stereotypes, such as negro or ostrogoto
collects words that have in general negative connotations, such as parassita or dilettante
contains terms relative to criminal acts or immoral actions, such as stupro or violento
The second approach aims to deal with the
dynamism of the informal language online trying to
capture new words relative to contexts defined in
each lexicon. Therefore, we use enriched versions
of the original lexica (described above), and
stylistic and linguistic information captured by means
of n-grams of words and characters as in the first
approach. The method for the expansion of a
given lexicon shares the idea of identifying new
words by considering their contextual similarity
with known words, as defined by some pre-trained
word embeddings. For its description, let assume
that L = fl1; : : : ; lmg is the initial lexicon of m
words, and W = f(w1; e(w1)); : : : ; (wn; e(wn))g
is the set of pre-trained word embeddings, where
each pair represents a word and its corresponding
embedding vector. This method aims to enrich the
lexicon with words strongly related to the context
from the original lexicon without being
necessarily associated to any particular word. Its idea is
to search for words having similar contexts to the
entire lexicon. This method has two main steps,
described below.</p>
      <p>Dictionary modeling. Firstly, we extract the
embedding e(li) for each word li 2 L; then, we
compute the average of these vectors to obtain a vector
describing the entire lexicon, e(L). We name this
vector the context embedding.</p>
      <sec id="sec-4-1">
        <title>Dictionary expansion. Using the cosine simi</title>
        <p>larity, we compare e(L) against the embedding
e(wi) of each wi 2 W ; then, we extract the
k most similar words to e(L), defining the set
EL = (w1; : : : ; wk). Finally, we insert the
extracted words into the original lexicon to build the
new lexicon, i.e., LE = L [ EL.</p>
        <p>
          Therefore, we carry out the experiments using
different pre-trained word embeddings for each
language: GloVe embeddings trained on 2
billion tweets
          <xref ref-type="bibr" rid="ref18">(Pennington et al., 2014)</xref>
          for English,
and word embeddings built on TWITA corpus6 for
Italian
          <xref ref-type="bibr" rid="ref18 ref3 ref5">(Basile and Novielli, 2014)</xref>
          . Finally, the
proposed expansion method is parametric and
requires a value for k, the number of words that are
going to extend the lexica. In particular, we use
k = 1000, 500 and 100.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>3.3 Experiments and Results</title>
        <p>To carry out the experiments, a SVM classifier
is employed with the radial basis function kernel
(RBF) using the following parameters: C = 5 and
= 0:1 for English and = 0:01 for Italian.
Considering the complexity of the target classification
for the Italian language due to imbalanced training
set (see Table 1), we used a Random Forest (RF)
classifier that aggregates the votes from different
6http://valeriobasile.github.io/twita/
about.html
decision trees to decide the final class of the tweet.</p>
        <p>The evaluation is performed using the test set
provided by the organizers of the AMI shared task.
For the competition, they use as evaluation
measures the Accuracy for Task A and the average of
F-score of both classes for Task B.</p>
        <p>Table 4 and Table 5 show the results obtained
in the competition compared with the baselines
provided by the organizers for each task.
Comparing the two approaches, in general AEL seems
to work better than MML. However, the
improvement of the results is very slight, especially for
Italian language. This soft variation is unexpected
considered the results obtained during the
experiments employing 10-fold cross validations. In
fact, AEL with enriched lexica using k equal 100
performed an Accuracy of 0.880. Moreover,
looking at Table 4, reporting the official results of the
AMI Task, only run 2 overcomes the baseline for
the detection of misogyny in English, and for this
run we used AEL approach excluding the
sentiment imbalance as feature. About the
identification of misogyny in Italian, the obtained results are
lower than provided baselines as well as the values
of F-score obtained in Task B for both languages
(see Table 5). Despite the usefulness of lexica for
a specific domain like misogyny, a lexicon-based
approach proves to be insufficient for this task.
Indeed, as the error analysis will confirm, misogyny,
as well as general hate speech, involves linguistic
devices such as humour, exclamations typical of
orality and contextual information that completes
the meaning transmitted by the tweet. Moreover,
the low values obtained also in Task B suggest
the necessity to implement dedicated approach for
each misogynistic category.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion and Conclusions</title>
      <p>This paper reports our participation in the AMI
shared task. The organizers provide also the gold
test set that helps us to understand better what are
the misclassified cases and the aspects that should
be considered in the next experiments.
Carrying out the error analysis, we notice that in both
datasets the content of URL affects the
transmitted information in the tweet (such as Right! As
they rape and butcher women and children !!!!!!
https://t.co/maEhwuYQ8B). The swear words are
often used also as exclamation without the aim to
offend (such as Volevo dire alla Yamamay che
tettona non sinonimo di curvy dato che di vita ha una
40, quindi confidence sta minchia.). Moreover,
despite the actual English corpus does not contain
several jokes, Italian misclassified tweets involve
humourous utterances (such as
@GrianneOhmsfor1 @BarbaraRaval A parte il fatto poi che
culona inchiavabile” e` il miglior giudizio politico
sentito sulla Merkel negli ultimi anni??”). In fact,
in general, humour, irony and sarcasm hinder the
correct classification of the texts, as we noticed
in English and Spanish corpora provided in the
IberEval framework. Participating in this shared
task gave us the opportunity to analyze and
compare multilingual datasets, and thus, to discover
and infer general aspects typical of hate speech
against women.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The work of Simona Frenda was partially funded
by the Spanish research project SomEMBED
TIN2015-71147-C2-1-P (MINECO/FEDER). We
also thank the support of CONACYT-Mexico
(projects FC-2410, CB-2015-01-257383).
7This run does not involve the sentiment imbalance
8This run involves the expansions of lexica with k = 100
Francesco Barbieri and Horacio Saggion. 2014.
Modelling irony in twitter. In Proceedings of the
StuEnglish
Run
baseline AMI
run 2
run 1
run 3
Italian
Run
baseline AMI
run 3
run 1
run 2
F-score
0.534
0.485
0.483
0.480</p>
      <p>Target
UBT+BoC
UBT+BoC
UBT+BoC
Target
UBT+BoC
UBT+BoC
UBT+BoC
F-score
0.440
0.414
0.414
0.411
Total
0.487
0.449
0.448
0.446
dent Research Workshop at the 14th Conference of
the European Chapter of the ACL.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Maria</given-names>
            <surname>Anzovino</surname>
          </string-name>
          , Elisabetta Fersini, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Automatic identification and classification of misogynistic language on twitter</article-title>
          .
          <source>In International Conference on Applications of Natural Language to Information Systems</source>
          , pages
          <fpage>57</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Baccianella</surname>
          </string-name>
          , Andrea Esuli, and
          <string-name>
            <given-names>Fabrizio</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Sentiwordnet 3.0: an enhanced lexical resource for sentiment analysis and opinion mining</article-title>
          .
          <source>In Lrec</source>
          , volume
          <volume>10</volume>
          , pages
          <fpage>2200</fpage>
          -
          <lpage>2204</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Pierpaolo</given-names>
            <surname>Basile</surname>
          </string-name>
          and
          <string-name>
            <given-names>Nicole</given-names>
            <surname>Novielli</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Uniba at evalita 2014-sentipolc task: Predicting tweet sentiment polarity combining micro-blogging, lexicon and semantic features</article-title>
          .
          <source>In Proceedings of EVALITA</source>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Bassignana</surname>
          </string-name>
          , Valerio Basile, and
          <string-name>
            <given-names>Patti</given-names>
            <surname>Viviana</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hurtlex: A multilingual lexicon of words to hurt</article-title>
          .
          <source>In Proceedings of CLiC-it, Turin</source>
          ,
          <fpage>10</fpage>
          -12
          <source>December</source>
          <year>2018</year>
          , CEUR.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Peter</given-names>
            <surname>Burnap and Matthew Leighton Williams</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Hate speech, machine classification and statistical modelling of information flows on twitter: Interpretation and communication for policy decision making</article-title>
          . Internet, Policy &amp; Politics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Ying</given-names>
            <surname>Chen</surname>
          </string-name>
          , Yilu Zhou, Sencun Zhu, and
          <string-name>
            <given-names>Heng</given-names>
            <surname>Xu</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Detecting offensive language in social media to protect adolescent online safety</article-title>
          .
          <source>In Privacy, Security, Risk and Trust (PASSAT)</source>
          , pages
          <fpage>71</fpage>
          -
          <lpage>80</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Fabio Del Vigna</surname>
            ,
            <given-names>Andrea</given-names>
          </string-name>
          <string-name>
            <surname>Cimino</surname>
            , Felice Dell'Orletta,
            <given-names>Marinella</given-names>
          </string-name>
          <string-name>
            <surname>Petrocchi</surname>
            , and
            <given-names>Maurizio</given-names>
          </string-name>
          <string-name>
            <surname>Tesconi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Hate me, hate me not: Hate speech detection on facebook</article-title>
          .
          <source>In Proceedings of ITASEC17.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Hugo</given-names>
            <surname>Jair</surname>
          </string-name>
          <string-name>
            <surname>Escalante</surname>
          </string-name>
          , Esau´ Villatoro-Tello, Sara E Garza,
          <string-name>
            <given-names>A Pastor</given-names>
            <surname>Lo</surname>
          </string-name>
          <article-title>´pez-</article-title>
          <string-name>
            <surname>Monroy</surname>
          </string-name>
          ,
          <article-title>Manuel Montes-y Go´mez, and Luis Villasen˜or-</article-title>
          <string-name>
            <surname>Pineda</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Early detection of deception and aggressiveness using profile-based representations</article-title>
          .
          <source>Expert Systems with Applications</source>
          ,
          <volume>89</volume>
          :
          <fpage>99</fpage>
          -
          <lpage>111</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Elisabetta</given-names>
            <surname>Fersini</surname>
          </string-name>
          , Maria Anzovino, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          . 2018a.
          <article-title>Overview of the task on automatic misogyny identification at ibereval</article-title>
          .
          <source>In Proceedings of Workshop IBEREVAL at 3rd SEPLN.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Elisabetta</given-names>
            <surname>Fersini</surname>
          </string-name>
          , Debora Nozza, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          . 2018b.
          <article-title>Overview of the evalita 2018 task on automatic misogyny identification (ami)</article-title>
          .
          <source>In Tommaso Caselli</source>
          , Nicole Novielli, Viviana Patti, and Paolo Rosso, editors,
          <source>Proceedings of the 6th evaluation campaign of Natural Language Processing and Speech tools for Italian (EVALITA'18)</source>
          , Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Thomas E</given-names>
            <surname>Ford and Christie Fitzgerald Boxer</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Sexist humor in the workplace: A case of subtle harassment</article-title>
          .
          <source>In Insidious Workplace Behavior</source>
          , pages
          <fpage>203</fpage>
          -
          <lpage>234</lpage>
          . Routledge.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Jesse</given-names>
            <surname>Fox</surname>
          </string-name>
          , Carlos Cruz, and Ji Young Lee.
          <year>2015</year>
          .
          <article-title>Perpetuating online sexism offline: Anonymity, interactivity, and the effects of sexist hashtags on social media</article-title>
          .
          <source>Computers in Human Behavior</source>
          ,
          <volume>52</volume>
          :
          <fpage>436</fpage>
          -
          <lpage>442</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Simona</given-names>
            <surname>Frenda</surname>
          </string-name>
          , Bilal Ghanem, and
          <string-name>
            <surname>Manuel</surname>
          </string-name>
          Montes-y Go´mez.
          <year>2018</year>
          .
          <article-title>Exploration of misogyny in spanish and english tweets</article-title>
          .
          <source>In Proceedings of Workshop IBEREVAL at 3rd SEPLN.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Njagi</given-names>
            <surname>Dennis Gitari</surname>
          </string-name>
          , Zhang Zuping, Hanyurwimfura Damien, and
          <string-name>
            <given-names>Jun</given-names>
            <surname>Long</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>A lexicon-based approach for hate speech detection</article-title>
          .
          <source>International Journal of Multimedia and Ubiquitous Engineering</source>
          ,
          <volume>10</volume>
          (
          <issue>4</issue>
          ):
          <fpage>215</fpage>
          -
          <lpage>230</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <article-title>Estefan´ıa Guzma´n Falco´n. 2018. Deteccio´n de lenguaje ofensivo en Twitter basada en expansio´n automa´tica de lexicones (tesis de maestr´ıa)</article-title>
          . Instituto Nacional de Astrof´ısica, O´ ptica y Electro´nica. Puebla, Me´xico.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Mehdad</surname>
          </string-name>
          and
          <string-name>
            <given-names>Joel</given-names>
            <surname>Tetreault</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Do characters abuse more than words</article-title>
          ?
          <source>In Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue</source>
          , pages
          <fpage>299</fpage>
          -
          <lpage>303</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Chikashi</given-names>
            <surname>Nobata</surname>
          </string-name>
          , Joel Tetreault, Achint Thomas,
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Mehdad</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Yi</given-names>
            <surname>Chang</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Abusive language detection in online user content</article-title>
          .
          <source>In Proceedings of the 25th international conference on WWW.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey</surname>
            <given-names>Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Glove: Global vectors for word representation</article-title>
          .
          <source>In Proceedings of EMNLP.</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Bailey</given-names>
            <surname>Poland</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Haters: Harassment, abuse, and violence online</article-title>
          . U of Nebraska Press.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Niloofar</given-names>
            <surname>Safi</surname>
          </string-name>
          <string-name>
            <surname>Samghabadi</surname>
          </string-name>
          , Suraj Maharjan, Alan Sprague, Raquel Diaz-Sprague, and
          <string-name>
            <given-names>Thamar</given-names>
            <surname>Solorio</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Detecting nastiness in social media</article-title>
          .
          <source>In Proceedings of ALW1.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>