<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Merging datasets for hate speech classification in Italian</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>(1) INESC TEC and (3) FEUP, University of Porto Rua Dr. Roberto Frias</institution>
          ,
          <addr-line>s/n 4200-465 Porto</addr-line>
          <country country="PT">PORTUGAL</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents an approach to the shared task HaSpeeDe within Evalita 2018. We followed a standard machine learning procedure with training, validation, and testing phases. We considered word embedding as features and deep learning for classification. We tested the effect of merging two datasets in the classification of messages from Facebook and Twitter. We concluded that using data for training and testing from the same social network was a requirement to achieve a good performance. Moreover, adding data from a different social network allowed to improve the results, indicating that more generalized models can be an advantage.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>ll manoscritto presenta un approccio per
la risoluzione dello shared task HaSpeeDe
organizzato all’interno di Evalita 2018.
La classificazione e` stata condotta con
caratteristiche del testo estratte con word
embedding e utilizzando algoritmi di deep
learning. Abbiamo voluto sperimentare
l’effetto dell’integrazione di messaggi di
Facebook e Twitter ha e abbiamo ottenuto
due risultati. 1) Addestrare modelli con un
dataset integrato migliora le performance
di classificazione in datasets provenienti
dai singoli social network suggerendo una
migliore capacita` di generalizzazione del
modello. 2) Tuttavia, utilizzare modelli
addestrati su datasets provenienti da un
social network per classificare messaggi
provenienti da un altro social network
comporta un peggioramento delle
performance indicando che e` indispensabile
includere nel train set messaggi dello stesso
social network che si e` interessati a
classificare nel test set.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        In the last few years, there is a growing attention to
the automatic detection of hate speech in text. This
appears as an answer to the increased spreading of
online abuse in social networks. Several
evaluation initiatives have been presenting different yet
related classification tasks, e.g. TRAC
        <xref ref-type="bibr" rid="ref15">(Kumar et
al., 2018)</xref>
        . Shared initiatives such as this, have
the advantage of promoting the development of
different but comparable solutions for the same
problem, within a short period of time. In this
paper, we describe the participation of the “Stop
PropagHate” team in the HaSpeeDe task within
Evalita 2018
        <xref ref-type="bibr" rid="ref23 ref3">(Bosco et al., 2018)</xref>
        .
      </p>
      <p>The goal of this task is to improve the
automatic classification of hate speech in Italian. More
specifically, there were three sub-tasks,
promoting the development of features that would work
independently of social network. For the task
HaSpeeDe-FB, only the Facebook dataset could
be used to train the model and classify Facebook
data; for HaSpeeDe-TW, only the Twitter dataset
could be used to classify Twitter data; and for the
Cross-HaSpeeDe, only the Facebook dataset could
be used to classify the Twitter and vice versa.</p>
      <p>In our approach, we focused on understanding
the effects of merging the two provided datasets.
As features, we used word embeddings and deep
learning for classification with a simple dense
neural network. In this paper, we present the details of
our approach, our results and conclusions.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>
        Previous research in the field of automatic
detection of hate speech can give us insight into how
to approach this problem. Two surveys
summarize previous research and conclude that the
approaches rely frequently on Machine Learning and
classification
        <xref ref-type="bibr" rid="ref10 ref11 ref13 ref14 ref19 ref2 ref20 ref24 ref27 ref6 ref7">(Schmidt and Wiegand, 2017;
Fortuna and Nunes, 2018)</xref>
        .
      </p>
      <p>
        Regarding the automatic classification of
messages, one first step is the gathering of training
data. Several studies published datasets
considering hate speech with different classification
systems
        <xref ref-type="bibr" rid="ref14 ref17 ref18 ref22 ref25 ref26 ref5 ref7">(Ross et al., 2017; Waseem and Hovy, 2016;
Davidson et al., 2017; Nobata et al., 2016; Jigsaw,
2018)</xref>
        . Although these could be useful datasets,
the annotated language is not Italian. Regarding
this language, the two existent datasets are used
in this task
        <xref ref-type="bibr" rid="ref21 ref23 ref8">(Del Vigna et al., 2017; Poletto et al.,
2017; Sanguinetti et al., 2018)</xref>
        .
      </p>
      <p>
        After data collection, one of the most important
steps when using classification is the process of
feature extraction
        <xref ref-type="bibr" rid="ref13 ref19 ref2 ref24 ref7">(Schmidt and Wiegand, 2017)</xref>
        .
Different methods are used, for instance word and
character n-grams
        <xref ref-type="bibr" rid="ref16 ref4">(Liu and Forss, 2014)</xref>
        ,
perpetrator characteristics
        <xref ref-type="bibr" rid="ref17 ref25 ref26 ref5">(Waseem and Hovy, 2016)</xref>
        ,
othering language
        <xref ref-type="bibr" rid="ref17 ref25 ref26 ref5">(Burnap and Williams, 2016)</xref>
        or
word embedings
        <xref ref-type="bibr" rid="ref9">(Djuric et al., 2015)</xref>
        . Regarding
the classification algorithms, the more common
are, for instance, SVM
        <xref ref-type="bibr" rid="ref8">(Del Vigna et al., 2017)</xref>
        or Random forests
        <xref ref-type="bibr" rid="ref16 ref4">(Burnap and Williams, 2014)</xref>
        .
Another popular approach, due to its good results,
is deep learning
        <xref ref-type="bibr" rid="ref13 ref13 ref19 ref19 ref2 ref2 ref24 ref24 ref26 ref7 ref7">(Yuan et al., 2016; Gamba¨ck and
Sikdar, 2017; Park and Fung, 2017)</xref>
        .
      </p>
      <p>
        Different studies proved that deep learning
algorithms outperform previous approaches. This
was the case when using character or
tokenbased n-grams with Recurrent Neural Network
Language Model (RNN)
        <xref ref-type="bibr" rid="ref17 ref18 ref25 ref26 ref5">(Mehdad and Tetreault,
2016)</xref>
        ; user behavioral characteristics with neural
network composed of multiple
Long-Short-TermMemory (LSTM)
        <xref ref-type="bibr" rid="ref13 ref19 ref2 ref24 ref7">(Park and Fung, 2017)</xref>
        ;
Convolutional Neural Networks (CNN), LSTM and
FastText
        <xref ref-type="bibr" rid="ref2">(Badjatiya et al., 2017)</xref>
        ; morpho-syntactical
features, sentiment polarity and word embedding
lexicons with LSTM
        <xref ref-type="bibr" rid="ref8">(Del Vigna et al., 2017)</xref>
        ;
users’ tendency towards racism or sexism with
RNN
        <xref ref-type="bibr" rid="ref20">(Pitsilis et al., 2018)</xref>
        ; abusive behavioral
norms, available metadata, patterns within the text
with RNN
        <xref ref-type="bibr" rid="ref12">(Founta et al., 2018)</xref>
        ; n-grams,
tfidf, POS, sentiment, misspellings, emojis, special
punctuation, capitalization, hashtags with CNN
and GRU
        <xref ref-type="bibr" rid="ref27">(Zhang et al., 2018)</xref>
        ; and word2vec with
Convolutional Neural Networks (CNN)
        <xref ref-type="bibr" rid="ref13 ref19 ref2 ref24 ref7">(Gamba¨ck
and Sikdar, 2017)</xref>
        .
      </p>
      <p>
        In this work, we propose an innovative
approach in hate speech detection by merging
different datasets, in the sequence of a previous
experiment
        <xref ref-type="bibr" rid="ref10 ref11">(Fortuna et al., 2018)</xref>
        . We merged two
datasets for aggression classification and the
results showed that, although training with similar
data is an advantage, adding data from different
platforms allowed slightly better results.
      </p>
      <p>Regarding the specificities of our approach in
this contest, the main research question of our
work concerns the effects of merging new datasets
on the performance of models for hate speech
classification. Accordingly with the previous study,
we hypothesize that merging datasets will lead to
a better performance. Additionally, we want to
investigate how models perform when only data
from different sources was used in the training.</p>
      <p>In the next sections, we present our
methodology and approach to this problem.
3
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Methodology</title>
      <sec id="sec-4-1">
        <title>Data</title>
        <p>
          The data proposed for this task results of
joining a collection of Facebook comments from
2016
          <xref ref-type="bibr" rid="ref8">(Del Vigna et al., 2017)</xref>
          with a Twitter
corpus developed in 2018
          <xref ref-type="bibr" rid="ref21 ref23">(Poletto et al., 2017;
Sanguinetti et al., 2018)</xref>
          . Both consist of a total
amount of 4,000 comments/tweets, randomly split
into development and test set, of 3,000 and 1,000
messages respectively. The data format is the
same with three tab-separated columns, each one
representing the ID of the message, the text and
the class (1 if the text contains hate speech, and 0
otherwise).
3.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Text pre-processing</title>
        <p>As a first step, we load the messages, remove the
retweet marker “RT” in case of the tweets, and also
the URL links present in the text.
3.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Feature extraction and classification</title>
        <p>
          We follow a methodology of classification with
training, testing and validation. Keeping 30% of
the data for validation allows us to estimate the
results we would achieve in the contest. We use
word embeddings and deep learning as presented
in previous literature
          <xref ref-type="bibr" rid="ref1 ref10 ref14 ref20 ref27 ref6">(Chollet and Allaire, 2018)</xref>
          .
We use the keras R package
          <xref ref-type="bibr" rid="ref1 ref6">(Allaire et al., 2018)</xref>
          and make our approach available in a public
repository1.
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>3.3.1 Word embedding</title>
        <p>In the procedure of feature extraction, we
vectorize the text. We start by tokenizing the
1https://github.com/StopPropagHate/
experiment_evalita_HaSpeeDe
data, considering only the top 10,000 words
in the dataset. Additionally, we consider only
the first 100 words of the tweets. We use
the functions text tokenizer, fit text tokenizer,
texts to sequences and pad sequences in our
extraction.</p>
      </sec>
      <sec id="sec-4-5">
        <title>3.3.2 Deep Learning</title>
        <p>For the classification we use 10 fold
crossvalidation and apply a simple dense neural
network. We use binary crossentropy for loss with
the rmsprop optimizer, we define the custom
metric F1, so that it would be in according to the
contest used metric. Regarding the model, we
instantiate an empty model and we customize it:
First we add an embedding layer where we
specify the input length (100, the maximum
length of the messages) and give the
dimensionality of the input data (dimensional space
of 10,000). We add a dropout of 0.25.</p>
        <p>We flatten the output and add a dense layer,
specified with 256 unit, with “relu” as a
parameter. We add a dropout of 0.25.</p>
        <p>We add a dense layer with just a single
neuron to serve as the output layer. Aiming for a
single output, we use a sigmoid activation.</p>
        <p>We use keras compile function to compile and
fit the model. We use batch size 128 and we tune
the number of epochs starting by using 10. We
also feed the model with the classes weights,
corresponding to the frequencies of the classes in the
training set. We average the F1 and loss results of
the 10 folds for each epoch. For the epoch number,
we kept the maximum number before overfitting
to happen (the results only improving in the
training set, but not in the test set). We save the final
model and apply it to the validation data, with the
function keras predict. We conduct a permutation
test in order to have a p-value associated to the F1.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Tasks and runs description</title>
      <p>We conduct three different experiments following
the procedure described in Section 3.</p>
      <p>Task HaSpeeDe-FB In the HaSpeeDe-FB run1,
we train and test with Facebook data. In the
HaSpeeDe-FB run2, we mix Facebook with the
Twitter provided data and see the effect in
predicting hate speech in Facebook.</p>
      <p>Task HaSpeeDe-TW We follow a similar
procedure, but we switched the roles of Facebook
and Twitter data. For theHaSpeeDe-TW run1 only
Twitter data is used. In a second run
HaSpeeDeTW run2, we mix data for training and use Twitter
for testing.</p>
      <p>Task Cross-HaSpeeDe This is a proposed
outof-domain task. In the Cross-HaSpeeDe-FB, only
the Facebook dataset can be used to classify
Twitter data. In the Cross-HaSpeeDe-TW, only the
Twitter dataset is used to classify Facebook data.
5</p>
    </sec>
    <sec id="sec-6">
      <title>Results and Discussion</title>
      <p>We separated the conditions with testing data from
Facebook from Twitter and we compared three
different conditions: the training data is from the
same social network (1), the training data is both
from and not from the social network (2), and the
training data is not from the social network (3).
5.1</p>
      <sec id="sec-6-1">
        <title>Results for Tuning and Validation</title>
        <p>For each of the runs in our experiment we tuned
the epoch parameter and we analyzed the average
of the 10 folds for each of the 10 epochs (Figure 1).</p>
        <p>The decided number of epochs for each run is
presented in the Table 1. We concluded that using
mixed data for training (Condition 2) has a better
performance (F1) than using data only from the
social network (Condition 1). Additionally, using
data only from other social network (Condition 3)
provided poor results. Finally, classifying
Facebook data was easier than Twitter data.</p>
        <p>C. system epoch
1 HaSpeeDe-FB run1 7
2 HaSpeeDe-FB run2 3
3 Cross-HaSpeeDe-TW 4
1 HaSpeeDe-TW run1 6
2 HaSpeeDe-TW run2 4
3 Cross-HaSpeeDe-FB 6</p>
        <p>F1
0.723
0.738
0.284
0.630
0.679
0.434
p-value
0.001
0.001
0.001
0.001
0.001
1
Regarding the contest results (Table 2), similarly
to the validation results we verified again that
using mixed data for training (Condition 2) is better.</p>
        <p>Also in this case we verified that using only data
from a different social network provided much
worse results (Condition 3). Opposing to the
validation results we found here that generally
classifying Facebook data was more difficult than
Twitter data.
1F0000....3468 ●●● ●●● ●●● ●●●● ●●●● ●●● ●● ●● ●● ●● da●●ta1Ftvra000al...ii446ndiantgion●● ● ●● ●● ● ● ● ● ● ● data 1F00..68 da●tatraining
0.3 ●● ●● ●● ● ● ● ● ● ● ● ●● tvraaliindia00nt..gi43on ●●●● ●●●● ●●●● ●●●● ●●●● ●● ●● ●● ●● ●● ● validation
lsso00..21 2 4 epoch 6● ● 8● ● 1●0 lsso000...210 2 4● ● epoch 6● ● 8● ● 1●0 lsso00..21 2 4 epoch 6● ● 8● ● 1●0
(d) HaSpeeDe-TW run1 (e) HaSpeeDe-TW run2 (f) Cross-HaSpeeDe-FB</p>
        <p>Regarding the main finding of this experiment,
the results show that in this contest adding new
data from a different social network brought
improved performance. However, in the scope of
this work it was not possible to investigate the
reasons for this. One possibility may be the increased
number of instances in the training when adding
new datasets. Also using data from a different
social network may bring less overfitting from
training with only a dataset.
6</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>Throughout our approach to this shared task, our
goal was to measure the effects of merging new
datasets on hate speech classification. Supported
by a previous experiment, we expected that adding
data would help the classification. Indeed, we
verified that merging datasets allowed us to have a
small improvement of the results.</p>
      <p>Complementary to this result, we tried the same
approach following the same method and idea, in
the Evalita 2018 AMI task. Merging datasets did
not help for misoginy classification. In this case,
we found that merging extra misogynistic or hate
speech data kept the mysoginy classification with
similar performance.</p>
      <p>The reason why merging datasets worked in one
case and not in the other remains unclear, and
requires exploration in future studies. Possible
variables interfering are the number of messages used
for training and also the number of distinct words
in the data.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This work was partially funded by the Google DNI
grant Stop PropagHate.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Allaire</surname>
          </string-name>
          , Francois Chollet, Yuan Tang, Daniel Falbel, Wouter Van Der Bijl, and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Studer</surname>
          </string-name>
          .
          <year>2018</year>
          . R interface to 'keras'.
          <source>Computer software manual](R package version 2.1.6)</source>
          . Retrieved from https://CRAN. R-project. org/package= keras.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Pinkesh</given-names>
            <surname>Badjatiya</surname>
          </string-name>
          , Shashank Gupta, Manish Gupta, and
          <string-name>
            <given-names>Vasudeva</given-names>
            <surname>Varma</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Deep learning for hate speech detection in tweets</article-title>
          .
          <source>In Proceedings of the 26th International Conference on World Wide Web Companion</source>
          , pages
          <fpage>759</fpage>
          -
          <lpage>760</lpage>
          . International World Wide Web Conferences Steering Committee.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <surname>Felice</surname>
            <given-names>DellOrletta</given-names>
          </string-name>
          , Fabio Poletto, Manuela Sanguinetti, and
          <string-name>
            <given-names>Maurizio</given-names>
            <surname>Tesconi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of the EVALITA 2018 Hate Speech Detection Task</article-title>
          . In Tommaso Caselli, Nicole Novielli, Viviana Patti, and Paolo Rosso, editors,
          <source>Proceedings of the 6th evaluation campaign of Natural Language Processing and Speech tools for Italian (EVALITA'18)</source>
          , Turin, Italy. CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Peter</given-names>
            <surname>Burnap and Matthew L. Williams</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Hate speech, machine classification and statistical modelling of information flows on Twitter: Interpretation and communication for policy decision making</article-title>
          .
          <source>In Proceedings of Internet, Policy &amp; Politics</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Pete</given-names>
            <surname>Burnap and Matthew L. Williams</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Us and them: identifying cyber hate on Twitter across multiple protected characteristics</article-title>
          .
          <source>EPJ Data Science</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ):
          <fpage>11</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Francois</given-names>
            <surname>Chollet</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Allaire</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Deep Learning with R</article-title>
          . Manning Publications Co.,
          <string-name>
            <surname>Greenwich</surname>
            ,
            <given-names>CT</given-names>
          </string-name>
          , USA, 1st edition.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Davidson</surname>
          </string-name>
          , Dana Warmsley,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Macy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ingmar</given-names>
            <surname>Weber</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Automated Hate Speech Detection and the Problem of Offensive Language</article-title>
          .
          <source>In Proceedings of ICWSM.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Fabio Del Vigna</surname>
            ,
            <given-names>Andrea</given-names>
          </string-name>
          <string-name>
            <surname>Cimino</surname>
            , Felice Dell'Orletta,
            <given-names>Marinella</given-names>
          </string-name>
          <string-name>
            <surname>Petrocchi</surname>
            , and
            <given-names>Maurizio</given-names>
          </string-name>
          <string-name>
            <surname>Tesconi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Hate me, hate me not: Hate speech detection on facebook</article-title>
          .
          <source>In Proceedings of the First Italian Conference on Cybersecurity</source>
          , pages
          <fpage>86</fpage>
          -
          <lpage>95</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Nemanja</given-names>
            <surname>Djuric</surname>
          </string-name>
          , Jing Zhou, Robin Morris, Mihajlo Grbovic, Vladan Radosavljevic, and
          <string-name>
            <given-names>Narayan</given-names>
            <surname>Bhamidipati</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Hate speech detection with comment embeddings</article-title>
          .
          <source>In Proceedings of the 24th International Conference on World Wide Web</source>
          , pages
          <fpage>29</fpage>
          -
          <lpage>30</lpage>
          .
          <year>ACM2</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Paula</given-names>
            <surname>Fortuna</surname>
          </string-name>
          and Se´rgio Nunes.
          <year>2018</year>
          .
          <article-title>A survey on automatic detection of hate speech in text</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>51</volume>
          (
          <issue>4</issue>
          ):
          <fpage>85</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Paula</given-names>
            <surname>Fortuna</surname>
          </string-name>
          , Jose´ Ferreira, Luiz Pires, Guilherme Routar, and Se´rgio Nunes.
          <year>2018</year>
          .
          <article-title>Merging datasets for aggressive text identification</article-title>
          .
          <source>In Proceedings of the First Workshop on Trolling, Aggression and Cyberbullying (TRAC-2018)</source>
          , pages
          <fpage>128</fpage>
          -
          <lpage>139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Antigoni-Maria</surname>
            <given-names>Founta</given-names>
          </string-name>
          , Despoina Chatzakou, Nicolas Kourtellis, Jeremy Blackburn, Athena Vakali, and
          <string-name>
            <given-names>Ilias</given-names>
            <surname>Leontiadis</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A unified deep learning architecture for abuse detection</article-title>
          . arXiv preprint arXiv:
          <year>1802</year>
          .00385.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <article-title>Bjo¨rn Gamba¨ck and Utpal Kumar Sikdar</article-title>
          .
          <year>2017</year>
          .
          <article-title>Using Convolutional Neural Networks to Classify Hatespeech</article-title>
          .
          <source>In Proceedings of the First Workshop on Abusive Language Online</source>
          , pages
          <fpage>85</fpage>
          -
          <lpage>90</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Jigsaw</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Toxic comment classification challenge identify and classify toxic online comments</article-title>
          . Available in https://www.kaggle.com/c/ jigsaw-toxic
          <article-title>-comment-classification-challenge, accessed last time in 23 May 2018</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Ritesh</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Atul</given-names>
            <surname>Kr</surname>
          </string-name>
          . Ojha, Shervin Malmasi, and
          <string-name>
            <given-names>Marcos</given-names>
            <surname>Zampieri</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Benchmarking Aggression Identification in Social Media</article-title>
          .
          <source>In Proceedings of the First Workshop on Trolling, Aggression and Cyberbulling (TRAC)</source>
          , Santa Fe, USA.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Shuhua</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Forss</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Combining n-gram based similarity analysis with sentiment analysis in web content classification</article-title>
          .
          <source>In International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management</source>
          , pages
          <fpage>530</fpage>
          -
          <lpage>537</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Mehdad</surname>
          </string-name>
          and
          <string-name>
            <given-names>Joel</given-names>
            <surname>Tetreault</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Do characters abuse more than words</article-title>
          ?
          <source>In Proceedings of the SIGdial 2016 Conference: The 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue</source>
          , pages
          <fpage>299</fpage>
          -
          <lpage>303</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Chikashi</given-names>
            <surname>Nobata</surname>
          </string-name>
          , Joel Tetreault, Achint Thomas,
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Mehdad</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Yi</given-names>
            <surname>Chang</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Abusive language detection in online user content</article-title>
          .
          <source>In Proceedings of the 25th International Conference on World Wide Web</source>
          , pages
          <fpage>145</fpage>
          -
          <lpage>153</lpage>
          . International World Wide Web Conferences Steering Committee.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Ji</given-names>
            <surname>Ho</surname>
          </string-name>
          Park and
          <string-name>
            <given-names>Pascale</given-names>
            <surname>Fung</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>One-step and Twostep Classification for Abusive Language Detection on Twitter</article-title>
          .
          <source>In Proceedings of the First Workshop on Abusive Language Online.</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Georgios K Pitsilis</surname>
            , Heri Ramampiaro, and
            <given-names>Helge</given-names>
          </string-name>
          <string-name>
            <surname>Langseth</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Detecting offensive language in tweets using deep learning</article-title>
          .
          <source>arXiv preprint arXiv:1801</source>
          .04433.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Poletto</surname>
          </string-name>
          , Marco Stranisci, Manuela Sanguinetti, Viviana Patti, and
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Hate speech annotation: Analysis of an italian Twitter corpus</article-title>
          .
          <source>In CEUR WORKSHOP PROCEEDINGS</source>
          , volume
          <year>2006</year>
          , pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . CEUR-WS.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Bjorn</given-names>
            <surname>Ross</surname>
          </string-name>
          , Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Wojatzki</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Measuring the reliability of hate speech annotations: The case of the european refugee crisis</article-title>
          .
          <source>arXiv preprint arXiv:1701</source>
          .
          <fpage>08118</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Fabio Poletto, Cristina Bosco, Viviana Patti, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Stranisci</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>An italian Twitter corpus of hate speech against immigrants</article-title>
          .
          <source>In Proceedings of LREC.</source>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Anna</given-names>
            <surname>Schmidt</surname>
          </string-name>
          and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Wiegand</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A survey on hate speech detection using natural language processing</article-title>
          .
          <source>SocialNLP</source>
          <year>2017</year>
          , page 1.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Zeerak</given-names>
            <surname>Waseem</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dirk</given-names>
            <surname>Hovy</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Hateful symbols or hateful people? predictive features for hate speech detection on Twitter</article-title>
          .
          <source>In Proceedings of NAACL-HLT</source>
          , pages
          <fpage>88</fpage>
          -
          <lpage>93</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Shuhan</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <surname>Xintao Wu</surname>
            , and
            <given-names>Yang</given-names>
          </string-name>
          <string-name>
            <surname>Xiang</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A two phase deep learning model for identifying discrimination from tweets</article-title>
          .
          <source>In International Conference on Extending Database Technology</source>
          , pages
          <fpage>696</fpage>
          -
          <lpage>697</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Ziqi</given-names>
            <surname>Zhang</surname>
          </string-name>
          , David Robinson,
          <string-name>
            <given-names>and Jonathan</given-names>
            <surname>Tepper</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Detecting hate speech on Twitter using a convolution-gru based deep neural network</article-title>
          .
          <source>In European Semantic Web Conference</source>
          , pages
          <fpage>745</fpage>
          -
          <lpage>760</lpage>
          . Springer.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>