<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UR NLP @ HaSpeeDe 2 at EVALITA 2020: Towards Robust Hate Speech Detection with Contextual Embeddings</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julia Hoffmann</string-name>
          <email>Julia1.Hoffmann@ur.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Udo Kruschwitz</string-name>
          <email>Udo.Kruschwitz@ur.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Regensburg</institution>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>8</lpage>
      <abstract>
        <p>We describe our approach to address Task A of the EVALITA 2020 Hate Speech Detection (HaSpeeDe2) challenge. We submitted two runs that are both based on contextual embeddings - which we had chosen due to their effectiveness in solving a wide range of NLP problems. For our baseline run we use stacked embeddings that serve as features in a linear SVM. Our second run is a simple ensemble approach of three SVMs with majority voting. Both approaches outperform the official baselines by a large margin, and the ensemble classifier in particular demonstrates robust performance on different types of test data coming 6th (out of 27 runs) for news headlines and 10th (out of 27) for Twitter feeds.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Hate speech in social media (and its automatic
detection) has become a major problem in recent
years. It can be generically defined as “language
that is used to express hatred towards a targeted
group or is intended to be derogatory, to
humiliate, or to insult the members of the group”
        <xref ref-type="bibr" rid="ref9">(Davidson et al., 2017)</xref>
        and is often based on aspects like
race, religion, ethnicity, and gender. The
problem is that what is considered acceptable for some
might not be for others. In addition to that, there is
a fine line between freedom of expression on the
one hand and censorship and illegal discrimination
on the other
        <xref ref-type="bibr" rid="ref24">(Zimmerman et al., 2018)</xref>
        . In fact,
this fine balance is reflected by the fundamental
human rights (as outlined in articles 19 and 20 of
        <xref ref-type="bibr" rid="ref21">(The United Nations, 1948)</xref>
        and
        <xref ref-type="bibr" rid="ref20">(The United
Nations General Assembly, 1966)</xref>
        which
simultane
      </p>
      <p>Copyright c 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
ously provide rights to freedom of expression and
prevent censorship and illegal discrimination. All
this contributes to making automatically detecting
hate speech a challenging task.</p>
      <p>Nevertheless, social media platforms such as
Twitter have defined clear guidelines prohibiting
the use of hateful behaviour.1 Accounts with such
contents can be reported and are subsequently
deleted. The challenge is to be able to detect such
content automatically with both high precision and
high recall.</p>
      <p>
        The EVALITA evaluation campaign introduced
a hate speech detection challenge applied to
Italian social media in 2018
        <xref ref-type="bibr" rid="ref6">(Bosco et al., 2018)</xref>
        . Its
success led to the continuation of the challenge in
2020, now called HaSpeeDe 2, which is split up
into three subtasks
        <xref ref-type="bibr" rid="ref18">(Sanguinetti et al., 2020)</xref>
        . This
report discusses our two runs that we submitted
to HaSpeeDe 2 Task A of EVALITA 2020
        <xref ref-type="bibr" rid="ref5">(Basile
et al., 2020)</xref>
        . We will first give some background
on the problem aimed at motivating our choice of
approach. We will then introduce our systems,
report results and discuss some findings. We will
also outline some scope for future developments.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>We will provide some background that should
motivate the system architectures we developed.
There are several aspects to be mentioned here.</p>
      <p>
        First of all, given the impressive advances in a
broad range of natural language processing tasks
using a transformer-based architecture
        <xref ref-type="bibr" rid="ref22">(Vaswani
et al., 2017)</xref>
        capturing contextual embeddings –
most prominently utilizing the various flavours of
BERT
        <xref ref-type="bibr" rid="ref10">(Devlin et al., 2019)</xref>
        – we decided to adopt
a transformer architecture as well. There are two
ways language models such as BERT could be
used – using pre-training and fine-tuning or just
feature-based without fine-tuning.
      </p>
      <p>1https://help.twitter.com/en/rules-and-policies/hatefulconduct-policy</p>
      <p>
        This leads us to the next design decision. The
winning team in the 2018 HaSpeeDe competition,
ItaliaNLP, submitted as one of their runs a SVM
with three different feature categories, namely raw
and lexical text, morpho-syntactic and lexicon
features, which performed extremely well in
particular when trained and tested on Twitter data
        <xref ref-type="bibr" rid="ref7">(Cimino et al., 2018)</xref>
        . Rather than designing an
end-to-end neural architecture that would be
finetuned on the available training data we therefore
opted for a simpler and slightly more transparent
architecture with an SVM backbone as our
classifier, i.e. the feature-based approach mentioned
above.
      </p>
      <p>
        Ensemble methods have repeatedly been shown
to outperform individual classifiers for a variety
of tasks including hate speech detection. For
example, an ensemble of ten simple neural
classifiers proposed by
        <xref ref-type="bibr" rid="ref24">(Zimmerman et al., 2018)</xref>
        outperformed a BERT-based approach on the
standard HatebaseTwitter benchmark dataset
        <xref ref-type="bibr" rid="ref12">(MacAvaney et al., 2019)</xref>
        . Other recent examples that
demonstrate the effectiveness of ensemble
methods for hate speech detection include
        <xref ref-type="bibr" rid="ref14 ref15 ref19 ref23 ref3 ref4">(Alonso et
al., 2020; Nourbakhsh et al., 2019; Seganti et al.,
2019; Zampieri et al., 2020; Badjatiya et al., 2017;
Park and Fung, 2017)</xref>
        . We should add that these
findings are not limited to the area of hate speech
detection as ensemble methods have a long history
in being successfully utilized in a broad range of
machine learning approaches, e.g.
        <xref ref-type="bibr" rid="ref13">(Molteni et al.,
1996)</xref>
        . Simple but effective ensemble approaches
have also been used for example in sentiment
classification of tweets, e.g.
        <xref ref-type="bibr" rid="ref11">(Hagen et al., 2015)</xref>
        , and
other social media classification tasks.
      </p>
      <p>Finally, given the task definition in which the
classifier was to be trained on social media data
but then tested on both social media and news
headlines we were aiming at an approach that
would have a robust performance across domains
rather than being tailored specifically to one type
of data.</p>
      <p>One additional motivation for our work is the
intention to develop approaches that can be
applied to different languages (we will get back to
that point when we outline future directions).</p>
      <p>We will now demonstrate how those motivating
considerations lead to the system architecture we
propose.</p>
    </sec>
    <sec id="sec-3">
      <title>System Architecture</title>
      <p>We submitted two runs of which the first one can
be considered our own baseline approach. We first
present both architectures at a conceptual level and
will go into the technical details when we discuss
the experimental setup in the next section. Our
runs are:</p>
      <p>Model 1: Stacked embeddings as features of
a linear SVM
Model 2: Ensemble of several SVMs with
different text representations – both contextual
embeddings and TF-IDF-based.</p>
      <p>Both models can be realised in many different
ways. The core idea, as motivated before, is to
experiment with transformer-based contextual
embeddings but to avoid fine-tuning and instead
deploy a traditional, more transparent approach of
SVM. The ensemble can consist of a variety of
different systems that can be aggregated in many
ways. In this paper (and as submitted) we treat
each system as equally important and use a simple
majority vote.</p>
      <p>
        Stacked embeddings have been shown to be
effective in NLP applications, e.g.
        <xref ref-type="bibr" rid="ref1 ref2">(Akbik et al.,
2018; Akbik et al., 2019)</xref>
        . Conceptually there is
some similarity to ensemble approaches in that
a combination of differently derived embedding
models turns out to be more effective than each
approach individually.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Model 1: Stacked embeddings + SVM</title>
        <p>Our own baseline model combines two different
document embeddings: transformer document and
document pool embeddings which are then fed
into a linear SVM to train a classifier. We keep
the architecture deliberately simple.</p>
        <p>
          There is a wide range of transformer-based
language models. One of our motivations was to
train a classifier that will generalise beyond a
specific domain but also has the potential to
generalise beyond a specific language. We therefore
opted for XLM-RoBERTa (XLM-R) that has been
shown to outperform alternative multilingual
models such as mBERT in various NLP tasks
          <xref ref-type="bibr" rid="ref8">(Conneau et al., 2020)</xref>
          . XLM-R is based on XLM
and RoBERTa. It is trained on data covering 100
languages in a very large (2TB) CommonCrawl.
Transformer document embeddings are obtained
from (the large version of) XLM-R. In addition
to that we use document pool embeddings which
consist of word embeddings using Flair
          <xref ref-type="bibr" rid="ref2">(Akbik et
al., 2019)</xref>
          . The exact experimental choices are
described further down.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Model 2: Ensemble of SVMs</title>
        <p>Our second system is an ensemble classifier
consisting of three SVMs each trained on a different
text representation, namely:</p>
        <p>Transformer document embeddings using
XLM-R</p>
        <sec id="sec-3-2-1">
          <title>Document pool embeddings</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>Straightforward TF-IDF.</title>
          <p>The first two of these are exactly the same as
we have seen in Model 1 except that they are not
stacked but fed into different classifiers. Again
we observe that the general setup is kept
simple to avoid overfitting for the specific problem at
hand thereby allowing more scope for future
experiments.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Setup</title>
      <p>We applied our systems to Task A - Hate Speech
Detection (Main Task).
4.1</p>
      <sec id="sec-4-1">
        <title>Data Sets</title>
        <p>Training and test data is briefly described here.</p>
        <p>Training Data Set: the training data set
consists of 6,839 tweets in total, 2,766 of them
classified as hate speech. The corpus has
three columns: tweet ID, text and the label
(0 = no hate speech, 1 = hate speech). Table
1 summarises these numbers.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Label</title>
        <p>0
1
Total</p>
      </sec>
      <sec id="sec-4-3">
        <title>Training Data Set</title>
        <p>4,073
2,766
6,839
Test Data Set: unlike training data which was
all Twitter feeds, there were two sets of test
data, the first one sampled from Twitter and
the second one from news headlines. The
Twitter test set has 1,263 entries in total, the
news test set 500. The two columns in both
sets are the ID and the text of the tweet and
news headlines, respectively. The classes 0
and 1 in the Twitter test set include 641 and
622 tweets respectively. In the news headline
test set 319 entries have the label 0, 181 the
label 1 (see Table 2).</p>
      </sec>
      <sec id="sec-4-4">
        <title>Label</title>
        <p>0
1
Total
In line with our overall aim of simplicity and
generalisibility (rather than tuning) we applied a
simple pre-processing pipeline that would apply to
both Twitter data as well as news headlines. There
are only small variations in the different
normalization steps as follows.</p>
        <p>
          For any embedding-based processing the text
was lower-cased and punctuation was removed so
that any input, be it tweet or news headline, would
be represented as a string of unpunctuated tokens.
For the calculation of our (sparse) TF-IDF
representation the text was tokenized and in addition to
that stopwords were removed. After that each
token was vectorized using TF-IDF. Figure 1 shows
an overview of the preprocessing.
All implementation was done in Python. For all
text and document embeddings we used flairNLP2.
Our SVMs were developed using scikit-learn
          <xref ref-type="bibr" rid="ref16">(Pedregosa et al., 2011)</xref>
          , and for the preprocessing
of the TF-IDF version and TF-IDF calculation we
used NLTK3 and scikit-learn.
        </p>
      </sec>
      <sec id="sec-4-5">
        <title>Stacked embeddings + SVM: as outlined, we</title>
        <p>
          use stacked embeddings composed of Transformer
Document and Document Pool Embeddings. The
Transformer Document Embeddings are obtained
using XLM-R. Document Pool Embeddings are
calculated using a mean-pooling over all word
embeddings. It consists of forward and backward
embeddings for the Italian language as provided by
flair
          <xref ref-type="bibr" rid="ref1">(Akbik et al., 2018)</xref>
          and as recommended. An
overview is given in Figure 2.
        </p>
        <p>Flair allows for the easy combination of
embeddings to create stacked embeddings – one for each
input text. These vectors (together with the labels)
are then used to train the SVM. Using grid-search
on the training data the most suitable parameter
settings were determined, and Table 3 specifies
the settings which were then used in the
submitted run.</p>
      </sec>
      <sec id="sec-4-6">
        <title>Parameter</title>
        <p>C
kernel
degree
gamma</p>
      </sec>
      <sec id="sec-4-7">
        <title>Value</title>
        <p>1.0
’linear’
3
1
2https://github.com/flairNLP/flair
3https://www.nltk.org</p>
        <p>Ensemble of SVMs: three different feature
representations are used to train one SVM each as
illustrated in Table 4. The first two incorporate the
same representations as already seen in Figure 2.</p>
      </sec>
      <sec id="sec-4-8">
        <title>Classifier</title>
        <p>SVM2.1
SVM2.2
SVM2.3</p>
      </sec>
      <sec id="sec-4-9">
        <title>Features</title>
        <p>Transformer Document Embeddings
Document Pool Embeddings</p>
        <p>TF-IDF</p>
        <p>Again we used grid-search for parameter tuning
(see Table 5).</p>
        <p>Input is run against each classifier, and through
majority voting over these three predictions the
final classification category is determined.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>We first present detailed results and then discuss
our findings and insights. We start with our
baseline approach and then move on to the
classifier ensemble. Macro-F1 is the official metric
for this competition. In addition to that we look
at Precision, Recall and F1 at category-level and
also include confusion matrices for each approach
(Model 1 and Model 2) and test set (Twitter data
and news headlines). There were 27 runs
submitted for each dataset and the official baseline was a
linear SVM with TF-IDF of word and char-grams.
5.1</p>
      <sec id="sec-5-1">
        <title>Model 1: Our Baseline</title>
        <p>Twitter Data: Training and testing on Twitter
data results in a Macro-F1 score of 0.7399 which
makes it into position 16 (out of 27). The official
task baseline is 0.7212. Details are displayed in
Table 6 and Figure 3.</p>
        <p>News Headlines: On the news headlines test
data we get a Macro-F1 of 0.6684 with official
baseline result of 0.6210 (rank 12). More details
are in Table 7 and Figure 4.
Twitter Data: Our ensemble approach gets a
Macro-F1 of 0.7599 (rank 10). More details are
included in Table 8 and Figure 5.</p>
        <p>News Headlines: On the news headlines test
data we get a Macro-F1 of 0.6984 with an official
baseline result of 0.6210 (rank 6). More details
can be found in Table 9 and Figure 6.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Metric</title>
        <p>Precision</p>
        <p>Recall</p>
        <p>F1
Our first observation we derive from the results is
that the ensemble approach we proposed for this
task does provide a robust and solid performance –
solid in that it scores well in the ranked list of
systems and robust in that it also ranks highly when
applied to out-of-domain data (coming 6th out of
27 submitted runs on data it had not been trained
on). Given the simplicity of our system
architecture and the composition of the official baseline
system we also note the superiority of
transformerbased contextual embeddings over bag-of-words
approaches (while this comes as no surprise it is
still worth pointing out). Moving from a
featurebased to a pre-training plus fine-tuning approach
will most certainly further push up the scores.</p>
        <p>Looking at the balance between precision and
recall, we find that both our approaches have a
tendency to return a fair number of false positives for
the Twitter data set. This could indicate that words
and phrases used to express hateful content is quite
common in social media even if it does not
actually represent hate speech. On the other hand, we
record a large proportion of false negatives when
classifying news headlines. This could be an
indicator of a more subtle way in which hate speech is
expressed in traditional news outlets.</p>
        <p>Generally speaking, both models perform better
on Twitter data than on news headlines – again an
insight that was to be expected due to the training
data. However, the fact that our approach managed
to score higher in the ranked list of systems for
data it was not trained on is a result that confirms
our initial assumptions – that using a corpus with
a very broad range of topics, styles and languages
as our core language model would help in making
the system transfer more easily to unseen input.</p>
        <p>
          This leads us to an area of future research.
While it would be possible to improve the
performance of our system by making the preprocessing,
the language model and any fine-tuning step match
more closely the expected test data – e.g. by using
AlBERTo, a BERT-based transformer trained on
Italian Twitter data
          <xref ref-type="bibr" rid="ref17">(Polignano et al., 2019)</xref>
          – we
are actually aiming at something else. As part of
the COURAGE research project4 we are exploring
ways to help teenagers manage social media
exposure by providing a virtual companion that would,
among other things, automatically identify
examples of hate speech, bullying or other toxic
content. Given this is a multi-national effort we are
interested in architectures that work for languages
including Italian, Spanish, German and English
with as little fine-tuning as possible. The ensemble
introduced here with its multilingual transformer
backbone turns out to be a step in that direction.
4https://www.upf.edu/web/courage
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>We presented a simple but effective architecture
to detect hate speech in Italian social media and
news headlines. Our ensemble-based architecture
relies on contextual embeddings trained on a large
multilingual corpus which we see as the basis for
the robustness of the approach. There is plenty
of room for further improvement and the results
we report here will serve as a benchmark in this
development.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>This work was supported by the project
COURAGE: A Social Media Companion
Safeguarding and Educating Students funded by the
Volkswagen Foundation, grant number 95564.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Akbik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Blythe</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Vollgraf</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Contextual string embeddings for sequence labeling</article-title>
          . In E. M.
          <string-name>
            <surname>Bender</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Derczynski</surname>
          </string-name>
          , and P. Isabelle, editors,
          <source>Proceedings of the 27th International Conference on Computational Linguistics</source>
          ,
          <string-name>
            <surname>COLING</surname>
          </string-name>
          <year>2018</year>
          ,
          <string-name>
            <given-names>Santa</given-names>
            <surname>Fe</surname>
          </string-name>
          , New Mexico, USA,
          <year>August</year>
          20-
          <issue>26</issue>
          ,
          <year>2018</year>
          , pages
          <fpage>1638</fpage>
          -
          <lpage>1649</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Akbik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bergmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Blythe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Rasul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schweter</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Vollgraf</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>FLAIR: An easy-to-use framework for state-of-the-art NLP</article-title>
          .
          <source>In Proceedings of the</source>
          <year>2019</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Demonstrations</article-title>
          , pages
          <fpage>54</fpage>
          -
          <lpage>59</lpage>
          , Minneapolis, Minnesota, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>P.</given-names>
            <surname>Alonso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Saini</surname>
          </string-name>
          , and G. Kova´cs.
          <year>2020</year>
          .
          <article-title>Hate Speech Detection Using Transformer Ensembles on the HASOC Dataset</article-title>
          . In A. Karpov and R. Potapova, editors,
          <source>Speech and Computer</source>
          , pages
          <fpage>13</fpage>
          -
          <lpage>21</lpage>
          , Cham. Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>P.</given-names>
            <surname>Badjatiya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gupta</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Varma</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Deep learning for hate speech detection in tweets</article-title>
          .
          <source>In Proceedings of the 26th International Conference on World Wide Web Companion</source>
          , pages
          <fpage>759</fpage>
          -
          <lpage>760</lpage>
          . International World Wide Web Conferences Steering Committee.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Croce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Di</given-names>
            <surname>Maro</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>EVALITA 2020: Overview of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Poletto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Tesconi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of the EVALITA 2018 hate speech detection task</article-title>
          . In T. Caselli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Novielli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          , and P. Rosso, editors,
          <source>Proceedings of the Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ), volume
          <volume>2263</volume>
          <source>of CEUR Workshop Proceedings. CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Cimino</surname>
          </string-name>
          , L. De Mattei, and
          <string-name>
            <given-names>F.</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Multi-task learning in deep neural networks at EVALITA 2018</article-title>
          . In T. Caselli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Novielli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          , and P. Rosso, editors,
          <source>Proceedings of the Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          ), volume
          <volume>2263</volume>
          <source>of CEUR Workshop Proceedings. CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chaudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wenzek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Guzma´n</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Unsupervised crosslingual representation learning at scale</article-title>
          . In D. Jurafsky,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Schluter</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. R</surname>
          </string-name>
          . Tetreault, editors,
          <source>Proceedings of ACL</source>
          , pages
          <fpage>8440</fpage>
          -
          <lpage>8451</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Davidson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Warmsley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. W.</given-names>
            <surname>Macy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I.</given-names>
            <surname>Weber</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Automated hate speech detection and the problem of offensive language</article-title>
          .
          <source>In Proceedings of ICWSM</source>
          <year>2017</year>
          , Montre´al, Que´bec, Canada, May
          <volume>15</volume>
          -18,
          <year>2017</year>
          , pages
          <fpage>512</fpage>
          -
          <lpage>515</lpage>
          . AAAI Press.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>BERT: Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In Proceedings of NAACL</source>
          , pages
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          , Minneapolis, Minnesota, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Hagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Bu¨chner, and</article-title>
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Webis: An ensemble for twitter sentiment detection</article-title>
          .
          <source>In Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval</source>
          <year>2015</year>
          ), pages
          <fpage>582</fpage>
          -
          <lpage>589</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>MacAvaney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Russell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goharian</surname>
          </string-name>
          , and
          <string-name>
            <given-names>O.</given-names>
            <surname>Frieder</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Hate speech detection: Challenges and solutions</article-title>
          .
          <source>PLoS ONE</source>
          ,
          <volume>14</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Molteni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Buizza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. N</given-names>
            <surname>Palmer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Petroliagis</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>The ECMWF ensemble prediction system: Methodology and validation</article-title>
          .
          <source>Quarterly Journal of the Royal Meteorological Society</source>
          ,
          <volume>122</volume>
          (
          <issue>529</issue>
          ):
          <fpage>73</fpage>
          -
          <lpage>119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Nourbakhsh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Vermeer</surname>
          </string-name>
          , G. Wiltvank, and
          <string-name>
            <surname>R. van der Goot.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>sthruggle at SemEval-2019 task 5: An ensemble approach to hate speech detection</article-title>
          .
          <source>In Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          , pages
          <fpage>484</fpage>
          -
          <lpage>488</lpage>
          , Minneapolis, Minnesota, USA, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Park</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Fung</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>One-step and two-step classification for abusive language detection on twitter</article-title>
          .
          <source>In Proceedings of The First Workshop on Abusive Language Online</source>
          , pages
          <fpage>41</fpage>
          -
          <lpage>45</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          , et al.
          <year>2011</year>
          .
          <article-title>Scikit-learn: Machine learning in python</article-title>
          .
          <source>Journal of machine learning research</source>
          ,
          <volume>12</volume>
          (Oct):
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Polignano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Basile</surname>
          </string-name>
          , M. de Gemmis, G. Semeraro, and
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Alberto: Italian BERT language understanding model for NLP challenging tasks based on tweets</article-title>
          . In R. Bernardi,
          <string-name>
            <given-names>R.</given-names>
            <surname>Navigli</surname>
          </string-name>
          , and G. Semeraro, editors,
          <source>Proceedings of the Sixth Italian Conference on Computational Linguistics</source>
          , Bari, Italy,
          <source>November 13-15</source>
          ,
          <year>2019</year>
          , volume
          <volume>2481</volume>
          <source>of CEUR Workshop Proceedings. CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Comandini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. Di</given-names>
            <surname>Nuovo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Frenda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stranisci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Caselli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I.</given-names>
            <surname>Russo</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>HaSpeeDe 2@EVALITA2020: Overview of the EVALITA 2020 Hate Speech Detection Task</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Seganti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sobol</surname>
          </string-name>
          , I. Orlova,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Staniszewski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Krumholc</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Koziel</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>NLPR@SRPOL at SemEval-2019 task 6 and task 5: Linguistically enhanced deep learning offensive sentence classifier</article-title>
          .
          <source>In Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          , pages
          <fpage>712</fpage>
          -
          <lpage>721</lpage>
          , Minneapolis, Minnesota, USA, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>The</given-names>
            <surname>United Nations General Assembly</surname>
          </string-name>
          .
          <year>1966</year>
          .
          <article-title>International covenant on civil and political rights</article-title>
          .
          <source>Treaty Series</source>
          ,
          <volume>999</volume>
          :
          <fpage>171</fpage>
          ,
          <string-name>
            <surname>December</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>The</given-names>
            <surname>United Nations</surname>
          </string-name>
          .
          <year>1948</year>
          .
          <article-title>Universal Declaration of Human Rights. The United Nations</article-title>
          , December.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I.</given-names>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Attention is all you need</article-title>
          .
          <source>In Proceedings of the 31st International Conference on Neural Information Processing Systems</source>
          , NIPS'
          <volume>17</volume>
          , pages
          <fpage>6000</fpage>
          -
          <lpage>6010</lpage>
          , USA. Curran Associates Inc.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Atanasova</surname>
          </string-name>
          , G. Karadzhov,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mubarak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Derczynski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitenis</surname>
          </string-name>
          , and
          <string-name>
            <surname>C</surname>
          </string-name>
          ¸ . C¸ o¨ltekin.
          <year>2020</year>
          . SemEval-2020
          <source>Task</source>
          <volume>12</volume>
          :
          <article-title>Multilingual Offensive Language Identification in Social Media (OffensEval</article-title>
          <year>2020</year>
          ). CoRR, abs/
          <year>2006</year>
          .07235.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Zimmerman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Kruschwitz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Fox</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Improving Hate Speech Detection with Deep Learning Ensembles</article-title>
          .
          <source>In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ), Miyazaki,
          <string-name>
            <given-names>Japan. European</given-names>
            <surname>Language Resources Association (ELRA).</surname>
          </string-name>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>