<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>YNU OXZ @ HaSpeeDe 2 and AMI : XLM-RoBERTa with Ordered Neurons LSTM for classification task at EVALITA 2020</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Xiaozhi Ou</string-name>
          <email>xiaozhiou88@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hongling Li</string-name>
          <email>honglingli66@126.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Yunnan University</institution>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. This paper describes the system that team YNU OXZ submitted for EVALITA 2020. We participate in the shared task on Automatic Misogyny Identification (AMI) and Hate Speech Detection (HaSpeeDe 2) at the 7th evaluation campaign EVALITA 2020. For HaSpeeDe 2, we participate in Task A - Hate Speech Detection and submitted two-run results for the news headline test and tweets headline test, respectively. Our submitted run is based on the pre-trained multilanguage model XLM-RoBERTa, and input into Convolution Neural Network and K-max Pooling (CNN + K-max Pooling). Then, an Ordered Neurons LSTM (ONLSTM) is added to the previous representation and submitted to a linear decision function. Regarding the AMI shared task for the automatic identification of misogynous content in the Italian language. We participate in subtask A about Misogyny &amp; Aggressive Behaviour Identification. Our system is similar to the one defined for HaSpeeDe and is based on the pre-trained multi-language model XLMRoBERTa, an Ordered Neurons LSTM (ON-LSTM), a Capsule Network, and a final classifier.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction and Background</title>
      <p>People use offensive contents in their social
media posts to degrade an individual or religion or
other organizations in many respects, the
identification of such social media posts is a necessity, a</p>
      <p>
        Copyright ⃝c 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
substantial amount of work has been done in
languages like English. However, hate speech and
offensive language identification in other language
scenario is still an area worth exploring. The latest
edition of EVALITA
        <xref ref-type="bibr" rid="ref6">(Caselli et al., 2018)</xref>
        hosted
the first Hate Speech (HS) detection in Social
Media (i.e. HaSpeeDe
        <xref ref-type="bibr" rid="ref24 ref5">(Bosco et al., 2018)</xref>
        ) task for
Italian, the HaSpeeDe 2 (Hate Speech Detection)
        <xref ref-type="bibr" rid="ref25">(Sanguinetti et al., 2020)</xref>
        shared task have been
organized within Evalita 2020 1. The ultimate goal
of HaSpeeDe 2 is to take a step further in the
state of the art of HS detection for Italian while
also exploring other side phenomena, the extent to
which they can be distinguished from HS, and
finally whether and how much automatic systems
are able to draw such conclusions. For AMI
        <xref ref-type="bibr" rid="ref11">(Elisabetta Fersini, 2020)</xref>
        , the second shared task at the
7th evaluation campaign EVALITA 2020
        <xref ref-type="bibr" rid="ref21 ref3 ref7">(Basile
et al., 2020)</xref>
        . Given the huge amount of
usergenerated content on the Web, and in particular on
social media, the problem of detecting, in order to
possibly limit the diffusion of hate speech against
women, is rapidly becoming fundamental
especially for the societal impact of the phenomenon,
it is very important to identify misogyny in social
media.
1.1
      </p>
      <sec id="sec-1-1">
        <title>Hate Speech (HaSpeeDe 2)</title>
        <p>In recent years, with the acceleration of
information dissemination, the identification of hate
speech and offense language has become a crucial
mission in multilingual sentiment analysis
fields and has attracted the attention of a large
number of industrial and academic researchers. From
an NLP perspective, much attention has been paid
to the topic of HS - together with all its
possible facets and related phenomena, such as
offensive/abusive language, and its identification. This
is shown by the proliferation, especially in the
last few years, of contributions on this topic (e.g.</p>
        <sec id="sec-1-1-1">
          <title>1http://www.evalita.it/2020/tasks</title>
          <p>
            Caselli et al. (2020), Jurgens et al. (2019), Fortuna
et al. (2019)), corpora and lexica (e.g. de Pelle
and Moreira (2017),
            <xref ref-type="bibr" rid="ref24">(Sanguinetti et al., 2018)</xref>
            ,
            <xref ref-type="bibr" rid="ref4">(Bassignana et al., 2018)</xref>
            ), dedicated
workshops, and shared tasks within national (GermEval
2, HASOC 3, IberLEF 4) and international
(SemEval 5) evaluation campaigns. Among them,
Gemeval2018 is about offensive language
recognition and aims to promote research on
offensive contents recognition in German language
microblogs. The best teams system is to train three
basic classifiers (maximum entropy and two
random forest sets) using five disjoint feature
sets and then used the maximum entropy
elementlevel classifier for final classification
            <xref ref-type="bibr" rid="ref1 ref20 ref26 ref4 ref6">(Montani and
Schu¨ller, 2018)</xref>
            . In the SemEval-2019 shared tasks
HatEval and OffensEval, HatEval is a
multilingual detection of hate speech against immigrants
and women on Twitter. Fermi team is the best
team of Hateval. It proposes an SVM model with
the RBF kernel and uses sentence embedding in
Google general sentence encoder as a function
            <xref ref-type="bibr" rid="ref15">(Indurthi et al., 2019)</xref>
            . OffensEval is about the
identification and classification of offensive language
in social media. The NULI team is the best
performing team, they use BERT-base without default
parameters
            <xref ref-type="bibr" rid="ref19 ref27">(Liu et al., 2019)</xref>
            . HASOC2019 is
proposed to identify hate speech and offensive
content in Indo-European languages. Its purpose is
to develop powerful technologies capable of
processing multilingual data and to develop a transfer
learning method that can utilize cross-lingual data.
The optimal system is a system based on ordered
neuron LSTM (ON-LSTM) and attention model
and adopts the K-folding approach for ensemble
            <xref ref-type="bibr" rid="ref27">(Wang et al., 2019)</xref>
            .
1.2
          </p>
        </sec>
      </sec>
      <sec id="sec-1-2">
        <title>Misogyny (AMI)</title>
        <p>Unfortunately, nowadays more and more incidents
of harassment against women have appeared and
misogynistic comments have been found in
social media, where misogynists hide behind by
anonymity security. Therefore, it is very important
to identify misogyny in social media. Pamungkas
et al. (2020) conducted extensive and in-depth
research on online misogyny, developed a
state-ofthe-art model for detecting misogyny in social
media and explored the feasibility of detecting
misog2https://projects.fzai.h-da.de/iggsa/germeval/
3https://hasocfire.github.io/hasoc/2020
4http://hitz.eus/sepln2019/
5http://alt.qcri.org/semeval2020/
yny in a multilingual environment. Aiming at
the TRAC-2 shared tasks of Aggression
Identification and Misogynistic Aggression Identification,
Samghabadi et al. (2020) propose an end-to-end
neural model using attention on top of BERT that
incorporates a multi-task learning paradigm to
address both the sub-tasks simultaneously. Arango
et al. (2019) discussed the implications for current
research and re-conduct experiments, a closer look
at model validation to give a more accurate
picture of the current state-of-the-art methods.
Recent investigations studied how the misogyny
phenomenon takes place, such as Farrell et al. (2019)
study this phenomenon by investigating the flow
of extreme language across seven online
communities on Reddit. Goenaga et al. (2018)
automatic misogyny identification using neural networks.
Automatic misogyny identification in Twitter has
been firstly investigated by Anzovino et al. (2018).
2
2.1</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Task and Data description</title>
      <sec id="sec-2-1">
        <title>Task description</title>
        <p>In this part, we describe one of the subtasks
HaSpeeDe 2 participating in EVALITA 2020. This
task introduces its novelty from three main
aspects (Language variety and test of time,
Stereotypical communication, Syntactic realization of HS).
We participated in Task A - Hate Speech Detection
(Main Task), a binary classification task aimed at
determining the presence or the absence of hateful
content in the text towards a given target (among
Immigrants, Muslims or Roma people).</p>
        <p>The AMI shared task proposes that
misogynous content in Italian is automatic identification
in Twitter. It is organized according to two main
subtasks, namely subtask A - Misogyny &amp;
Aggressive Behaviour Identification and subtask B
Unbiased Misogyny Identification. We participate
in subtask A, the system must recognize whether
the text is misogyny, and if it is misogyny, it must
also recognize whether it expresses an aggressive
attitude.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Data description</title>
        <p>HaSpeeDe 2 task organizer provides a new
HS training dataset (binary task) based on
Twitter data, accompanied by a test set including both
in-domain and out-of-domain data (tweets + news
headlines), as well as from different time periods.
The HaSpeeDe 2020 new training set already
contains the Twitter dataset of HaSpeeDe 2018. The
new dataset contains a total of 6,839 tweets (label
0 means NOT HS, label 1 means HS), of which
HS contains 2,766, NOT HS contains 4,703, the
tweets headlines test set contains 1,263 tweets, and
the news headlines test set contains 500 elements.
In the experimental run, the data we recommend
for this task is the result of combining the
Facebook dataset (training set + test set) of HaSpeeDe
2018 with the new training set of HaSpeeDe 2020,
this is to analyze the influence of out-of-domain
texts in the training set. The two contain a total of
10,839 comments/tweets.</p>
        <p>The AMI organizer provided a raw dataset
(5,000 tweets) as the training set for participants in
subtask A, the raw dataset is a balanced dataset of
tweets manually labeled according to two levels:
• Misogynous: defines if a tweet is
misogynous or not misogynous. Label 0 means Not
misogynous tweet, label 1 means
Misogynous tweet.
• Aggressiveness: denotes the subject of the
misogynistic tweet (misogynous tweet is
label 1). Label 0 means Non-aggressive tweet,
label 1 means Aggressive tweet. Not
misogynous tweet (misogynous tweet is label 0) are
labeled as 0 by default.</p>
        <p>For the test set (1,000 tweets) for subtask A
provided by the AMI organizer, only the annotations
on the “misogynous” and “aggressiveness” fields
in the raw dataset will consider.</p>
        <p>As shown in Figure 1, we use stratified
sampling technology (StratifiedKFold), using
StratifiedKFold cross-validation instead of ordinary
kfold cross-validation to evaluate a classifier. The
reason is that StratifiedKFold can utilize stratified
sampling to divide, which can ensure that the
proportion of each category in the generated training
set and validation set is consistent with the
original training set so that the generated data
distribution disorder will not occur. In the experiment, we
used 5-fold stratified sampling. For the HaSpeeDe
2 training set (Merged dataset), each of which
included the randomly sampled training set (8,671)
and validation set (2,168). For the AMI training
set (raw dataset), each of which included the
randomly sampled training set (4,000) and validation
set (1,000).
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Description of the system</title>
      <p>In this part, we introduce our final submission
system. Figure 2 shows the overall framework
of the system we submitted to HaSpeeDe 2 Task
A. We use the pre-trained multi-language model
XLM-RoBERTa. We discover the limitations of
BERT’s pooler output (P O) and obtained rich
semantic information by extracting the hidden state
(The last four hidden layers) of XLM-RoBERTa,
which is used as input for Convolution Neural
Network and K-max Pooling (CNN + K-max
Pooling). Then, we input the output of (CNN + K-max
Pooling) into the Ordered Neurons LSTM
(ONLSTM). Finally, we concatenate the P O and
output of ON-LSTM ON-LSTM together and pass it
through the Linear layer and Softmax for the final
classification.</p>
      <p>
        Figure 3 shows the overall framework of the
system we submitted to AMI subtask A. We
use the pre-trained multi-language model
XLMRoBERTa. We first get pooler output (P O) and
obtained rich semantic information by extracting
the hidden state (The last four hidden layers) of
XLM-RoBERTa, which is input into Ordered
Neurons LSTM (ON-LSTM). Then, we input the
output of ON-LSTM into Capsule Network.Finally,
we concatenate the P O and output of Capsule
togetherand through the Linear layer and Softmax
for the final classification.
Early work in the field of cross-language
understanding has proved the effectiveness of
multilingual masked language model (MLM) in
crosslanguage understanding, but models such as
XLM
        <xref ref-type="bibr" rid="ref12 ref16 ref17 ref18 ref19 ref2 ref27">(Lample and Conneau, 2019)</xref>
        and
Multilingual BERT
        <xref ref-type="bibr" rid="ref10">(Devlin et al., 2018)</xref>
        (pre-trained on
Wikipedia) are still limited in learning useful
representations of low resource languages.
XLMRoBERTa
        <xref ref-type="bibr" rid="ref8">(Conneau et al., 2020)</xref>
        shows that the
performance of cross-language transfer tasks can
be significantly improved by using the large-scale
multi-language pre-training model. It can be
understood as a combination of XLM and
RoBERTa. It is trained on 2.5 TB of newly created clean
CommonCrawl data in 100 languages. Because
the training of the model in this task must make
full use of the whole sentence content to extract
useful semantic features, which may help to
deepen the understanding of the sentence and reduce
the impact of noise on speech. Therefore, we use
XLM-RoBERTa in this work.
      </p>
      <p>
        In the classification task, the original output of
XLM-RoBERTa is obtained through the last
hidden state of the model. However, the output
usually does not summarize the semantic content of the
input. Recent studies have shown that
abundant semantic information features are learned by the
top hidden layer of BERT
        <xref ref-type="bibr" rid="ref16">(Jawahar et al., 2019)</xref>
        ,
which we call the semantic layer. In our opinion,
the same is true of XLM-RoBERTa. Therefore, in
order to make the model obtain more abundant
semantic information features, we propose the
system as shown in Figure 2 for HaSpeeDe 2 Task A.
Firstly, we get P O. Secondly, we extract the
hidden state of the last four layers of XLM-RoBERTa
and input them into CNN and K-max Pooling.
Then, input into ON-LSTM. For AMI subtask A,
we propose the system as shown in Figure 3.
Firstly, we get P O. Secondly, we extract the hidden
state of the last four layers of XLM-RoBERTa and
input them into ON-LSTM. Then, input into
Capsule.
3.2
      </p>
      <sec id="sec-3-1">
        <title>CNN and K-max Pooling</title>
        <p>As shown in Figure 2, we input the extracted
hidden states of the last four layers of
XLMRoBERTa into CNN and K-max Pooling for
convolution operations to obtain multiple feature
maps. The specific operation: a sentence contains
L words, each of which has a dimension of d after
the embedding layer, and the representation of the
sentence is formed by splicing the L words to form
a matrix of L ∗ d. There are several convolution
kernels in the convolutional layer, the size of which
is N ∗ d, and N is the filter window size. The
convolution operation is to apply a convolution kernel
to create a new feature in a matrix that is spliced
by words. Its formula is as follows:</p>
        <p>Cl = f (w ∗ x(l : L + N − 1) + b)
(1)
where l represents the lth word, Cl is the feature,
w is the convolution kernel, b is the bias term, and
f is a nonlinear function. After the convolution
operation of the whole sentence, a feature map is
obtained, which is a vector of size L + N - 1.</p>
        <p>Another important idea of CNN is pooling. The
pooling layer is usually connected behind the
convolution layer. The purpose of introducing it is to
simplify the output of the convolutional layer and
perform dimensionality reduction on the features
of the Filter to form the final feature. Here is the
K-max Pooling operation, which takes the value
of the scores in Top K among all the feature
values, and retains the original order of these feature
values, that is, by retaining some feature
information for subsequent use. Obviously, K-max
Pooling can express the same type of feature multiple
times, that is, can express the intensity of a certain
type of feature; in addition, because the relative
order of these Top K eigenvalues is preserved, it
should be said that it retains part of the position
information. However, this location information
is only the relative order between features, not
absolute location information.
3.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Ordered Neurons LSTM</title>
        <p>
          For HaSpeeDe 2, as shown in Figure 2, we input
the output of CNN and K-max pooling into
ONLSTM. For AMI, as shown in Figure 3, We input
the extracted hidden states of the last four layers
of XLM-RoBERTa into ON-LSTM. ON-LSTM is
a new variant of LSTM, which sorts the neurons
in a specific order, allowing the hierarchical
structure (tree structure) to be integrated into the LSTM
to express richer information. The gate structure
and output structure of ON-LSTM are still similar
to the original LSTM. The difference is that the
update mechanism from cbt to ct is different. The
formula is as follows
          <xref ref-type="bibr" rid="ref26">(Shen et al., 2018)</xref>
          :
fet = →−cs(sof tmax(Wfext + Ufeht 1 + bf ) (2)
e
iet = ←c−s(sof tmax(Weixt + Ueiht 1 + bi)
e
wt = fet ◦ iet
ct =wt ◦ (ft ◦ ct 1 + it ◦ cbt) + (fet − wt)
◦ ct 1 + (iet − wt) ◦ cbt
(3)
(4)
(5)
        </p>
        <p>Among them, −→cs and ←c−s are cumsum()
operations in the right and left directions, respectively.
the newly introduced fet and iet represent the
master forget gate and master input gate respectively.
wt represents a vector where the intersection part
is 1 and the rest is all 0. In this way, the high-level
information remains a considerable long distance,
while the low-level information may be updated at
each step of input, thereby embedding the
hierarchical structure through information grading.
3.4</p>
      </sec>
      <sec id="sec-3-3">
        <title>Capsule Network</title>
        <p>
          As shown in Figure 3, we input the output of
ONLSTM into Capsule. In the deep learning
model, spatial patterns are aggregated at a lower
level, which helps to represent higher-level concepts.
We use the Capsule Network
          <xref ref-type="bibr" rid="ref22">(Sabour et al., 2017)</xref>
          to enhance the models feature extraction
capabilities, spatial insensitivity methods are inevitably
limited by the abundant text structure (such as
saving the location of words, semantic
information, grammatical structure, etc.), difficult to
effectively encode, and lack of text expression
ability. The Capsule network effectively improved this
disadvantage by using neuron vectors instead of
individual neuron nodes of traditional neural
networks to train this new neural network in the
dynamic routing way. The Capsule’s parameter
update algorithm is routing-by-agreement, a
lowerlevel capsule prefers to send its output to
higherlevel capsule whose activity vectors have a big
scalar product with the prediction coming from the
lower-level capsule. The calculation formula of
the Capsule is as follows:
        </p>
        <p>Vj =
∥ Sj ∥2</p>
        <p>Sj
1+ ∥ Sj ∥2 ∥ Sj ∥
Sj =
∑ Cij u^jji;
i
u^jji = Wij ui
(6)
(7)
where Vj is the vector output of capsule j and
Sj is its total input, prediction vectors u^jji is by
multiplying the output ui of a capsule in the layer
below by a weight matrix Wij , the Cij are
coupling coefficients that are determined by the
iterative dynamic routing process.</p>
        <p>The most fundamental difference between the
Capsule network and the traditional artificial
neural network lies in the unit structure of the
network. For traditional neural networks, the
calculation of neurons can be divided into the following
three steps: 1. Perform a scalar weighted
calculation on the input. 2. Sum the weighted input
scalars. 3. Nonlinearization from scalar to the
scalar. For the Capsule, its calculation is divided
into the following four steps: 1. Do matrix
multiplication on the input vector. 2. Scalar weighting
of the input vector. 3. Sum the weighted vector.
4. Vector-to-vector nonlinearization. The biggest
difference between the Capsule network and the
traditional neural network is the unit output. The
output of the traditional neural network is a
value, while the output of the Capsule network is a
vector, which can contain abundant features and is
more interpretable.</p>
      </sec>
      <sec id="sec-3-4">
        <title>For the XLM-RoBERTa, we use XLM</title>
        <p>RoBERTa-base6 pre-trained model, which
contains 12 layers. We use Binary cross-entropy,
Adam optimizer with a learning rate of 5e-5.
The batch size is set to 32 and the max sequence
length is set to 80. We extract the hidden layer
state of XLM-RoBERTa by setting the
output hidden States is true. The model is trained in
8 epochs with a dropout rate of 0.1.</p>
      </sec>
      <sec id="sec-3-5">
        <title>For the Convolution Neural Network,we use</title>
        <p>2D convolution (nn.Conv2d7). The size of the
convolution kernel is set to (3,4,5) and the
number of convolution kernels is set to 256.</p>
        <p>For the ON-LSTM, we set the hidden units to
128 and num levels to 16.</p>
      </sec>
      <sec id="sec-3-6">
        <title>For the Capsule Network, we set num capsule</title>
        <p>to 10, dim capsule to 16, routings to 4.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results and Discussion</title>
      <sec id="sec-4-1">
        <title>Task</title>
      </sec>
      <sec id="sec-4-2">
        <title>HaSpeeDe</title>
      </sec>
      <sec id="sec-4-3">
        <title>Tweets News AMI subtask A</title>
      </sec>
      <sec id="sec-4-4">
        <title>Our Score</title>
      </sec>
      <sec id="sec-4-5">
        <title>Macro F1</title>
        <p>put of BERT is P O. In the same way, we just put
P O as the output of XLM-RoBERTa.The results
are shown in Table 2. We can see that the results
are not good when only P O is used as the output
of XLM-RoBERTa. We think that just using P O
as the output will lose some effective semantic
information. So we think that deep and abundant
semantic features are effective for this work. We
extract the hidden state of XLM-RoBERTa and we
also discover that the performance of the model
improves with the increase of the semantic layer.</p>
        <p>Table 3 shows the performance of our model at
different semantic layers. Table 4 shows our results
on the test set.
In our experiment, we find the limitations of P O
for sentiment analysis of hate text in Italian
languages. In the classification task, the original
outIn this work, we have similar tasks as discussed in
Section 4.1, and we consider the influence of P O
6https://huggingface.co/xlm-roberta-base for identifying misogyny content. We conduct
ex7https://pytorch.org/docs/stable/generated/torch.nn.Conv2d periments on the AMI subtask A base on the
mod</p>
        <sec id="sec-4-5-1">
          <title>AMI subtask A</title>
        </sec>
        <sec id="sec-4-5-2">
          <title>The last four hidden states of XLM-RoBERTa</title>
        </sec>
        <sec id="sec-4-5-3">
          <title>News</title>
        </sec>
      </sec>
      <sec id="sec-4-6">
        <title>Not Hate Hate</title>
        <sec id="sec-4-6-1">
          <title>Tweets</title>
        </sec>
      </sec>
      <sec id="sec-4-7">
        <title>Not Hate Hate</title>
        <p>P
el in HaSpeeDe 2, and in order to improve the
performance, we propose a new method base on this
model. Table 5 shows the comparative
experimental data of the CNN + K-max Pooling + ON-LSTM
method and the ON-LSTM + Capsule method.
Table 6 shows the results of our new model for
AMI subtask A on the test set. Run 1 only extracts
the last four hidden layer states of XLM-RoBERTa
and inputs them into ON-LSTM, then through the
Capsule Network, and finally performs
classification (without using P O). Run 2 is to concatenate
the output of the Capsule Network with the
obtained P O and input it to the classifier for final
classification (using P O). We think that
concatenate the P O and the hidden layer will retain richer
semantic information and show excellent results.</p>
        <sec id="sec-4-7-1">
          <title>Base on XLM-RoBERTa model</title>
          <p>(The validation set of 1-fold)</p>
        </sec>
        <sec id="sec-4-7-2">
          <title>Method</title>
          <p>CNN + K-max Pooling + ON-LSTM
(HaSpeeDe 2 Model)
ON-LSTM + Capsule
(AMI model)</p>
        </sec>
        <sec id="sec-4-7-3">
          <title>Macro F1</title>
          <p>0.786
In the experiment, we find the limitation of
only using pooler output as the XLM-RoBERTa’s
output. To obtain deeper and more abundant
semantic features, we extract the hidden layer
s</p>
        </sec>
      </sec>
      <sec id="sec-4-8">
        <title>Run 1 (without using P O)</title>
        <p>Run 2 (using P O)
tate of XLM-RoBERTa. The result shows that it
is helpful to improve the performance of
XLMRoBERTa to obtain more abundant semantic
information features by extracting the hidden state of
XLM-RoBERTa. We test the effects of using the
external dataset (Merged dataset) and not using
the external dataset (raw dataset). Our conclusion
is that using data from the same social network for
training and test is a necessary condition for good
performance. In addition, adding data from
different social networks can improve results.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Maria</given-names>
            <surname>Anzovino</surname>
          </string-name>
          , Elisabetta Fersini, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Automatic identification and classification of misogynistic language on twitter</article-title>
          .
          <source>In International Conference on Applications of Natural Language to Information Systems</source>
          , pages
          <fpage>57</fpage>
          -
          <lpage>64</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Ayme</surname>
            ´ Arango, Jorge Pe´rez, and
            <given-names>Barbara</given-names>
          </string-name>
          <string-name>
            <surname>Poblete</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Hate speech detection is not as easy as you may think: A closer look at model validation</article-title>
          .
          <source>In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , pages
          <fpage>45</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Evalita 2020: Overview of the 7th evaluation campaign of natural language processing and speech tools for italian</article-title>
          .
          <source>In Valerio Basile</source>
          , Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Bassignana</surname>
          </string-name>
          , Valerio Basile, and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hurtlex: A multilingual lexicon of words to hurt</article-title>
          .
          <source>In 5th Italian Conference on Computational Linguistics</source>
          , CLiC-it
          <year>2018</year>
          , volume
          <volume>2253</volume>
          , pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . CEUR-WS.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <surname>Dell'Orletta Felice</surname>
            , Fabio Poletto, Manuela Sanguinetti, and
            <given-names>Tesconi</given-names>
          </string-name>
          <string-name>
            <surname>Maurizio</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of the evalita 2018 hate speech detection task</article-title>
          .
          <source>In EVALITA 2018-Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</source>
          , volume
          <volume>2263</volume>
          , pages
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          . CEUR.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Tomasso</given-names>
            <surname>Caselli</surname>
          </string-name>
          , Nicole Novielli, Viviana Patti, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Sixth evaluation campaign of natural language processing and speech tools for italian: Final workshop</article-title>
          (evalita
          <year>2018</year>
          ).
          <source>In EVALITA 2018. CEUR Workshop Proceedings (CEUR-WS.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Tommaso</given-names>
            <surname>Caselli</surname>
          </string-name>
          , Valerio Basile, Jelena Mitrovic´,
          <string-name>
            <given-names>Inga</given-names>
            <surname>Kartoziya</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Granitzer</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>I feel offended, dont be abusive! implicit/explicit messages in offensive and abusive language</article-title>
          .
          <source>In Proceedings of The 12th Language Resources and Evaluation Conference</source>
          , pages
          <fpage>6193</fpage>
          -
          <lpage>6202</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Alexis</given-names>
            <surname>Conneau</surname>
          </string-name>
          , Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Unsupervised cross-lingual representation learning at scale</article-title>
          .
          <source>In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Rogers Prates de Pelle and Viviane P Moreira</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Offensive comments in the brazilian web: a dataset and baseline results</article-title>
          .
          <source>In Anais do VI Brazilian Workshop on Social Network Analysis and Mining. SBC.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso Elisabetta Fersini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Debora</given-names>
            <surname>Nozza</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Overview of the evalita 2020 automatic misogyny identification (ami) task</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA 2020), Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Tracie</given-names>
            <surname>Farrell</surname>
          </string-name>
          , Miriam Fernandez, Jakub Novotny, and
          <string-name>
            <given-names>Harith</given-names>
            <surname>Alani</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Exploring misogyny across the manosphere in reddit</article-title>
          .
          <source>In Proceedings of the 10th ACM Conference on Web Science</source>
          , pages
          <fpage>87</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Paula</given-names>
            <surname>Fortuna</surname>
          </string-name>
          ,
          <source>Joao Rocha da Silva</source>
          , Leo Wanner,
          <source>Se´rgio Nunes</source>
          , et al.
          <year>2019</year>
          .
          <article-title>A hierarchically-labeled portuguese hate speech dataset</article-title>
          .
          <source>In Proceedings of the Third Workshop on Abusive Language Online</source>
          , pages
          <fpage>94</fpage>
          -
          <lpage>104</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Iakes</given-names>
            <surname>Goenaga</surname>
          </string-name>
          , Aitziber Atutxa, Koldo Gojenola, Arantza Casillas, Arantza D´ıaz de Ilarraza, Nerea Ezeiza, Maite Oronoz, Alicia Pe´rez, and
          <string-name>
            <surname>Olatz</surname>
          </string-name>
          Perez-de Vin˜aspre.
          <year>2018</year>
          .
          <article-title>Automatic misogyny identification using neural networks</article-title>
          .
          <source>In IberEval@ SEPLN</source>
          , pages
          <fpage>249</fpage>
          -
          <lpage>254</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Vijayasaradhi</given-names>
            <surname>Indurthi</surname>
          </string-name>
          , Bakhtiyar Syed, Manish Shrivastava, Nikhil Chakravartula, Manish Gupta, and
          <string-name>
            <given-names>Vasudeva</given-names>
            <surname>Varma</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Fermi at semeval-2019 task 5: Using sentence embeddings to identify hate speech against immigrants and women in twitter</article-title>
          .
          <source>In Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          , pages
          <fpage>70</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Ganesh</given-names>
            <surname>Jawahar</surname>
          </string-name>
          , Benot Sagot, and
          <string-name>
            <given-names>Djam</given-names>
            <surname>Seddah</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>What does bert learn about the structure of language? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>David</given-names>
            <surname>Jurgens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Eshwar</given-names>
            <surname>Chandrasekharan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Libby</given-names>
            <surname>Hemphill</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A just and comprehensive strategy for using nlp to address online abuse</article-title>
          . arXiv preprint arXiv:
          <year>1906</year>
          .01738.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Guillaume</given-names>
            <surname>Lample</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alexis</given-names>
            <surname>Conneau</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Crosslingual language model pretraining</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Ping</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Wen</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Liang</given-names>
            <surname>Zou</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Nuli at semeval-2019 task 6: Transfer learning for offensive language detection using bidirectional transformers</article-title>
          .
          <source>In Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          , pages
          <fpage>87</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Joaquın</given-names>
            <surname>Padilla Montani</surname>
          </string-name>
          and Peter Schu¨ller.
          <year>2018</year>
          . Tuwienkbs at germeval 2018:
          <article-title>German abusive tweet detection</article-title>
          .
          <source>In 14th Conference on Natural Language Processing KONVENS</source>
          , volume
          <year>2018</year>
          , page 45.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Endang</given-names>
            <surname>Wahyu</surname>
          </string-name>
          <string-name>
            <surname>Pamungkas</surname>
          </string-name>
          , Valerio Basile, and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Misogyny detection in twitter: a multilingual and cross-domain study</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>57</volume>
          (
          <issue>6</issue>
          ):
          <fpage>102360</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Sara</given-names>
            <surname>Sabour</surname>
          </string-name>
          , Nicholas Frosst, and
          <string-name>
            <given-names>Geoffrey E</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Dynamic routing between capsules</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Niloofar</given-names>
            <surname>Safi</surname>
          </string-name>
          <string-name>
            <given-names>Samghabadi</given-names>
            , Parth Patwa, PYKL Srinivas, Prerana Mukherjee,
            <surname>Amitava Das</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Thamar</given-names>
            <surname>Solorio</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Aggression and misogyny detection using bert: A multi-task approach</article-title>
          .
          <source>In Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying</source>
          , pages
          <fpage>126</fpage>
          -
          <lpage>131</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Fabio Poletto, Cristina Bosco, Viviana Patti, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Stranisci</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>An italian twitter corpus of hate speech against immigrants</article-title>
          .
          <source>In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Gloria Comandini, Elisa Di Nuovo, Simona Frenda, Marco Stranisci, Cristina Bosco, Tommaso Caselli, Viviana Patti, and
          <string-name>
            <given-names>Irene</given-names>
            <surname>Russo</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>HaSpeeDe 2@EVALITA2020: Overview of the EVALITA 2020 Hate Speech Detection Task</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Yikang</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Shawn</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Sordoni</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Aaron</given-names>
            <surname>Courville</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Ordered neurons: Integrating tree structures into recurrent neural networks</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Bin</given-names>
            <surname>Wang</surname>
          </string-name>
          , Yunxia Ding, Shengyan Liu, and
          <string-name>
            <given-names>Xiaobing</given-names>
            <surname>Zhou</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Ynu wb at hasoc 2019: Ordered neurons lstm with attention for identifying hate speech and offensive language</article-title>
          .
          <source>In FIRE (Working Notes)</source>
          , pages
          <fpage>191</fpage>
          -
          <lpage>198</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>