<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>YNU qyc at MeO endEs@IberLEF 2021:The XLM-RoBERTa and LSTM for Identifying O ensive Tweets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>hi Qu[</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Information Science and Engineering Yunnan University</institution>
          ,
          <addr-line>Yunnan</addr-line>
          ,
          <country country="CN">P.R. China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Our team (qu) participate in the classi cation tasks of MeO endEs@IberLEF 2021. The task aims to distinguish o ensive Spanish in online communication. We only participate in subtask 3 which aims to identify the o ensive tweets (non-contextual Mexican Spanish). To solve the problem, we mainly put forward the model (XLM-R combined with LSTM). In our model, the F1 score is 0.6676 and the precision is 0.7433.</p>
      </abstract>
      <kwd-group>
        <kwd>o ensive</kwd>
        <kwd>XLM-RoBERTa</kwd>
        <kwd>LSTM</kwd>
        <kwd>K-folding ensemble</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>With the popularization of the Internet, online communication becomes a
important part of life. While social networking brings great convenience to human
beings, it also has huge risks and hidden dangers [1]. One risk arises from online
comments (o ensive), which may hurt netizens, or even cause long-term harm
to the victims, leading them to depression and suicide [2]. Therefore, o ensive
comments can be detected and analyzed which have a positive e ect on every
Internet user. Many social media and technology companies are studying the
identi cation of o ensive speech in the network, and the most important thing
is how to improve the e ectiveness of identi cation [3].</p>
      <p>The main task of MeO endEs@IberLEF 2021 [4] is to analyze and detect
o ensive Spanish on social networks, hoping to nd a better solution to identify
o ensive Spanish and its categories, and explore the impact of metadata on the
identi cation task [5]. The mission is divided into four subtasks and two di erent
Spanish languages (universal Spanish and Mexican Spanish). The goal of subtask
1 is to classify generic Spanish comments (non-contextual) into multiple
categories (OFP: O ensive, target is a person. OFG: O ensive, target is a group of
people or collective. NOM: Non-O ensive, but with inadequate language. NO:
non-o ensive). The goal of subtask 2 is also to do multiple classi cations of
comments, but the di erence is that subtask 2 provide metadata. The goal of
subtask 3 is to classify Mexican Spanish tweets (non-context) into o ensive and
non-o ensive tweets. Subtask 4 is also a dichotomy for Mexican Spanish tweets,
but the di erence between subtask 4 and subtask 3 is that subtask 4 will
provide metadata. Metadata associated with tweets include date, retweet count,
bear count, reply status, quote status and so on [6].</p>
      <p>In this competition, we only take part in subtask 3 (Non-contextual binary
classi cation for Mexican Spanish). For subtask 3, we mainly use the XLM-R
model combined with LSTM (Long-Short Term Memory). The XLM-R model is
a combination of the XLM model and RoBERTa model. Relatively speaking, the
new model increases the number of model's languages and expands the number
of train set, which greatly improve the performance of the downstream tasks. At
the same time, the LSTM network [7] can further improve the model accuracy.
Finally, we use the k-fold ensemble method to get higher scores.</p>
      <p>The structure of the rest in this paper is as follows: The second part is the
related research work on the recognition and classi cation of o ensive language.
The third part is the relevant description of the data set and the construction
of our model. The fourth part is the summary and analysis of the experimental
results.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>The identi cation and classi cation of o ensive language in social media have
always existed. As early as 2009, Yin [8] identi ed and classi ed the harassment
behaviors that appeared in Web2.0. The author's idea was to extract the
features of N-Grams, TF-IDF, semantic and other features from the text at rst,
and classi ed the text by SVM (Support Vector Machine). This method achieved
excellent results at that time. For now, it has to be admitted that traditional
machine learning methods still play a key role in the recognition and classi
cation of o ensive language. However, the traditional machine learning method
also has a major drawback, which relies excessively on features and model
parameters selection. With the development of Deep Learning, more and more
neural network models make outstanding achievements in the classi cation of
o ensive language. At rst, after Kim [9] applied CNN (Convolutional Neural
Networks) to text classi cation, many text recognition models related to CNN
are produced in the NLP ( Natural Language Processing</p>
      <p>dict.cnki.net). In the eld of text detection and recognition, we have to talk
about the LSTM model. RNN (Recurrent Neural Network) [10] plays an
important role in the eld of NLP, but it has the shortcoming of long-term dependence,
and LSTM is designed for RNN's shortcomings [11]. Chakrabarty's, Zhang's and
Srivastava`s experiments prove the signi cant contribution of the LSTM model
in text classi cation [12].</p>
      <p>In 2018, Google introduced a new NLP model: the BERT model [13]. The
BERT model is excellent in many elds of NLP. Subsequently, the BERT model
is also used in the recognition and classi cation of o ensive language in social
networking. A large number of facts prove that the e ect of BERT model is
indeed better than the general model. But in the cross-lingual realm, BERT still
has many disadvantages.</p>
      <p>Although the BERT model can be trained in multiple languages, it also
has this irreparable disadvantage: di erent languages cannot learn from each
other and communicate with each other. The XLM o ers its solution for this
shortcoming. By training di erent kinds of languages under the same model, the
model can absorb the information from di erent languages. [14].</p>
      <p>
        RoBERTa's improvement over traditional BERT is mainly re ected in the
following three aspects: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) BERT uses static masking. RoBERTa adopts the
dynamic masking. (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) RoBERTa cancels the NSP task relative to BERT. (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
Compared with BERT, RoBERTa increase the batch size [15].
3
3.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methodology and Data</title>
      <sec id="sec-3-1">
        <title>Data description</title>
        <p>In subtask 3, we use the data set provided by MeO endEs@IberLEF 2021 that
is collected from Mexican-Spanish data on Twitter, and the data are manually
tagged (o ensive and non-o ensive). There are training data (5060 pieces),
validation data (76 pieces) and test data (2183 pieces). The ratio of o ensive data
items and non-o ensive data items is about 5:2, and the data distribution is
uneven.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>XLM-R and LSTM Model</title>
        <p>
          Although Spanish is the fourth most spoken language in the world, there are
relatively few resources about it. So we need to use cross-lingual model for this
task. As shown in Figure1 below, the XLM-R model+LSTM model is adopted
in this paper. The experimental steps of this model are as follows:
(
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) Encode.We add [CLS] and [SEP] for subsequent classi cation tasks and
separation between sentences. (Language embeddings+Position embeddings+Token
embeddings)
        </p>
        <p>
          (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) After the text data is encoded and input to the 12 hidden layers and the
hidden state sequence output by the last layer is obtained (last hidden state).
        </p>
        <p>
          (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) Input the hidden state sequence to LSTM and GRU and we can get the
output (H N, H AVERAGE, H MAX).
        </p>
        <p>
          (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) We input P O (Pooler output), H n, H average, and H max into the
classi er. The H n is the second output returned by GRU: the hidden state of
the last time step. H AVERAGE is the average-pooling of GRU output. H MAX
is the max-pooling of GRU output.
        </p>
        <p>At the same time, we also use the comparison model, only using the XLM-R
model, and not combined with LSTM. The experimental steps of this model are
as follows: Firstly, we get pooler output (P O) from input data through hidden
layer calculation. P O is a two-dimensional vector. Secondly, we transform P O
into a three-dimensional vector and input it into Bi-GRU to get the result of a
three-dimensional vector. Finally, we transform the result of three-dimensional
vector into two-dimensional vector and input it into the classi er and we can
get the labels of tweets. The results of the two comparative experiments (the
parameters of the two models are consistent) in the development set are shown
in the following table 1. As shown in the following table, the model combined
with LSTM has more excellent performance than the other.
In this article, the training data set only have 5,060 pieces. And the amount of
data is relatively small, so we use the K-folding ensemble method to improve the
performance of the model. The idea of the K-folding ensemble method comes
from k-folding cross-validation. We divide the data set into K in total, use K-1
for training, and use the remaining one as a validation set. Repeat K times in
total to ensure that each piece of data can become a validation set. Finally, we
will accumulate these K results to nd the average value as the nal result. The
K-folding ensemble method is shown in Figure 2.</p>
        <p>The purpose of using the K-fold ensemble method is to extract di erent
features of all data as much as possible in the case of existing training data, which
can e ectively improve the generalization ability of the model. We use validation
set to test the e ect of k-fold ensemble method. In the Table 1, compared with
the method without k-fold ensemble method, the test result in F1 score increased
by 0.0624.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiment and Results</title>
      <p>4.1</p>
      <sec id="sec-4-1">
        <title>Data preprocessing</title>
        <p>
          We observe the data set and nd that the data downloaded from Twitter are
not processed and looked very messy. In order to enable the model to obtain
data features and get a good training model, we further process the data. The
speci c process is as follows: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) All user names are replaced by "#username".
Under normal circumstances, we usually think that usernames do not contain
emotions, so to prevent the in uence of usernames on model training, we
replace all usernames. (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) All abbreviations are expanded. (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) There are emojis in
unprocessed sentences. Computers usually cannot recognize these emojis. The
emojis are replaced with the corresponding words by emotion lexicons [16] to
express the corresponding emotion. (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) Change all uppercase to lowercase, which
can unify the standard and facilitate model training. (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ) Remove stop words and
repeated words.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Experiment setting</title>
        <p>The platform used in this experiment: Intel CORE i7 CPU (16G), hard disk
1T, NVIDIA RTX 3080Ti. The operating system is Windows 10. The editor
is Pychrom2020. The framework of Deep Learning is PyTorch. The setting of
hyper-parameters has a great in uence on the performance of the model. The
following Table 2 shows the settings of the experimental hyper-parameters:
The evaluation criteria of subtask 3 are precision, recall ,and F1 score. In this
mission, the partial results and the baseline shown in Table 3 below.</p>
        <p>
          In the Table 3, the results of team vic gomez, team qu and the baseline of
task 3 are included. The team vic gomez win the rst place in the subtask 3.
And our team-name is qu. In the competition subtask 3, the ideal results were
not achieved. We think there may be two main reasons for the poor results: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
The data set is unbalanced. In the training data set, there are only 5060 items,
and the ratio of o ensive data to non-o ensive data is 5:2. This model has poor
adaptability to unbalanced training data. Because the linear classi er used in
this model is biased to most classes, it causes the deviation of the model. (
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
The parameters' design of the model are unreasonable. The results of the model
in the train set are excellent, but the results in the test set are not the best. So
the model may cause over- t problem.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper, we mainly propose a model (XLM-R and LSTM) for o ensive
detection and classi cation in Mexican Spanish (non-contextual). And the
kfold ensemble method is adopted to improve the generalization ability of the
model. We think that there is still room for improvement in model performance.
In the next research, we will further improve the optimization model (solve the
problems caused by data imbalance by over-sampling and under-sampling), and
improve the e ciency of recognition and classi cation.
13. Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep
bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805
(2018)
14. Lample, G., Conneau, A.: Cross-lingual language model pretraining. arXiv preprint
arXiv:1901.07291 (2019)
15. Zampieri, M., Malmasi, S., Nakov, P., Rosenthal, S., Farra, N., Kumar, R.:
Semeval-2019 task 6: Identifying and categorizing o ensive language in social
media (o enseval). arXiv preprint arXiv:1903.08983 (2019)
16. Majumder, P., Patel, D., Modha, S., Mandl, T.: Overview of the hasoc track at re
2019: Hate speech and o ensive content identi cation in indo-european languages.
In: Working Notes of FIRE 2019 - Forum for Information Retrieval Evaluation
(2019)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Whittaker</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kowalski</surname>
            ,
            <given-names>R.M.:</given-names>
          </string-name>
          <article-title>Cyberbullying via social media</article-title>
          .
          <source>Journal of school violence 14(1)</source>
          ,
          <volume>11</volume>
          {
          <fpage>29</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Pasi</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piwowarski</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Azzopardi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <source>[lecture notes in computer science] advances in information retrieval</source>
          volume
          <volume>10772</volume>
          |
          <article-title>| deep learning for detecting cyberbullying across multiple social media platforms 10</article-title>
          .1007/978-3-
          <fpage>319</fpage>
          -76941-7
          <source>(Chapter 11)</source>
          ,
          <volume>141</volume>
          {
          <fpage>153</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Badjatiya</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varma</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Deep learning for hate speech detection in tweets</article-title>
          .
          <source>In: Proceedings of the 26th international conference on World Wide Web companion</source>
          . pp.
          <volume>759</volume>
          {
          <issue>760</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Montes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aragon</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agerri</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alvarez-Carmona</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alvarez Mellado</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carrillo-de Albornoz</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Chiruzzo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freitas</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            <given-names>Adorno</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Gutierrez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Jimenez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.M.</given-names>
            ,
            <surname>Lima</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Plaza-de Arco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.M.</given-names>
            ,
            <surname>Taule</surname>
          </string-name>
          , M. (eds.):
          <source>Proceedings of the Iberian Languages Evaluation Forum (IberLEF</source>
          <year>2021</year>
          ) (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Plaza-del-Arco</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Casavantes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Jair</given-names>
            <surname>Escalante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Martin-Valdivia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.T.</given-names>
            ,
            <surname>Montejo-Raez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Montes-</surname>
          </string-name>
          y-Gomez,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Jarqu</surname>
          </string-name>
          n-Vasquez,
          <string-name>
            <surname>H.</surname>
          </string-name>
          ,
          <article-title>Villasen~or-</article-title>
          <string-name>
            <surname>Pineda</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Overview of the MeO endEs task on o ensive text detection at IberLEF 2021</article-title>
          .
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <issue>0</issue>
          ) (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Mehdad</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tetreault</surname>
          </string-name>
          , J.:
          <article-title>Do characters abuse more than words?</article-title>
          <source>In: Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue</source>
          . pp.
          <volume>299</volume>
          {
          <issue>303</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hochreiter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9(8)</source>
          ,
          <volume>1735</volume>
          {
          <fpage>1780</fpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xue</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hong</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davison</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kontostathis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Detection of harassment on web 2.0</article-title>
          .
          <source>Proceedings of the Content Analysis in the WEB 2</source>
          ,
          <issue>1</issue>
          {
          <issue>7</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Convolutionalneuralnetworksforsentence classi cation</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. Gamback,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Sikdar</surname>
          </string-name>
          , U.K.:
          <article-title>Using convolutional neural networks to classify hatespeech</article-title>
          .
          <source>In: Proceedings of the rst workshop on abusive language online</source>
          . pp.
          <volume>85</volume>
          {
          <issue>90</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , Y.,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Cn-hit-mi</article-title>
          . t at semeval
          <article-title>-2019 task 6: O ensive language identi cation based on bilstm with double attention</article-title>
          .
          <source>In: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          . pp.
          <volume>564</volume>
          {
          <issue>570</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khurana</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Detecting aggression and toxicity using a multi dimension capsule network</article-title>
          .
          <source>In: Proceedings of the Third Workshop on Abusive Language Online</source>
          . pp.
          <volume>157</volume>
          {
          <issue>162</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>