<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>WLV-RIT at HASOC-Dravidian-CodeMix-FIRE2020: Ofensive Language Identification in Code-switched YouTube Comments</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tharindu Ranasinghe</string-name>
          <email>T.D.RanasingheHettiarachchige@wlv.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sarthak Gupte</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcos Zampieri</string-name>
          <email>marcos.zampieri@rit.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ifeoma Nwogu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Rochester Institute of Technology</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Wolverhampton</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>This paper describes the WLV-RIT entry to the Hate Speech and Ofensive Content Identification in Indo-European Languages (HASOC) shared task 2020. The HASOC 2020 organizers provided participants with annotated datasets containing social media posts of code-mixed in Dravidian languages (Malayalam-English and Tamil-English). We participated in task 1: Ofensive comment identification in Code-mixed Malayalam Youtube comments. In our methodology, we take advantage of available English data by applying cross-lingual contextual word embeddings and transfer learning to make predictions to Malayalam data. We further improve the results using various fine tuning strategies. Our system achieved 0.89 weighted average F1 score for the test set and it ranked 5ℎ place out of 12 participants.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;ofensive language identification</kwd>
        <kwd>hate speech</kwd>
        <kwd>text classification</kwd>
        <kwd>code-switching</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Ofensive content is pervasive in social media putting users of various platforms at risk [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
The pervasiveness of such content motivated the development of several systems capable of
identifying ofensive posts in a number of languages [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. Once identified, these posts can be
then set aside for human moderation or deleted from online platforms mitigating risks to their
users [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Recent studies have addressed many types of ofensive content such as online abuse [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ],
aggression [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], cyberbullying [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ], and hate speech [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. International workshops and
competitions such as HatEval 2019 [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], OfensEval 2019 and 2020 [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ], co-located with
SemEval, have been organized in the last two years attracting a large number of participants.
Most high performing system in these competitions used neural networks and contextual word
embeddings such as BERT [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        In this paper we describe the WLV-RIT entry to the the HASOC 2020 shared task which
featured Malayalam-English code-switched data. Building on the experience of recent high
performing models submitted to the OfensEval competitions, we use a transformer-based
architecture described in detail in Section 3. The HASOC 2020 code-switching dataset is a
particularly challenging one for state-of-the-art ofensive language detection systems and we make
use of transfer learning techniques that have been recently applied to project predictions from
English to resource-poorer languages with great success [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Task description and Datasets</title>
      <p>
        The goal of this task is to identify the ofensive language of the code-mixed dateset of
comments/posts in Dravidian Languages (Tamil-English and Malayalam-English) collected from
social media. Each comment/post is annotated with an ofensive language label at the
comment/post level. The dataset has been collected from YouTube comments [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. We participated
in task 1 which is a message-level label classification task; given a YouTube comment in
Codemixed (Mixture of Native and Roman Script) Tamil and Malayalam, systems have to classify
whether a post is ofensive or not-ofensive. To the best of our knowledge, this is the first
dataset to be released for ofensive language detection in Dravidian Code-Mixed text [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        In addition to the dataset provided by the organisers we also used an English Ofensive
Language Identification Dataset (OLID) [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] used in the SemEval-2019 Task 6 (OfensEval) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
for transfer learning experiments which are describing in Section 3. OLID is arguably one of
the most popular ofensive language datasets. It contains manually annotated tweets with the
following three-level taxonomy and labels:
      </p>
      <sec id="sec-2-1">
        <title>A: Ofensive language identification - ofensive vs. non-ofensive;</title>
        <p>B: Categorization of ofensive language - targeted insult or thread vs. untargeted profanity;</p>
      </sec>
      <sec id="sec-2-2">
        <title>C: Ofensive language target identification - individual vs. group vs. other.</title>
        <p>
          We adopted the transfer learning strategy similar to previous recent work [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. We believe that
the flexibility provided by the hierarchical annotation model of OLID allows us to map OLID
level A (ofensive vs. non-ofensive ) to labels in the HASOC Malayalam-English dataset.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methods</title>
      <p>The methodology applied in this work is divided in two parts. Subsection 3.1 describes
traditional machine learning applied to this task and in Subsection 3.2 we describe the transformer
models used.</p>
      <p>
        The motivation behind our methodology is the recent success that the transformers had in
wide range of NLP tasks like language generation [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], sequence classification [
        <xref ref-type="bibr" rid="ref19 ref20">19, 20</xref>
        ], word
similarity [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], named entity recognition [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], question and answering [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] etc. The main idea
of the methodology is that we train a classification model with several transformer models
inorder to identify ofensive texts. However, the transformer models are known to be resource
intensive requiring fairly large datasets [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], therefore, we also experimented with several
traditional machine learning models
      </p>
      <sec id="sec-3-1">
        <title>3.1. Traditional Machine Learning Methods</title>
        <p>
          In the first part of the methodology, we used traditional machine learning models. We
experimented with three models; Multinomial Naive Bayes [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], Support Vector Machines [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], and
Random Forest [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ]. The models take an input vector created using Bag-of-words and outputs
a label, either ofensive or non-ofensive. The models for Multinomial Naive Bayes, SVM and
Random Forest were implemented using the Scikit-learn [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ].
        </p>
        <p>
          Data Preprocessing We performed three preprocessing techniques; removing punctuations,
removing emojis and lemmatising the English words. This was done with the use of the NLTK
(Natural Language Toolkit) library [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ] in Python.
        </p>
        <p>Hyper Parameter Optimisation Optimisation of hyper parameters was performed on SVM
and random forest only. For SVM, the hyper parameters fine-tuned were alpha, random state
and max iteration, where alpha represents regularisation, random state is used for shufling
of the data and max iteration denotes number of passes through the training data which is
also known as epochs. Optimal values achieved were alpha=0.001, random state=5, max
iteration=15. For random forest, only one hyper parameter was used which is n-estimator that
denotes number of decision trees created. Optimal value achieved for number of trees was 500.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Transformers Models</title>
        <p>
          As the second part of the methodology, we used Transformer models. Transformer
architectures have been trained on general tasks like language modelling and then can be fine-tuned
for classification tasks [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]. They take an input of a sequence and outputs the representation
of the sequence. The sequence has one or two segments that the first token of the sequence is
always [CLS] which contains the special classification embedding and another special token
[SEP] is used for separating segments. For text classification tasks, Transformer models take
the final hidden state h of the first token [CLS] as the representation of the whole sequence
[
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]. A simple softmax classifier is added to the top of the transformer model to predict the
probability of a class as shown in Equation 1 where W is the task-specific parameter matrix.
The architecture diagram of the classification is shown in Figure 1
 ( |h) =  
( h)
(1)
Transformers We experimented two pretrained transformer models; BERT [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and
XLMROBERTA [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] We used the HuggingFace’s implementation of the transformer models [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ]
and the pre-trained models available in the HuggingFace model repository.1 These models
were used mainly considering their support to Malayalam language. For BERT we used the
BERT multilingual model (BERT-M) and for XLM-ROBERTA (XLM-R) we used the
XLM-RLarge model. Both models support 104 languages including Malayalam. The interesting fact
about XLM-R is that it is very compatible in monolingual benchmarks while achieving best
results in cross-lingual benchmarks at the same time [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ].
        </p>
        <p>
          1HuggingFace model repository - https://huggingface.co/models
Transfer Learning The main idea of the transfer learning strategy is that we train a
classiifcation model on a resource rich language, typically English, using a transformer model and
perform transfer learning on a less resource language. We trained the classification model on
the first level of OLID [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ] and then we save the weights of the transformer model as well as
the softmax layer. We use this saved weights from English to initialise the weights when we are
training the classification model for Malayalam. This strategy has improved the performance
of diferent languages with less resources for ofensive language identification such as Hindi,
Bengali etc [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Therefore we experimented with this strategy to see whether it improves the
results for Malayalam too. According to the recent research, cross-lingual transformers have
slight edge when using this transfer-learning strategy [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>
          Data Preprocessing The data preprocessing for this task was kept fairly minimal to make it
portable for other languages too. We only followed one data preprocessing technique;
converting emojis to text. Emojis are found to play a key role in expressing emotions in the context of
social media [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ]. But, we cannot assure the existence of embeddings for emojis in pretrained
models. Therefore as a preprocessing step, we converted emojis to text. For this conversion
we used the Python libraries demoji 2 and emoji 3. demoji returns a normal descriptive text
and emoji returns a specifically formatted text. For an example, the conversion of , is ‘slightly
smiling face’ using demoji and ‘:slightly_smiling_face:’ using emoji. Considering that demoji
returns a normal text, we used demoji to convert the emojis to text.
        </p>
        <p>
          Fine-tuning To improve the models, we experimented diferent fine-tuning strategies:
majority class self-ensemble, average self-ensemble, language modelling, which are described
below. These fine tuning strategies have shown promising results in recent shared tasks [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ].
2demoji repository - https://github.com/bsolomon1124/demojis
3emoji repository - https://github.com/carpedm20/emoji
1. Self-Ensemble (SE) - Self-ensemble is found as a technique which results better
performance than the performance of a single model [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ]. In this approach, same model
architecture is trained or fine-tuned with diferent random seeds or train-validation splits.
Then the output of each model is aggregated to generate the final results. As the
aggregation methods, we analysed majority-class and average in this research. The number
of models used with self-ensemble will be denoted by  .
        </p>
        <p>• Majority-class SE (MSE) - As the majority class, we computed the mode of the classes
predicted by each model. Given a data instance, following the softmax layer, a
model predicts probabilities for each class and the class with highest probability is
taken as the model predicted class.
• Average SE (ASE) - In average SE, final probability of class  is calculated as the
average of probabilities predicted by each model as in Equation 2 where h is the
ifnal hidden state of the [CLS] token. Then the class with highest probability is
selected as the final class.</p>
        <p>( |ℎ) =</p>
        <p>∑ =1   ( |ℎ)

(2)
2. Language Modelling (LM) - As language modelling, we retrained the transformer model
on task dataset before fine-tuning it for the downstream task; text classification. This
training is took place according with the model’s initial trained objective. Following this
technique model understanding on the task data can be improved.</p>
        <p>Implementation</p>
        <p>We used a Nvidia Tesla K80 GPU to train the models. We mainly fine tuned
the learning rate and number of epochs of the classification model manually to obtain the best
results for the validation set. We obtained 1 −
5 as the best value for learning rate and 3 as
the best value for number of epochs for all the languages. Training for English language took
around 1 hour while training for Malayalam took around 30 minutes.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Evaluation</title>
      <p>In this section, we report the experiments we conducted and their results. As informed by the
task organisers, we used Weighted Average F1 score to measure the model performance. We
also report Precision, Recall and F1 score for each class label as well the Macro F1 score in the
results tables. Results in Tables 1 - 5 are computed on validation dataset. Finally, in Section 4.1
we report the results provided by organisers to our models, for the test set.
learning outperformed other models. Also we could notice that transfer learning improved
both models; BERT and XLM-R.</p>
      <p>Model
Random Forest
Linear SVM
Mult. Naive Bayes
The self ensemble methods were experimented using all the transformer models and obtained
results are summarised in Tables 3 and 4. In most experiments, ASE has given a higher F1 than
MSE and it improved the results over the default settings. With that fine tuning strategy too
XLM-R with transfer learning outperformed all the other models.</p>
      <p>The language modeling fine tuning strategy were experimented using all the transformer
models and obtained results are summarised in Table 5. These experimented were done on top
of ASE fine tuning strategy since it provided better results than the default settings. Results
show that language modeling clearly improved the results. In fact, the best result from our
experiments were shown when XLM-R model with transfer learning fine tuned with ASE and
language modeling.</p>
      <sec id="sec-4-1">
        <title>4.1. Submission Results</title>
        <p>Considering the evaluation results on the validation set, we selected the fine-tuned XLM-R(TL)
model with ASE + language modeling as our oficial submission to the HASOC task. According
to the results provided by the organisers, our best model has scored 0.89 weighted average F1
score on the test set and ranked 5ℎ out of 12 participants.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Analysis</title>
      <p>In addition to the experiments described in this paper, we carried out a qualitative analysis on
the dataset to find interesting patterns and observations. In the training data out of 3,200 tweets
only 567 were labelled ofensive and the remaining 2,633 were labelled as not-ofensive. The
use of English words were minimal although there are many tweets which are in Malayalam
language but written in Roman script. When analysing the tweets labelled as ofensive, we
observed that there are many tweets in the dataset which are actually not-ofensive but labelled
as ofensive. Free English translations of some examples include:
(1) Spent 4 years proclaiming to be a Royal Mech.
(2) There are 25k dislikes from Ikka (Mammooty) fans, you are free to unlike and cry.
(3) Nice, looks like a TV drama series from SuryaTV (a Malayalam channel).
(4) Have you no shame defaming a reputed hospital?
We observed that between 20% and 25% of the tweets which are labelled as ofensive are similar
to the example shown above which has certainly impacted the performance of the models.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>
        In this paper we have presented the system submitted by the WLV-RIT team to the HASOC
2020 - Ofensive Language Identification - Dravidian Code Mix Task 1 at FIRE 2020.
Following a recent study [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], we have shown that the XLM-R with transfer learning is the most
successful transformer model from several transformer models we experimented. It should be
noted that the traditional machine learning models comes very close to the performance of
the transformer models. We have shown that the best traditional machine learning algorithm
we experimented; Random Forest outperforms the majority of our transformer model based
experiments. This can be due to properties of the dataset or due to the fact that a low-resource
language like Malayalam is under represented in multilingual pre-trained models. With several
ifne tuning strategies, XLM-R with transfer learning provides the best result for the validation
set. Finally, our approach achieved 5ℎ place in the leaderboard for the test set.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>We would like to thank the HASOC organizers for running this interesting shared task and for
replying promptly to all our inquiries. We further thank the anonymous reviewers for their
insightful feedback.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bauman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Toomey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Walker</surname>
          </string-name>
          ,
          <article-title>Associations among bullying, cyberbullying, and suicide in high school students</article-title>
          ,
          <source>Journal of adolescence 36</source>
          (
          <year>2013</year>
          )
          <fpage>341</fpage>
          -
          <lpage>350</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Ojha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <article-title>Evaluating aggression identification in social media</article-title>
          ,
          <source>in: Proceedings of TRAC</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitenis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          , T. Ranasinghe, Ofensive Language Identification in Greek,
          <source>in: Proceedings of LREC</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Risch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Krestel</surname>
          </string-name>
          ,
          <article-title>Delete or not Delete? Semi-Automatic Comment Moderation for the Newsroom</article-title>
          ,
          <source>in: Proceedings of TRAC</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Nobata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tetreault</surname>
          </string-name>
          , A. Thomas,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Mehdad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <article-title>Abusive language detection in online user content</article-title>
          ,
          <source>in: Proceedings of WWW</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Founta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Djouvas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chatzakou</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Leontiadis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Blackburn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Stringhini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vakali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sirivianos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kourtellis</surname>
          </string-name>
          ,
          <article-title>Large scale Crowdsourcing and Characterization of Twitter Abusive Behavior</article-title>
          ,
          <source>in: Proceedings of ICWSM</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Ojha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          , Benchmarking Aggression Identification in Social Media,
          <source>in: Proceedings of TRAC</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chelmis</surname>
          </string-name>
          , D.-S. Zois,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <article-title>Mining patterns of cyberbullying on twitter</article-title>
          ,
          <source>in: Proceedings of ICDMW</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chelmis</surname>
          </string-name>
          , D.-S. Zois,
          <article-title>Cyberbullying detection on instagram with optimal online feature selection</article-title>
          ,
          <source>in: Proceedings of ASONAM</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          , Detecting Hate Speech in Social Media,
          <source>in: Proceedings of the International Conference Recent Advances in Natural Language Processing (RANLP)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>467</fpage>
          -
          <lpage>472</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B.</given-names>
            <surname>Mathew</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Illendula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Saha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sarkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <article-title>Temporal efects of unmoderated hate speech in gab</article-title>
          , arXiv preprint arXiv:
          <year>1909</year>
          .
          <volume>10966</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fersini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nozza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M. R.</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          , Semeval
          <article-title>-2019 task 5: Multilingual detection of hate speech against immigrants and women in twitter</article-title>
          ,
          <source>in: Proceedings of SemEval</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Farra</surname>
          </string-name>
          , R. Kumar, SemEval
          <article-title>-2019 Task 6: Identifying and Categorizing Ofensive Language in Social Media (OfensEval)</article-title>
          ,
          <source>in: Proceedings of SemEval</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Atanasova</surname>
          </string-name>
          , G. Karadzhov,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mubarak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Derczynski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitenis</surname>
          </string-name>
          , Ç. Çöltekin, Semeval-2020 task 12:
          <article-title>Multilingual ofensive language identification in social media</article-title>
          (ofenseval
          <year>2020</year>
          ), in: Proceedings of SemEval,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: Proceedings of NAACL</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ranasinghe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <article-title>Multilingual ofensive language identification with crosslingual embeddings</article-title>
          ,
          <source>in: Proceedings of EMNLP</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          , P. B,
          <string-name>
            <surname>S. KP</surname>
          </string-name>
          , T. Mandl,
          <article-title>Overview of the track on "HASOC-Ofensive Language Identification- DravidianCodeMix"</article-title>
          ,
          <source>in: Proceedings of FIRE</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Farra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>Predicting the Type and Target of Ofensive Posts in Social Media</article-title>
          ,
          <source>in: Proceedings of NAACL</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ranasinghe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hettiarachchi</surname>
          </string-name>
          , BRUMS at SemEval-2020
          <source>Task</source>
          <volume>12</volume>
          :
          <article-title>Transformer based Multilingual Ofensive Language Identification in Social Media</article-title>
          ,
          <source>in: Proceedings of SemEval</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ranasinghe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hettiarachchi</surname>
          </string-name>
          , BRUMS at HASOC 2019:
          <article-title>Deep Learning Models for Multilingual Hate Speech and Ofensive Language Identification</article-title>
          ,
          <source>In Proceedings of FIRE</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>H.</given-names>
            <surname>Hettiarachchi</surname>
          </string-name>
          , T. Ranasinghe, Brums at semeval
          <article-title>-2020 task 3: Contextualised embeddings for predicting the (graded) efect of context in word similarity</article-title>
          ,
          <source>in: Proceedings of SemEval</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>C.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Er</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Bond:
          <article-title>Bert-assisted opendomain named entity recognition with distant supervision</article-title>
          ,
          <source>in: Proceedings of SIGKDD</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>W.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>End-to-end open-domain question answering with BERTserini</article-title>
          ,
          <source>in: Proceedings of NAACL</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>A. M. Kibriya</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Pfahringer</surname>
          </string-name>
          , G. Holmes,
          <article-title>Multinomial naive bayes for text categorization revisited</article-title>
          ,
          <source>in: Australasian Joint Conference on Artificial Intelligence</source>
          , Springer,
          <year>2004</year>
          , pp.
          <fpage>488</fpage>
          -
          <lpage>499</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>C.</given-names>
            <surname>Cortes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vapnik</surname>
          </string-name>
          ,
          <article-title>Support-vector networks</article-title>
          ,
          <source>Machine learning 20</source>
          (
          <year>1995</year>
          )
          <fpage>273</fpage>
          -
          <lpage>297</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>A.</given-names>
            <surname>Liaw</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiener</surname>
          </string-name>
          , et al.,
          <article-title>Classification and regression by randomforest</article-title>
          ,
          <source>R news 2</source>
          (
          <year>2002</year>
          )
          <fpage>18</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          , et al.,
          <article-title>Scikit-learn: Machine learning in python</article-title>
          ,
          <source>the Journal of machine Learning research 12</source>
          (
          <year>2011</year>
          )
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bird</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Klein</surname>
          </string-name>
          , E. Loper,
          <article-title>Natural language processing with Python: analyzing text with the natural language toolkit, "</article-title>
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          ,
          <source>Inc."</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>How to fine-tune bert for text classification?</article-title>
          , in: M.
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Liu</surname>
          </string-name>
          , Y. Liu (Eds.),
          <source>Chinese Computational Linguistics</source>
          , Springer International Publishing,
          <year>2019</year>
          , pp.
          <fpage>194</fpage>
          -
          <lpage>206</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chaudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wenzek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Guzmán</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Unsupervised cross-lingual representation learning at scale</article-title>
          , arXiv preprint arXiv:
          <year>1911</year>
          .
          <volume>02116</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>T.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Debut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chaumond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Delangue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cistac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Louf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Funtowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Brew</surname>
          </string-name>
          ,
          <article-title>Huggingface's transformers: State-of-the-art natural language processing</article-title>
          , ArXiv abs/
          <year>1910</year>
          .03771 (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Farra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>Predicting the type and target of ofensive posts in social media</article-title>
          ,
          <source>in: Proceedings of NAACL</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>H.</given-names>
            <surname>Hettiarachchi</surname>
          </string-name>
          , T. Ranasinghe,
          <article-title>Emoji powered capsule network to detect type and target of ofensive posts in social media</article-title>
          ,
          <source>in: Proceedings of RANLP</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>H.</given-names>
            <surname>Hettiarachchi</surname>
          </string-name>
          , T. Ranasinghe, Infominer at wnut
          <article-title>-2020 task 2: Transformer-based covid-19 informative tweet extraction</article-title>
          , in: Proceedings of W-NUT,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>Improving bert fine-tuning via self-ensemble and selfdistillation</article-title>
          , arXiv preprint arXiv:
          <year>2002</year>
          .
          <volume>10345</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>