<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>A. Kalaivani);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Ofensive language detection in English, Hindi, and Marathi languages</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Adaikkan Kalaivani</string-name>
          <email>kalaivania@ssn.edu.in</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Durairaj Thenmozhi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of CSE, Sri Sivasubramaniya Nadar College of Engineering</institution>
          ,
          <addr-line>Kalavakkam, TamilNadu</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Information and Communication Engineering, Anna University</institution>
          ,
          <addr-line>Chennai</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Research Centre, Department of CSE, Sri Sivasubramaniya Nadar College of Engineering</institution>
          ,
          <addr-line>Kalavakkam, TamilNadu</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>Hate speech and ofensive language are phenomena that spread with the rising popularity of social media forums. Automatic detection of such content is crucial for predicting conflicts among social communities and blocking inappropriate content from social media forums. This paper aims to describe our team SSN_NLP_MLRG submission to HASOC 2021: Hate speech and ofensive language detection in English and Indo-Aryan language, where we explore diferent models to perform the subtask1 includes subtask A: To detect the comments is Hate speech and ofensive (HOF) or NOT and subtask B: To classify the HOF comments into profanity (PRFN), Hate speech (HATE), Ofensive (OFFN) in English, Hindi language and subtask A in Marathi language. The experiments cover diferent learning techniques that include machine learning, transfer learning, and Multilingual pre-trained models. Our best models are Roberta for English subtask A, BERT for English subtask B, and MBERT for the Hindi subtask A, Hindi subtask B, and Marathi subtask A. Our team achieved the macro-averaged F1 scores of 0.7919, 0.7320, 0.8223, 0.6242, and 0.5110 in the English subtask A, Hindi subtask A, Marathi subtask A, English subtask B, and Hindi subtask B, respectively.</p>
      </abstract>
      <kwd-group>
        <kwd>Transfer learning</kwd>
        <kwd>Code-Mixed language</kwd>
        <kwd>Machine learning</kwd>
        <kwd>Language modeling</kwd>
        <kwd>Low-resource language</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Social media is a vast online communication forum that enables the public to express themselves
easily, at times, anonymously. While expressing their opinion oneself is a right of humans
that is cherished, inducing and spreading ofensive content towards another social community
is an abuse of this liberty [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Therefore, social media forums and other means of online
communication platforms have begun to play a larger role in hate and ofensive crimes. Many
online social media forums such as Twitter, Facebook, Instagram, and YouTube consider hate
speech and ofensive content harmful and have the policy to remove such content. Due to the
societal concern and how widespread ofensive content is becoming on the Internet, there is a
strong motivation to detect hate speech and ofensive content in social media forums. Hate
speech1 defines the attacks against someone or group community, based on these attributes as
CEUR
Workshop
Proceedings
race, gender, ethnicity, religion, sexual orientation, age, physical or mental disability, and others.
Ofensive content 2 is a language that could seriously ofend an individual or group based on
their age, religious or political beliefs, marital or parental status, sexual orientation, physical
features, national origin, or disability.
      </p>
      <p>
        Hindi is an Indo-Aryan language with the oficial languages of India and spoken chiefly in
North India. Marathi is an Indo-Aryan language spoken predominantly by Marathi people of
the Maharashtra state in India and oficial and co-oficial language in the Maharashtra and Goa
states of Western India. Code-mixed language is a phenomenon that combines one or more
languages and also native language written in roman script. The detection and categorization
of hate speech and ofensive language in the indirect comments [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] of the code-mixed are
challenging tasks not only in the English language. Therefore, there is an open research area in
the field of a code-mixed multilingual community such as Hindi, Marathi languages, etc.
      </p>
      <p>
        The HASOC 2019 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and 2020 competitions aim to train the systems capable of detecting
hate speech and ofensive content in social media forums for the English, Hindi, and German
languages. In the HASOC 20193, the organizers ofered sub-tasks A, B, C for English, Hindi,
and sub-tasks A, B for German languages. Furthermore, whether the content is NOT or HOF
(sub-task A), what are the characteristics of the HOF content (sub-task B), and who is the target
of the HOF message (sub-task C). In the HASOC 20204, they organized the two tasks for English,
German, and Hindi Languages, namely subtask A: To identify the given comment is HOF or
NOT and subtask B: To categorize the HOF comment into ofensive, hate, and profanity.
      </p>
      <p>This paper presents our approaches to HASOC-2021. We have participated in subtask1 shared
task consisting of English subtask A, English subtask B, Hindi subtask A, Hindi subtask B, and
Marathi subtask A. The goal of subtask A is to identify and detect the social media comments are
hate speech and ofensive (HOF) or NOT. The subtask B aims to categorize the characteristics of
the HOF content into hate, ofensive, and Profanity. We used the machine learning algorithms,
BERT, MBERT, ALBERT, RoBERTa, DistilBERT model with ktrain library, ULMFiT to adapt
and fine-tune the system. We used the NLTK library for pre-processing the training set and
testing set for all three languages. The paper outlines as follows. The survey of relevant works
describes in Section 2. The details of the experiment data and technique of our models are in
Sections 3 and 4. Section 5 describes the analysis part of the experiment results. Finally, Section
6 shows the concluded work and discusses further work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        The OfensEval 2019 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is the shared task to identify the ofensive content in the English
language. Most of the teams employed BERT with diferent parameters to detect ofensive
language. In the SemEval 2019 shared task, the researchers used machine learning approaches,
bi-directional LSTM models to classify ofensive language [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Mostly, the researchers have
adapted and fine-tuned the BERT, GPT, and ULMFiT models for detecting ofensive language
in the shared task. The OfensEval 2020 [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ] is the Shared Task on Multilingual Ofensive
      </p>
      <sec id="sec-2-1">
        <title>2https://www.lawinsider.com/dictionary/ofensive-content 3https://hasocfire.github.io/hasoc/2019/dataset.html 4https://hasocfire.github.io/hasoc/2020/dataset.html</title>
        <p>Language Identification for the English, Greek, Danish, Turkish, Arabic languages. Most teams
used con-textualized Transformers, ELMo embeddings, BERT, RoBERTa, and the multilingual
mBERT to detect and categorize the ofensive language in five diferent languages.</p>
        <p>
          For the low-resource language, the authors used the cross-lingual data augmentation
technique for the input context [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. In the Gemeval2018 shared task [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], They obtained the optimal
solution by using the maximum entropy meta-level classifier model to identify the micro-posts
of ofensive language in the German language. In the HatEval [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], the top team used the
SVM model to detect the Twitter comments is hate speech against women and immigrants
in multilingual language. The majority of teams used deep learning techniques in that LSTM
model to detect hate speech and ofensive language in the shared task of HASOC 2019.
        </p>
        <p>
          In the HASOC 2020, the best-performed teams have used the variants of the BERT
transformers model [
          <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
          ] to identify and categorize the hate speech and ofensive language in
the English, German, and Hindi (code-mixed) languages. From the observation, most of the
researchers were used machine learning techniques, deep learning techniques, variation of
pre-trained transformer models to detect hate speech and ofensive language. We are still
dealing with the issue of recognizing ofensive content in low-resource languages, as well as
the problem of handling an imbalanced dataset in various code-mixed languages. This problem
opens new research in diferent low-resource languages other than English. HASOC 2021 5
shared task organizers provide the resource for the English, Hindi (code-mixed), and Marathi
languages [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experiment data</title>
      <p>This section presents the task description, data pre-processing techniques in the shared task
HASOC 2021.</p>
      <sec id="sec-3-1">
        <title>3.1. Data description</title>
        <p>The organizers ofer two shared tasks, namely subtask1 and subtask2. In subtask1, they provided
the HASOC 2021 dataset of the English, Hindi (code-mixed), and Marathi languages. In subtask2,
the organizers provided the Hindi code-mixed conversation dataset. Our team SSN_NLP_MLRG
participated in the subtask1 for all three languages. Table 1 presents the annotated tweets
for the English, Marathi, Hindi HASOC 2021 dataset. For training and testing the system, the
English dataset has 3832 and 1281 posts. The Hindi code-mixed dataset contains 4594 posts,
1532 posts for training and testing the system. The Marathi code-mixed dataset consists of 1863
posts for the train system and 625 comments for testing the model system. Table 2 shows the
statistics of the dataset for all three languages.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Task description</title>
        <p>
          The shared task of HASOC 2021 is to identify the hate speech and ofensive content in English
and Indo-Aryan languages [
          <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
          ]. The English language includes two tasks, namely subtasks
        </p>
        <sec id="sec-3-2-1">
          <title>5https://hasocfire.github.io/hasoc/2021/dataset.html</title>
          <p>found the little bastard now the fun begins
my first time seeing report about this is very heart breaking</p>
          <p>technically that is still turning back the clock dick head
india has got the worst finance minister and health minister ever
HOF
NOT
HOF
HOF
A and B as same as Hindi language. The Marathi language ofers one task, namely subtask A.
Subtask A is a binary text classification task that focuses the systems able to classify the given
social media comments into two classes, namely, HOF and NOT.</p>
          <p>Hate and Ofensive content (HOF): The social media comments contain harassment,
profane, insults, threatening words.</p>
          <p>Non-Hate and Ofensive (NOT): The social media comments do not include hate and
ofensive content.</p>
          <p>Subtask B is a multi-text classification that focuses the systems able to classify the given
online comments into three classes, namely HATE, OFFN, PRFN.</p>
          <p>Hate speech (HATE): The social media comments which contain hate words.
Ofensive (OFFN): The social media posts which contain ofensive content.</p>
          <p>Profane (PRFN): The social media posts contain profanity words.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Data pre-processing</title>
        <p>
          The data pre-processing is to clean the social media comments from the unnecessary noisy
content is present in the given dataset and transform it into a coherent form, which can be
portable for English, Hindi code-mixed, and Marathi languages. We used the NLTK6 for data
cleaning, data duplication from the HASOC 2021 dataset [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. First, we remove @ symbol with a
string denoted as user-id because it does not have any meaningful expressions. Next, we remove
the hashtag with a text as the user’s name because it afects the performance of our model. The
example of data pre-processing like “@AjeebBharti @BeraJaykrishna @khan_nainam @Policy
@Twitter Prove it !!! What evidence you have bloody hell prove kar” and after that “prove
it what evidence you have bloody hell prove kar”. After that, we removed the punctuation,
numerals, symbols, URLs, and emojis and then converted the upper case text into small case text.
Finally, we replaced the misspelling ofense words and string with * into appropriate matched
words presented in the collected vocabulary words.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>This section presents the experimental analysis of the various methods used for the validation
process.</p>
      <sec id="sec-4-1">
        <title>4.1. BERT-based model and Language Model</title>
        <p>
          We have experimented with various pre-trained models, namely BERT uncased, DistilBERT
base-uncased, ALBERT (Albert-base-v2), RoBERTa base, ULMFiT language modeling, and
Machine learning techniques for subtask A of English language. We used the BERT, ALBERT
pre-trained models and machine learning techniques for subtask B of the English language. For
the Hindi language subtask A and B, We used the multi-cased BERT transformers to adapt and
ifne-tune the system to classify the hate speech and ofensive content from the given dataset.
For the Marathi language subtask A, we used the multi-cased BERT (MBERT) for the binary text
classification task. For the validation process of the system, we take 25% of the data from the
training dataset for the three languages. We used the above-mentioned pre-trained models with
the ktrain7 library that is useful to build the system using machine learning, neural network,
and deep learning techniques. We have analyzed the training system to set the various batch
size to 6, 32 and learning rates as 2e-5, 3e-5, and the epochs to 6, 7, 9, and 10. We used the
ULMFiT[
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] framework in that Average-SGD Weight-Dropped LSTM (AWD-LSTM) architecture
model to predict the hate speech and ofensive content and their characteristics for the English
language dataset. Table 3 presents the validation results for the BERT-based models of the three
languages.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Machine Learning Techniques</title>
        <sec id="sec-4-2-1">
          <title>For the machine learning techniques, we have conducted two experiments.</title>
          <p>4.2.1. Experiment 1
In the first experiment, we have used the following models, namely support vector machine
classifier (SVM), Naive Bayes classifier (NB), random forest classifier (RF), and Extreme gradient
boosting ensemble classifier (XGB), and used to predict the hate speech and ofensive content
in the given English dataset. We used the sci-kit learn library for the implementation of the
machine learning classifiers. For using Term frequency-inverse document frequency (TF-IDF)
vectorization, we extracted the Ngram, character level, word-level features from the given
dataset. For using the sklearn CountVectorizer, we build vocabulary for known words and also
tokenize the collected data. FastText is the pre-trained vector for 157 languages trained on
Common crawl and Wikipedia. We used the FastText pre-trained word embedding vectors for
the English language, namely Wikipedia Tamil vectors (wiki.ta.vec).
4.2.2. Experiment 2
In the second experiment, we used the Gensim library for vector embeddings. Gensim is the
fastest library for training the system of vector embedding. We experiment with the logistic
regression, Multinomial Naive Bayes (NB), Random forest, and Linear Support vector machine
(SVC) models to predict the system. We have utilized the genism for pre-processing and
lemmatized the training dataset for this experiment 2. We have utilized the word cloud for
categorized the training dataset. We have used a Document to vector transformer (Doc2vec)
and text to TFIDF transformer for extracted the features by using the genism model. Table 4
presents the validation results for the machine learning models of the English language.</p>
          <p>Finally, we have used the MBERT model to predict the hate speech and ofensive content and
got a macro F1-score of 0.8223 with the epochs to 10 and the learning rate as 2e-5 for the Marathi
subtask A. We got macro F1-scores of 0.7320, 0.511 of the Hindi subtask A, Hindi subtask B with
the epochs to 10, and the learning rate as 2e-5 for the MBERT model. For English subtask A,
We got a macro F1 score of 0.7919 for the RoBERTa model with the 07 epochs and the learning
rate as 3e-5. We got a macro F1 score of 0.624 with the 09 epochs and the learning rate as 2e-5
for the BERT model of the English subtask B.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Result Analysis</title>
      <p>This section presents the evaluation of our model. For instance, we used the evaluation metrics
like precision, recall, macro-averaged F1-score. We have submitted our best model after
comparing the performance of our methods for English, Marathi, and Hindi code-mixed languages. The
HASOC 2021 organizers provided the test data for English subtask A, English subtask B, Hindi
subtask A, Hindi subtask B, and Marathi subtask A. From the performance of the validation
system, the RoBERTa model achieved an accuracy of 0.83 and Precision, Recall, and macro
F1-score of 0.81, 0.80, and 0.80, when compared with the performance of the other machine
learning approaches and pre-trained language models. The F1-score for the Not-ofensive
comments and hate speech ofensive comments for the RoBERTa model is 0.74, 0.87 respectively.
Therefore, the RoBERTa model performs well than other models for the English subtask A. The
accuracy of the English subtask B is 0.66, and the F1-score for the OFFN, HATE, PRFN, and
NONE comments for the BERT model are 0.54, 0.73, 0.41, and 0.77.</p>
      <p>From the observation, the BERT model performs well than other approaches of machine
learning techniques for the English subtask B. For the Hindi language, the MBERT model
achieved an accuracy of 0.79 for subtask A, 0.69 for subtask B, and F1-score of HOF and NOT
comments are 0.65, 0.86 for the subtask A, and the F1-score for the OFFN, HATE, PRFN, and
NONE comments for the subtask B task are 0.83, 0.19, 0.40 and 0.43 respectively. For Marathi
Language, MBERT achieved an accuracy of 0.88, the macro F1-score is 0.87, and the F1-score of
HOF and NOT comments are 0.91 and 0.83. We adapted and fine-tuned the BERT-based model
to build and predict the data and its characteristics for all the languages. Our team submitted
three runs for the English subtask A and one runs for the subtasks of the other languages.</p>
      <p>The final results 8 of our team for the three languages are present in Table 5. Our team
SSN_NLP_MLRG submission got the 19ℎ , 12ℎ , 25ℎ , 9ℎ , 19ℎ rank in the shared task for English
subtask A, English subtask B, Hindi subtask A, Hindi subtask B, Marathi subtask A respectively.</p>
      <sec id="sec-5-1">
        <title>8https://hasocfire.github.io/hasoc/2021/results.html</title>
        <p>We classify the performance of the model for all the three languages by using the confusion
matrix are presented in the Figure 1 for English subtask A for the RoBERTa model, Figure 2
for the English subtask B for the BERT model, Figure 3 for the Hindi subtask A for the MBERT
model, Figure 4 for the Hindi subtask B for the MBERT model, Figure 5 for the Marathi subtask
A for the MBERT model. From the confusion matrix, we noticed that many test cases were
classified as HOF comments by the RoBERTa model for the English subtask A. For English A,
the Precision, Recall, and F1-score for the HOF and NOT comments are 0.81, 0.89, 0.85, and
0.79, 0.66, 0.62 respectively. For English B, the F1-score for the OFFN, HATE, PRFN, and NONE
comments are 0.72, 0.46, 0.75, and 0.56. For Hindi A, the Precision, Recall, and F1-score for the
HOF and NOT comments are 0.81, 0.87, 0.84, and 0.69, 0.58, 0.63 respectively. For Hindi B, the
F1-score for the OFFN, HATE, PRFN, and NONE comments are 0.83, 0.40, 0.50, and 0.31. For
Marathi A, the Precision, Recall, and F1-score for the HOF and NOT comments are 0.89, 0.88,
0.88, and 0.75, 0.77, 0.76 respectively. Overall, the hate speech and ofensive comments perform
well by Bert-based models for all three languages.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This paper presents the team submitted runs for the hate speech and ofensive language
identification for the HASOC 2021 subtask1 shared task in the Forum for Information Retrieval
Evaluation (FIRE) 2021. We experimented with diferent approaches such as machine learning
techniques, pre-trained BERT-based models. The results show the RoBERTa models perform
well than the other BERT-based models and the machine learning approaches for the English
subtask A. The BERT uncased performs well in the English subtask B. MBERT performs well in
Hindi subtask A, Hindi subtask B, and Marathi subtask A. Based on the evaluation, the Overall
BERT-based model performs well for the three languages. Our team submission had a macro
F1-score of 0.8223, for the Marathi subtask A, macro F1-score of 0.7320, 0.511 for the Hindi
subtask A and Hindi subtask B code-mixed language, and macro F1-score of 0.7919, 0.624 for the
English subtask A and English subtask B. For future work, we will handle the sarcastic feature
and imbalanced dataset to avoid misclassification and extend this work into other low-resource
languages.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>MacAvaney</surname>
          </string-name>
          , H.
          <string-name>
            <surname>-R. Yao</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Russell</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Goharian</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Frieder</surname>
          </string-name>
          ,
          <article-title>Hate speech detection: Challenges and solutions</article-title>
          ,
          <source>PloS one 14</source>
          (
          <year>2019</year>
          )
          <article-title>e0221152</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kalaivani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Thenmozhi</surname>
          </string-name>
          ,
          <article-title>Sentimental Analysis using Deep Learning Techniques</article-title>
          ,
          <source>International Journal of Recent Technology and Engineering (IJRTE) 7</source>
          (
          <year>2019</year>
          )
          <fpage>600</fpage>
          -
          <lpage>606</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kalaivani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Thenmozhi</surname>
          </string-name>
          ,
          <article-title>Sarcasm Identification and Detection in Conversion Context using BERT</article-title>
          ,
          <source>in: Proceedings of the Second Workshop on Figurative Language Processing</source>
          , Association for Computational Linguistics, Online,
          <year>2020</year>
          , pp.
          <fpage>72</fpage>
          -
          <lpage>76</lpage>
          . URL: https://www. aclweb.org/anthology/2020.figlang-
          <volume>1</volume>
          .10.
          <article-title>doi:1 0 . 1 8 6 5 3 / v 1 / 2 0 2 0 . f i g l a n g - 1 . 1 0 .</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Modha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mandlia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <article-title>Overview of the HASOC Track at FIRE 2019: Hate Speech and Ofensive Content Identification in Indo-European Languages</article-title>
          ,
          <source>in: Proceedings of the 11th Forum for Information Retrieval Evaluation</source>
          , FIRE '19,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2019</year>
          , p.
          <fpage>14</fpage>
          -
          <lpage>17</lpage>
          . URL: https://doi.org/10.1145/3368567.3368584.
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 3 3 6 8 5 6 7 . 3 3 6 8 5 8 4 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Farra</surname>
          </string-name>
          , R. Kumar, SemEval
          <article-title>-2019 Task 6: Identifying and Categorizing Ofensive Language in Social Media (OfensEval</article-title>
          ),
          <year>2019</year>
          .
          <article-title>a r X i v : 1 9 0 3 . 0 8 9 8 3</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Thenmozhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Senthil</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sharavanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chandrabose</surname>
          </string-name>
          ,
          <article-title>SSN_NLP at SemEval2019 Task 6: Ofensive Language Identification in Social Media using Traditional and Deep Machine Learning Approaches</article-title>
          ,
          <source>in: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Minneapolis, Minnesota, USA,
          <year>2019</year>
          , pp.
          <fpage>739</fpage>
          -
          <lpage>744</lpage>
          . URL: https://www.aclweb.org/anthology/S19-2130. doi:
          <article-title>1 0 . 1 8 6 5 3 / v 1 / S 1 9 - 2 1 3 0</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Atanasova</surname>
          </string-name>
          , G. Karadzhov,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mubarak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Derczynski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitenis</surname>
          </string-name>
          , Çağrı Çöltekin, SemEval-2020
          <source>Task</source>
          <volume>12</volume>
          :
          <article-title>Multilingual Ofensive Language Identification in Social Media (OfensEval</article-title>
          <year>2020</year>
          ),
          <year>2020</year>
          .
          <article-title>a r X i v : 2 0 0 6 . 0 7 2 3 5</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kalaivani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Thenmozhi</surname>
          </string-name>
          , SSN_NLP_MLRG at SemEval-2020
          <source>Task</source>
          <volume>12</volume>
          :
          <article-title>Ofensive Language Identification in English, Danish, Greek Using BERT and Machine Learning Approach</article-title>
          , in: Proceedings of the Fourteenth Workshop on Semantic Evaluation, International Committee for Computational Linguistics,
          <source>Barcelona (online)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>2161</fpage>
          -
          <lpage>2170</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .semeval-
          <volume>1</volume>
          .
          <fpage>287</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>McCann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Keskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          , XLDA:
          <article-title>Cross-Lingual Data Augmentation for Natural Language Inference and Question Answering</article-title>
          , CoRR abs/
          <year>1905</year>
          .11471 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1905</year>
          .11471.
          <article-title>a r X i v : 1 9 0 5 . 1 1 4 7 1</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Siegel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ruppenhofer</surname>
          </string-name>
          ,
          <article-title>Overview of the GermEval 2018 Shared Task on the Identification of Ofensive Language</article-title>
          , in: In Proceedings of GermEval,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fersini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nozza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Rangel Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , M. Sanguinetti, SemEval
          <article-title>-2019 Task 5: Multilingual Detection of Hate Speech Against Immigrants and Women in Twitter</article-title>
          ,
          <source>in: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Minneapolis, Minnesota, USA,
          <year>2019</year>
          , pp.
          <fpage>54</fpage>
          -
          <lpage>63</lpage>
          . URL: https://www.aclweb.org/anthology/S19-2007.
          <article-title>doi:1 0 . 1 8 6 5 3 / v 1 / S 1 9 - 2 0 0 7</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kalaivani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Thenmozhi</surname>
          </string-name>
          , SSN_NLP_
          <article-title>MLRG@HASOC-FIRE2020: Multilingual Hate Speech and Ofensive Content Detection in Indo-European Languages using ALBERT</article-title>
          , in: P. Mehta,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          , M. Mitra (Eds.), Working Notes of FIRE 2020 -
          <article-title>Forum for Information Retrieval Evaluation, Hyderabad</article-title>
          , India,
          <source>December 16-20</source>
          ,
          <year>2020</year>
          , volume
          <volume>2826</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>188</fpage>
          -
          <lpage>194</lpage>
          . URL: http: //ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2826</volume>
          /
          <fpage>T2</fpage>
          -12.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Modha</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Kumar</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <article-title>Overview of the HASOC Track at FIRE 2020: Hate Speech and Ofensive Language Identification in Tamil, Malayalam, Hindi, English and German, in: Forum for Information Retrieval Evaluation</article-title>
          ,
          <string-name>
            <surname>FIRE</surname>
          </string-name>
          <year>2020</year>
          ,
          <article-title>Association for Computing Machinery</article-title>
          , New York, NY, USA,
          <year>2020</year>
          , p.
          <fpage>29</fpage>
          -
          <lpage>32</lpage>
          . URL: https: //doi.org/10.1145/3441501.3441517.
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 3 4 4 1 5 0 1 . 3 4 4 1 5 1 7 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gaikwad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ranasinghe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Homan</surname>
          </string-name>
          ,
          <article-title>Cross-lingual Ofensive Language Identification for Low Resource Languages: The Case of Marathi</article-title>
          ,
          <source>in: Proceedings of RANLP</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Modha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. K.</given-names>
            <surname>Shahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Madhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Satapara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ranasinghe</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Zampieri, Overview of the HASOC Subtrack at FIRE 2021: Hate Speech and Ofensive Content Identification in English and Indo-Aryan Languages and Conversational Hate Speech</article-title>
          , in: FIRE 2021:
          <article-title>Forum for Information Retrieval Evaluation, Virtual Event</article-title>
          ,
          <fpage>13th</fpage>
          -17th
          <source>December</source>
          <year>2021</year>
          , ACM,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Modha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. K.</given-names>
            <surname>Shahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Madhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Satapara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schäfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ranasinghe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nandini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Jaiswal</surname>
          </string-name>
          ,
          <article-title>Overview of the HASOC subtrack at FIRE 2021: Hate Speech and Ofensive Content Identification in English and Indo-Aryan Languages</article-title>
          , in: Working Notes of FIRE 2021 -
          <article-title>Forum for Information Retrieval Evaluation</article-title>
          ,
          <string-name>
            <surname>CEUR</surname>
          </string-name>
          ,
          <year>2021</year>
          . URL: http://ceur-ws.org/.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bird</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Klein</surname>
          </string-name>
          , E. Loper,
          <article-title>Natural language processing with Python: Analyzing text with the natural language toolkit, ”</article-title>
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          , Inc.”,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kalaivani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Thenmozhi</surname>
          </string-name>
          , SSN_NLP_MLRG@
          <string-name>
            <surname>Dravidian-CodeMix-FIRE2020</surname>
          </string-name>
          :
          <article-title>Sentiment Code-Mixed Text Classification in Tamil and Malayalam using ULMFiT.</article-title>
          ,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>528</fpage>
          -
          <lpage>534</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>