<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the track on HASOC-Ofensive Language Identification-DravidianCodeMix</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bharathi Raja Chakravarthi</string-name>
          <email>bharathi.raja@insight-centre.org</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anand Kumar M</string-name>
          <email>m_anandkumar@nitk.edu.in</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John P. McCrae</string-name>
          <email>john.mccrae@insight-centre.org</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>B Premjith</string-name>
          <email>b_premjith@cb.amrita.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>K.P Soman</string-name>
          <email>kp_soman@amrita.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Mandl</string-name>
          <email>mandl@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Computational Engineering and Networking (CEN), Amrita School of Engineering</institution>
          ,
          <addr-line>Coimbatore, Amrita Vishwa Vidyapeetham</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Information Technology, National Institute of Technology Karnataka Surathkal</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Information Science, University of Hildesheim</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Insight SFI Research Centre for Data Analytics, Data Science Institute, National University of Ireland</institution>
          ,
          <addr-line>Galway</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present the results and main findings of the HASOC-Ofensive Language Identification on code mixed Dravidian languages. The task featured two tasks. Task 1 is about ofensive language identification in Malayalam language where the comment were written in both native script and Latin script. Task 2 is about ofensive language identification in Tamil and Malayalam languages where the comments were written in Latin script (non-native script). For both the task, given a comment the participants should develop a system to classify the text into ofensive or not-ofensive. In total 96 participants participated and 12 participants submitted the papers. In this paper, we present the task, data, the results and discuss the system submission and methods used by participants.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Dravidian languages</kwd>
        <kwd>Tamil</kwd>
        <kwd>Malayalam</kwd>
        <kwd>Ofensive language identification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        considered as under-resourced [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. Malayalam and Tamil have their own scripts for writing such as
Malayalam script and Tamil script. However, social media users uses Latin script to write it in online.
Therefore, most of the social media comments/posts in Malayalam and Tamil are available in Latin
script or in code-mixed form [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. The research in ofensive language identification from social media
comments/posts is still in its infancy, particularly in code-mixed Indian languages. The unavailability
of gold standard corpus and limited scientific study in the area motivates the introduction of
organizing this shared task. Shared tasks such as Task-6 in SemEval-2019 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], Task-12 in SemEval-2020 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
and Hate Speech and Ofensive Content Identification in Indo-European Languages (HASOC)-2019
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] inspired us to organize this event for Dravidian languages.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Task Description</title>
      <p>The goal of this task is to identify ofensive language from a code-mixed dataset of comments/posts
in Dravidian Languages (Malayalam-English and Tamil-English) collected from social media. The
comment/post may contain more than one sentence but the average sentence length of the corpora
is 1. Each comment/post is annotated with ofensive language label at the comment/post level. The
task-1 dataset also has class imbalance problems depicting real-world scenarios. The participants
were provided with development, training and test dataset.</p>
      <p>Task1:</p>
      <p>This is a message-level label classification task. Given a YouTube comment in code-mixed
Malayalam, systems have to classify it into ofensive or not-ofensive.</p>
      <p>Task2:</p>
      <p>This is a message-level label classification task. Given a tweet or Youtube comments in Tanglish
and Manglish (Tamil and Malayalam using written using Roman Characters), systems have to classify
it into ofensive or not-ofensive.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset Description</title>
      <p>
        3.1. Task 1 data
For the Task 1, we downloaded data from YouTube comments. The comments were downloaded from
movie trailers during 2019. All the comment from those movie trailers were downloaded using a
YouTube comment scrapper 1. We utilized these comments to make a dataset for ofensive language
identification classification dataset. The dataset contains all types of code-mixing such as mixing the
scripts of Malayalam script and Latin script, mixing at the word level, mixing at inter-sentential and
intra-sential [
        <xref ref-type="bibr" rid="ref12">12, 13, 14</xref>
        ].
3.2. Task 2 data
The Tamil code-mixed dataset for Task 2 was collected from the Twitter tweets and comments on
the Helo App. We have considered only the comments/posts in the Latin characters. Malayalam
dataset for Task 2 has collected from YouTube comments. The training dataset for the Tanglish and
Manglish used for the Task 2 contained 4000 comments. In Tamil, 2997 comments were collected from
Twitter and 1003 are from Helo App. Out of 4000 comments, 1980 comments are ofensive and 2020
comments are not ofensive. The test dataset consists of 940 comments on which 475 are ofensive
1https://github.com/egbertbouman/youtube-comment-downloader
and 465 are not ofensive. Malayalam training set contains 1953 ofensive comments and 2047
notofensive comments, whereas the test set consist of 512 ofensive and 488 not-ofensive comments.
The comments were annotated manually and verified by experts. The comments were annotated
with two tags - OFF (ofensive comment) and NOT (not-ofensive comment). The dataset was given
to the participants in CSV format.
      </p>
      <p>The baseline system for Task 2 for Manglish and Tanglish data used Support Vector Machine (SVM)
classifier [15] with Term Frequency - Inverse Document Frequency (TD-IDF) features. This system
didn’t undergo any preprocessing and analyzed the data at the word level. The SVM classifier was
trained using a linear kernel and a regularization parameter (C) value = 1. Comments are evenly
distributed in both classes, and hence the baseline model didn’t consider any approach to make classifier
a balanced one. The baseline model for Malayalam achieved the F-score of 0.68, precision score of
0.69 and the recall score of 0.68 and Tamil achieved the F-score of 0.89, precision and recall of 0.89.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation</title>
      <p>All teams were allowed a total of 3 submission per Task. Participants had to submit prediction for
Tasks in language of their choosing. A submission’s final score for each Task was computed as the
highest score of the submission. The systems were evaluated on precision, recall and F1-score. This
takes into account the varying degrees of importance of each class in the dataset. We used a
classification report tool from Scikit learn2.</p>
      <p>Precision =</p>
      <p>Recall =
 
 
 
+  
  +</p>
      <p>Precision ∗ Recall
F-Score = 2 ∗ Precision + Recall
(1)
(2)
(3)</p>
    </sec>
    <sec id="sec-5">
      <title>5. Methodology</title>
      <p>We received a total of 12 submissions for Task 1 for Malayalam, 13 submission for Task 2 for
Malayalam, and 12 submission for Task 2 for Tamil. The systems were evaluated based on F1 scores and
a rank list was prepared. Table 1, Table 2 and Table 3 show the rank lists of Malayalam and Tamil
tasks. We briefly describe below the methodologies used by the participating teams who submitted
the paper.</p>
      <p>• SivaSai@BITS [16]: The authors proposed a novel and flexible approach of selective translation
and transliteration to be able to reap better results out of fine-tuning and ensemble multilingual
transformer networks like XLM-RoBERTa and mBERT.
• CENmates [19]: The participants used TF-IDF vectors along with character level n-grams as
features to the proposed system for system development. They developed and evaluated four
systems consisting of logistic regression, XGBoost, long short-term memory networks, and
attention networks. They noted that simple TD-IDF approach with character level n-gram
features using machine learning classifiers was producing good results almost the same as the
results obtained with deep learning-based classifiers.
2https://scikit-learn.org/stable/modules/generated/sklearn.metrics.classification_report.html</p>
      <sec id="sec-5-1">
        <title>TeamName</title>
      </sec>
      <sec id="sec-5-2">
        <title>SivaSai@BITS [16]</title>
      </sec>
      <sec id="sec-5-3">
        <title>IIITG-ADBU [17]</title>
      </sec>
      <sec id="sec-5-4">
        <title>CFILT-IITBOMBAY</title>
      </sec>
      <sec id="sec-5-5">
        <title>SSNCSE-NLP [18]</title>
      </sec>
      <sec id="sec-5-6">
        <title>CENMates [19]</title>
        <p>NIT-AI-NLP [20]
YUN [21]
Zyy1510 [22]</p>
      </sec>
      <sec id="sec-5-7">
        <title>Gauravarora [23]</title>
        <p>WLV-RIT [24]</p>
      </sec>
      <sec id="sec-5-8">
        <title>Kjdong( only not) Ajees [25]</title>
      </sec>
      <sec id="sec-5-9">
        <title>TeamName</title>
      </sec>
      <sec id="sec-5-10">
        <title>CENmates [19]</title>
      </sec>
      <sec id="sec-5-11">
        <title>SivaSai [16]</title>
      </sec>
      <sec id="sec-5-12">
        <title>KBCNMUJAL [26]</title>
      </sec>
      <sec id="sec-5-13">
        <title>IIITG-ABDU [17]</title>
      </sec>
      <sec id="sec-5-14">
        <title>SSNCSE-NLP [18]</title>
      </sec>
      <sec id="sec-5-15">
        <title>Gauravarora [23] CFILT [27] NITP [20] Ajees [25]</title>
      </sec>
      <sec id="sec-5-16">
        <title>Baseline YUN [21] Zyy1510 [22] CUSAT [28]</title>
        <p>• CUSATNLP [28]: The participants did extensive preprocessing and removed the hashtags, URLs,
and emojis. They used the simple LSTM layers, a recurrent dropout (0.2). The system results
were very low compared to other participants.
• KBCNMUJAL [26]: The participants as well used the common preprocessing techniques to
simplify the text messages in addition to that they also remove the special tags. They performed
elaborate analysis with classical machine learning model with char n-gram and word n-gram as
features. They also reported results for ensemble models. They ranked second for Malayalam
and third for Tamil in Task 2.
• Gauravarora [23]: Participants proposed pre-training ULMFiT on synthetically generated
codemixed data, generated by modelling code-mixed data generation as a Markov process using
Markov chains. Their model achieved 0.88 weighted F1-score for code-mixed Tamil-English</p>
      </sec>
      <sec id="sec-5-17">
        <title>TeamName</title>
      </sec>
      <sec id="sec-5-18">
        <title>SivaSaiBITS [16]</title>
      </sec>
      <sec id="sec-5-19">
        <title>SSNCSE-NLP [18]</title>
      </sec>
      <sec id="sec-5-20">
        <title>Gauravarora [23]</title>
      </sec>
      <sec id="sec-5-21">
        <title>KBCNMUJAL [26]</title>
      </sec>
      <sec id="sec-5-22">
        <title>IIITG-ADBU [17]</title>
        <p>Zyy1510 [22]</p>
      </sec>
      <sec id="sec-5-23">
        <title>CENmates [19]</title>
        <p>CFILT [27]
YUN [21]
NIT-AI-NLP [20]</p>
      </sec>
      <sec id="sec-5-24">
        <title>Baseline</title>
        <p>Ajees [25]
• SSNCSE [18]: The participants used char n-gram, TFIDF and fine-tuned BERT in combination
with machine learning models such as MLP, Random Forest and Naive Bayes.
• NITP-AI-NLP [20]: The authors explored deep learning models such as attention-based Long
Short Term Memory (LSTM), Convolution Neural Network (CNN), and machine learning
models such as support vector machine, Logistic regression, Random forest, and Naive Bayes. They
have hown that the use of character N-gram Term Frequency-Inverse Document Frequency
(TF-IDF) features plays a promising role in identifying ofensive social media posts.
• YUN [21]: The authors proposed an ensemble model which makes full use of the information
of rich sequential patterns. More precisely, the proposed model contains a self-attention based
on the BiLSTM and the sub-word representation learning. Experimental results of their model
on the Malayalam-English of Task 1, Tamil-English and Malayalam-English of Task 2 have
achieved the F1 values of 0.93, 0.85 and 0.67, respectively, and ranked 3rd, 5th, 9th, respectively.
• Zyy1510 [22]: ensemble model combines with diferent models to improve the F-1 value of
the framework. The ensemble model is a combination of a BiLSTM (Bidirectional LSTM), an
LSTM+Convolution, and a CNN (Convolution Neural Network) model. The proposed model
have achieved an F-1 of 0.93 (ranked 3 rd ) in Malayalam-English of task1, and F-1 of 0.87 (ranked
3 rd ) and 0.67 (ranked 9 th ) in Tamil-English and Malayalam-English of task2, respectively
• Ajees [25]: Mainly three types of machine learning models were experimented using two types
of word embedding techniques. The first architecture was a simple MLP classifier using two
hidden layers. The second one was a combination of CNN-BiLSTM network with four
convolutional layers. And the third one was a BiLSTM stack with two hidden layers. In order to
represent the individual words in the comments, a simple CountVectorizer as well as BERT was
used. CountVectorizer is used to convert text data to a vector of token counts. It also provides
the pre-processing of text data before generating the vector representation. This capability
makes it highly flexible towards feature representation for text processing. On the other hand,
BERT is a pre-trained NLP model that provides contextualized word embeddings. Static word
embedding techniques provide the same vector for polysemous words without proper
consideration of their context. Whereas the dynamic embedding techniques like ELMo and BERT can
consider the context of words before generating their embeddings. BERT word representations
take the entire input sentence into the equation for calculating the word embeddings.
• WLV-RIT [24]: applying cross-lingual contextual word embeddings and transfer learning to
make predictions to Malayalam data. They further improve the results using various fine tuning
strategies.
• IIITG-ADBU [17] paper presents the results obtained by our SVM and XLM-RoBERTa based
classifiers in the shared task.
• CFILT[27] presents an ensemble of multilingual BERT models for this task and devise a novel
training strategy involving data augmentation using random transliteration. They achieve an
F-score of 0.95 for hate speech and ofensive content detection on the Malayalam code-mixed
YouTube comments test data in task 1. In task 2, they achieve F-scores of 0.86 and 0.72
respectively for hate speech and ofensive content detection on Tamil and Malayalam code-mixed
Twitter test data.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Results and Discussion</title>
      <p>In Malayalam Task 1, teams SivaSai@BITS and IIITG-ADBU shared the first position with an F1-score
of 0.95. These two systems achieved precision and recall score of 0.95. Teams from CFILT-IITBOMBAY
and SSNCSE-NLP achieved the second position with an F-score of 0.94. The top four teams attained
Fscore higher than 0.90. The diference between the evaluation scores of top teams is minuscule. Table
1 presents the results of the Malayalam Task 1. Teams placed in the first and second position utilized
transformer-based model for classification of YouTube comments into OFF and NOT. Transliteration
of Romanized text into the native script is also found to be efective in this method. Another important
fact visible from the results is the Support Vector Machine classifiers with TF-IDF features also reach
top positions. Other systems submitted to the task use deep learning models using Bidirectional
LSTM, LSTM, CNN and ULMFiT.</p>
      <p>In Malayalam Task 2, CENmates reached the first position with an F-score of 0.78. Teams
SivaSai@BITS and KBCNMUJAL bagged the second place, and their F-score was 0.77. The scores of the
top five teams are close. Team CENmates used TF-IDF features with character n-gram as features for
classification using machine learning algorithms. SivaSai@BITS used the same approach followed
for Malayalam Task 1 for this task also. Team KBCNMUJAL used character and word n-grams
features for with machine learning classifiers. Other teams used transformer-based models and Deep
Learning-based models. Table 2 presents the result of Malayalam Task 2.</p>
      <p>In Tamil Task 2, team SivaSai@BITS placed in the first position with an F-score of 0.90. They used
the transformer-based model for this task also. Team SSNCSE-NLP grabbed the second position, and
they scored an F-score of 0.88. They used TF-IDF with character n-gram features for classification.
Gauravarora, who also came second position followed a pre-trained ULMFiT for the classification.
Three teams reached third position with an F-score of 0.87. These three teams used entirely diferent
features and classifiers for the prediction task. Team KBCNMUJAL uses character n-gram and word
ngram features for representing text with ensemble models for classification. Team IIITG-ADBU used
the same model used for other tasks for this task also. Team Zyy1510 used an ensemble of BiLSTM,
LSTM+Convolution and a Convolution for the classification of social media texts into OFF and NOT.
Table 3 presents the results of Tamil Task 2.</p>
      <p>When we analyse the models submitted to the Tasks, most of them used either transformer-based
models or conventional machine learning classifier with TF-IDF features. The performance of the deep
learning models such as Bidirectional LSTM, LSTM, and CNN was not up to the mark.
Transformerbased models used BERT for generating the embedding.</p>
      <p>Figure 1 shows the box-plots of the performance of the systems for Malayalam Task 1, Malayalam
Task 2 and Tamil Task 2, respectively. From the Figures, it is clear that Median of the F-score of all
the submissions is close to the top score in Malayalam Task 1 and Malayalam Task 2.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>We presented the overview and results of the shared task on HASOC-Ofensive language detection in
Dravidian languages. A wide range of systems were evaluated on two task in two languages relying
on a thoroughly annotated dataset. The task setup provided an opportunity to test models on
codemixed Dravidian languages with mixed scripts. NLP for code-mixed texts is very challenging and
under-resourced setting also make it even more challenging. We found that many systems were
based on transformers and pre-trained embedding based systems. We that this shared task makes a
lasting contribution to the research in Dravidian languages.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This publication is the outcome of the research supported in part by a research grant from Science
Foundation Ireland (SFI) under Grant Number SFI/12/RC/2289_P2 (Insight_2), co-funded by the
European Regional Development Fund as well as by the EU H2020 programme under grant agreements
825182 (Prêt-à-LLOD), and Irish Research Council grant IRCLA/2017/129 (CARDAMOM-Comparative
Deep Models of Language for Minority and Historical Languages).</p>
      <p>We would like to thank Mrs.Sharmila Devi V, PhD scholar, Karpagam College of Engineering,
Coimbatore for sharing the task 2 Tamil code-mixed ofensive dataset.
guage Technologies for Under-resourced languages (SLTU) and Collaboration and Computing
for Under-Resourced Languages (CCURL), European Language Resources association, Marseille,
France, 2020, pp. 177–184. URL: https://www.aclweb.org/anthology/2020.sltu-1.25.
[13] B. R. Chakravarthi, V. Muralidaran, R. Priyadharshini, J. P. McCrae, Corpus creation for
sentiment analysis in code-mixed Tamil-English text, in: Proceedings of the 1st Joint Workshop on
Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration and
Computing for Under-Resourced Languages (CCURL), European Language Resources
association, Marseille, France, 2020, pp. 202–210. URL: https://www.aclweb.org/anthology/2020.sltu-1.
28.
[14] A. Hande, R. Priyadharshini, B. R. Chakravarthi, KanCMD: Kannada codemixed dataset for
sentiment analysis and ofensive language detection, in: Proceedings of the Third Workshop
on Computational Modeling of People’s Opinions, Personality, and Emotions in Social Media,
Barcelona, Spain, 2020.
[15] K. Soman, R. Loganathan, V. Ajay, Machine learning with SVM and other kernel methods, PHI</p>
      <p>Learning Pvt. Ltd., 2009.
[16] S. Sai, Y. Sharma, Siva@HASOC-Dravidian-CodeMix-FIRE-2020: Multilingual Ofensive Speech</p>
      <p>Detection in Code-mixed and Romanized Text, in: FIRE (Working Notes), 2020.
[17] A. Baruah, K. A. Das, F. A. Barbhuiya, K. Dey,
IIITG-ADBU@HASOC-Dravidian-CodeMixFIRE2020: Ofensive Content Detection in Code-Mixed Dravidian Text, in: FIRE (Working
Notes), 2020.
[18] N. N. Balaji, B. Bharathi, SSNCSE-NLP@HASOC-Dravidian-CodeMix- FIRE2020: Ofensive
Language Identification on Multilingual Code Mixing Text, in: FIRE (Working Notes), 2020.
[19] V. P V, P. Ramanan, R. Devi G, CENMates@HASOC-Dravidian-CodeMix-FIRE2020: Ofensive
Language Identification on Code-mixed Social Media Comments, in: FIRE (Working Notes),
2020.
[20] S. Kumar, Abhinav adn Saumya, J. P. Singh,
NITP-AI-NLP@HASOC-Dravidian-CodeMixFIRE2020: A Machine Learning Approach to Identify Ofensive Languages from Dravidian
CodeMixed Text, in: FIRE (Working Notes), 2020.
[21] K. Dong, YUN@HASOC-Dravidian-CodeMix-FIRE2020: A Multi-component Sentiment
Analysis Model for Ofensive Language Identification, in: FIRE (Working Notes), 2020.
[22] Y. Zhu, X. Zhou, Zyy1510@HASOC-Dravidian-CodeMix-FIRE2020: An Ensemble Model for</p>
      <p>Ofensive Language Identification, in: FIRE (Working Notes), 2020.
[23] G. Arora, Gauravarora@HASOC-Dravidian-CodeMix- FIRE2020: Pre-training ULMFiT on
Synthetically Generated Code-Mixed Data for Hate Speech Detection, in: FIRE (Working Notes),
2020.
[24] T. Ranasinghe, M. Zampieri, WLV-RIT @ HASOC 2020: Ofensive Language Identification in</p>
      <p>Code-switched Texts, in: FIRE (Working Notes), 2020.
[25] A. A P, Ajees@HASOC-Dravidian-CodeMix-FIRE2020, in: FIRE (Working Notes), 2020.
[26] V. Pathak, M. Joshi, P. Joshi, M. Mundada, T. Joshi,
KBCNMUJAL@HASOC-Dravidian-CodeMixFIRE2020: Using Machine Learning for Detection of Hate Speech and Ofensive Codemix Social
Media text, in: FIRE (Working Notes), 2020.
[27] P. Singh, P. Bhattacharyya, CFILT IIT Bombay@HASOC-Dravidian-CodeMix FIRE 2020:
Assisting ensemble of transformers with random transliteration, in: FIRE (Working Notes), 2020.
[28] S. Renjit, CUSAT-NLP@HASOC-Dravidian-CodeMix-FIRE2020: Identifying Ofensive Language
from Manglish Tweets, in: FIRE (Working Notes), 2020.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Jose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sherly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>A survey of current datasets for code-switching research</article-title>
          ,
          <source>in: 2020 6th International Conference on Advanced Computing and Communication Systems (ICACCS)</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P. V.</given-names>
            <surname>Veena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. P.</given-names>
            <surname>Soman</surname>
          </string-name>
          ,
          <article-title>An efective way of word-level language identification for code-mixed facebook comments using word-embedding via character-embedding</article-title>
          ,
          <source>in: 2017 International Conference on Advances in Computing, Communications and Informatics (ICACCI)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>1552</fpage>
          -
          <lpage>1556</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICACCI.
          <year>2017</year>
          .
          <volume>8126062</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Rajasekaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>McGuinness</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. E.O</given-names>
            <surname>'Connor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Bilingual lexicon induction across orthographically-distinct under-resourced Dravidian languages</article-title>
          ,
          <source>in: Proceedings of the Seventh Workshop on NLP for Similar Languages, Varieties and Dialects</source>
          , Barcelona, Spain,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vegupatti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Named entity recognition for code-mixed Indian corpus using meta embedding</article-title>
          ,
          <source>in: 2020 6th International Conference on Advanced Computing and Communication Systems (ICACCS)</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>WordNet gloss translation for under-resourced languages using multilingual neural machine translation</article-title>
          ,
          <source>in: Proceedings of the Second Workshop on Multilingualism at the Intersection of Knowledge Bases and Machine Translation, European Association for Machine Translation</source>
          , Dublin, Ireland,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          . URL: https://www.aclweb.org/anthology/W19-7101.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <article-title>Leveraging orthographic information to improve machine translation of under-resourced languages</article-title>
          ,
          <source>Ph.D. thesis, NUI Galway</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Comparison of Diferent Orthographies for Machine Translation of Under-Resourced Dravidian Languages</article-title>
          ,
          <source>in: 2nd Conference on Language, Data and Knowledge (LDK</source>
          <year>2019</year>
          ), volume
          <volume>70</volume>
          of OpenAccess Series in Informatics (OASIcs),
          <source>Schloss Dagstuhl-Leibniz-Zentrum fuer Informatik</source>
          , Dagstuhl, Germany,
          <year>2019</year>
          , pp.
          <volume>6</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          :
          <fpage>14</fpage>
          . URL: http://drops.dagstuhl.de/opus/volltexte/2019/10370. doi:
          <volume>10</volume>
          .4230/OASIcs.LDK.
          <year>2019</year>
          .
          <volume>6</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stearns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jayapal</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. S</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zarrouk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Multilingual multimodal machine translation for Dravidian languages utilizing phonetic transcription</article-title>
          ,
          <source>in: Proceedings of the 2nd Workshop on Technologies for MT of Low Resource Languages, European Association for Machine Translation</source>
          , Dublin, Ireland,
          <year>2019</year>
          , pp.
          <fpage>56</fpage>
          -
          <lpage>63</lpage>
          . URL: https://www.aclweb.org/anthology/W19-6809.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Farra</surname>
          </string-name>
          , R. Kumar, SemEval
          <article-title>-2019 task 6: Identifying and categorizing ofensive language in social media (OfensEval)</article-title>
          ,
          <source>in: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Minneapolis, Minnesota, USA,
          <year>2019</year>
          , pp.
          <fpage>75</fpage>
          -
          <lpage>86</lpage>
          . URL: https://www.aclweb.org/anthology/ S19-2010. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>S19</fpage>
          -2010.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Atanasova</surname>
          </string-name>
          , G. Karadzhov,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mubarak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Derczynski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitenis</surname>
          </string-name>
          , Ç. Çöltekin, Semeval-2020 task 12:
          <article-title>Multilingual ofensive language identification in social media</article-title>
          (ofenseval
          <year>2020</year>
          ), arXiv preprint arXiv:
          <year>2006</year>
          .
          <volume>07235</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Modha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mandlia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <source>Overview of the HASOC track at FIRE</source>
          <year>2019</year>
          :
          <article-title>Hate speech and ofensive content identification in Indo-European languages</article-title>
          ,
          <source>in: Proceedings of the 11th Forum for Information Retrieval Evaluation</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>14</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          , N. Jose,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sherly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>A sentiment analysis dataset for code-mixed Malayalam-English</article-title>
          ,
          <source>in: Proceedings of the 1st Joint Workshop on Spoken Lan-</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>