<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Sentiment Analysis on Multilingual Code Mixing Text Using BERT-BASE: participation of IRLab@IIT(BHU) in Dravidian-CodeMix and HASOC tasks of FIRE2020</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anita Saroj</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sukomal Pal</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indian Institute of Technology (BHU)</institution>
          ,
          <addr-line>Varanasi</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper discusses our participation in the “Sentiment Analysis in Dravidian-CodeMix”, DravidianCodeMix and “Hate Speech and Ofensive Content Identification in Indo-European Languages”- FIRE 2020 tasks of identifying subjective opinions or reactions on a given topic. Several techniques are applied for sentiment analysis including the recent word embeddings-based methods. BERT, Word2Vec, and ELMo are currently among the most promising and ready-to-use word embedding methods that can convert words into meaningful vectors. We used the BERT_BASE model for sentiment classification of Dravidian-CodeMix data and for HASOC task, our team submitted systems for all the two sub-tasks in three languages - Hindi, English, and German with BERT-based system. We report our approach and results which are promising.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;BERT</kwd>
        <kwd>Classification</kwd>
        <kwd>Dravidian</kwd>
        <kwd>CodeMix</kwd>
        <kwd>Sentiment Analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>and English-Malayalam and build models for classifying “posts” into positive, negative, mixed
feelings and unknown_state, and other_language1][.</p>
      <p>
        In a related note, while the social media helps quickly spread the information in society,
it has also become hotbeds for hate speech and ofensive content. Hate speech is defined as
“language that is used to express anger towards a targeted group or is meant to be derogatory,
to humiliate, or to insult the members of the group [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].” People convicted of using hate speech
are often enforced to face large fines and even imprisonment. These laws should spread to the
internet and social media, leading many sites to create their own provisions against hate speech.
Both Twitter and Facebook have responded to criticisms for not doing enough to prevent hate
speech on their sites by instituting policies to prevent their platforms for attacks on people
based on characteristics like race, ethnicity, gender, and sexual orientation threats of violence
towards others. The HASOC task at FIRE 2020 provides another platform to build and test
robust automatic hate speech and ofensive content identification systems where content could
range from political and religious to caste and gender or any other issues that potentially divide
and polarise a society. It is, therefore, imperative to use such systems that can help automatically
iflter out hate and ofensive contents so that they do not spread in the community. HASOC
FIRE 2020 sets up the task divided into two subtasks: hate speech language identification and
automatic categorization of hate speech types.
      </p>
      <p>The rest of the paper is organized as follows. Sectio2ndescribes the existing work on
Codemix and HASOC content identification respectively. Section 4 focuses on the methodology. In
Section 5 we describe the experiments. Section6 discusses the results of the task. Section7
concludes, discussing the lessons learned.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>There exist a few relevant past literature of research work done in sentiment analysis for the
CodeMix language.</p>
      <p>Garain et al. (2020) 3[] used feature extraction algorithms in conjunction with traditional
machine learning algorithms such as the Support Vector Regression and Grid Search to solve
the Hindi-English codemix sentiment analysis and garnered an1-Fscore of 66.2%. Nguyen and
Dogruoz (2013) [4] analyzed Turkish-Dutch posts from an online chat forum and compared
dictionary-based methods with language models, adding logistic regression and linear-chain
CRF. They achieved the best result with word-level accuracy9o7f.6% and post-level classification
0.89%. Yadav and Chakraborty (2020) 5[] used multilingual and cross-lingual embeddings to
eficiently transfer knowledge from monolingual text to code-mixed text for sentiment analysis
of the code-mixed text. They achieve an F1-score of 0.58 (without a parallel corpus) an0d.62
(with the parallel corpus) on the same benchmark in a zero-shot way compared 0to.68 F1-score
in supervised settings.</p>
      <p>
        The development of social media has witnessed the proliferation of two separate but
connected phenomena: first, they helped to create a more open and connected world, and second,
they contributed to the spread of hate speech and rude behavior6[]. Previous work on
hatespeech language detection and related phenomena has seen diferent system architectures’
deployment with varying performance levels. We can observe three significant waves of the
system: discrete linear models, neural networks, pre-trained language model7s, 8[, 9, 10, 11].
Linear models (like SVM, logistic regression, or ensemble models) are very competitive and
robust methods to identify of-line successfully. An ofensive language that, in many cases,
improves more complex approaches based on neural networks1[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Task definitions</title>
      <p>Although Dravidian CodeMix and HASOC were simultaneously run in FIRE 2020, they attempted
to focus on two related but diferent aspects with clearly distinct task definitions. Below we
describe them.</p>
      <sec id="sec-3-1">
        <title>3.1. Dravidian Code-mix Task</title>
        <p>The task is primarily of sentiment detection from code-mixed social media text. We attempt
the message-level polarity classification task1[3]. Given a YouTube comment, the system must
classify it into positive, negative, neutral, mixed feelings, or not in intended languages. The
track consists of multi-class classification
tasks:• Positive state: Positive state shows the positive opinion, thoughts or attitude of the
person, such as relaxing, happiness, thanking to someone, forgiveness.
• Negative state: Negative state shows the negative opinion, thoughts, or attitude of the
person, such as sad, angry, worried, and violent.
• Mixed feelings: Mixed feelings show the speaker is experiencing both positive and
negative emotions, such as comparing two products.
• Unknown state: There is no clear or implicit indicator of the emotional state of the speaker.</p>
        <p>Examples are liking or membership or asking questions about the release date or the film
dialogue.
• Other language: For Malayalam and Tamil, if the sentence does not contain Malayalam
and Tamil, it is not Malayalam and not Tamil.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. HASOC Task</title>
        <p>The task of identifying hate and ofensive posts were done in two diferent languages: Hindi and
German. Each language has two subtasks. A brief description of each work is given below14][.
• Sub-task A: Identifying Hate, ofensive and profane content. Sub-task A is a
coarsegrained binary classification. Teams are required to classify tweets into two classes: Hate
and Ofensive (HOF) and Non- Hate and ofensive (NOT).</p>
        <p>(NOT) Non Hate-Offensive: The post does not contain any hate speech, profane,
ofensive content.</p>
        <p>(HOF) Hate and Offensive: Post contains Hate, ofensive, and profane content.
• S u b - t a s k B : There is no clear or implicit indicator of the emotional state of the speaker.</p>
        <p>Examples are liking or membership or asking questions about the release date or the film
dialogue.</p>
        <p>H A T E : Posts under this class contain Hate speech content.</p>
        <p>O F F N : Posts under this class contain ofensive content.</p>
        <p>P R F N : These posts contain profane words.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>This section describes the general model and architecture commonly followed for identification
of sentiments in both CodeMix and HASOC tasks. and then segregates them according to the
requirement of the tasks.</p>
      <p>• Preprocessing: We first convert all the texts into lowercase. We also replace all the links
with the word “URL” and all the numbers with the word “number”, remove the leading
and trailing white-spaces, and replace multiple white-spaces between words with a single
whitespace. We remove all the punctuation symbols using a pre-initialized string, string
punctuation available in the string library. After removing non-letters from the data, all
tokens are lemmatized.
• Model Architecture: BERT stands for Bi-directional Encoder Representations from
Transformers [15]. Figure 1 shows the architecture. It is based on a transformer and reads
the input from both directions at once. It uses two training approaches, namely Masked
Language Model and Next Sentence Prediction. In our approach, it is a pre-trained model,
so fine-tuning is done to use it for a specific task.</p>
      <p>We used Rectified Linear Units (ReLU), Sigmoid activation function, and Adam algorithm
as an optimizer. Input data is converted in the form of input representation of BERT.
In the training phase, the model receives a pair of sentences as input and learns to
predict whether the second sentence in the pair is the subsequent sentence in the original
document. The model diferentiates between two sentences during training by adding
a CLS token at the beginning of the first sentence and a SEP token at the end of each
sentence. Approximately15 percent of the words in the input are masked. The input
data is converted into a combination of token embedding plus sentence embedding, the
transformer positional embedding. The transformer encoder reads the input embedding.</p>
      <p>The final embedding is fed into our model which makes the predictions.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments</title>
      <p>In this section, we present a BERT_BASE experiments on Dravidian-CodeMix and HASOC data.</p>
      <sec id="sec-5-1">
        <title>5.1. Dataset</title>
        <sec id="sec-5-1-1">
          <title>5.1.1. Dravidian-CodeMix</title>
          <p>The dataset is created from YouTube comments and shared by the task organizers in a
tabseparated format for a code-mixed dataset for Tamil1[6] and Malayalam1[7]. The dataset
was collected from the comments from the Tamil and Malayalam trailers of movies released in
2019. The sentences written in code-mixed Tamil-English (Tanglish) and Malayalam-English
(Manglish). Figure2 shows the distribution of training instances of the task.
5.1.2. HASOC
HASOC provides a platform for multilingual research on the identification of problematic
content and a data challenge. There are 2 sub-tasks for each language such as English, German
and Hindi, altogether over 10.000 annotated tweets from Twitter. There are five columns in
the CSV file for each language, i.e., tweet_id, text, task1, task2, and ID. Tabl1eshows the data
distribution of the task.</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Experimental Settings</title>
        <p>For both the tasks, we used the released BERT_BASE pre-trained model (Un-cased:12-layer,
768-hidden, 12-heads, 110M parameters). Based on the maximum size of the sentence in the
dataset, performed padding of the sentence to forge the sentence of equal measure. We use
recurrent dropout of 0.5 and Sigmoid as an activation function. The dropout layer, with a rate
of 0.2, is used to avoid overfitting of the model and set the number of epochs to 10. The initial
learning rate is2e−5, and the batch size is8. We used a Softmax activation function at the output
layer. We use the Adam optimizer and categorical cross-entropy loss function for training.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Results</title>
      <sec id="sec-6-1">
        <title>6.1. Dravidian-CodeMix</title>
        <p>The metric for evaluating the systems was as follows. The organizers used F-scores across the
positive, negative, unknown_state, mixed_feelings, and the other_language. The final ranking
was based on the average F-score. Our submitted system garnered an F-score o0f.59% for
English-Tamil and0.60% for English-Malayalam. The detailed results are shown in Ta3b.lWee
compare our result with the team SRJ result, which is rank1. We found that our submitted
result is0.06, 0.14 (F-score) low compared to rank1 for English-Tamil, English-Malayalam,
respectively.
6.2. HASOC</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>The metric for evaluating the systems was as follows. The organizers used F-scores across the
sub-task A and sub-task B. The final ranking was based on the average F-score. Our submitted
system garnered an F-score of0.5028% for Hindi sub-task A and0.3840% for German sub-task
A. The detailed results are shown in Tabl3e.</p>
      <p>In this task, we attempted to explore the sentiment analysis of code-mixed English-Tamil and
English-Malayalam data while participating in Dravidian-CodeMix - F2I0R2E0. Our system
was based on the BERT_BASE model. The system, when evaluated by the organizers, garnered
an  1-score of 0.59 for English-Tamil and0.6 for English-Malayalam. There was a choice of
developing an unconstrained approach, but we only used the provided data to train the system.
As a future work, we would like to use other embeddings: ELMo, BER_TLARGE, and other
deep learning models. We attempted to explore hate speech and ofensive content identification
in the Indo-European language while participating in HASOC - FIRE 2020. Our system was
based on the BERT_Base model for multi-class classification. The system, when evaluated by
the organizers, achieved</p>
      <p>1
-score of 0.50 for Hindi. There was a choice of developing an unconstrained approach, but here
also we restricted ourselves only to the provided data for training. As future work, we want to
work on multitask learning with other embeddings.
[3] A. Garain, S. K. Mahata, D. Das, Junlp@ semeval-2020 task 9: Sentiment analysis of
hindi-english code mixed data, arXiv preprint arXiv:2007.12561 (2020).
[4] D. Nguyen, A. S. Doğruöz, Word level language identification in online multilingual
communication, in: Proceedings of the 2013 Conference on Empirical Methods in Natural
Language Processing, Association for Computational Linguistics, Seattle, Washington,
USA, 2013, pp. 857–862. URL: https://www.aclweb.org/anthology/D13-108.4
[5] S. Yadav, T. Chakraborty, Unsupervised sentiment analysis for code-mixed data, arXiv
preprint arXiv:2001.11384 (2020).
[6] D. Colla, T. Caselli, V. Basile, J. Mitrovic, M. Granitzer, Grupato at semeval-2020 task
12: Retraining mbert on social media and fine-tuned ofensive language models, in:
Proceedings of the International Workshop on Semantic Evaluation (SemEval), 2020.
[7] A. Cimino, L. De Mattei, F. Dell’Orletta, Multi-task learning in deep neural networks at
evalita 2018, Proceedings of the 6th evaluation campaign of Natural Language Processing
and Speech tools for Italian (EVALITA’18) (2018) 86–95.
[8] M. Karan, J. Šnajder, Cross-domain detection of abusive language online, in: Proceedings
of the 2nd workshop on abusive language online (ALW2), 2018, pp. 132–137.
[9] P. Liu, W. Li, L. Zou, Nuli at semeval-2019 task 6: Transfer learning for ofensive language
detection using bidirectional transformers, in: Proceedings of the 13th International
Workshop on Semantic Evaluation, 2019, pp. 87–91.
[10] J. Mitrović, B. Birkeneder, M. Granitzer, nlpup at semeval-2019 task 6: a deep neural
language model for ofensive language detection, in: Proceedings of the 13th International
Workshop on Semantic Evaluation, 2019, pp. 722–726.
[11] Z. Waseem, D. Hovy, Hateful symbols or hateful people? predictive features for hate
speech detection on twitter, in: Proceedings of the NAACL student research workshop,
2016, pp. 88–93.
[12] J. P. Montani, P. Schüller, Tuwienkbs at germeval 2018: German abusive tweet detection,
in: 14th Conference on Natural Language Processing KONVENS, volume 2018, 2018, p. 45.
[13] B. R. Chakravarthi, R. Priyadharshini, V. Muralidaran, S. Suryawanshi, N. Jose, J. P. Sherly,
Elizabeth McCrae, Overview of the track on Sentiment Analysis for Dravidian Languages
in Code-Mixed Text, in: Working Notes of the Forum for Information Retrieval Evaluation
(FIRE 2020). CEUR Workshop Proceedings. In: CEUR-WS. org, Hyderabad, India, 2020.
[14] T. Mandl, S. Modha, G. K. Shahi, A. K. Jaiswal, D. Nandini, D. Patel, P. Majumder, J. Schäfer,
Overview of the HASOC track at FIRE 2020: Hate Speech and Ofensive Content
Identification in Indo-European Languages), in: Working Notes of FIRE 2020 - Forum for
Information Retrieval Evaluation, CEUR, 2020.
[15] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional
transformers for language understanding, arXiv preprint arXiv:1810.04805 (2018).
[16] B. R. Chakravarthi, V. Muralidaran, R. Priyadharshini, J. P. McCrae, Corpus creation for
sentiment analysis in code-mixed Tamil-English text, in: Proceedings of the 1st Joint
Workshop on Spoken Language Technologies for Under-resourced languages (SLTU)
and Collaboration and Computing for Under-Resourced Languages (CCURL), European
Language Resources association, Marseille, France, 2020, pp. 202–210. URhLt:tps://www.
aclweb.org/anthology/2020.sltu-1.2.8
[17] B. R. Chakravarthi, N. Jose, S. Suryawanshi, E. Sherly, J. P. McCrae, A sentiment analysis
dataset for code-mixed Malayalam-English, in: Proceedings of the 1st Joint Workshop on
Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration
and Computing for Under-Resourced Languages (CCURL), European Language Resources
association, Marseille, France, 2020, pp. 177–184. URLh:ttps://www.aclweb.org/anthology/
2020.sltu-1.25.
Data distribution of HASOC task</p>
      <p>Language</p>
      <p>Hindi
English
German</p>
      <p>Training</p>
      <p>Testing
2963
3794
2452
663
526
526</p>
      <sec id="sec-7-1">
        <title>Team name</title>
      </sec>
      <sec id="sec-7-2">
        <title>Model</title>
        <p>BERT_BASE</p>
      </sec>
      <sec id="sec-7-3">
        <title>Hindi</title>
        <p>Sub-task A
0.5028</p>
      </sec>
      <sec id="sec-7-4">
        <title>German</title>
        <p>Sub-task A
0.3840
Classifier result of HASOC dataset at F1 Macro average</p>
        <p>Precision
0.59
0.64</p>
        <p>Rank
7
1
Classifier result of Dravidian-CodeMix dataset at Precision, Recall, F-score and Accuracy in %.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Muralidaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          , N. Jose,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Sherly</surname>
          </string-name>
          ,
          <string-name>
            <surname>Elizabeth</surname>
            <given-names>McCrae</given-names>
          </string-name>
          ,
          <article-title>Overview of the track on Sentiment Analysis for Dravidian Languages in Code-Mixed Text</article-title>
          ,
          <source>in: Proceedings of the 12th Forum for Information Retrieval Evaluation</source>
          ,
          <source>FIRE '20</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Davidson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Warmsley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Macy</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Weber</surname>
          </string-name>
          ,
          <article-title>Automated hate speech detection and the problem of ofensive language</article-title>
          ,
          <source>arXiv preprint arXiv:1703.04009</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>English-Tamil Recall</surname>
          </string-name>
          F-score
          <source>0.61 0.59 0.67 0.65 Precision 0.68 0</source>
          .74 English
          <string-name>
            <surname>-Malayalam Recall</surname>
          </string-name>
          F-score
          <source>0.6 0.6 0.75 0</source>
          .
          <fpage>74</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>