<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IRLab@IITBHU@Dravidian-CodeMix-FIRE2020: Sentiment Analysis for Dravidian Languages in Code-Mixed Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Supriya Chanda</string-name>
          <email>supriyachanda.rs.cse18@itbhu.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sukomal Pal</string-name>
          <email>spal.cse@itbhu.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indian Institute of Technology (BHU)</institution>
          ,
          <addr-line>Varanasi</addr-line>
          ,
          <country country="IN">INDIA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the IRlab@IITBHU system for the Dravidian-CodeMix - FIRE 2020: Sentiment Analysis for Dravidian Languages pairs Tamil-English (TA-EN) and Malayalam-English (ML-EN) in Code-Mixed text. We submitted three models for sentiment analysis of code-mixed TA-EN and MA-EN datasets. Run-1 was obtained from the BERT and Logistic regression classifier, Run-2 used the DistilBERT and Logistic regression classifier, and Run-3 used the fastText model for producing the results. Run-3 outperformed Run-1 and Run-2 for both the datasets. We obtained an  1-score of 0.58, rank 8/14 in TA-EN language pair and for ML-EN, an  1-score of 0.63 with rank 11/15.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Code Mixed</kwd>
        <kwd>Malayalam</kwd>
        <kwd>Tamil</kwd>
        <kwd>BERT</kwd>
        <kwd>fastText</kwd>
        <kwd>Sentiment Analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Internet and digitization enabled people express their views, sentiments, opinions through blog posts,
online forums, product review websites, and diferent social media. Millions of people from diferent
linguistic and cultural backgrounds use social networking sites like Facebook, Twitter, LinkedIn, and
YouTube to express their emotions, opinions, and share views on diferent issues that matter in their
lives. As a large number of Indian users can speak multiple languages proficiently (at least two: native
languages like Malayalam, Tamil, Hindi, and English), an unplanned switching between languages
often happens unconsciously. Even though many languages have their own scripts, social media users
often use non-native scripts, usually Roman script, because of socio-linguistics reasons. This
phenomenon is called code-mixing and is defined as “the embedding of linguistic units such as phrases,
words and morphemes of one language into an utterance of another language" (Myers-Scotton[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]).
Code-mixed data is generally observed in a place of informal communication like social media. The
data can be easily extracted from social media sources using diferent APIs. Sentiment analysis (SA)
on social media text has become an important research task in academia and industry in the past
two decades. SA helps understand people’s opinion from movie/product reviews, and thus help take
decision to improve customer satisfaction through advertisement and marketing.
      </p>
      <p>
        The shared task [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] here aims to identify sentiment polarity of the code-mixed data of YouTube
comments in Dravidian Language pairs (Malayalam-English [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and Tamil-English [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]) collected from
social media. In the past few years, there have been multiple attempts to process code-mixed data, and
a shared task on sentiment analysis of code-mixed Indian languages[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] was organized in ICON 2017.
However, the freely available data apart from Hindi-English and Bengali-English are still limited in
Indian languages, although some other languages like English-Spanish and Chinese-English datasets
are available for research.
      </p>
      <p>The rest of the paper is organized as follows. Section 2 describes the dataset, pre-processing and
processing techniques. In Section 3, we report our results and analysis. Finally we conclude in Section
4.</p>
    </sec>
    <sec id="sec-2">
      <title>2. System Description</title>
      <sec id="sec-2-1">
        <title>2.1. Datasets</title>
        <p>The Dravidian-CodeMix shared task1 organizers provided a dataset that consists of 15,744
TamilEnglish and 6,739 Malayalam-English YouTube video comments. The statistics of training,
development, and test data corpus collection and their class distribution are shown in Table 1. Here, each
comment is annotated by six (for ML-EN) and eleven (for TA-EN) independent annotators. An
interannotator agreement score of 0.6 with Krippendorf’s alpha is obtained for the Tamil-English dataset,
and score of 0.8 with Krippendorf’s alpha for the Malayalam-English dataset. Some comment
examples from the training dataset (Tamil-English) are shown in Table 2. The dataset provided sufers
from general problems of social media data, particularly code-mixed data. The sentences are short
with lack of well-defined grammatical structures, and many spelling mistakes.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Data Pre-processing</title>
        <p>The YouTube comment dataset used in this work is already labelled into five categories: Positive,
Negative, Mixed_feelings, unknown_state and not-Tamil or not-Malayalam. Our pre-processing of
comments includes the following steps:
• Removal of extended words: number of words which have one or more contiguous repeating
characters 2
• Removal of exclamations and other punctuation
• Removal of non-ASCII characters, all the emoticons, symbols, numbers, special characters.</p>
        <sec id="sec-2-2-1">
          <title>1https://dravidian-codemix.github.io/2020/index.html 2https://github.com/SupriyaChanda/Dravidian-CodeMix-FIRE2020</title>
          <p>Sample comments from dataset(Tamil-English)
Ena da bgm ithu yuvannnnnnnnnn rocksssssss
Kola gaadula iruka... Thalaivaaaaaaaa waiting layea veri aaguthey
Wow wow wow... Thalaivaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa....... proud to be every Indian...
&lt;3 thanks to shankar sir and holl team...</p>
          <p>Nenu ee movie chusanu super movie
Super. 1 like is equivalent to 100 likes.</p>
          <p>Category
Positive
Negative
Mixed_feelings
not-Tamil
unknown_state</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Word Embedding</title>
        <p>Word embedding is arguably the most widely known technology in the recent history of NLP. It
captures the semantic property of a word. We use bert-base-uncased and distilbert-base-uncased
pre-trained models3 to get a vector as an embedding for the sentence that we can use for
classification. Apart from these two pre-trained models, we experiment with other pre-trained models like
bert-base-multilingual-uncased, bert-base-multilingual-cased.</p>
        <p>• BERT: Bidirectional Encoder Representations from Transformers (BERT)[7] is a technique for
NLP pre-training developed by Google. BERT is pre-trained on a large corpus of unlabelled
text, including the entire Wikipedia (that is 2,500 million words!) and the Book Corpus (800
million words). BERT-Base uncased has 12 layers (transformer blocks), 12 attention heads, and
110 million parameters.
• DistilBERT: DistilBERT[8] is a smaller version of BERT developed and open-sourced by the
team at HuggingFace. It is a lighter and faster version of BERT that roughly matches its
performance. DistilBERT also compares surprisingly well to BERT on downstream tasks while having
about half and one third the number of parameters.
• fastText: fastText, developed by Facebook, combines certain concepts introduced by the NLP
and ML communities, representing sentences with a bag-of-words and n-grams using subword
information and sharing them across classes through a hidden representation. fastText[9] can
learn vector representations of out-of-vocabulary words, which is useful for our dataset that
contains Malayalam and Tamil words in Roman script.</p>
        <p>After pre-processing our data and transforming all the comments into vector, we implement our
classification algorithms and construct our training models. We used the multinomial logistic
regression 4 with the fastText embeddings for unigrams, bigrams, and trigrams present along with diferent
learning rates and epochs. we got the maximum  1 score on fastText text classification model with
-wordNgrams= 1, learning rate = 0.1 and epochs = 10.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results and Analysis</title>
      <p>We use scikit-learn5 machine learning package for the implementation. A Macro  1 score was
used to evaluate every system. Macro  1 score of the overall system was the average of  1 scores of</p>
      <sec id="sec-3-1">
        <title>3https://huggingface.co/transformers/pretrained_models.html 4https://fasttext.cc/docs/en/supervised-tutorial.html 5http://scikit-learn.org</title>
        <p>Precision, recall,  1-score, and support for all experiment on Tamil-English test data
Precision, recall,  1-scores, and support for all experiment on Malayalam-English test data
the individual classes. Table 3 shows our oficial performances as shared by the organizers vis-a-vis
the best performing team. Table 4 and Table 5 report our results on Tamil-English and
MalayalamEnglish dataset respectively. We select three models that performed well during the validation phase
over others which was also in the oficial results (shown in Table 3).
and submit them for final prediction of the test dataset. We observe that fastText gives better  1 scores
In the training data, there are some ambiguous samples. Some examples are given below.
• The Tamil-English sentence Srk fan plz dislike tha video is labeled as Positive, when the sentence
has negative sentiment word like dislike.
• The Tamil-English sentence Wow wow wow... Thalaivaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa.......
proud to be every Indian... &lt;3 thanks to shankar sir and holl team... is labeled as mixed_feelings,
when there is many positive words like wow, proud, thanks.</p>
        <p>Our models were trained on this ambiguous data, and we could not verify the correctness of
labelling as we do not have knowledge of Tamil or Malayalam languages. Inconsistency of the labellings,
if any, might have worsened the results on test data. Another aspect is very small sentence length.
That might also be the reason why fastText unigram gave better results than n-grams, word n-grams
were not able to capture the sentiment of a sentence.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>This study reports performance of our system for the shared task on Sentiment Analysis for
Dravidian Languages in Code-Mixed Text in Dravidian-CodeMix - FIRE 2020. We conducted a number of
experiments on a real-world code-mixed YouTube comments dataset involving a few embedding
techniques: fastText, BERT, and DistilBERT. We find that fastText outperforms other techniques on this
task. However, there are room for improvement. In the future, we plan to use other pre-trained models
with necessary fine-tuning. We also plan to explore multilingual embeddings for the languages.
[7] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT: Pre-training of Deep Bidirectional
Transformers for Language Understanding, Proceedings of the 2019 Conference of the North (2019).
doi:10.18653/v1/n19-1423.
[8] V. Sanh, L. Debut, J. Chaumond, T. Wolf, DistilBERT, a distilled version of BERT: smaller, faster,
cheaper and lighter, 2019. arXiv:1910.01108.
[9] T. Mikolov, E. Grave, P. Bojanowski, C. Puhrsch, A. Joulin, Advances in Pre-Training Distributed
Word Representations, in: Proceedings of the International Conference on Language Resources
and Evaluation (LREC 2018), 2018.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Myers-Scotton</surname>
          </string-name>
          ,
          <article-title>Common and Uncommon Ground: Social and Structural Factors in Codeswitching, Language in Society 22 (</article-title>
          <year>1993</year>
          )
          <fpage>475</fpage>
          -
          <lpage>503</lpage>
          . URL: http://www.jstor.org/stable/ 4168471.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Muralidaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          , N. Jose,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Sherly</surname>
          </string-name>
          ,
          <string-name>
            <surname>Elizabeth</surname>
            <given-names>McCrae</given-names>
          </string-name>
          ,
          <article-title>Overview of the track on Sentiment Analysis for Davidian Languages in CodeMixed Text</article-title>
          ,
          <source>in: Proceedings of the 12th Forum for Information Retrieval Evaluation</source>
          ,
          <source>FIRE '20</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Muralidaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          , N. Jose,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Sherly</surname>
          </string-name>
          ,
          <string-name>
            <surname>Elizabeth</surname>
            <given-names>McCrae</given-names>
          </string-name>
          ,
          <article-title>Overview of the track on Sentiment Analysis for Davidian Languages in CodeMixed Text, in: Working Notes of the Forum for Information Retrieval Evaluation (FIRE 2020)</article-title>
          . CEUR Workshop Proceedings. In: CEUR-WS. org, Hyderabad, India,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          , N. Jose,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sherly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>A sentiment analysis dataset for code-mixed Malayalam-English, in: Proceedings of the 1st Joint Workshop on Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration and Computing for Under-Resourced Languages (CCURL), European Language Resources association</article-title>
          , Marseille, France,
          <year>2020</year>
          , pp.
          <fpage>177</fpage>
          -
          <lpage>184</lpage>
          . URL: https://www.aclweb.org/anthology/2020.sltu-
          <volume>1</volume>
          .
          <fpage>25</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Muralidaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Corpus creation for sentiment analysis in code-mixed Tamil-English text</article-title>
          ,
          <source>in: Proceedings of the 1st Joint Workshop on Spoken Language Technologies</source>
          for
          <article-title>Under-resourced languages (SLTU) and Collaboration and Computing for Under-Resourced Languages (CCURL), European Language Resources association</article-title>
          , Marseille, France,
          <year>2020</year>
          , pp.
          <fpage>202</fpage>
          -
          <lpage>210</lpage>
          . URL: https://www.aclweb.org/anthology/2020.sltu-
          <volume>1</volume>
          .
          <fpage>28</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B. G.</given-names>
            <surname>Patra</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
          </string-name>
          ,
          <source>Sentiment Analysis of Code-Mixed Indian Languages: An Overview of SAIL Code-Mixed Shared Task @ICON-2017</source>
          ,
          <year>2018</year>
          . arXiv:
          <year>1803</year>
          .06745.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>