<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>N. N. A. Balaji); bharathib@ssn.edu.in (B. Bharathi);
bhuvanaj@ssn.edu.in (J. Bhuvana)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>SSNCSE_NLP@Dravidian-CodeMix-FIRE2020: Sentiment Analysis for Dravidian Languages in Code-Mixed Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nitin Nikamanth Appiah Balaji</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>B. Bharathi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. Bhuvana</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of CSE, Sri Siva Subramaniya Nadar College of Engineering</institution>
          ,
          <addr-line>Tamil Nadu</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Social media has become a place for expressing people's emotions, love, and hatred. With the ubiquitous availability of the internet entertains individuals to spend more time and express their feelings good or bad openly on social media platforms. In this study, we compare and analyze the methods for comment-level text polarity classification task using the Dravidian-CodeMix-FIRE2020 data-set. We contrast machine learning models with features extracted from techniques such as TF, TFIDF, BERT, fastText, and LSTM. The TF and the BERT embedding gave the best results compared. Our models scored F1 scores of 0.61 and 0.71 for the Tamil-English and the Malayalam-English tasks respectively.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Machine Learning</kwd>
        <kwd>NLP</kwd>
        <kwd>Code-mixed text</kwd>
        <kwd>Sentiment analysis</kwd>
        <kwd>BERT embeddings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>by unsupervised training on a large corpus containing diferent languages. The BERT model
produced results on par with the count vectorization model.</p>
      <p>This work is an account of the submissions made to the Dravidian-Codemix challenge
[10, 11, 12]. The subsequent sections are arranged as follows: The Section 2 explaining the
data-set distribution and preprocessing steps. The Section 3 details the experimental setup and
the various feature that was trialed for the task. The Section 4 provides a subjective analytical
comparison of the performance of various models in the test-data. Finally the Section 5 briefs
the objectives.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Data Description</title>
      <p>The data-set consists of YouTube comments with message-level sentiment polarity labels. The
classes for the tasks include positive, negative, neutral, mixed emotions, or if the comment
is not in the intended language for label. This becomes a multi-class classification task. The
individual comments contain an average of sentence length 1, making it more balanced and
easier to analyze. But there is some imbalance in the classes as it simulates the real-life scenario.
The data-set distribution among the training, development, and test sets are described in table
1. More detailed description on the data-set is provided in [13] and [14] for the Tamil-English
and Malayalam-English tasks respectively.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Experiments</title>
      <p>The experimental structure for the task can be expressed in two stages - the feature extraction
stage and the classifier stage. Techniques such as Count vectorization, TFIDF vectorization,
BERT, fastText are analyzed for the feature extraction stage and diferent classifiers such as
Logistic regression, Multi Layer Perceptron, Naive Bayes, and Random Forest classifiers are
compared. The machine learning algorithms, count, and TFIDF vectorization are studied using
the sci-kit learn 1 implementations. The sentence transformers [15] implementation for the
multilingual BERT and the pymagnitude [16] implementation for fastText is considered. The
Keras version of the LSTM was used for this study.</p>
      <p>The features extracted from the first stage are used to train the machine learning models
in the second stage and their performance is compared using the F1 score and the accuracy
score. The metrics methods of the sci-kit learn package is used to measure the performance.
The implementation with experimented and selected hyper-parameters are available in the link
2.</p>
      <sec id="sec-3-1">
        <title>3.1. Count And TFIDF Vectorization</title>
        <p>The content of the text comments is a mix of various languages, their syntax, and the
interchanging between diferent symbols. It becomes hard to capture the coherent intensity of the
comments with the extant pre-trained models. So bag of words based word and char count
models are deployed and analyzed by varying the n-gram range. The n-gram range of 2-3 and
1-5 gave the optimal result on the dev-set for Tamil-English and Malayalam-English sub-tasks
respectively. The Term Frequency Inverse Document frequency model helps to give a lesser
weight-age to the banal words in the corpus. This technique emphasizes more on the unique
terms in the corpus than the repeated words, rendering a better model.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Bidirectional LSTM</title>
        <p>Bidirectional Long short-term memory network is trained using word one-hot embedding.
This structure helps to learn the semantic syntax of the mixing of diferent languages and
consolidates structures from the comments. An LSTM could get relations from long distance
in the sequence. A multi-input, single-output RNN network is constructed with a single layer
1024 dimension hidden biLSTM layer and a dense layer. The words are converted into one-hot
vectors of 150-time frames and 100 dimensions for each word. This vector is fed to the LSTM
network. The network is trained for 7 epochs with a batch size of 128 and a learning rate of
0.001.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Multilingual Embedding Models</title>
        <p>The YouTube comments selected for the study contains text from English and Dravidian
languages coalesced together. This becomes a major problem to consider when applying
monolingual pre-trained models and fine-tuning for this particular task. But with the option of
pre-trained models in an unsupervised manner on a large collection of languages, it is possible
to fine-tune such multilingual models for fitting well for the Code-mix application.</p>
        <p>As the fastText and the BERT multilingual models [17] had shown fruitful results, it is
considered for this experiment. For FastText, a fixed length of 300 dimension vector is generated
by averaging the word-wise vectors of the entire sentence with the pymagnitude implementation
[16]. The Tamil specific and Malayalam specific pre-trained models from the FastText
multilanguage resources is used 3. Similarly a 512 dimension vector is generated by the BERT
(distiluse2https://github.com/nikamanthab/SSN_NLP-FIRE2020/tree/master/Codemix
3https://fasttext.cc/docs/en/crawl-vectors.html
base-multilingual-cased) pre-trained model from the SentenceTransformers implementation
[15]. The extracted features are used to train a classification model. The Multi Layer Perceptron
with 0.001 learning rate trained for 25 iterations shined better when compared to the Random
Forest classifier or the Naive Bayes classifier.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Observations</title>
      <p>The code-mix data-set presents a new challenge of applying alternating symbols and syntax
from majorly two diferent languages - English and a Dravidian language. Due to this very
reason, the primitive pre-trained fastText model showed relatively poorer results than the
models trained from scratch. But in contrast, the multilingual BERT which is an attention-based
transformer model, much convoluted from trained on a larger corpus produced comparable
results to that of the TFIDF and count vectorization models. The biLSTM didn’t perform on par
with the vectorization models but showed better performance than the fastText model.</p>
      <p>Out of all the analyzed models the count vectorization and the BERT model produced the
best performance for the Tamil-English corpus. For the Malayalam-English corpus the TFIDF
and count vectorization techniques generated equal performance models. The performance of
each model on the dev-set is presented in Table 2 and the results on the test-set is presented in
Table 3.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>Social media platforms are growing rapidly and entrench more and more people in taking part
in these platforms. Even though some opinions may be acceptable by one, it may be hurting for
others, so it becomes necessary to devise an efective model for the sentiment polarity detection
for the text comments. In this study, we have analyzed a variety of feature extraction techniques
and conclude that the Count, TFIDF based vectorization, and multilingual BERT technique
performs well on code-mix polarity labeling task. With these features, we reach a weighted F1
score of 0.61 for the Tamil-English task and 0.71 for the Malayalam-English tasks respectively.
6:1–6:14. URL: http://drops.dagstuhl.de/opus/volltexte/2019/10370. doi:10.4230/OASIcs.</p>
      <p>LDK.2019.6.
[8] B. R. Chakravarthi, M. Arcan, J. P. McCrae, WordNet gloss translation for under-resourced
languages using multilingual neural machine translation, in: Proceedings of the Second
Workshop on Multilingualism at the Intersection of Knowledge Bases and Machine
Translation, European Association for Machine Translation, Dublin, Ireland, 2019, pp. 1–7. URL:
https://www.aclweb.org/anthology/W19-7101.
[9] B. R. Chakravarthi, R. Priyadharshini, B. Stearns, A. Jayapal, S. S, M. Arcan, M. Zarrouk, J. P.</p>
      <p>McCrae, Multilingual multimodal machine translation for Dravidian languages utilizing
phonetic transcription, in: Proceedings of the 2nd Workshop on Technologies for MT of
Low Resource Languages, European Association for Machine Translation, Dublin, Ireland,
2019, pp. 56–63. URL: https://www.aclweb.org/anthology/W19-6809.
[10] B. R. Chakravarthi, R. Priyadharshini, V. Muralidaran, S. Suryawanshi, N. Jose, E. Sherly,
J. P. McCrae, Overview of the track on Sentiment Analysis for Dravidian Languages in
Code-Mixed Text, in: Working Notes of the Forum for Information Retrieval Evaluation
(FIRE 2020). CEUR Workshop Proceedings. In: CEUR-WS. org, Hyderabad, India, 2020.
[11] B. R. Chakravarthi, R. Priyadharshini, V. Muralidaran, S. Suryawanshi, N. Jose, E. Sherly,
J. P. McCrae, Overview of the track on Sentiment Analysis for Dravidian Languages in
Code-Mixed Text, in: Proceedings of the 12th Forum for Information Retrieval Evaluation,
FIRE ’20, 2020.
[12] B. R. Chakravarthi, Leveraging orthographic information to improve machine translation
of under-resourced languages, Ph.D. thesis, NUI Galway, 2020.
[13] B. R. Chakravarthi, V. Muralidaran, R. Priyadharshini, J. P. McCrae, Corpus creation for
sentiment analysis in code-mixed Tamil-English text, in: Proceedings of the 1st Joint
Workshop on Spoken Language Technologies for Under-resourced languages (SLTU)
and Collaboration and Computing for Under-Resourced Languages (CCURL), European
Language Resources association, Marseille, France, 2020, pp. 202–210. URL: https://www.
aclweb.org/anthology/2020.sltu-1.28.
[14] B. R. Chakravarthi, N. Jose, S. Suryawanshi, E. Sherly, J. P. McCrae, A sentiment analysis
dataset for code-mixed Malayalam-English, in: Proceedings of the 1st Joint Workshop on
Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration
and Computing for Under-Resourced Languages (CCURL), European Language Resources
association, Marseille, France, 2020, pp. 177–184. URL: https://www.aclweb.org/anthology/
2020.sltu-1.25.
[15] N. Reimers, I. Gurevych, Making monolingual sentence embeddings multilingual using
knowledge distillation, arXiv preprint arXiv:2004.09813 (2020). URL: http://arxiv.org/abs/
2004.09813.
[16] A. Patel, A. Sands, C. Callison-Burch, M. Apidianaki, Magnitude: A fast, eficient universal
vector embedding utility package, in: Proceedings of the 2018 Conference on Empirical
Methods in Natural Language Processing: System Demonstrations, 2018, pp. 120–126.
[17] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional
transformers for language understanding, arXiv preprint arXiv:1810.04805 (2018).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Jose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sherly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>A survey of current datasets for code-switching research</article-title>
          ,
          <source>in: 2020 6th International Conference on Advanced Computing and Communication Systems (ICACCS)</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vegupatti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Named entity recognition for code-mixed Indian corpus using meta embedding</article-title>
          ,
          <source>in: 2020 6th International Conference on Advanced Computing and Communication Systems (ICACCS)</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>A survey of orthographic information in machine translation</article-title>
          , arXiv e-prints (
          <year>2020</year>
          ) arXiv-
          <fpage>2008</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Rani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Goswami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Fransen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>A comparative study of diferent state-of-the-art hate speech detection methods for HindiEnglish code-mixed data</article-title>
          ,
          <source>in: Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying, European Language Resources Association (ELRA)</source>
          , Marseille, France,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Verma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Buitelaar</surname>
          </string-name>
          ,
          <article-title>A dataset for troll classification of Tamil memes</article-title>
          ,
          <source>in: Proceedings of the 5th Workshop on Indian Language Data Resource and Evaluation (WILDRE-5)</source>
          ,
          <source>European Language Resources Association (ELRA)</source>
          , Marseille, France,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Improving Wordnets for Under-Resourced Languages Using Machine Translation</article-title>
          ,
          <source>in: Proceedings of the 9th Global WordNet Conference, The Global WordNet Conference 2018 Committee</source>
          ,
          <year>2018</year>
          . URL: http://compling.hss. ntu.edu.sg/events/2018-gwc/pdfs/GWC2018_paper_
          <fpage>16</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Comparison of Diferent Orthographies for Machine Translation of Under-Resourced Dravidian Languages</article-title>
          ,
          <source>in: 2nd Conference on Language, Data and Knowledge (LDK</source>
          <year>2019</year>
          ), volume
          <volume>70</volume>
          of OpenAccess Series in Informatics (OASIcs),
          <source>Schloss Dagstuhl-Leibniz-Zentrum fuer Informatik</source>
          , Dagstuhl, Germany,
          <year>2019</year>
          , pp.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>