<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Parameswari_faith_nagaraju@Dravidian-CodeMix- FIRE: A machine-learning approach using n-grams in sentiment analysis for code-mixed texts: A case study in Tamil and Malayalam</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>ParameswariKrishnamurthy</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Faith Varghese</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>NagarajuVuppala</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Independent Researcher</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Hyderabad</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>eBhashasetu Language Services Pvt Ltd</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Sentiment analysis is a fast growing research positioned to uncover the underlying meaning of a text by categorizing it into diferent levels. This paper is an attempt to decode the deeply entangled code-mixed Malayalam and Tamil datasets and classify its interlined meaning at five various levels. Along with the corpus creation, [1] propose a five-level classification for Malayalam and Tamil code-mixed datasets. In this paper, we follow the five-level annotated datasets and aim to solve the classification problem by implementing unigram and bigram knowledge with a Multinomial Naive Bayes model. Our model scores an F1-score of 0.55 for Tamil and 0.48 for Malayalam.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Sentiment Analysis</kwd>
        <kwd>Code-mixed texts</kwd>
        <kwd>n-gram</kwd>
        <kwd>a Multinomial Naive Bayes model</kwd>
        <kwd>Tamil</kwd>
        <kwd>Malayalam</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Sentiments are constructed using elements to express positive or negative sentiments and
sentiment analysis thus detects the opinion of the sentence/document and classifies it into
positive, negative or neutral. 2[
        <xref ref-type="bibr" rid="ref3">, 3</xref>
        ] try to address four diferent problems predominating in
this research community, namely, subjectivity classification, word sentiment classification,
document sentiment classification, and opinion extraction. The automated process of discerning
or monitoring the opinions about a given subject, not only assists us in training the machine to
associate certain inputs with the corresponding outputs but also spots the keywords to assess
the stance of the consumer, to scan its polarity. The proliferation of commercial applications
has been one of the major reasons for the flourishing of sentiment analysis in the industrial field
as well. This provides a strong motivation for research on Tamil and Malayalam code-mixed
data in sentiment analysis and ofers many challenging research problems, which would have
been tough to address, otherwise. Presently, sentiment analysis is the cynosure of social media
research. Tamil and Malayalam, a widely used language in social media in diferent domains
needs a robust sentiment analysis. Starting from the assessment of marketing the success of an
ad campaign or new product launch, to determining the versions of a product or service that
are popular, and even identifying the demographics of people’s likes and dislikes particularly, is
a contribution much needed in these languages. Hate speech detection is another inevitable
tantamount achievement of this domain. In this paper, an n-gram knowledge-base trained with
a Multinomial Naive Bayes model is employed to analyse code-mixed texts in sentiment analysis
for Tamil and Malayalam.
      </p>
      <p>
        Tamil and Malayalam, the major Dravidian languages, are agglutinative i.e. word may
contain multiple morphemes attached to the stem with distinct morpheme boundaries to form
a multimorphemic word [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4, 5, 6</xref>
        ]. Understanding certain linguistic features which are encoded
as inflection and derivation is crucial in sentiment classification. According to a KPMG report
(2017)1, Tamil and Malayalam users of the internet are 42% and 27% respectively. Social media
platforms have millions of Indian language users and it is evident that due to bilingualism and
multilingualism in India, the code-mixed use of language is a common factor.
      </p>
      <p>
        The majority among the vast group of social media users with bilingual or multilingual
proficiency prefer to adopt code-mixing as it conveys the concept in the most simplest and acceptable
fashion [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. It is a normal tendency that the internet users opt for Roman transliteration over
the native scripts when they comment in social media. In some cases, users mix more than a
language in their usage resulting in code-mixing. Similar to this concept another form we
witness is the usage of combination of more than a script in expressing the idea. The conventional
approach of feature based classification may not assist us in yielding an optimum prediction
as the code-mixed language structure and spelling cannot be predefined. Implementation of
sentiment analysis in Indian languages is so recent that its emergence is noted by the end of
the first decade of the 21st century and however, extensive research has not been reported in
Tamil or Malayalam code-mixed corpora.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Brief Survey</title>
      <p>
        A few eforts on extracting sentiments in Tamil and Malayalam are reported here.
• The research by [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] focus on domain-specific sentence-level mood extraction from
Malayalam text using two methods of sentiment analysis viz., machine learning method and
semantic orientation method. The task is carried out using a semantic orientation method
using pointwise mutual information retrieval algorithm.
• the research article “SentiMa - Sentiment Extraction for Malayalam” b1y0][ propounds a
rule-based approach for opinion analysis from Malayalam movie reviews. The rule-based
approach that has been suggested by the researcher for extracting the opinion analysis in
Malayalam is the Negation-Rule that has claimed to have achieved 85% accuracy.
• Authors in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] employed sentiment analysis on the tweets in three languages, namely,
Hindi, Bengali, and Tamil for datasets collected from twitter over a period of three months.
The purpose of this shared task was to classify the collected tweets into positive, negative
and neutral polarity. From the submissions of six teams, maximum accuracy attained
1https://assets.kpmg/content/dam/kpmg/in/pdf/2017/04/Indian-languages-Defining-Indias-Internet.pdf
were 43.2 %, 55.67 %, and 39.28 % for Bengali, Hindi and Tamil respectively and to achieve
this accuracy the teams had employed supervised classification algorithms such as Naïve
Bayes, Multinomial Naïve Bayes, Support Vector Machines and Decision Tree.
• Authors in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] propose a N-gram model for opinion classification of Tamil tweets. 7418
Tamil unicode tweets were manually annotated by various domain experts for this study
and it aimed at three level classification. 61.29% is the overall accuracy estimated by the
model in this approach.
• The lexicon-based method is used by 1[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] uses which is also efectively practiced [14, 15,
16] for Malayalam as it is crucial to identify functional features along with the lexical
categories, as certain linguistic features are encoded as inflection and derivation . For
this study, 87347 Malayalam unicode sentences with political content were selected from
news web resources in the period between 1st August 2018 and 30th September 2018 and
the feature-based classification assisted to achieve a f-score of 0.9290.
• In order to address the decision problem of code-mixed Tamil and Malayalam texts in
sentiment analysis, [17] [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [18], [19] built a corpus of Tamil-English and Malayalam-English
by manually annotating 15,744 and 6,738 YouTube comments respectively. They employ
ifve-level classification for annotation: Positive, Negative, Neutral, not-Malayalam or
not-Tamil and Mixed feelings. The annotation derived at an agreement of Krippendorf’s
alpha 0.6 and 0.89 for Tamil and Malayalam respectively.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Trained Dataset: An Analysis</title>
      <p>Dravidian-CodeMix Forum for Information Retrieval Evaluation (FIRE) 2020 releases the datasets
of Tamil with 11335 sentences and Malayalam 4851 sentences. An overview of the categories
present in the dataset is given below:</p>
      <sec id="sec-3-1">
        <title>Category</title>
      </sec>
      <sec id="sec-3-2">
        <title>Tamil Category Percentage</title>
      </sec>
      <sec id="sec-3-3">
        <title>Malayalam</title>
      </sec>
      <sec id="sec-3-4">
        <title>Positive 7627</title>
        <p>Negative 1448
Mixed_feelings 1283
Unknown_state 609
Not-Tamil/Malayalam 368</p>
        <p>As seen in Table1, the dataset covering the Positive category is higher than any other
categories both in Tamil and Malayalam. The initial hypothesis is that the class-imbalance in
the training data may adversely afect the model accuracy. Furthermore, it is to be also noted
that building a balanced corpus from diverse YouTube comments is a cumbersome task. Hence,
we adopt n-gram modelling to examine whether imbalanced data with code-mix for sentiment
analysis can be handled.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Technique used</title>
      <p>The algorithm used in the present research in sentiment analysis for code-mixed data is described
in the following stages.</p>
      <sec id="sec-4-1">
        <title>1. Preprocessing 2. N-Gram modelling 3. Conversion into weighted features and machine learning</title>
        <sec id="sec-4-1-1">
          <title>4.1. Preprocessing</title>
          <p>The input data is preprocessed which involve dataset cleaning, removing stop words,
punctuations, numbers and non-unicode characters. As the dataset is given in Roman script, the text is
converted into lowercase.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.2. N-gram modelling</title>
          <p>From the pre-processed trained dataset, bigrams and unigrams are identified with highest
probabilistic sentiment category. The unigram and bigram database of Tamil are 18,195 and
58,129 respectively. Similarly in Malayalam 12,356 and 28,291 respectively. Each sentence is
converted into a unigram and bigram model based on the maximal match.</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>4.3. Conversion into weighted features and machine learning</title>
          <p>There are many techniques in machine learning that can be used to categorize data. The
present task used Term Frequency- Inverse Document Frequency (TF-IDF) approach 2[0].
TFIDF is a numerical statistic that shows the relevance or importance of a word/keyword to a
document in a collection of corpus. After experimenting with the Bag of words approach, we
observed that TF-IDF gives better results. We have used it as an initial step is to convert the
data(training and testing) into numerals. These numerals are feature vectors, i.e each vector has
its own importance. Since we are dealing with the classification of texts into positive, negative,
unknown_state, mixed_feelings, not Malayalam, not Tamil, Multinomial Naive Bayes (NB)
Algorithm [20] is used to classify text/comments into categories. This model works eficiently
when there are multiple categories with the combination of TF-IDF features. Multinomial NB
model and trains each sentence with its given category as shown in the flowchart-12</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results and Discussion</title>
      <p>The training dataset of Tamil includes 3149 sentences and Malayalam consists of 1348 sentences.
An overview of the categories present in the testing dataset is given below:</p>
      <p>2The code can be accessed here: https://github.com/nagaraju291990/sentimentAnalysis</p>
      <sec id="sec-5-1">
        <title>The table3 provides the results of the dataset given in table2-.</title>
      </sec>
      <sec id="sec-5-2">
        <title>Language</title>
      </sec>
      <sec id="sec-5-3">
        <title>Precision Recall F-score</title>
        <p>Tamil 0.55
Malayalam 0.53
0.66
0.51
0.55
0.48</p>
        <p>In the case of identifying positive category, our model performs better due to the datasize.
Whereas in identifying other categories, the failure occurred. Similarly, as noticed there are
cases in which the annotated data has some errors too:</p>
        <p>For instance, in Malayalam as seen in sentence (1), it is actually annotated as Not-Malayalam,
however it expresses Positive. Similarly, sentence (2) is tagged as Negative, however it denotes
Positive.</p>
        <p>(1) ml_sen_468:
P r o u d t o b e a v y p i n k k a r i . . . . Not-malayalam
‘Proud to be a Vypin lady’</p>
        <p>(2) ml_sen_1319:
I t h i n u a p u r a t h o r u t r a i l e r i l l a Negative
‘There is no trailer beyond this’</p>
        <p>Similarly, for instance in Tamil sentence (3) is tagged as positive, but it is not-Tamil and
sentence (4) expresses negative feeling, however tagged as Positive.</p>
        <p>(3) ta_sent_22:
I a m s i m b u f a n s l i k e d h a n u s h a c t i n g Positive
’I am the fan of Simbu and like Dhanush acting’</p>
        <p>(4) ta_sent_29:
I n u m n a s a g u r a t h u k u l l a i n t h a m a t h i r i y e t h a n a k o d u m a i y a p a k a p o r a n o o t h e r i l a y e e Positive
‘’I dont know how many inhuman act I will witness like this before I die”</p>
        <p>Another inevitable reason is the n-gram model and its size. Higher the model learns the
context the better the prediction would be; and the major factor that resulted in imprecise labelling
is the shortfall of the n-gram training in the given category. Also there are inconsistencies in
the data regarding the distribution of categories. For example positive categorised data amounts
to 65% while the remaining categories each fall under 15% or much less percentage.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>The possibility of sentiment analysis in Tamil and Malayalam code-mixed texts by applying the
combination of unigram and bigram approach is discussed in this paper. The model accuracy
can be improved with the increasing in the database of the unigrams and bigrams. On the
other hand Multinomial NB machine learning model works well when we need to train the
dataset with more than two sentiments while other models work well with Boolean sentiment
classification. To further improve the model, spelling normalization of code-mixed data and
integrating linguistic methods would be adopted in the future study. Although this requires a
lot of data and accurately human classified sentences to improve the accuracy of the machine
learning model, the distribution of sentiments equally in the train dataset is another important
aspect for machine learning model not to be biased to a particular sentiment.
malayalam political texts on social media, Unpublished Mphil dissertation: University of
Hyderabad (2018).
[14] P. D. Turney, M. L. Littman, Unsupervised learning of semantic orientation from a
hundredbillion-word corpus, arXiv preprint cs/0212012 (2002).
[15] M. Taboada, J. Brooke, M. Tofiloski, K. Voll, M. Stede, Lexicon-based methods for sentiment
analysis, Computational linguistics 37 (2011) 267–307.
[16] P. Chesley, B. Vincent, L. Xu, R. K. Srihari, Using verbs and adjectives to automatically
classify blog sentiment, Training 580 (2006) 233.
[17] B. R. Chakravarthi, V. Muralidaran, R. Priyadharshini, J. P. McCrae, Corpus creation for
sentiment analysis in code-mixed Tamil-English text, in: Proceedings of the 1st Joint
Workshop on Spoken Language Technologies for Under-resourced languages (SLTU)
and Collaboration and Computing for Under-Resourced Languages (CCURL), European
Language Resources association, Marseille, France, 2020, pp. 202–210. URhLt:tps://www.
aclweb.org/anthology/2020.sltu-1.2.8
[18] B. R. Chakravarthi, R. Priyadharshini, V. Muralidaran, S. Suryawanshi, N. Jose, J. P. Sherly,
Elizabeth McCrae, Overview of the track on Sentiment Analysis for Davidian Languages
in Code-Mixed Text, in: Working Notes of the Forum for Information Retrieval Evaluation
(FIRE 2020). CEUR Workshop Proceedings. In: CEUR-WS. org, Hyderabad, India, 2020.
[19] B. R. Chakravarthi, R. Priyadharshini, V. Muralidaran, S. Suryawanshi, N. Jose, J. P. Sherly,
Elizabeth McCrae, Overview of the track on Sentiment Analysis for Davidian Languages in
Code-Mixed Text, in: Proceedings of the 12th Forum for Information Retrieval Evaluation,
FIRE ’20, 2020.
[20] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel,
P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher,
M. Perrot, E. Duchesnay, Scikit-learn: Machine learning in Python, Journal of Machine
Learning Research 12 (2011) 2825–2830.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          , N. Jose,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sherly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>A sentiment analysis dataset for code-mixed Malayalam-English, in: Proceedings of the 1st Joint Workshop on Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration and Computing for Under-Resourced Languages (CCURL), European Language Resources association</article-title>
          , Marseille, France,
          <year>2020</year>
          , pp.
          <fpage>177</fpage>
          -
          <lpage>184</lpage>
          . URLh:ttps://www.aclweb.org/anthology/ 2020.sltu-
          <volume>1</volume>
          .
          <fpage>25</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Sentiment analysis and opinion mining (series synthesis lectures on human language technologies)</article-title>
          . vol.
          <volume>16</volume>
          , San Mateo, CA, USA: Morgan (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tan</surname>
          </string-name>
          , X. Cheng,
          <article-title>A survey on sentiment detection of reviews</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>36</volume>
          (
          <year>2009</year>
          )
          <fpage>10760</fpage>
          -
          <lpage>10773</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <article-title>Leveraging orthographic information to improve machine translation of under-resourced languages</article-title>
          ,
          <source>Ph.D. thesis, NUI Galway</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Comparison of Diferent Orthographies for Machine Translation of Under-Resourced Dravidian Languages</article-title>
          ,
          <source>in: 2nd Conference on Language, Data and Knowledge (LDK</source>
          <year>2019</year>
          ), volume
          <volume>70</volume>
          oOf penAccess Series in Informatics (OASIcs),
          <source>Schloss Dagstuhl-Leibniz-Zentrum fuer Informatik</source>
          , Dagstuhl, Germany,
          <year>2019</year>
          , pp.
          <volume>6</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          :
          <fpage>14</fpage>
          . URL: http://drops.dagstuhl.de/opus/volltexte/2019/1037.0doi:
          <fpage>10</fpage>
          .4230/OASIcs. LDK.
          <year>2019</year>
          .
          <volume>6</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stearns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jayapal</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. S</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zarrouk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Multilingual multimodal machine translation for Dravidian languages utilizing phonetic transcription</article-title>
          ,
          <source>in: Proceedings of the 2nd Workshop on Technologies for MT of Low Resource Languages, European Association for Machine Translation</source>
          , Dublin, Ireland,
          <year>2019</year>
          , pp.
          <fpage>56</fpage>
          -
          <lpage>63</lpage>
          . URL: https://www.aclweb.org/anthology/W19-680.
          <fpage>9</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Jose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sherly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>A survey of current datasets for code-switching research</article-title>
          ,
          <source>in: 2020 6th International Conference on Advanced Computing &amp; Communication Systems (ICACCS)</source>
          ,
          <year>India</year>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vegupatti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Named entity recognition for code-mixed Indian corpus using meta embedding</article-title>
          ,
          <source>in: 2020 6th International Conference on Advanced Computing &amp; Communication Systems (ICACCS)</source>
          ,
          <year>India</year>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N.</given-names>
            <surname>Mohandas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Nair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Govindaru</surname>
          </string-name>
          ,
          <article-title>Domain specific sentence level mood extraction from malayalam text</article-title>
          ,
          <source>in: 2012 International Conference on Advances in Computing and Communications</source>
          , IEEE,
          <year>2012</year>
          , pp.
          <fpage>78</fpage>
          -
          <lpage>81</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Nair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Jayan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rajeev</surname>
          </string-name>
          , E. Sherly,
          <article-title>Sentiment analysis of malayalam film review using machine learning techniques</article-title>
          ,
          <source>in: 2015 international conference on advances in computing, communications and informatics (ICACCI)</source>
          , IEEE,
          <year>2015</year>
          , pp.
          <fpage>2381</fpage>
          -
          <lpage>2384</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B. G.</given-names>
            <surname>Patra</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Prasath</surname>
          </string-name>
          ,
          <article-title>Shared task on sentiment analysis in indian languages (sail) tweets- an overview</article-title>
          ,
          <source>in: International Conference on Mining Intelligence and Knowledge Exploration</source>
          , Springer,
          <year>2015</year>
          , pp.
          <fpage>650</fpage>
          -
          <lpage>655</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>N.</given-names>
            <surname>Ravishankar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Shriram</surname>
          </string-name>
          ,
          <article-title>Grammar rule-based sentiment categorisation model for classification of tamil tweets</article-title>
          ,
          <source>International Journal of Intelligent Systems Technologies and Applications</source>
          <volume>17</volume>
          (
          <year>2018</year>
          )
          <fpage>89</fpage>
          -
          <lpage>97</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F. T.</given-names>
            <surname>Varghese</surname>
          </string-name>
          ,
          <article-title>A computational implementation of opinion analysis: a case study of</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>