<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>XLM-RoBERTa for Sarcasm Detection in Dravidian Code-Mixed Languages</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Supriya Chanda</string-name>
          <email>supriyachanda.rs.cse18@itbhu.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Krishna Tewari</string-name>
          <email>krishnatewari.rs.cse24@itbhu.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arjun Mukherjee</string-name>
          <email>arjunmukherjee.rs.cse23@itbhu.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sukomal Pal</string-name>
          <email>spal.cse@itbhu.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Retrieval Lab (IReL), Dept. of CSE, Indian Institute of Technology (BHU) Varanasi</institution>
          ,
          <country country="IN">INDIA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Social Media</institution>
          ,
          <addr-line>Code-Mixed, BERT, ChatGPT, Sarcasm, Sentiment Analysis, Tamil, Malayalam</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Sarcasm refers to a form of irony where what is meant is actually said in an opposite manner usually in a mocking or humorous form, it could be either verbal or written. We very often come across multiple sarcastic comments which are code-mixed in various social-media platforms. In order to obtain insights from the textual data available or encountered upon, we would need a system to identify the sentiments behind the text and detect sarcasm. In this paper, we present a solution submitted for the shared task titled 'Sarcasm Identification of Dravidian Languages Tamil and Malayalam, which was organized by Dravidian CodeMix 2024 at the Forum for Information Retrieval Evaluation (FIRE) 2024. This paper explores an approach to sarcasm detection, leveraging the BERT (Bidirectional Encoder Representations from Transformers) and a supplementary layer of neural networks for precise classification into two distinct classes: sarcastic and non-sarcastic comments. It also uses ChatGPT for the same and performs a comparative study between GPT and BERT-based models. Our experiment demonstrates that our model efectively detects sarcastic comments, achieving an  1 score of 0.74 for both the Tamil-English and Malayalam-English code-mixed datasets, in contrary to GPT which can just achieve the  1 score of 0.64 for the above mentioned datasets. This score reflects a reasonable overall performance and places us at the third position in the ranking for Malayalam-English language pairs and at the first position in the ranking for Tamil-English language pairs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In the fast pacing age of technology, we often come across certain captivating linguistic puzzles which
are one of the fields humans have expertise into through tone or context and emotional cues, but for
machines it remains a formidable task [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Among these puzzles, sarcasm remains as one of the most
complicated linguistic problem. Sarcasm involves expressing thoughts in a manner that conceals the
true intentions of the speaker, often infused with a dose of mockery or humor, serving as a linguistic
tool to convey sentiments in a subtle manner [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>In the domain of sentiment analysis, sarcasm plays a very significant role in understanding of data.
Thus, accurately detecting sarcasm becomes a crucial aspect. In a world driven by technology, the boom
in social media users has been unfolding exponentially, with a staggering 60% of the global population
actively participating on these platforms, dedicating an average of 2 hours and 24 minutes daily to
their online engagements (as reported by smartinsights 1). Individuals can freely express their views
across wide range of subjects, events, personalities, products and a lot more domains, generating an
astounding volume of data real-time at a staggering rate of 328.77 million terabytes daily; all because
of the panorama provided by social media platforms. A considerable portion of this corpus is
codemixed, i.e. a phenomenon of linguistic wherein individuals very often combine intricacies of diferent
languages.</p>
      <p>ceur-ws.org</p>
      <p>
        Code-mixing can very often be identified as the pinnacle of utmost experience users have of varied
and diverse linguistic background. Through a peep hole of highlights we can observe the skills users
possess in weaving together diferent languages to enhance communication, adroitly switching between
tongues to convey their feelings [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In the globe of social media platforms, the amalgamation of
multiple languages whilst online interaction, imparting comments, posts or messages; is a common
occurrence and shouldn’t be a matter of surprise [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The analysis of text not emerged in its native form
presents an incremental layer of complexity in the versified area of natural language processing.
      </p>
      <p>
        The significance of sentiment analysis application on the aforesaid data is tremendous. It not only
provides valuable insights on various fields such as product description, market research, social trends
and customer feedback but can also counter the spread of hate speech on social media platforms [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
thus protecting the mental well being of individuals. Henceforth, the ability to derive intelligence in
user queries wrapped up with sarcasm opens up the road not taken to provide relevant information and
responses to the users.
      </p>
      <p>This shared task focuses on precise identification and determination of sarcasm within a
codemixed dataset comprising of comments and posts in Tamil-English and Malayalam-English. It helps
to dive deeper on how sarcasm is used in mixed-language conversations sourced from social-media.
Deciphering the subtleties of communication happening digitally can be achieved by gaining a deeper
insight into how sarcasm works in the ever- revolving globe of online interactions.</p>
      <p>In this paper, we applied a method that leverages XLM-Roberta and ChatGPT to enhance its capability
in identifying sarcasm and determining sentiment polarity within code-mixed comments and posts
written in Tamil-English and Malayalam-English, which are commonly encountered on social media
platforms, and also did a comparative study between both the approaches determining which one works
more eficiently on code-mixed data.</p>
      <p>The rest of the paper is structured as follows. Section 2 provides a concise overview of prior research
in this field. In Section 3, we delve into the datasets we utilized for our investigation. Section 4
elaborates on our computational methodologies, model specifications, and the techniques we employed
for evaluation. Next, we present our results and conduct a comprehensive analysis in Section 5. Finally,
we conclude in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>Code-mixing in languages has been the ongoing subject of extensive research over quite a long period
of time equi-validating it to around decades. However, the analysis of code-mixed text, particularly
in the area of social media, represents a comparatively new frontier or domain in the field of Natural
Language Processing(NLP).</p>
      <p>
        In recent years, we have explored various text processing tasks on code-mixed data, focusing on
diferent language pairs such as Bengali-English, Hindi-English, and Dravidian language
pairs—specifically Tamil-English, Malayalam-English, and Kannada-English. Our research has covered word-level,
sentence-level, and sentence-pair-level tasks. For instance, word-level language identification [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is
one of the foundational tasks that facilitates downstream processing of code-mixed data. We have
observed that incorporating a language identification module improves the performance of sentiment
analysis tasks [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. Additionally, our findings show that meta-embeddings outperform pretrained
word embeddings [9]. Another key sentence-level task we addressed is hate speech identification [ 10].
Recently, we also explored the novel task of information retrieval on code-mixed data [11]. However,
despite the advancements made, these tasks remain challenging due to the complexities inherent in
processing code-mixed content.
      </p>
      <p>Identification and detection of sarcasm has always been a significant task at downstream in the
domain of NLP which has always attracted researchers. Multiple eforts and contributions have been
made to solve this challenging problem and one such notable approach in this regard is the use of
IndicBERT for detecting sarcasm in social media text, as proposed by Amir et al [12]. This model
focuses on capturing contextual information and identifying sarcasm. Continuing further Wicana et
al. [13] dived into diferent machine learning methodologies to detect sarcasm. Their work ofered a
wide perspective on the latest techniques and associated dificulties in deploying that model to identify
and detect sarcasm. They explored a range of neural network-based classification structures, including
models like subword-level LSTM, Hierarchical LSTM, BERT, mBERT, XLM RoBERTa, LSTM, GRU, and
XLNet.</p>
      <p>Without a slightest doubt IndicBERT has proven itself to be very eficient in understanding nuances
and language-specific characteristics of Indian languages [ 14]. In order to deal with diferent intertwined
threads if sarcasm detection, we can see researchers employing a variety of techniques. For instance,
an attention-based BiLSTM model, combined with a feature-rich Convolutional Neural Network (CNN)
approach [15], has been utilized. It is essential to note that while sarcasm and hate speech are related,
they are not the same, and they demand distinct approaches.</p>
      <p>
        Eforts to identify sarcasm in social media content have lead to pathway of various innovative
approaches thus far. One such approach is employing prompt-based ChatGPT and its comparison
in multilingual sarcasm detection [16]. Additionally, Hegde et al. [17] investigated using the same
for Tamil and Malayalam code-mixed texts. While other approach in this domain is the use of the
multilingual XLM-RoBERTa with CNN and BiLSTM for sarcasm detection [18]. Furthermore, Agrawal
et al. [19] swam through the pool of emotional transitions to improve sarcasm detection, re-iterating
the dynamic nature of emotional signals in detection of sarcasm. As sarcasm goes beyond just written
text. Pandey and Vishwakarma [20] managed to handle the challenge of multi-modal sarcasm detection
in videos. They dealt with deep learning approaches to efectively use sensory inputs, visuals and
audio for sarcasm detection and identification. The organizers held a similar task last year, where we
used the mBERT model and achieved good results for both language pairs [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. From previous years
ifndings we can see Transformer-based language models have been pivotal in advancing language
comprehension [21, 22].
      </p>
      <p>These studies provide a glimpse into how evolving is the area of sarcasm detection and how a
wide spectrum of approaches and techniques are being applied, each ofering significant insights,
advancements and contribution in this field.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset</title>
      <p>The dataset that has been provided by the organizers [23] have served as a valuable resource for our
research, having comprised of several code mixed comments in Tamil-English and Malayalam-English,
sourced from social media. We often find that comments or post are of multiple sentences, but the
dataset mainly has an average sentence length of one. The realm of dataset encouraged us to investigate
how sarcasm is manifested in code-mixed contexts in social media. It includes development, training
and test datasets of comments in Tamil-English and Malayalam-English, having various code-mixed and
linguistic characteristics which provide a rich foundation for our research in detection and identification
of sarcasm. Table 1 provides a summary of the dataset statistics for both language pairs. It includes
details such as the total number of samples, the distribution across diferent classes or labels, and the
proportion of each language in the code-mixed data. These statistics give an overall view of the dataset
composition, helping to understand the balance and diversity within the dataset. Table 2 presents
example sentences from the dataset for both language pairs. The table includes the original sentence,
and any associated labels or annotations (such as Sarcastic or Non-sarcastic) to demonstrate the variety
of samples within the dataset.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <sec id="sec-4-1">
        <title>4.1. Preprocessing</title>
        <p>In the data preprocessing phase, we ensured to conduct several essential text preprocessing steps
to refine the dataset. We targeted in removing hashtags, punctuation marks, URLs, numbers and
mentions that does not have a great significance. Emojis were removed with their corresponding text
representations. Also, any extra spaces were removed to ensure a clean and consistent text corpus
for analysis. All our experiments were conducted after performing thorough preprocessing on the
code-mixed datasets.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Model Architecture</title>
        <p>In our research, we have used the bert-based-multilingual (XLM-RoBERTa) pre-trained model with
some fine tuning and ChatGPT to create a basic foundation of our task. In the recent trends, we have
found ChatGPT providing a tremendous boom in linguistics. So we have also employed ChatGPT for
our task of sarcasm prediction in code-mixed data, we utilized the model to predict labels on the test
dataset. A prompt-based approach was employed, where each test sentence was presented to ChatGPT
with a carefully crafted prompt instructing the model to classify the sentence as either “Sarcastic” or
“Non-sarcastic.” The prompt was designed to ensure that ChatGPT understood the task in the context of
code-mixed language data, taking into account the complexities of identifying sarcasm in multilingual
social media content. XLM-RoBERTa is built on transformer architecture, which involves self-attention
mechanism in both ends of encoding and decoding. These models are pre-trained on vast multilingual
text corpora and have a track-record of delivering exceptional performance when fine-tuned with
respect to the assigned tasks for downstream. For our objective of identifying sarcasm in code-mixed
language, we opted for BERT (Bidirectional Encoder Representations from Transformers) model, with a
focus in multilingual variant known as XLM-RoBERTa trained in 100 diferent languages.</p>
        <p>We have submitted two approaches prediction files. In the first approach, we used prompt-based
ChatGPT on the test dataset to classify them into Sarcastic or Not-Sarcastic. In the second approach,
we used XLM-RobERTa where we took a special token as input which propagated through the layers,
applying self-attention mechanisms and forwarding the output to the next consequent layer (see Figure
1). The output from final layer was fed into Neural Network which classified the component into two
categories: Sarcastic or Not-Sarcastic.</p>
        <p>S
e
n
t
e
n
c
e</p>
        <p>X
L
M
R
o
B
E
R
T
a
768
512
256
dropout
0.3
2
1</p>
        <p>Softmax</p>
        <p>In the training phase, we carefully tuned specific hyperparameters to guide the learning process
efectively. These hyperparameters were optimized to ensure the model’s proficiency in sarcasm
detection. We used a batch size of 16, a learning rate of 2e-5, and the AdamW optimizer to manage the
weight updates. The model was trained for a maximum of 2 epochs, as we observed overfitting beyond
this point using an early stopping method. This early stopping mechanism monitored the validation
loss and prompted the model to exit training after 2 epochs when overfitting became apparent. The loss
function employed was binary cross-entropy, given the binary nature of the sarcasm classification task.
These hyperparameter settings allowed the models to achieve optimal performance without overfitting,
ensuring accurate detection of sarcasm in code-mixed data.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results and Discussion</title>
      <p>This section is meant to deep dive into comprehensive evaluation of ChatGPT’s and our proposed models
performance on both datasets: Tamil-English and Malayalam-English, as part of Sarcasm Detection
and Identification in code-mixed Dravidian Languages. The performance of our proposed models
is examined using a range of evaluation metrics, with a primary focus on macro-averaged  1-score,
accuracy, recall and weighted average  1-score. The test data provided to us by the organizers served
as the foundation of our model evaluation.</p>
      <p>Our methodology involved fine-tuning our model based on the training and validation datasets,
ensuring it was well-prepared for the subsequent test data. Upon submission of our prediction file,
for our first approach which included employing ChatGPT, we achieved an  1 Score of 0.64 for both
language pairs and for our second approach which involved XLM-RoBERTa we achieved an  1 Score
of 0.74 again for both language pairs. This score reflects a reasonable overall performance and places
us at the third position in the ranking for Malayalam-English language pairs and at the first position
in the ranking for Tamil-English language pairs. Table 3 and 4 display the performance of the test
outcomes for our proposed model and top scored team for Malayalam-English and Tamil-English
language respectively [24]. Table 5 and 6 show the class wise classification report for both language
pairs on test data.</p>
      <p>While our system demonstrated commendable accuracy, it’s worth noting that other competing
teams surpassed us in both Precision and Recall, which ultimately influenced our  1 score and final
ranking especially in Malayalam-English, also, our model could reach this F1 score in 2 epochs. This
outcome encourages further refinement of our approach to enhance our model’s precision and recall,
aiming for even more competitive results in future endeavors. Figure 2 display the confusion matrices
of two language pairs based on our submission. The confusion matrices indicate that XLM-RoBERTa
consistently outperforms ChatGPT across both language pairs. For the Tamil-English dataset,
XLMRoBERTa achieved higher accuracy in classifying both non-sarcastic and sarcastic instances, with a
significantly lower number of false negatives (524 vs. 926 for ChatGPT) and better sarcasm detection
(985 true positives vs. 846 for ChatGPT). Similarly, in the Malayalam-English dataset, XLM-RoBERTa
demonstrated stronger performance, with fewer misclassifications in both categories and a notably
lower false-negative rate (256 vs. 400 for ChatGPT). The performance gap is more pronounced in the
Malayalam-English pair, where ChatGPT struggles to detect sarcasm, achieving only 225 true positives
compared to 312 for XLM-RoBERTa. These results suggest that XLM-RoBERTa is more efective in
capturing the nuances of code-mixed sarcasm, providing better generalization and robustness across
diverse language pairs.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>In this research, we have tried to solve the intricate task of identifying sarcasm in code-mixed comments
and/or posts, especially in Tamil-English and Malayalam-English languages, extracted from the world
of social media. Our dive into the waves of sentiment analysis definitely ensures us of the growing
significance of user direct or indirect opinions/expressions in the context of enhancing strategies, be
it business, marketing, government or other. In our experimentation, we tried to harness the power
of ChatGPT, which yielded a  1 score of 0.64 and the pre-trained multilingual BERT model, which
produced an outstanding  1 score of 0.74. This achievement shined the limelight on our approach
in capturing the nuances of sarcasm in code-mixed data. Despite of impressive accuracy, we surely
acknowledge the competitive environment where other teams excelled in Precision and Recall, afecting
our  1 score and final ranking, especially in Malayalam-English. Gaining insights from this, we are
all geared up for refining our methodology further for both our approaches, especially fine-tuning the
approach where we use ChatGPT and steadfast in our mission of enhancing precision and recall, with
the aim of achieving more competitive results.</p>
      <p>(a) ChatGPT for Tamil-English</p>
      <p>(b) XLM-RoBERTa for Tamil-English
(c) ChatGPT for Malayalam-English
(d) XLM-RoBERTa for Malayalam-English</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>The author(s) have not employed any Generative AI tools.
pretrained model and word-level language tag, Natural Language Processing (2024) 1–23. doi:10.
1017/nlp.2024.30.
[9] S. Chanda, R. Singh, S. Pal, Is meta embedding better than pre-trained word embedding to perform
sentiment analysis for dravidian languages in code-mixed text?, Working Notes of FIRE (2021).
[10] S. Chanda, S. Sheth, S. Pal, Coarse and fine-grained conversational hate speech and ofensive
content identification in code-mixed languages using fine-tuned multilingual embedding, in:
Forum for Information Retrieval Evaluation (Working Notes)(FIRE). CEUR-WS. org, 2022, pp.
502–512.
[11] S. Chanda, S. Pal, The efect of stopword removal on information retrieval for code-mixed data
obtained via social media, SN Computer Science 4 (2023) 494.
[12] S. Amir, B. C. Wallace, H. Lyu, P. Carvalho, M. J. Silva, Modelling context with user embeddings
for sarcasm detection in social media, in: Proceedings of the 20th SIGNLL Conference on
Computational Natural Language Learning, Association for Computational Linguistics, Berlin, Germany,
2016, pp. 167–177. URL: https://aclanthology.org/K16-1017. doi:10.18653/v1/K16- 1017.
[13] S. G. Wicana, T. Y. Ibisoglu, U. Yavanoglu, A review on sarcasm detection from machine-learning
perspective, 2017 IEEE 11th International Conference on Semantic Computing (ICSC) (2017)
469–476. URL: https://api.semanticscholar.org/CorpusID:16074739.
[14] K. Jain, A. Deshpande, K. Shridhar, F. Laumann, A. Dash, Indic-transformers: An analysis of
transformer language models for indian languages, 2020. arXiv:2011.02323.
[15] D. K. Jain, A. Kumar, G. Garg, Sarcasm detection in mash-up language using soft-attention
based bi-directional lstm and feature-rich cnn, Appl. Soft Comput. 91 (2020) 106198. URL: https:
//api.semanticscholar.org/CorpusID:216439240.
[16] M. A. Galal, A. H. Yousef, H. H. Zayed, W. Medhat, Arabic sarcasm detection: An enhanced
ifne-tuned language model approach, Ain Shams Engineering Journal 15 (2024) 102736. URL:
https://doi.org/10.1016/j.asej.2024.102736.
[17] N. N, V. V, A. Hegde, H. L. Shashirekha, Learning models with text augmentation for sarcasm
detection in malayalam and tamil code-mixed texts, in: Forum for Information Retrieval Evaluation
(Working Notes)(FIRE). CEUR-WS. org, 2023, pp. 287–298.
[18] M. Rahguouy, H. B. Giglou, T. Rahgooy, C. D. Seals, Null at semeval-2022 task 6: Intended sarcasm
detection using stylistically fused contextualized representation and deep learning, in: Proceedings
of the 16th International Workshop on Semantic Evaluation, 2022, pp. 862–870.
[19] A. Agrawal, A. An, M. Papagelis, Leveraging transitions of emotions for sarcasm detection,
Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in
Information Retrieval (2020). URL: https://api.semanticscholar.org/CorpusID:220729631.
[20] A. Pandey, D. K. Vishwakarma, Multimodal sarcasm detection (msd) in videos using deep
learning models, in: 2023 International Conference in Advances in Power, Signal, and Information
Technology (APSIT), 2023, pp. 811–814. doi:10.1109/APSIT58554.2023.10201731.
[21] N. Sripriya, T. Durairaj, K. Nandhini, B. Bharathi, K. K. Ponnusamy, C. Rajkumar, P. K. Kumaresan,
R. Ponnusamy, C. Subalalitha, B. R. Chakravarthi, Findings of shared task on sarcasm identification
in code-mixed dravidian languages, FIRE 2023 16 (2023) 22.
[22] B. R. Chakravarthi, N. Sripriya, B. Bharathi, K. Nandhini, S. C. Navaneethakrishnan, T. Durairaj,
R. Ponnusamy, P. K. Kumaresan, K. K. Ponnusamy, C. Rajkumar, Overview of the shared task
on sarcasm identification of dravidian languages (malayalam and tamil) in dravidiancodemix, in:
Forum of Information Retrieval and Evaluation FIRE-2023, 2023, pp. 231–239.
[23] B. R. Chakravarthi, N. Sripriya, B. Bharathi, K. Nandhini, S. Chinnaudayar Navaneethakrishnan,
T. Durairaj, R. Ponnusamy, P. K. Kumaresan, K. K. Ponnusamy, C. Rajkumar, Overview of the shared
task on sarcasm identification of Dravidian languages (Malayalam and Tamil) in DravidianCodeMix,
in: Forum of Information Retrieval and Evaluation FIRE - 2023, 2023, pp. 231–239.
[24] B. R. Chakravarthi, S. N, B. B, N. K, T. Durairaj, R. Ponnusamy, P. K. Kumaresan, K. K. Ponnusamy,
C. Rajkumar, Overview of sarcasm identification of dravidian languages in
dravidiancodemix@fire2024, in: Forum of Information Retrieval and Evaluation FIRE - 2024, DAIICT, Gandhinagar, 2024.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Sentiment analysis and subjectivity</article-title>
          ,
          <source>in: Handbook of Natural Language Processing</source>
          , Second Edition, Chapman and Hall/CRC,
          <year>2010</year>
          , pp.
          <fpage>627</fpage>
          -
          <lpage>666</lpage>
          . URL: http://www.crcnetbase.com/doi/abs/10. 1201/
          <fpage>9781420085938</fpage>
          -
          <lpage>c26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P. B.</given-names>
            <surname>Aditya Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Carman</surname>
          </string-name>
          ,
          <article-title>Automatic sarcasm detection: A survey</article-title>
          ,
          <source>ACM Comput. Surv</source>
          .
          <volume>50</volume>
          (
          <year>2017</year>
          )
          <volume>73</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>73</lpage>
          :
          <fpage>22</fpage>
          . URL: https://doi.org/10.1145/3124420.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chanda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Misha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <article-title>Sarcasm detection in tamil and malayalam dravidian code-mixed text</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>336</fpage>
          -
          <lpage>343</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Utsab</surname>
          </string-name>
          <string-name>
            <given-names>Barman</given-names>
            ,
            <surname>Amitava Das</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Foster</surname>
          </string-name>
          ,
          <article-title>Code mixing: A challenge for language identification in the language of social media</article-title>
          ,
          <source>in: Proceedings of the First Workshop on Computational Approaches to Code Switching@EMNLP</source>
          <year>2014</year>
          , Doha, Qatar, October
          <volume>25</volume>
          ,
          <year>2014</year>
          , Association for Computational Linguistics,
          <year>2014</year>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>23</lpage>
          . URL: https://doi.org/10.3115/v1/
          <fpage>W14</fpage>
          -3902.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chanda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <article-title>Hate Content Identification in Code-mixed Social Media Data, Chapman</article-title>
          and Hall/CRC,
          <year>2025</year>
          , pp.
          <fpage>225</fpage>
          -
          <lpage>247</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chanda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Misha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <article-title>Advancing language identification in code-mixed tulu texts: Harnessing deep learning techniques</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>223</fpage>
          -
          <lpage>230</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chanda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <article-title>Irlab@ iitbhu@ dravidian-codemix-fire2020: Sentiment analysis for dravidian languages in code-mixed text</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>535</fpage>
          -
          <lpage>540</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chanda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <article-title>Sentiment analysis of code-mixed dravidian languages leveraging</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>