<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of Memotion 3: Sentiment and Emotion Analysis of Codemixed Hinglish Memes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shreyash Mishra</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>S Suryavardan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Megha Chakraborty</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Parth Patwa</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anku Rani</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aman Chadha</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aishwarya Reganti</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amitava Das</string-name>
          <email>amitava@mailbox.sc.edu</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amit Sheth</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manoj Chinnakotla</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Asif Ekbal</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Srijan Kumar</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IIIT Sri City</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>India</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Microsoft</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Amazon AI</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Georgia Tech</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>IIT Patna</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Stanford University</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of South Carolina</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Analyzing memes on the internet has emerged as a crucial endeavor due to the impact this multi-modal form of content wields in shaping online discourse. Memes have become a powerful tool for expressing emotions and sentiments, possibly even spreading hate and misinformation, through humor and sarcasm. In this paper, we present the overview of the Memotion 3 shared task, as part of the DeFactify 2 workshop at AAAI-23. The task released an annotated dataset of Hindi-English code-mixed memes based on their Sentiment (Task A), Emotion (Task B), and Emotion intensity (Task C). Each of these is defined as an individual task and the participants are ranked separately for each task. Over 50 teams registered for the shared task and 5 made final submissions to the test set of the Memotion 3 dataset. CLIP, BERT modifications, ViT etc. were the most popular models among the participants along with approaches such as Student-Teacher model, Fusion, and Ensembling. The best final F1 score for Task A is 34.41, Task B is 79.77 and Task C is 59.82.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Memes</kwd>
        <kwd>codemixed</kwd>
        <kwd>multimodal</kwd>
        <kwd>Hindi-English</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The term meme is derived from the Ancient Greek word mimema, meaning imitated thing, which
comes from the verb mimeisthai, meaning to mimic. The term was coined by Richard Dawkins,
a British evolutionary biologist, in his book The Selfish Gene [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Dawkins used the concept of a
meme to explain the spread of cultural phenomena and ideas, drawing parallels between memes
and genes. Dawkins provided examples of memes such as melodies, catchphrases, fashion, and
arch-building technology in his book. Interestingly, the word meme is a self-describing term,
also known as autological, meaning that it is a meme itself.
      </p>
      <p>
        The reach of online content is vast and has the capacity to reach a massive viewership. Memes
propagate through multiple mediums including social media, messaging, apps, and emails. An
article on memes quotes "When you plant a fertile meme in my mind, you literally, parasitize
my brain" [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This quote indicates the impact of memes on a viewer’s mind. It spreads to others
who find it motivating, ofensive, amusing, or relatable in some ways [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Memes are used to
convey an opinion. When a person comes across a meme that they find relatable, they often
share it with others who they think would find relatable. Diferent people interpret memes
diferently which means what can be ofensive for one person might not be for other. This
diference in interpretation is a result of cultural, gender, and demographic diferences.
      </p>
      <p>
        The memes can also be used to spread ofensive and hateful content [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Identifying and
halting the dissemination of hateful memes is an arduous undertaking for both human beings
and AI models, as it requires comprehending the intricate subtleties and socio-political contexts
that form their interpretation. The deficiencies of current hate speech moderation techniques
highlight the pressing need to enhance the efectiveness of automated hate speech detection.
Given their subtlety and multi-modal nature, memes pose a more challenging issue to address
compared to only text.
      </p>
      <p>
        In this paper, we present the findings of the Memotion 3 shared task, where participants
were provided with an annotated dataset of 10k memes [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and were tasked with detecting
the sentiment, emotion and emotion intensity of the meme. Unlike the previous iteration of
memotion [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ] which provided English memes, current iteration studies codemixed
HindiEnglish (Hinglish) memes.
      </p>
      <p>The paper is organized as follows: we describe the related work and the task in section 2 and
3 respectively; Section 4 contains the details of the dataset we collected for Memotion analysis:
Memotion 3.0; followed by a brief description of baseline models and their results in 5. We
conclude and mention some possible future works in section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Sentiment and emotion analysis: There has been significant research on sentiment analysis
for text over many years [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Work on sentiment analysis using ML methods like SVM, logistic
regression, random forest, XGBoost, k-nearest neighbor has been done in [
        <xref ref-type="bibr" rid="ref10 ref11 ref9">9, 10, 11</xref>
        ]. Works
which use DL methods include [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12, 13, 14</xref>
        ]. For detailed surveys on sentiment analysis in social
media, please refer to [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ].
      </p>
      <p>
        The HaHa shared task provides a dataset for humor detection on social media [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Öhman
et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] release an English dataset to detect eight emotions like joy, sadness, disgust etc. A
comprehensive survey of textual emotion detection is provided in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]
      </p>
      <p>
        Hatespeech detection: It is important to detect hatespeech to keep social media safe
for everyone including minorities. Towards this goal, researchers have curated and released
annotated datasets [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. The ofenseval shared task [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] at SemEval 2019 releases an annotated
dataset of 14 English tweets to detect the type and target of ofensive language. The HatEval
shared task releases English and Spanish datasets to detect hate towards women and immigrants.
Patwa et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] conduct a shared task on Hindi hostile tweets detection. The TRAC workshop
series [
        <xref ref-type="bibr" rid="ref23 ref24 ref25">23, 24, 25</xref>
        ] conducts multiple shared tasks to detect aggression and misogyny in English,
Hindi, and Bengali datasets.
      </p>
      <p>
        Methods to detect hatespeech in text include CNNs and RNNS [
        <xref ref-type="bibr" rid="ref26 ref27 ref28 ref29">26, 27, 28, 29</xref>
        ], Bert-like
models [
        <xref ref-type="bibr" rid="ref30 ref31 ref32">30, 31, 32</xref>
        ], incorporating linguistic characteristics [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] etc.
      </p>
      <p>
        Codemixed Language Processing : Codemixed language processing is a challenging task
because the informal mixing of 2 or more languages and the proliferation of unique number
of ways to write the same word [
        <xref ref-type="bibr" rid="ref34 ref35">34, 35</xref>
        ]. The Sentimix task [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ] at semeval 2020 focused
on sentiment analysis of Hinglish and Spanish-English tweets. [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ] organized a shared task
on detecting ofense in 3 codemixed dravidian languages - Tamil [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ], Malayalam[
        <xref ref-type="bibr" rid="ref39">39</xref>
        ] and
Kannada[
        <xref ref-type="bibr" rid="ref40">40</xref>
        ]. Methods explored to tackle codemixing include graph convolutional netowrks
[
        <xref ref-type="bibr" rid="ref41">41</xref>
        ], BERT based models [
        <xref ref-type="bibr" rid="ref42 ref43">42, 43</xref>
        ], modifying loss function [
        <xref ref-type="bibr" rid="ref44 ref45">44, 45</xref>
        ] or positional embeddings
[
        <xref ref-type="bibr" rid="ref46">46</xref>
        ] to incorporate codemixing etc.
      </p>
      <p>
        Multimodal analysis : Although most of the existing research focuses on unimodal (text)
analysis, the use of multi-modal content likes memes and videos is fast increasing. Multimodal
datatsets having text and image, or videos are useful for tasks like image captioning, hatespeech
detection, emotion analysis in videos, sentiment analysis [
        <xref ref-type="bibr" rid="ref47 ref48 ref49 ref50">47, 48, 49, 50</xref>
        ] etc among other tasks.
The Hateful Memes Challenge dataset [
        <xref ref-type="bibr" rid="ref51">51</xref>
        ] and the multiof dataset [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] address hatespeech and
ofense detection in memes. However, they are binary classification task and the memes are in
English whereas memotion 3 has multi-class and multi-label tasks on code-mixed data. There
have been very few works on code-mixed meme analysis. [
        <xref ref-type="bibr" rid="ref52">52</xref>
        ] conduct a shared task to detect
trolling in Tamil codemixed memes whereas [
        <xref ref-type="bibr" rid="ref53">53</xref>
        ] release a dataset to detect hate in codemixed
Bengali memes. Popular methods for multimodal learning include image-text joint embeddings
[
        <xref ref-type="bibr" rid="ref54 ref55 ref56">54, 55, 56</xref>
        ] and transformer based models [
        <xref ref-type="bibr" rid="ref57 ref58">57, 58</xref>
        ].
      </p>
      <p>
        Previous iteration of Memotion : Memotion 1 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and Memotion 2[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] shared task released
datasets of 10k memes each [
        <xref ref-type="bibr" rid="ref59 ref6">6, 59</xref>
        ]. These datasets were annotated on the same tasks as
Memotion 3. However, both these datasets only focus on English memes, where as in memotion
3, we focus on Hinglish codemixed memes. Methods like ensembling [
        <xref ref-type="bibr" rid="ref60 ref61 ref62">60, 61, 62</xref>
        ] and bert-like
models [
        <xref ref-type="bibr" rid="ref63 ref64 ref65">63, 64, 65</xref>
        ] were common across memotion 1 and Memotion 2.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Task Details</title>
      <p>3.1. Tasks
Memotion 3 is the 3rd iteration of the Memotion task. The challenge consists of three sub-tasks:
1. Task A - Sentiment analysis of memes: Given a meme, the system is supposed to
classify the meme’s sentiment as positive, negative or neutral.
2. Task B - Overall emotion analysis of memes: This task’s goal is to pinpoint certain
emotions connected to a given meme. Whether a meme is humorous, sarcastic, ofensive,
or motivating should be indicated by the system. There are multiple categories in which
a meme can fit.
3. Task C - Classifying the intensity of meme emotions: The task is to determine the
degree to which a particular emotion is being expressed. The ranking of these emotions
is as follows:
a) Humour: Not funny, funny, very funny and hilarious
b) Sarcasm: Not Sarcastic, little sarcastic, very sarcastic and extremely sarcastic
c) Ofensive: Not ofensive, slightly ofensive, very ofensive and hateful ofensive
d) Motivation: Not motivational, motivational</p>
      <sec id="sec-3-1">
        <title>3.2. Dataset</title>
        <p>
          The tasks were conducted on the Memotion 3 dataset [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. It consists of Hindi-English codemixed
memes, which were collected from selenium based web crawler and they were gathered from
various public platforms like Reddit, Google Images, etc and annotated manually. The dataset
consists of 10,000 meme images divided into a train-val-test split of 7000-1500-1500. Each
meme is annotated for its sentiment, emotion and intensity of emotion. Images also have their
corresponding OCR text extracted with the help of Google Vision APIs and their respective
URLs. For more details of the dataset, please refer to [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Evaluation</title>
        <p>As mentioned previously, there are three tasks. Scoring is done for each task separately, and
separate leaderboards are generated. For each task, we use weighted average F1 score to measure
the performance of a model. The participants had access to only train and validation set. They
were asked to submit a maximum of 3 submissions on the test set for each task, the best of
which was selected as part of the leaderboard.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Baselines</title>
        <p>
          For multi-modal data, it is crucial to take into account both the visual and textual properties,
particularly in the case of memes where the context can only be recorded using a mix of
both elements. For textual features, we employ a multilingual form of BERT, namely
HinglishBERT (BERT-base-multilingual-cased) [
          <xref ref-type="bibr" rid="ref66">66</xref>
          ], which is tuned on Hinglish data. The trained
Vision Transformer (ViT) model provides the visual properties [
          <xref ref-type="bibr" rid="ref67">67</xref>
          ]. The Hinglish-BERT
embedding is concatenated with the pooled output from the ViT model. After passing through an
MLP, the combined features are then categorised in a final classification layer. For more details
about the baseline, please refer [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Participating systems</title>
      <p>There were 47 team registrations for the task in the Memotion 3.0 task page, of which 5 teams
made submissions for the final test set of the dataset. The results for all three tasks are given in
the following section and an overview of the 4 teams that presented their description papers
are provided below.</p>
      <p>
        wentaorub [
        <xref ref-type="bibr" rid="ref68">68</xref>
        ] use CLIP [
        <xref ref-type="bibr" rid="ref69">69</xref>
        ] for individual text and image embeddings, before concatenating
them and passing them through multi-headed attention layers for classification. They also use
the OSCAR [
        <xref ref-type="bibr" rid="ref70">70</xref>
        ] model in this approach and an ensemble of their models are used for the final
submission. Datasets such as Facebook Hateful memes [
        <xref ref-type="bibr" rid="ref51">51</xref>
        ], MMHS150k [
        <xref ref-type="bibr" rid="ref49">49</xref>
        ] etc. are used for
pre-training. This architecture helped them achieve the best performance in Task B and C.
      </p>
      <p>
        NYCU_TWO [
        <xref ref-type="bibr" rid="ref71">71</xref>
        ] propose a two model pipelines, namely Coopoerative Teaching Model
(CTM) for task A and Cascaded Emotion Classifier (CEC) for task B and C. A fusion of
multimodal embeddings from pre-trained Swin-Transformer [
        <xref ref-type="bibr" rid="ref72">72</xref>
        ] and CLIP are passed to the CTM
and CEC pipelines. CEC helps leverage task C predictions for task B by jointly training the
model. This team attained the best results in 3 out of the 4 labels in Task B.
      </p>
      <p>
        NUAA-QMUL-AIIT [
        <xref ref-type="bibr" rid="ref73">73</xref>
        ] refer to their approach as Squeeze-and-Excitation Fusion or
SEFusion. The textual features from pre-trained RoBERTa [
        <xref ref-type="bibr" rid="ref74">74</xref>
        ] and visual features from CLIP-ViT
[
        <xref ref-type="bibr" rid="ref75">75</xref>
        ] are fused to obtain multi-modal embeddings. The fusion is the SEFusion module, which
uses a learned activation of the squeezed features, allowing for weighted fusion of multi-modal
embeddings. This approach led to NUAA-QMUL-AIIT being the 1st ranked team in Task A.
      </p>
      <p>
        CUFE used a LightGBM [
        <xref ref-type="bibr" rid="ref76">76</xref>
        ] classifier for classification on every individual emotion or
label in all Tasks. The inputs to the classifier were pre-trained Hinglish-DistilBERT for text
embeddings and ResNet18 [77] for image embeddings. Other features, such as occurrence
of characters, word count etc. from text and number of faces in the memes (extracted using
Facenet’s [78] multitask cascaded CNN), were also used. CUFE obtained the highest score in 2
labels in Task B and 1 label in Task C.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>
        The performance in Task A i.e. sentiment classification is presented in Table 1. Our proposed
baseline achieved a score of 33.28%. Out of the five final submissions, four teams managed to
surpass the baseline. The top two teams, namely, NUAA-QMUL-AIIT [
        <xref ref-type="bibr" rid="ref73">73</xref>
        ] and NYCU_TWO
[
        <xref ref-type="bibr" rid="ref71">71</xref>
        ] improved on the baseline by 1.1% and 0.9% respectively. The diference in F1 scores for this
task is minimal across all participants.
      </p>
      <p>
        Table 2 shows the leaderboard for Task B. Three teams out-performed the baseline for this
task and top performing team wentaorub [
        <xref ref-type="bibr" rid="ref68">68</xref>
        ] improves on the final score by 5.4%. Team CUFE
performs lower than the baseline score by 3.4%, however, they present the highest score on the
wentorub
NYCU_TWO
NUAA-QMUL-AIIT
BASELINE
CUFE
      </p>
      <p>
        CSECU-DSG
Ofensive sub-task and on the Sarcasm sub-task, jointly with NYCU_TWO [
        <xref ref-type="bibr" rid="ref71">71</xref>
        ]. wentaorub
[
        <xref ref-type="bibr" rid="ref68">68</xref>
        ], NYCU_TWO, and NUAA-QMUL-AIIT achieved the same score on Motivation, which is 4%
higher than the baseline. NYCU_TWO also scored the highest in the Humour sub-task, improving
on the baseline by 4.5%. Based on these F1 scores, we can deduce that the Motivation class is
the easiest to detect, similar to Memotion 2.0 [
        <xref ref-type="bibr" rid="ref59">59</xref>
        ], despite some teams performing poorly on
this class. However, in this iteration, the Ofensive class is the hardest to detect, instead of the
Sarcasm class. No single team performs the best on all the classes.
      </p>
      <p>S</p>
      <p>H</p>
      <p>
        As shown in the leaderboard for Task C in Table 3, all the participating teams outperform
the baseline in this task. The minimum improvement in F1 score is 0.8% by team CUFE and
the maximum improvement in the final score is 7.5% by wentaorub [
        <xref ref-type="bibr" rid="ref68">68</xref>
        ]. The submission
by wentaorub exceeds all other teams in the sub-tasks Humour, Ofensive and Motivation.
NUAA-QMUL-AIIT [
        <xref ref-type="bibr" rid="ref73">73</xref>
        ] matched the highest score of wentaorub in the Motivation category,
while CUFE achieved the highest score in the Sarcasm sub-task. The Motivation class has
significantly higher F1 scores than other classes, due to it having only 2 intensities whereas
other classes have 4 intensities each. The best overall score is only 59.82%, which shows the
dificulty of the task.
      </p>
      <p>
        A major observation when comparing performances with the previous iteration of the task
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] is that the performance on Task A is much worse in Memotion 3. Overall Scores in Tasks B
and C have remained about the same compared to Memotion 2.
      </p>
      <p>For each individual task, we analyze the memes from the test set that all participants made
wrong predictions on. For Task A, there are 121 memes where all participants mis-classified the
label, out of which 66 memes were true negative sentiment followed closely by true positive
memes. For task B, there are 231 such memes, majority of which belong to "humor" and
"sarcasm" class. Finally, for Task C, 421 memes are mis-classified by all the systems - most
memes mis-classified by all teams are "Very Funny", "Very Sarcastic", "Slightly Ofensive". Some
such examples for Task A, B and C are shown in figures 1, 2, and 3 respectively. Further, we
note that most of such dificult examples have code-mixed text.</p>
      <p>
        As for the overall performances, only two teams - NUAA-QMUL-AIIT and NYCU_TWO [
        <xref ref-type="bibr" rid="ref71 ref73">73, 71</xref>
        ]
- perform better than the baseline in all tasks.
      </p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Future Work</title>
      <p>In this paper, we summarize the approaches used by the participants for the Memotion 3 task
and analyze the results. Due to the multi-modal nature of the dataset, all teams use a pre-trained
image and text embedding models. However, each team presents a novel model pipeline. The
highest scores achieved in Task A, Task B and Task C of Memotion 3.0 are 34.41%, 79.77% and
59.82% respectively, which shows there is significant room for improvement. On analysis of the
results and the mis-classified examples on the test set, we find that "Sarcasm" and "Humour"
are dificult to identify, especially in code-mixed memes.</p>
      <p>While we address Hind-English code-mixed memes in this paper, future work could include
exploring other languages/language pairs. A unified baseline model to analyze memes in
multiple languages could also be an interesting possibility.
tree, in: Neurips, 2017.
[77] K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: CVPR,
2016.
[78] F. Schrof, D. Kalenichenko, J. Philbin, FaceNet: A unified embedding for face recognition
and clustering, in: CVPR, 2015.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Dawkins</surname>
          </string-name>
          ,
          <article-title>The selfish gene</article-title>
          ,
          <source>Granada Publishing Lim</source>
          ,
          <year>1979</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Marwick</surname>
          </string-name>
          , Memes, Contexts
          <volume>12</volume>
          (
          <year>2013</year>
          )
          <fpage>12</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Akhther</surname>
          </string-name>
          ,
          <article-title>Internet memes as form of cultural discourse: A rhetorical analysis on facebook</article-title>
          ,
          <year>2018</year>
          . doi:
          <volume>10</volume>
          .31234/osf.io/sx6t7.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Buitelaar</surname>
          </string-name>
          ,
          <article-title>Multimodal meme dataset (MultiOFF) for identifying ofensive content in image and text</article-title>
          , in: TRAC,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chadha</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chinnakotla</surname>
          </string-name>
          , et al.,
          <article-title>Memotion 3: Dataset on sentiment and emotion analysis of codemixed hindi-english memes</article-title>
          ,
          <source>arXiv preprint arXiv:2303.09892</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bhageria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Scott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. PYKL</given-names>
            ,
            <surname>A. Das</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          , et al.,
          <article-title>SemEval-2020 task 8: Memotion analysis- the visuo-lingual metaphor!</article-title>
          , in: SemEval,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ramamoorthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gunti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ekbal</surname>
          </string-name>
          , et al.,
          <article-title>Findings of memotion 2: Sentiment and emotion analysis of memes</article-title>
          , in: Proceedings of De-Factify: Workshop on Multimodal Fact Checking and Hate Speech Detection, ceur,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Perelygin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Potts</surname>
          </string-name>
          ,
          <article-title>Recursive deep models for semantic compositionality over a sentiment treebank</article-title>
          ,
          <source>in: EMNLP</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>A. I. Saad</surname>
          </string-name>
          ,
          <article-title>Opinion mining on us airline twitter data using machine learning techniques</article-title>
          , in: 2020 16th international computer engineering conference (ICENCO), IEEE,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Alzyout</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Bashabsheh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Najadat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alaiad</surname>
          </string-name>
          ,
          <article-title>Sentiment analysis of arabic tweets about violence against women using machine learning</article-title>
          ,
          <source>in: 12th ICICS</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Prabhakar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Santhosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Krishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sudhakar</surname>
          </string-name>
          ,
          <article-title>Sentiment analysis of us airline twitter data using new adaboost approach, (IJERT) 7 (</article-title>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Kokab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Asghar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Naz</surname>
          </string-name>
          ,
          <article-title>Transformer-based deep learning models for the sentiment analysis of social media data</article-title>
          ,
          <source>Array</source>
          <volume>14</volume>
          (
          <year>2022</year>
          )
          <fpage>100157</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S. G.</given-names>
            <surname>Tesfagergish</surname>
          </string-name>
          , J. Kapočiu¯tė-Dzikienė,
          <string-name>
            <given-names>R.</given-names>
            <surname>Damaševičius</surname>
          </string-name>
          ,
          <article-title>Zero-shot emotion detection for semi-supervised sentiment analysis using sentence transformers and ensemble learning</article-title>
          ,
          <source>Applied Sciences</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>K. L. Tan</surname>
            ,
            <given-names>C. P.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K. M.</given-names>
          </string-name>
          <string-name>
            <surname>Lim</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. S. M. Anbananthen</surname>
          </string-name>
          ,
          <article-title>Sentiment analysis with ensemble hybrid deep learning model</article-title>
          ,
          <source>IEEE Access 10</source>
          (
          <year>2022</year>
          )
          <fpage>103694</fpage>
          -
          <lpage>103704</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>L.</given-names>
            <surname>Yue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zuo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <article-title>A survey of sentiment analysis in social media</article-title>
          ,
          <source>Knowledge and Information Systems</source>
          <volume>60</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>K.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bhattacharyya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bag</surname>
          </string-name>
          ,
          <article-title>A survey of sentiment analysis from social media data</article-title>
          ,
          <source>IEEE Transactions on Computational Social Systems</source>
          <volume>7</volume>
          (
          <year>2020</year>
          )
          <fpage>450</fpage>
          -
          <lpage>464</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chiruzzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Castro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Etcheverry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Garat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Prada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rosá</surname>
          </string-name>
          , Overview of haha at iberlef 2019:
          <article-title>Humor analysis based on human annotation</article-title>
          ., in: IberLEF@ SEPLN,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>E.</given-names>
            <surname>Öhman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pàmies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kajava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tiedemann</surname>
          </string-name>
          ,
          <article-title>Xed: A multilingual dataset for sentiment analysis</article-title>
          and
          <source>emotion detection</source>
          ,
          <year>2020</year>
          . arXiv:
          <year>2011</year>
          .01612.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Acheampong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wenyu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Nunoo-Mensah</surname>
          </string-name>
          ,
          <article-title>Text-based emotion detection: Advances, challenges, and opportunities</article-title>
          , Engineering Reports (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Waseem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hovy</surname>
          </string-name>
          ,
          <article-title>Hateful symbols or hateful people? predictive features for hate speech detection on Twitter</article-title>
          , in: NAACL,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          , et al.,
          <article-title>Semeval-2019 task 6: Identifying and categorizing ofensive language in social media (ofenseval</article-title>
          ), arXiv:
          <year>1903</year>
          .
          <volume>08983</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bhardwaj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Guptha</surname>
          </string-name>
          , G. Kumari,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. PYKL</given-names>
            ,
            <surname>A. Das</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ekbal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Akhtar</surname>
          </string-name>
          , T. Chakraborty,
          <article-title>Overview of constraint 2021 shared tasks: Detecting english covid-19 fake news and hindi hostile posts</article-title>
          ,
          <source>in: Combating Online Hostile Posts in Regional Languages during Emergency Situation</source>
          , Springer International Publishing,
          <year>2021</year>
          , pp.
          <fpage>42</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Ojha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <article-title>Benchmarking aggression identification in social media</article-title>
          ,
          <source>in: TRAC workshop</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Ojha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <article-title>Evaluating aggression identification in social media</article-title>
          ,
          <source>in: TRAC workshop</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ratan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Nandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. N.</given-names>
            <surname>Devi</surname>
          </string-name>
          , et al.,
          <source>The ComMA dataset v0</source>
          .
          <article-title>2: Annotating aggression and bias in multilingual social media discourse</article-title>
          ,
          <source>in: LREC</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>B.</given-names>
            <surname>Gambäck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U. K.</given-names>
            <surname>Sikdar</surname>
          </string-name>
          ,
          <article-title>Using convolutional neural networks to classify hate-speech</article-title>
          ,
          <source>in: Proceedings of the first workshop on abusive language online</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>85</fpage>
          -
          <lpage>90</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <article-title>Inf-hateval at semeval-2019 task 5: Convolutional neural networks for hate speech detection against women and immigrants on twitter</article-title>
          ,
          <source>in: SemEval</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>K.</given-names>
            <surname>Winter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kern</surname>
          </string-name>
          , Know-center at semeval
          <article-title>-2019 task 5: multilingual hate speech detection on twitter using cnns</article-title>
          ,
          <source>in: Semeval</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pykl</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Mukherjee</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Pulabaigari</surname>
          </string-name>
          ,
          <article-title>Hater-O-genius aggression classification using capsule networks</article-title>
          ,
          <source>in: Proceedings of the 17th International Conference on Natural Language Processing (ICON)</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Mazari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Boudoukhani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Djefal</surname>
          </string-name>
          ,
          <article-title>Bert-based ensemble learning for multi-aspect hate speech detection</article-title>
          , Cluster
          <string-name>
            <surname>Computing</surname>
          </string-name>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Samghabadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pykl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Solorio</surname>
          </string-name>
          ,
          <article-title>Aggression and misogyny detection using bert: A multi-task approach</article-title>
          , in: Proceedings of the second workshop on trolling,
          <source>aggression and cyberbullying</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>J.</given-names>
            <surname>Risch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Krestel</surname>
          </string-name>
          ,
          <article-title>Bagging BERT models for robust aggression identification</article-title>
          ,
          <source>in: Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>S.</given-names>
            <surname>Nagar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Barbhuiya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Dey</surname>
          </string-name>
          ,
          <article-title>Towards more robust hate speech detection: using social context and user data</article-title>
          ,
          <source>Social Network Analysis and Mining</source>
          <volume>13</volume>
          (
          <year>2023</year>
          )
          <fpage>47</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>A.</given-names>
            <surname>Laddha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hanoosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <article-title>Understanding chat messages for sticker recommendation in messaging apps</article-title>
          ,
          <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
          <volume>34</volume>
          (
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .1609/aaai.v34i08.
          <fpage>7019</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>A.</given-names>
            <surname>Laddha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hanoosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <article-title>Large scale multilingual sticker recommendation in messaging apps</article-title>
          ,
          <source>AI</source>
          Magazine
          <volume>42</volume>
          (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .1609/aaai.12023.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          , G. Aguilar,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pandey</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. PYKL</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gambäck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Solorio</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
          </string-name>
          , SemEval
          <article-title>-2020 task 9: Overview of sentiment analysis of code-mixed tweets</article-title>
          ,
          <source>in: Proceedings of the Fourteenth Workshop on Semantic Evaluation</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          , et al.,
          <article-title>Findings of the shared task on ofensive language identification in Tamil, Malayalam, and Kannada</article-title>
          ,
          <source>in: Proceedings of the First Workshop on Speech and Language Technologies for Dravidian Languages</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Muralidaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>Corpus creation for sentiment analysis in code-mixed tamil-english text</article-title>
          , arXiv:
          <year>2006</year>
          .
          <volume>00206</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          , N. Jose,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sherly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <article-title>A sentiment analysis dataset for code-mixed malayalam-english</article-title>
          , arXiv preprint arXiv:
          <year>2006</year>
          .
          <volume>00210</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <article-title>KanCMD: Kannada CodeMixed dataset for sentiment analysis and ofensive language detection</article-title>
          ,
          <source>in: Workshop on Computational Modeling of People's Opinions, Personality, and Emotion's in Social Media</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dowlagar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mamidi</surname>
          </string-name>
          ,
          <article-title>Graph convolutional networks with multi-headed attention for code-mixed sentiment analysis</article-title>
          ,
          <source>in: Proceedings of the First Workshop on Speech and Language Technologies for Dravidian Languages</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>J.</given-names>
            <surname>Risch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Stoll</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ziegele</surname>
          </string-name>
          , R. Krestel, hpidedis at germeval 2019:
          <article-title>Ofensive language identification using a german bert model</article-title>
          .,
          <source>in: KONVENS</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>D.</given-names>
            <surname>Tula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Potluri</surname>
          </string-name>
          , S. Ms,
          <string-name>
            <given-names>S.</given-names>
            <surname>Doddapaneni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sahu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sukumaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          , Bitions@DravidianLangTech-EACL2021:
          <article-title>Ensemble of multilingual language models with pseudo labeling for ofence detection in Dravidian languages</article-title>
          ,
          <source>in: Proceedings of the First Workshop on Speech and Language Technologies for Dravidian Languages</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>D.</given-names>
            <surname>Tula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Shreyas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Reddy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sahu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Doddapaneni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Potluri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sukumaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <article-title>Ofence detection in dravidian languages using code-mixing index-based focal loss</article-title>
          ,
          <source>SN Computer Science</source>
          <volume>3</volume>
          (
          <year>2022</year>
          ).
          <source>doi:10.1007/s42979-022-01190-1.</source>
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ma</surname>
          </string-name>
          , L.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Hao</surname>
          </string-name>
          , XLP at SemEval
          <article-title>-2020 task 9: Cross-lingual models with focal loss for sentiment analysis of code-mixing language</article-title>
          , in: Semeval,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Kandukuri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Manduru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
          </string-name>
          ,
          <article-title>Pesto: Switching point based dynamic and relative positional encoding for code-mixed languages (student abstract)</article-title>
          ,
          <source>AAAI</source>
          (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .1609/aaai.v36i11.
          <fpage>21587</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Flaxman</surname>
          </string-name>
          ,
          <article-title>Multimodal sentiment analysis to explore the structure of emotions</article-title>
          , in: KDD,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>R.</given-names>
            <surname>Jha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kolla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bhagat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>S. Pal,</given-names>
          </string-name>
          <article-title>Image2tweet: Datasets in Hindi and English for generating tweets from images</article-title>
          ,
          <source>in: Proceedings of the 18th International Conference on Natural Language Processing (ICON)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>R.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gibert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Karatzas</surname>
          </string-name>
          ,
          <article-title>Exploring hate speech detection in multimodal publications</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1910</year>
          .03814.
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          [50]
          <string-name>
            <given-names>B.</given-names>
            <surname>Zadeh</surname>
          </string-name>
          , et al.,
          <article-title>Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph</article-title>
          ,
          <source>in: ACL</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          [51]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kiela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Firooz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Goswami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ringshia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Testuggine</surname>
          </string-name>
          ,
          <article-title>The hateful memes challenge: Detecting hate speech in multimodal memes</article-title>
          ,
          <source>Neurips</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          [52]
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryawanshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <article-title>Findings of the shared task on troll meme classification in Tamil, in: Speech and Language Technologies for Dravidian Languages</article-title>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          [53]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hossain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Sharif</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>M. Hoque, MUTE: A multimodal dataset for detecting hateful memes</article-title>
          ,
          <source>in: Proceedings of the 2nd AACL Student Research Workshop</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          [54]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xie</surname>
          </string-name>
          , L. Liu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Learning text-image joint embedding for eficient cross-modal retrieval with deep feature engineering</article-title>
          ,
          <source>ACM Transactions on Information Systems</source>
          <volume>40</volume>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>27</lpage>
          . URL: https://doi.org/10.1145%2F3490519. doi:
          <volume>10</volume>
          .1145/3490519.
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          [55]
          <string-name>
            <given-names>V.</given-names>
            <surname>Krishna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ramamoorthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chadha</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
          </string-name>
          , Imaginator:
          <article-title>Pre-trained image+text joint embeddings using word-level grounding of images</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2305</volume>
          .
          <fpage>10438</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          [56]
          <string-name>
            <given-names>N.</given-names>
            <surname>Gunti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ramamoorthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
          </string-name>
          ,
          <article-title>Memotion analysis through the lens of joint embedding (student abstract)</article-title>
          ,
          <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
          <volume>36</volume>
          (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .1609/aaai.v36i11.
          <fpage>21616</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref57">
        <mixed-citation>
          [57]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Batra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1908</year>
          .02265.
        </mixed-citation>
      </ref>
      <ref id="ref58">
        <mixed-citation>
          [58]
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yatskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-J. Hsieh</surname>
            ,
            <given-names>K.-W.</given-names>
          </string-name>
          <string-name>
            <surname>Chang</surname>
          </string-name>
          ,
          <article-title>Visualbert: A simple and performant baseline for vision</article-title>
          and language,
          <year>2019</year>
          . arXiv:
          <year>1908</year>
          .03557.
        </mixed-citation>
      </ref>
      <ref id="ref59">
        <mixed-citation>
          [59]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ramamoorthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gunti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ekbal</surname>
          </string-name>
          , et al.,
          <article-title>Memotion 2: Dataset on sentiment and emotion analysis of memes</article-title>
          ,
          <source>in: Proceedings of De-Factify: Workshop on Multimodal Fact Checking and Hate Speech Detection</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref60">
        <mixed-citation>
          [60]
          <string-name>
            <given-names>K. N.</given-names>
            <surname>Phan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.-S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.-J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.-H.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <source>Little flower at memotion 2</source>
          .
          <article-title>0 2022: Ensemble of multi-modal model using attention mechanism in memotion analysis</article-title>
          ,
          <source>in: Proceedings of De-Factify: Workshop on Multimodal Fact Checking and Hate Speech Detection</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref61">
        <mixed-citation>
          [61]
          <string-name>
            <given-names>T.</given-names>
            <surname>Morishita</surname>
          </string-name>
          , G. Morio,
          <string-name>
            <given-names>S.</given-names>
            <surname>Horiguchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ozaki</surname>
          </string-name>
          , T. Miyoshi, Hitachi at SemEval
          <article-title>-2020 task 8: Simple but efective modality ensemble for meme emotion recognition</article-title>
          ,
          <source>in: SemEval</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref62">
        <mixed-citation>
          [62]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Xu</surname>
          </string-name>
          , Guoym at SemEval
          <article-title>-2020 task 8: Ensemble-based classification of visuo-lingual metaphor in memes</article-title>
          , in: SemEval,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref63">
        <mixed-citation>
          [63]
          <string-name>
            <given-names>G.-A.</given-names>
            <surname>Vlad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.-E.</given-names>
            <surname>Zaharia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.-C.</given-names>
            <surname>Cercel</surname>
          </string-name>
          , et al., Upb at semeval
          <article-title>-2020 task 8: Joint textual and visual modeling in a multi-task learning architecture for memotion analysis</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <year>2009</year>
          .02779.
        </mixed-citation>
      </ref>
      <ref id="ref64">
        <mixed-citation>
          [64]
          <string-name>
            <given-names>T. T.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. T.</given-names>
            <surname>Pham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. D.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , et al.,
          <source>Hcilab at memotion 2</source>
          .
          <article-title>0 2022: Analysis of sentiment, emotion and intensity of emotion classes from meme images using single and multi modalities</article-title>
          ,
          <source>in: Proceedings of De-Factify: Workshop on Multimodal Fact Checking and Hate Speech Detection</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref65">
        <mixed-citation>
          [65]
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <source>Amazon pars at memotion 2</source>
          .
          <article-title>0 2022: Multi-modal multi-task learning for memotion 2.0 challenge</article-title>
          , Proceedings of De-Factify: Workshop on Multimodal Fact Checking and Hate Speech Detection (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref66">
        <mixed-citation>
          [66]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bhange</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kasliwal</surname>
          </string-name>
          , Hinglishnlp:
          <article-title>Fine-tuned language models for hinglish sentiment detection</article-title>
          , arXiv preprint arXiv:
          <year>2008</year>
          .
          <volume>09820</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref67">
        <mixed-citation>
          [67]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Beyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          , et al.,
          <article-title>An image is worth 16x16 words: Transformers for image recognition at scale</article-title>
          , arXiv:
          <year>2010</year>
          .
          <volume>11929</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref68">
        <mixed-citation>
          [68]
          <string-name>
            <given-names>W.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kolossa</surname>
          </string-name>
          , wentaorub at Memotion 3:
          <article-title>Ensemble learning for multi-modal meme classification</article-title>
          ,
          <source>in: Proceedings of De-Factify 2: Workshop on Multimodal Fact Checking and Hate Speech Detection, CEUR</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref69">
        <mixed-citation>
          [69]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          , et al.,
          <article-title>Learning transferable visual models from natural language supervision</article-title>
          , in: ICML,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref70">
        <mixed-citation>
          [70]
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , et al.,
          <article-title>Oscar: Object-semantics aligned pre-training for vision-language tasks</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <year>2004</year>
          .06165.
        </mixed-citation>
      </ref>
      <ref id="ref71">
        <mixed-citation>
          [71]
          <string-name>
            <given-names>Y.-C.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.-D. Wang</surname>
          </string-name>
          , T.-Y. Ou, W.-C.
          <article-title>Peng, NYCU_TWO at Memotion 3: Good foundation, good teacher, then you have good meme analysis</article-title>
          ,
          <source>in: Proceedings of De-Factify 2: Workshop on Multimodal Fact Checking and Hate Speech Detection, CEUR</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref72">
        <mixed-citation>
          [72]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <article-title>Swin transformer: Hierarchical vision transformer using shifted windows</article-title>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2103</volume>
          .
          <fpage>14030</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref73">
        <mixed-citation>
          [73]
          <string-name>
            <given-names>X.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ma</surname>
          </string-name>
          , A. Zubiaga,
          <article-title>NUAA-QMUL-AIIT at Memotion 3: Multi-modal fusion with squeeze-and-excitation for internet meme emotion analysis</article-title>
          ,
          <source>in: Proceedings of De-Factify 2: Workshop on Multimodal Fact Checking and Hate Speech Detection, CEUR</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref74">
        <mixed-citation>
          [74]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach (</article-title>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref75">
        <mixed-citation>
          [75]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Puigcerver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          , et al.,
          <article-title>A large-scale study of representation learning with the visual task adaptation benchmark</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <year>1910</year>
          .04867.
        </mixed-citation>
      </ref>
      <ref id="ref76">
        <mixed-citation>
          [76]
          <string-name>
            <given-names>G.</given-names>
            <surname>Ke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Finley</surname>
          </string-name>
          , et al.,
          <article-title>Lightgbm: A highly eficient gradient boosting decision</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>