<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Thinking, Fast and Slow: From the Speed of FastText to the Depth of Retrieval Augmented Large Language Models For Humour Classification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jana Viktória Kováčiková</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marek Šuppa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cisco Systems</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Comenius University</institution>
          ,
          <addr-line>Bratislava</addr-line>
          ,
          <country country="SK">Slovakia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>In this paper, we present the submission of our team NaiveNeuron to the JOKER 2024 task competition on Automatic Humour Analysis. Our first approach involves utilizing a fastText classifier for Task 2. Our subsequent approaches then incorporate Retrieval-Augmented Generation (RAG) with Large Language Models (LLMs) and adapt it for Humour Classification. By utilizing fastText, known for its eficiency, along with the depth and contextual understanding provided by RAG-enhanced LLMs, we demonstrate significant potential for improvements in humour detection accuracy. Although such systems may not be able to obtain the best possible accuracy as of yet, due to their simplicity we view them as highly potent baselines and believe they may be a very solid starting point for future research in this area.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Humour classification</kwd>
        <kwd>fastText</kwd>
        <kwd>Large Language Model</kwd>
        <kwd>Retrieval Agumented Generation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        For humour detection tasks, researchers have mainly explored using transformer models, such as
BERT. For instance, the study [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] evaluates the performance of diferent BERT variants in detecting
humour by fine-tuning them on humour-specific datasets. Their findings indicate that BERT-based
models significantly outperform traditional methods. Similarly, the study [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] uses BERT to generate
sentence embeddings, which are then processed by a neural network for binary humour classification.
Additionally, some studies have explored multimodal approaches to humour detection, integrating
textual, visual, and auditory cues to enhance the model’s understanding of humour. The paper [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
addresses multiple classification tasks, including humuor detection, sarcasm detection, and sentiment
analysis to analyze memes. Similar to our task, it employs multiclass humour classification. There
are also humour-classification-related works based on fastText models, such as the paper [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] which
addresses the problem of humour detection in Hindi-English code-mixed tweets.
      </p>
      <p>
        When it comes to the use of Large Language Models (LLMs) for humour detection, the work closes
to ours would be [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], in which the authors also provided LLMs with prompts and asked them to classify
the input. To the best of our knowledge, however, they did not make use of Retrieval Agumented
Generation (RAG) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], making our work the first to make use of it in the context of humour detection.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Datasets</title>
      <sec id="sec-3-1">
        <title>3.1. Datasets for Task 2</title>
        <p>• id: a unique identifier,
• text: humorous text.</p>
        <p>
          The JOKER 2024 datasets [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] for Task 2 consisted of three files. The train and the test data (
joker-2024task2-classification-train-input.json and joker-2024-task2-classification-test.json ) were provided in JSON
formats with the following fields:
        </p>
        <p>The train labels were provided in the format of JSON qrels file (
joker-2024-task2-classification-trainqrels.json) with the following fields:
• id: a unique identifier from the input file,
• class: class identifier for each humorous fenomena.</p>
        <p>There were 6 possible classes: IR (irony), SC (sarcasm), EX (exaggeration), AID (incongruity-absurdity),
SD (self-deprecating) and WS (wit-surprise). The train data consisted of 1742 samples and the test data
consisted of 6642 samples.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methods</title>
      <sec id="sec-4-1">
        <title>4.1. Using FastText Classifier</title>
        <p>Task 2 is a multiclass classification task with the aim to automatically classify text according to the
following classes: irony, sarcasm, exaggeration, incongruity-absurdity, self-deprecating and wit-surprise.
For Task 2, we employed the fastText library.</p>
        <p>Our approach began with partitioning the original training dataset,
joker-2024-task2-classificationtrain-input.json, and its corresponding labels, joker-2024-task2-classification-train-qrels.json , into training
(70%), validation (15%), and test (15%) sets.</p>
        <p>
          Prior to training the fastText classifier [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], we reformatted the data into a text file acceptable for
fastText. We then conducted a hyperparameter optimization for the classifier, experimenting with
various values for the following hyperparameters:
• minCount: minimum number of word occurrences,
• wordNgrams: maximum length of word n-gram,
• minn: minimum length of character n-gram,
• maxn: maximum length of character n-gram,
• lr: learning rate.
        </p>
        <p>We tested all possible combinations of hyperparameter values as shown in Table 1. Additionally, we
set the following parameters to fixed values: epoch=75 (number of training epochs) and dim=100 (size
of word vectors). We evaluated the model’s precision, recall, accuracy, and F1-score for each
hyperparameter combination using the validation set. Based on these metrics, the optimal hyperparameter
configuration was determined to be minCount=5, wordNgrams=1, minn=1, maxn=4, and lr=1.</p>
        <p>We also explored the impact of various preprocessing methods on the input data. Table 2 provides an
overview of the preprocessing methods tested and their corresponding validation set accuracies. Our
analysis concluded that using the original data without any preprocessing yielded the best results.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Large Language Models and Retrieval Augmented Generation</title>
        <p>Our aim with Large Language Models (LLMs) is twofold. First we aim to study how capable are they
of classifying the input text into one of the six categories in a zero-shot setup – that is, without being
provided any specific example. To evaluate this sort of performance we utilize the following prompt:
# Task
You a r e a t e x t c l a s s i f i e r t h a t c l a s s i f i e s E n g l i s h t e x t t o t h e
f o l l o w i n g c l a s s e s :
− IR ( i r o n y ) : I r o n y r e l i e s on a gap between t h e l i t e r a l meaning and
t h e i n t e n d e d meaning , c r e a t i n g a humorous t w i s t or r e v e r s a l .
− SC ( sarcasm ) − Sarcasm i n v o l v e s u s i n g i r o n y t o mock , c r i t i c i z e , or
convey contempt .
− EX ( e x a g g e r a t i o n ) − E x a g g e r a t i o n i n v o l v e s m a g n i f y i n g or
o v e r s t a t i n g something beyond i t s normal or r e a l i s t i c p r o p o r t i o n s .
− AID ( i n c o n g r u i t y − a b s u r d i t y ) − I n c o n g r u i t y r e f e r s t o t h e u n e x p e c t e d
or c o n t r a d i c t o r y e l e m e n t s t h a t a r e combined i n a humorous way
and A b s u r d i t y i n v o l v e s p r e s e n t i n g s i t u a t i o n s , e v e n t s , or i d e a s
t h a t a r e i n h e r e n t l y i l l o g i c a l , i r r a t i o n a l , or n o n s e n s i c a l .
− SD ( s e l f − d e p r e c a t i n g ) − S e l f − d e p r e c a t i n g humour i n v o l v e s making
fun o f o n e s e l f or h i g h l i g h t i n g one ’ s own f l a w s , weaknesses , or
e m b a r r a s s i n g s i t u a t i o n s i n a l i g h t h e a r t e d manner .
− WS ( wit − s u r p r i s e ) − Wit r e f e r s t o c l e v e r , quick , and i n t e l l i g e n t
humour and S u r p r i s e i n humour i n v o l v e s i n t r o d u c i n g u n e x p e c t e d
e le m en t s , t w i s t s , or p u n c h l i n e s t h a t c a t c h t h e a u d i e n c e o f f guard
.</p>
        <p>For any i n p u t , o u t p u t ONLY t h e u p p e r c a s e d c l a s s from t h e l i s t above .</p>
        <p>In the second case we aim to see to what extent the performance of the model improves when its
prompt includes a list of examples of both the input as well as its correct classification. To implement
this approach the examples, which are obtained from the the training set based input text, are added
towards the end of the aforementioned prompt in the format</p>
        <p>Input: {text}\nOutput: {class}\n\n,
where class is the correct output class.</p>
        <p>All of the LLM-based experiments were done with the temperature being set to 0 to ensure their
reproducibility.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <sec id="sec-5-1">
        <title>5.1. FastText-based Classification</title>
        <p>The final fastText classifier with parameters minCount=5, wordNgrams=1, minn=1, maxn=4, and lr=1,
dim=100, and epoch=200, was trained on a combined set of training and validation data (85% of the
original dataset joker-2024-task2-classification-train-input.json and
joker-2024-task2-classification-trainqrels.json) and achieved a test accuracy of 0.5725 on the remaining 15% of the original data. This
trained model was then used to predict labels for the test dataset, joker-2024-task2-classification-test.json .
The predictions were saved in naiveneuron_task2_fasttext.json and submitted.</p>
        <p>While the fastText classifier demonstrated moderate performance with a test accuracy of 0.5725, there
is significant potential for improvement by employing more advanced models such as large language
models (LLMs). On one hand, the fastText classifier had the advantage of being straightforward and
easy to implement. On the other hand, transformer-based models like BERT (Bidirectional Encoder
Representations from Transformers) could potentially ofer enhanced capabilities in understanding
context and semantic nuances.</p>
        <p>Accuracy</p>
        <p>Precision
Model
GPT-3.5-Turbo
GPT-4o
GPT-4
GPT-3.5-Turbo + RAG
GPT-4o + RAG
GPT-4 + RAG
LLama3 70b + RAG
GPT-4 + RAG UAE
LLama3 70b + RAG UAE</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. LLM-based Classification</title>
        <p>The results of the LLM-based classification, which was done across multiple models including
GPT-3.5Turbo, GPT-4, GPT-4o as well as LLama-3, can be seen in Table 3. All of the models were evaluated on
the same test split as the fastText evaluation used.</p>
        <p>As the results suggest, the zero-shot performance of LLMs on this task (the first three rows of Table 3)
leaves quite a bit to be desired, which is a phenomenon we can observe across multiple models.</p>
        <p>Involving Retrieval Augmented Generation (RAG) does indeed help, with the best
obtained performance tying that of the fastText model. All of the RAG models used the
mixedbread-ai/mxbai-embed-large-v11 embedding model which features 335M parameters. We
also experiment with using WhereIsAI/UAE-Large-V12 which is of similar size but has yielded better</p>
        <sec id="sec-5-2-1">
          <title>1https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1</title>
          <p>2https://huggingface.co/WhereIsAI/UAE-Large-V1
results in independent benchmarks. As the final two rows of Table 3 suggests, however, its inclusion in
the overall pipeline did not lead to an increase in performance.</p>
          <p>The Large Language Models (LLMs) were also able to uncover various issues with the actual evaluation
(test set) data. In various cases all three of the models considered (that is GPT-3.5-Turbo, GPT-4 as well
as the LLama3 model) refused to respond with a direct classification, despite being instructed to do so.</p>
          <p>One such example was the sample with the ID 6471 with the content Where PR = Public Reactions..
To illustrate how dificult it was for the models to respond to this input, we provide their output below
verbatim:
GPT-3.5-Turbo The input is not a valid example of any of the given classes. Please provide a text that can be classified
as either IR, SC, EX, AID, SD, or WS.</p>
          <p>GPT-4 Sorry, but I’m not sure what you mean by "PR = Public Reactions." Could you please provide more
context or clarify your question?
Llama3 I’m ready when you are! Please provide the input text, and I’ll classify it into one of the following
classes:
- IR (irony)
- SC (sarcasm)
- EX (exaggeration)
- AID (incongruity-absurdity)
- SD (self-deprecating)
- WS (wit-surprise)</p>
          <p>Go ahead and provide the input!</p>
          <p>This highlights the inherent limitations of LLMs and their combination with RAG: as their output
space is not limited, they might end up responding in free text which won’t match one of the desired
classes, especially when the input is indeed dificult to classify. In this particular case, however, we
believe it would be best to remove this particular item from the test set due to its significant ambiguity,
as it is not clear whether it does indeed fall into one of the aforementioned classes.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>In this study, we explored the use of the fastText and LLM-based classifiers for fine-grained sentiment
analysis of short humorous texts as part of Task 2 in JOKER 2024. Despite the fastText classifier
achieving a moderate test accuracy of 0.5725, which was only tied by the evaluated LLM-based systems,
our findings indicate potential for improvement. The simplicity and ease of use of the fastText model
were advantageous; however, its limitations in capturing the complex nuances of humour were evident.
Similarly, we only used LLMs in zero-shot and/or few-shot manner, without changing their parameters
at all, which was only able to yield a classification of the same accuracy as the fastText model. This
suggests that using classification models that utilize pre-trained Language Models as well as actually
updating the parameters of Large Language Models that are being used might be necessary for obtaining
higher performance.</p>
      <p>Despite that, we believe that thanks to their simplicity we believe that the proposed models have the
potential to serve as potent yet easy to setup baseline for any future exploration in humour classification.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <sec id="sec-7-1">
        <title>This research was partially supported by grant APVV-21-0114.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Peyrard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Borges</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Gligorić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>West</surname>
          </string-name>
          ,
          <article-title>Laughing heads: Can transformers detect what makes a sentence funny?</article-title>
          ,
          <source>arXiv preprint arXiv:2105.09142</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>I.</given-names>
            <surname>Annamoradnejad</surname>
          </string-name>
          , G. Zoghi,
          <article-title>Colbert: Using bert sentence embedding for humor detection</article-title>
          ,
          <source>arXiv preprint arXiv:2004.12765 1</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Chauhan</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. S R</surname>
          </string-name>
          , A. Ekbal,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharyya</surname>
          </string-name>
          ,
          <article-title>All-in-one: A deep attentive multi-task learning framework for humour, sarcasm, ofensive, motivation, and sentiment on memes</article-title>
          , in: K.
          <string-name>
            <surname>-F. Wong</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Knight</surname>
          </string-name>
          , H. Wu (Eds.),
          <source>Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing</source>
          , Association for Computational Linguistics, Suzhou, China,
          <year>2020</year>
          , pp.
          <fpage>281</fpage>
          -
          <lpage>290</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .aacl-main.
          <volume>31</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Sane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tripathi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Sane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mamidi</surname>
          </string-name>
          ,
          <article-title>Deep learning techniques for humor detection in Hindi-English code-mixed tweets</article-title>
          , in: A.
          <string-name>
            <surname>Balahur</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Klinger</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Hoste</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Strapparava</surname>
          </string-name>
          , O. De Clercq (Eds.),
          <source>Proceedings of the Tenth Workshop on Computational Approaches</source>
          to Subjectivity,
          <article-title>Sentiment and Social Media Analysis, Association for Computational Linguistics</article-title>
          , Minneapolis, USA,
          <year>2019</year>
          , pp.
          <fpage>57</fpage>
          -
          <lpage>61</lpage>
          . URL: https://aclanthology.org/W19-1307. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W19</fpage>
          -1307.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Dsilva</surname>
          </string-name>
          ,
          <article-title>Augmenting Large Language Models with Humor Theory To Understand Puns</article-title>
          ,
          <source>Ph.D. thesis</source>
          , Purdue University Graduate School,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Perez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Piktus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Petroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Karpukhin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Küttler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          , W.-t. Yih,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rocktäschel</surname>
          </string-name>
          , et al.,
          <article-title>Retrieval-augmented generation for knowledge-intensive nlp tasks</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>9459</fpage>
          -
          <lpage>9474</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          , A.-G. Bosser,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Thomas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M. P.</given-names>
            <surname>Preciado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          ,
          <article-title>Clef 2024 joker lab: Automatic humour analysis</article-title>
          , in: N.
          <string-name>
            <surname>Goharian</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Tonellotto</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>He</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Lipani</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>McDonald</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Macdonald</surname>
          </string-name>
          , I. Ounis (Eds.),
          <source>Advances in Information Retrieval</source>
          , Springer Nature Switzerland, Cham,
          <year>2024</year>
          , pp.
          <fpage>36</fpage>
          -
          <lpage>43</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , T. Mikolov,
          <article-title>Bag of tricks for eficient text classification</article-title>
          ,
          <source>arXiv preprint arXiv:1607.01759</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>