<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Grenoble, France
* Corresponding author.</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Overview of the CLEF 2024 JOKER Task 2: Humour Classification According to Genre and Technique</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Victor Manuel Palma Preciado</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Grigori Sidorov</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liana Ermakova</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anne-Gwenn Bosser</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tristan Miller</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adam Jatowt</string-name>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Austrian Research Institute for Artificial Intelligence (OFAI)</institution>
          ,
          <addr-line>Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science, University of Manitoba</institution>
          ,
          <addr-line>Winnipeg</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>École Nationale d'Ingénieurs de Brest, Lab-STICC CNRS UMR 6285</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Instituto Politécnico Nacional (IPN), Centro de Investigación en Computación (CIC)</institution>
          ,
          <addr-line>Mexico City</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Université de Bretagne Occidentale</institution>
          ,
          <addr-line>HCTI</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Innsbruck</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This paper presents details of Task 2 of the JOKER-2024 track, which was held as part of the 15th Conference and Labs of the Evaluation Forum (CLEF 2024). The JOKER-2024 aims to foster progress in diferent humour processing techniques. In JOKER-2024 Task 2, participants aim to classify sentences in English that use a specific humour technique or genre. In this paper, we present the data used for this task and review the participants' results.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;humour</kwd>
        <kwd>humour classification</kwd>
        <kwd>humour genre</kwd>
        <kwd>humour technique</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Task 1 Humour-aware information retrieval</title>
      </sec>
      <sec id="sec-1-2">
        <title>Task 2 Humour classification according to genre and technique</title>
      </sec>
      <sec id="sec-1-3">
        <title>Task 3 Translation of puns from English to French.</title>
        <p>In this paper, we present Task 2 of the JOKER whose main objective is to classify textual humour
according to its genre or technique. Task 2 was the most popular JOKER task this year with 18 teams
submitting 54 runs over 103 runs submitted to the track in total.</p>
        <p>
          For the general presentation of the JOKER 2024 edition, refer to the overview paper [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. For the more
detailed discussions of Task 1 on retrieving [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and Task 3 on translating [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] humourous texts, refer to
the respecting Task overview papers.
        </p>
        <p>The rest of the paper is organised as follows. Section 2 describes the task, the data collection, and the
evaluation metrics. Section 3 provides an overview of the participants’ approaches. Section 4 presents
and discusses the participants’ results on the train and test data as well as the analysis of the results per
class. Section 5 concludes the paper.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Task Description</title>
      <p>In this section, we explain the CLEF 2024 JOKER track’s Task 2 on classifying humorous texts.</p>
      <p>For the purposes of this task, we constructed a humour taxonomy based on integrating various existing
humour classifications and taxonomies used in the literature and in diferent, separate corpora covering
particular aspects of humour. All texts in the corpus were analysed by a professional specialising in
humour research, who annotated each text with a human classification label according to the following
humour technique classification:
IR: Irony relies on a gap between the literal meaning and the intended meaning, creating a humorous
twist or reversal.</p>
      <sec id="sec-2-1">
        <title>SC: Sarcasm involves using irony to mock, criticise, or convey contempt.</title>
        <p>EX: Exaggeration involves magnifying or overstating something beyond its normal or realistic
proportions.</p>
        <p>AID: Incongruity/Absurdity refers (in the case of incongruity) to the unexpected or contradictory
elements that are combined in a humorous way and (for absurdity) involve presenting situations,
events, or ideas that are inherently illogical, irrational, or nonsensical.</p>
        <p>SD: Self-deprecating humour involves making fun of oneself or highlighting one’s own flaws,
weaknesses, or embarrassing situations in a lighthearted manner.</p>
        <p>WS: Wit/Surprise refers (in the case of wit) to clever, quick, and intelligent humour, and (for surprise)
to introducing unexpected elements, twists, or punchlines that catch the audience of guard.
Thus, the humour classification of Task 2 is a classification where the goal is to identify in a target text
the particular technique used for generating humour.</p>
        <p>
          The data for this task is a mixture of existing corpora on irony and sarcasm detection [
          <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
          ] and on
COVID-19 humour [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], our JOKER corpus 2023 [
          <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
          ], and jokes retrieved from public humour sites
according to the predefined categories selected in a balanced manner. An example data instance is given
below:
Sentence “Finally figured out the reason I look so bad in photos. It’s my face. ”
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Humour technique Self-deprecating (SD)</title>
        <p>The details on the amount of data for each class in the training and test sets are given in Table 1. The
Table provides evidence that the test and the train data come from the same distribution. There are
1,715 sentences in the training set all labelled as either SC, EX, WS, SD, AID, or IR (as discussed above).
The test data consists of 6,642 unlabelled texts that contain one of the earlier described types of humour.
From these texts, 722 were used for the evaluation.</p>
        <p>Runs for the task are evaluated according to standard metrics for classification, namely:
• Precision - the ratio of true positive predictions (correctly identified positive instances) to the total
number of positive predictions made by the classifier (both true positives and false positives).
• Recall - the ability of a classifier to identify all relevant instances in a dataset. It is the ratio of
true positive predictions (correctly identified positive instances) to the total number of actual
positive instances (both true positives and false negatives).
• F1 - the harmonic mean of precision and recall.</p>
        <p>• Accuracy - the ratio of the number of correct predictions to the total number of instances.</p>
        <p>We report precision, recall, and F1 scores per class, as well as macro average (averaging the unweighted
mean per class) and weighted average (averaging the support-weighted mean per class).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Participants’ submissions and approaches</title>
      <p>In this section, we detail the approaches to humour classification as deployed by participants of the
track.</p>
      <p>Task 2 proved to be the most popular task at JOKER-2024, with 54 submissions – see Table 2. We
observed a great variety of approaches, ranging from state-of-the-art large language models (LLMs) to
more classic probabilistic ones. The list below explains the approaches used by each participating team
and briefly introduces their work; we also provide citations to the respective system description papers
for those teams that submitted them to the workshop proceedings.</p>
      <p>The AB&amp;DPV team [9] submitted a total of seven runs, opting to use embeddings with the help
of Word2Vec. To develop their results, they used the Multilayer Perceptron (MLP), Random Forest,
Decision Tree, and Gaussian Naive Bayes classifiers.</p>
      <p>The CodeRangers team [11] submitted a total of two runs. The team used BERT-uncased and
RoBERTa by fine-tuning them on the provided data for Task 2. They observed slightly higher accuracy
for RoBERTa than for BERT during their experiments on the training data.</p>
      <p>The CYUT team [10] submitted a total of three runs. RoBERTa was first fine-tuned, using an 80/20
split of the dataset we shared to train and validate models. GPT-4 was used with zero-shot prompting and
chain-of-thought prompting. Attempts to classify among all the classes at once proved too challenging
for the model. Therefore hierarchical categories were created and a four-step classification method was
employed (using either binary or three-way classification for a given step). After first discriminating
AID and WS on the one hand from IR, SC, SD, and EX on the other, further steps were used to classify
down the grouping hierarchy. Llama 3-8b was fine-tuned on a single GPU, which was made possible
through four-bit quantisation with QLoRa.</p>
      <p>The DadJokers team [12] submitted a total of three runs. The team’s first classification approach
is done using BERT base uncased and the second attempt for classification uses a traditional
machine learning model, the Random Forest classifier. The authors applied TFIDFVectorizer and
SentenceTransformer as preprocessing steps.</p>
      <p>The Dajana&amp;Kathy team submitted a single run. Their approach involves the use of TF–IDF and
BERT embeddings with a variety of models such as SVM, Random Forest, LSTM, and Transformers. A
similar approach was taken by the Frane team.</p>
      <p>The Jokester team [13] submitted a single run. They combined several classifiers available through
the Scikit library: a voting classifier weighted the results obtained by an SVC and a stack of Random
Forest, Decision Tree, Gradient Boosting, and Logistic Regression. We do not report their results as all
texts predicted as belonging to the SD class.</p>
      <p>The HumourInsights team [14] submitted one run. Although this team reported using a variety of
classical approaches for the classification task, they chose to submit only their best model. They used
TF–IDF to extract features for later use in diferent methods. They employed boosting methods such as
ADA and Gradient with mixed results, but these did not reach the accuracy obtained with KNN and
Random Forest.</p>
      <p>The NLPalma team [15] made three submissions for the classification task using two diferent
approaches: one with more classical classifiers and the other with a more well-known model in the
BERT-like lineage.</p>
      <p>The PunDerstand team [19] submitted a total of four runs. The authors employed the DeBERTa
model which, after fine-tuning, gave rise to two runs, one on a raw, unprocessed dataset and one on
balanced data. The latter was ensured by an undersampling strategy. In another run, they used GPT-4o,
the most recent large language model developed by OpenAI. Few-shot prompting was employed with
one example for each class of humour. The team also provided a run with manually guided annotation.</p>
      <p>The Tomislav&amp;Rowan team [20] submitted a total of three runs. After preprocessing the text,
TF–IDF was used for vectorising it. Three models were trained on the data: Logistic Regression, Naïve
Bayes, and an SVM.</p>
      <p>The Petra&amp;Regina team [18] submitted a total of one run. Processed data was vectorised using
TF–IDF and class labels were encoded using a linear regression algorithm.</p>
      <p>The UAms team [21] submitted a single run. This team chose to use a BERT classifier trained on 90%
of the training data, with taking special precautions against overfitting.</p>
      <p>The ORPAILLEUR team [17] submitted a total of nine runs. The team explored the potential of
advanced LLMs within a consistent methodological framework. They employed a four-bit quantised
version of three LLMs: Llama2-7b1, Mistral-7b2, and Llama3-8b. The final hidden state of the last token
was used as input to the feed-forward layer with the softmax function to get the class probabilities. The
team also explored QLoRa adapters.</p>
      <p>The NaiveNeuron team [16] submitted three runs through iterations of diferent LLMs, including
various instances of GPT-3.5, GPT-4, and GPT-4.0 Plus RAG. They obtained the best results with
GPT4+RAG using a 70/15/15 split for testing. The experimentation of this team was not limited to GPT
models; they also used fastText and Llama 3. However, they achieved more favorable results with
zero-shot and few-shot classification using GPT-RAG.</p>
      <p>The Arampatzis team submitted eight runs for this task. The team has experimented with the
following approaches: XLNet, Multilayer Perceptron, BERT, RoBERTa, DistilBERT, DeBERTa, Electra,
AlBERT.</p>
      <p>Finally, the RubyAiYoungTeam team submitted a single run, without providing details of their
approach.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>This section details the results of the CLEF 2024 JOKER Task 2 on Humour Classification.</p>
      <p>A total of 18 teams submitted 54 runs for Task 2. This was the most popular JOKER task this year
which might be explained by the variety of classification models. Participants used mostly LLMs and
traditional classifiers although some teams experimented with fine-tuned models and diferent setups.
Some runs had problems with additional classes in their predictions (e.g., “ERROR”). We lfitered out
these predictions, which explains some diferences in the numbers of instances in the runs.</p>
      <sec id="sec-4-1">
        <title>4.1. Test results</title>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Train results</title>
        <p>To check for a possible overfitting problem, we report training data results for each run in Tables 5
and 6. Table 5 presents the results the training data in terms of accuracy, macro and weighted-average
precision, recall, and F1, while Table 6 presents the results for each run and each class in terms of
precision, recall, and F1. Random Forest, LLAMA, and Mistral models achieved 100% accuracy on the
training data, while on the test data, Random Forest did not score well, suggesting the existence of an
overfitting problem. GPT-4 obtained very low scores both on test and training data. However, LLMs
showed much better generalisation capacity on unseen data than traditional models.</p>
        <p>We conducted various analyses to better understand the peculiarities of the dataset and the results of
each team. The following two subsections present some figures and discussion focussing on performance
across classes of data and across the various system archetypes used by the participants.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Performance per class</title>
        <p>As was shown already in Table 1 above, the dataset is unbalanced with respect to its class distribution.
Therefore, the relevant model performance metrics for this type of dataset tend to be macro and weighted
averages.</p>
        <p>From Table 7 it can be observed that the classes with the best performance over all submissions are
generally those that occur most frequently in the test set. Overall, the best average classification across
diferent runs was for class AID, and the worst, despite being the third largest in terms of number of
texts, was class EX.</p>
        <p>The following five subsections examine system performance on the individual classes, with tables
presenting the best, worst, median, and mean (“average”) system performance in terms of precision,
recall, and F-score. We excluded four incomplete runs from the computation of the average and median
values. All scores are reported in percentage.
4.3.1. SD: Self-deprecating humour
From Table 8 we can observe a very noticeable diference between the median and the highest result
for the SD class.
4.3.2. WS: Wit/Surprise humour
For Wit/Surprise (see Table 9), there isn’t such a pronounced diference between the median, average,
and the best result, although we note that the LLMs lead with the best results. It’s worth noting that
the other models on average are not far behind. Where there is a very marked disparity is in the lowest
performance compared to the others. In this case, the BERT-like model was not able to understand this
type of class well. Of course, this doesn’t mean that BERT-like models are incapable of competing, as
the table only shows the worst result. It’s clear that other models can perform quite well, especially
considering that the median is 42%.
4.3.3. EX: Exaggeration humour
Exaggeration (see Table 10) is a class that shows a very similar trend to the previous ones, with a rather
high disparity between the worst and the best performance. It is possible that these results are linked to
the way the models’ learning was handled.
4.3.4. IR, SC: Irony and Sarcasm humour
The cases of Irony and Sarcasm (see Tables 11 and 12) are peculiar. Perhaps owing to their semantic
similarity, both classes show a similar pattern of results, with quite close medians and averages.
Furthermore, the fact that the model with the best performance in both is an LLM (specifically, Mistral)
supports the argument that LLMs outperform more traditional models. It is worth noting that the
ORPAILLEUR team’s focus on using LLMs was quite rewarding, as they proved superior in all cases.
4.3.5. AID: Incongruity/Absurdity
The results for Incongruity/Absurdity (see Table 13) are by far the best among all the classes, with even
the classical methods achieving good results in the worst case. As expected, the LLMs again lead with a
very high score of 90% in F-score.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Performance by Model</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Conclusions</title>
      <p>This paper provides an overview of Task 2 from the CLEF 2024 JOKER track aiming at classifying
humorous texts based on specific humour techniques or genres. We constructed a reusable test collection
of 2,437 manually annotated short humouristic texts. A total of 18 teams submitted 54 runs to the
JOKER Task 2 on humour classification being the most popular JOKER task this year.</p>
      <p>In our evaluation campaign, we observed the best results from the ORPAILLEUR [17] team, whose use
of LLMs obtained the highest overall accuracy. We generally found state-of-the-art LLM-based models
to clearly outperform traditional classifiers. However, the overall and per-class scores of all participants
show significant room for improvement. These results indicate that the semantics and pragmatics of
humour is still challenging for LLMs despite their significant recent advances. Our evaluation setup
itself could also be improved in terms of the quality and quantity of data, or perhaps of providing
individual classification tasks for each technique and genre; these are options that we are exploring for
future iterations of the JOKER track.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This project has received a government grant managed by the National Research Agency under the
program “Investissements d’avenir” integrated into France 2030, with the Reference ANR-19-GURE-0001.
This Lab would not have been possible without the great support of numerous individuals; we would
like to thank in particular Jaap Kamps for his valuable comments and suggestions.
ACM SIGIR Conference on Research and Development in Information Retrieval, Association for
Computing Machinery, New York, NY, 2023, pp. 2796–2806. doi:10.1145/3539618.3591885.
[9] D. P. Varadi, A. Bartulović, JOKER 2024 by AB&amp;DPV: From ‘LOL’ to ‘MDR’ using AI models to
retrieve and translate puns, in: G. Faggioli, N. Ferro, P. Galuscakova, A. García Seco de Herrera
(Eds.), Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[10] S.-H. Wu, Y.-F. Huang, T.-Y. Lau, Humour classification by fine-tuning LLMs: CYUT at CLEF 2024
JOKER Lab subtask humour classification according to genre and technique, in: G. Faggioli, N. Ferro,
P. Galuscakova, A. García Seco de Herrera (Eds.), Working Notes of CLEF 2024 – Conference and
Labs of the Evaluation Forum, 2024.
[11] S. Narayanan, J. J, S. V, CLEF 2024 JOKER Task 3 : Using RoBERTa and BERT-uncased for humour
classification according to genre and technique, in: G. Faggioli, N. Ferro, P. Galuscakova, A. García
Seco de Herrera (Eds.), Working Notes of CLEF 2024 – Conference and Labs of the Evaluation
Forum, 2024.
[12] M. Saipranav, J. Sridharan, G. N. G, CLEF 2024 JOKER Task 3: Using BERT and random forest
classifier for humor classification according to genre and technique, in: G. Faggioli, N. Ferro,
P. Galuscakova, A. García Seco de Herrera (Eds.), Working Notes of CLEF 2024 – Conference and
Labs of the Evaluation Forum, 2024.
[13] H. Baguian, H. N. Ashley, JOKER Track @ CLEF 2024: The Jokesters’ approaches for retrieving,
classifying, and translating wordplay, in: G. Faggioli, N. Ferro, P. Galuscakova, A. García Seco de
Herrera (Eds.), Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[14] R. Subramanian, V. Sivaraman, B. B, Humour insights at JOKER 2024 Task 2: Humour classification
according to genre and technique, in: G. Faggioli, N. Ferro, P. Galuscakova, A. García Seco de
Herrera (Eds.), Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[15] V. M. Palma Preciado, C. Palma Preciado, G. Sidorov, NLPalma Joker 2024: Yet, no humor with
humorousness – Task 2 humour classification according to genre and technique, in: G.
Faggioli, N. Ferro, P. Galuscakova, A. García Seco de Herrera (Eds.), Working Notes of CLEF 2024 –
Conference and Labs of the Evaluation Forum, 2024.
[16] J. V. Kováčiková, M. Šuppa, Thinking, fast and slow: From the speed of FastText to the depth of
retrieval augmented large language models for humour classification, in: G. Faggioli, N. Ferro,
P. Galuscakova, A. García Seco de Herrera (Eds.), Working Notes of CLEF 2024 – Conference and
Labs of the Evaluation Forum, 2024.
[17] P. Epron, G. Guibon, M. Couceiro, CLEF 2024 JOKER Task 2: Leveraging large language models
for humor genre classification, in: G. Faggioli, N. Ferro, P. Galuscakova, A. García Seco de Herrera
(Eds.), Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[18] R. Elagina, P. Vučić, Convergential approach in machine learning for efective humour analysis
and translation, in: G. Faggioli, N. Ferro, P. Galuscakova, A. García Seco de Herrera (Eds.), Working
Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[19] R. R. Dsilva, N. Bhardwaj, CLEF JOKER 2024 Task 2: Who’s laughing now? Humor classification
by genre and technique, in: G. Faggioli, N. Ferro, P. Galuscakova, A. García Seco de Herrera (Eds.),
Working Notes of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[20] R. Mann, T. Mikulandric, CLEF 2024 JOKER Tasks 1–3: Humour identification and classification,
in: G. Faggioli, N. Ferro, P. Galuscakova, A. García Seco de Herrera (Eds.), Working Notes of CLEF
2024 – Conference and Labs of the Evaluation Forum, 2024.
[21] L. Buijs, M. Cazemier, E. Schuurman, J. Kamps, University of Amsterdam at the CLEF 2024 Joker
Track, in: G. Faggioli, N. Ferro, P. Galuscakova, A. García Seco de Herrera (Eds.), Working Notes
of CLEF 2024 – Conference and Labs of the Evaluation Forum, 2024.
[22] S. S. Balaji, S. N. N.S.R, S. S, Vayam Solve Kurmaha @ CLEF 2024: Task 2: Humor classification
according to genre and technique using BERT embeddings and transformers, in: G. Faggioli, N. Ferro,
P. Galuscakova, A. García Seco de Herrera (Eds.), Working Notes of CLEF 2024 – Conference and
Labs of the Evaluation Forum, 2024.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bosser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M.</given-names>
            <surname>Palma-Preciado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          ,
          <article-title>Overview of CLEF 2024 JOKER track: Automatic humor analysis</article-title>
          , in: L.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>G. Q.</given-names>
          </string-name>
          <string-name>
            <surname>Philippe Mulhem</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Schwab</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Soulier</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. M. D. Nunzio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Galuščáková</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Fifteenth International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ), Lecture Notes in Computer Science, Springer,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          , A.-G. Bosser,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          ,
          <article-title>Overview of the CLEF 2024 JOKER Task 1: Humouraware information retrieval</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuscakova</surname>
          </string-name>
          , A. G. Seco de Herrera (Eds.),
          <source>Working Notes of the Conference and Labs of the Evaluation Forum (CLEF</source>
          <year>2024</year>
          ), CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          , A.-G. Bosser,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          ,
          <article-title>Overview of the CLEF 2024 JOKER Task 3: Translation of puns from English to French</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuscakova</surname>
          </string-name>
          , A. G. Seco de Herrera (Eds.),
          <source>Working Notes of the Conference and Labs of the Evaluation Forum (CLEF</source>
          <year>2024</year>
          ), CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>I. Abu</given-names>
            <surname>Farha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. V.</given-names>
            <surname>Oprea</surname>
          </string-name>
          , S. Wilson, W. Magdy, SemEval
          <article-title>-2022 Task 6: iSarcasmEval, intended sarcasm detection in English and Arabic</article-title>
          ,
          <source>in: Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022)</source>
          , Association for Computational Linguistics,
          <year>2022</year>
          , pp.
          <fpage>802</fpage>
          -
          <lpage>814</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          .semeval-
          <volume>1</volume>
          .
          <fpage>111</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Frenda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pedrani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Cignarella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Panizzon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Marco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Scarlini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bernardi</surname>
          </string-name>
          ,
          <article-title>EPIC: Multi-perspective annotation of a corpus of irony, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics</article-title>
          , volume
          <volume>1</volume>
          , Association for Computational Linguistics,
          <year>2023</year>
          , pp.
          <fpage>13844</fpage>
          -
          <lpage>13857</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>774</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N. R.</given-names>
            <surname>Bogireddy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suresh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rai</surname>
          </string-name>
          ,
          <string-name>
            <surname>I'</surname>
          </string-name>
          <article-title>m out of breath from laughing! I think? A dataset of COVID-19 humor and its toxic variants</article-title>
          ,
          <source>in: Companion Proceedings of the ACM Web Conference</source>
          <year>2023</year>
          ,
          <article-title>Association for Computing Machinery</article-title>
          , New York, NY,
          <year>2023</year>
          , pp.
          <fpage>1004</fpage>
          -
          <lpage>1013</lpage>
          . doi:
          <volume>10</volume>
          .1145/ 3543873.3587591.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-G.</given-names>
            <surname>Bosser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M.</given-names>
            <surname>Palma Preciado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          , Overview of JOKER - CLEF
          <article-title>-2023 track on automatic wordplay analysis</article-title>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Giachanou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Aliannejadi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Vlachos</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction</source>
          , volume
          <volume>14163</volume>
          , Springer Nature Switzerland, Cham,
          <year>2023</year>
          , pp.
          <fpage>397</fpage>
          -
          <lpage>415</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -42448-9_
          <fpage>26</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-G.</given-names>
            <surname>Bosser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          , The JOKER Corpus:
          <article-title>English-French parallel data for multilingual wordplay recognition</article-title>
          ,
          <source>in: SIGIR '23: Proceedings of the 46th International</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>