<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Overview of JOKER 2023 Automatic Wordplay Analysis Task 2 - Pun Location and Interpretation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Liana Ermakova</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tristan Miller</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anne-Gwenn Bosser</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Victor Manuel Palma Preciado</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Grigori Sidorov</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adam Jatowt</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Austrian Research Institute for Artificial Intelligence (OFAI)</institution>
          ,
          <addr-line>Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>École Nationale d'Ingénieurs de Brest, Lab-STICC CNRS UMR 6285</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Instituto Politécnico Nacional (IPN), Centro de Investigacion en Computacion (CIC)</institution>
          ,
          <addr-line>Mexico City</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Université de Bretagne Occidentale</institution>
          ,
          <addr-line>HCTI</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Innsbruck</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>8</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>This paper presents an overview of Task 2 of the JOKER-2023 track on automatic wordplay analysis. The goal of the JOKER track series is to bring together linguists, translators, and computer scientists to foster progress in the automatic interpretation, generation, and translation of wordplay. Task 2 is focussed on pun location and interpretation. Automatic pun interpretation is important for advancing natural language understanding, enabling humor generation, aiding in translation and cross-linguistic understanding, enhancing information retrieval, and contributing to the field of computational creativity. In this overview, we present the general setup of the shared task we organized as part of the CLEF-2023 evaluation campaign, the participants' approaches, and the quantitative results.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;wordplay</kwd>
        <kwd>puns</kwd>
        <kwd>computational humour</kwd>
        <kwd>wordplay interpretation</kwd>
        <kwd>wordplay detection</kwd>
        <kwd>pun location</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>natural language understanding, enabling humour generation, aiding in translation and
crosslinguistic comprehension, enhancing information retrieval, and contributing to the field of
computational creativity.</p>
      <p>Interpreting a pun involves two steps: recognizing the word or phrase carrying multiple
meanings (pun location), and then identifying those meanings (pun interpretation). Pun
interpretation systems need to identify potential sources of ambiguity in the context and narrow
down the possible interpretations.</p>
      <p>
        Pun comprehension, when done by humans, involves recognizing that there is a pun, and
then understanding it, thereby finding humour in the unexpected or clever connection between
the diferent meanings of the words involved. Automatic pun interpretation refers to the use of
computational techniques and algorithms to automatically analyze and understand puns without
human intervention. Some systems analyze the words involved in the pun, including their
meanings and relationships with other words based on lexical resources such as dictionaries,
thesauri, or semantic networks such as WordNet [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. Systems can consider the surrounding
context of the pun to gain a better understanding of the intended meaning. This can involve
analyzing the broader text or discourse in which the pun appears, including syntactic and
semantic features.
      </p>
      <p>Humour is an important aspect of interpersonal interactions and our social behavior. Humour
can depend on subjective factors, which makes its automatic processing challenging. Thus,
dealing with humour, even in its written form, becomes a rather complex task even if at first
sight the problem may seem trivial. Wordplay processing is a specific task in a broader area of
automatic humour processing which involves detection, classification, generation, or translation
of humour.</p>
      <p>The paper discusses two distinct subtasks of JOKER-2023: Task 2.1 on pun location in English,
French, and Spanish; and Task 2.2 on pun interpretation in English. Each subtask is presented
individually, covering various aspects such as the objectives, data collection process, evaluation
metrics, approaches used by the participants, and the corresponding results. In total, nine teams
submitted 64 runs in total for Task 2; the breakdown across teams, languages, and subtasks is
given in Table 1.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task 2.1: Pun location</title>
      <sec id="sec-2-1">
        <title>2.1. Task description</title>
        <p>Pun location (Task 2.1) is a finer-grained version of pun detection. The goal is to identify the
words that carry the double meaning in a text which is known a priori to contain a pun. The
double meaning here produces the humorous efect of wordplay. For example, the first of the
following sentences contains a pun where the word propane evokes the similar-sounding word
profane, although the latter is not included in the sentence explicitly, while the second sentence
contains a pun exploiting two distinct meanings of the word interest:
• (1) When the church bought gas for their annual barbecue, proceeds went from the sacred
to the propane.</p>
        <p>• (2) I used to be a banker but I lost interest.</p>
        <p>
          Note that for the pun detection task which is the basis of Task 1 of JOKER-2023 Track, the
correct answer for these two instances would be “true”. Now, for the pun location task, the
correct answers are respectively “propane” and “interest”. System performance is reported in
terms of accuracy for this subtask.
2.2. Data
The pun location data is drawn from the positive examples of JOKER Task 1 [
          <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
          ], with each
text being accompanied by an annotation that reproduces the word being punned upon, as
described above. These positive examples for the pun detection task were short jokes in English,
French, and Spanish with single puns. A detailed description of the English and French pun
detection and location datasets can be found in our SIGIR 2023 resource paper [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>The statistics on data per language are given in Table 2. These statistics present the actual
numbers used for the assessment, representing the efective figures for both the test and training
data. Nevertheless, when providing files to the participants, we included the training data
within the input of the test file. Incorporating the training data within the dataset, for which
participants were required to make predictions, enables a comprehensive evaluation of the
systems’ performance on both the training and testing data. By incorporating the training
data, we can assess the systems’ ability to generalize to unseen test data by observing their
performance on familiar training examples.</p>
        <p>The test and training data sets were provided to participants as JSON or delimited text files
with fields containing the text of the punning joke and a unique ID. For training data for the pun
location task, there is an additional field reproducing the pun word. System output is expected
as a JSON or delimited text file with fields for the run ID, text ID, the pun word, and a boolean
lfag indicating whether the run is manual or automatic.
Input format. The base data is provided in JSON and CSV formats with the following fields:
id a unique identifier
text the text of the instance of wordplay</p>
        <sec id="sec-2-1-1">
          <title>Input example:</title>
          <p>[{"id":"en_135",
"text":"Cleopatra was the Pharaohs one of all."},
{"id":"en_226",
"text":"At a flower show the first prize is often a bloom ribbon."}
]
Qrels. We provide training data as JSON or TSV qrels files with the following fields:
id a unique identifier from the input file
location the portion of the text containing the wordplay</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Example of a qrel file:</title>
          <p>[{"id":"en_135","location":"Pharaohs"},
{"id":"en_226",
"location":"bloom"}
]
Output format. Systems were expected to submit their results in a TREC-style JSON or TSV
ifle with the following fields:
run_id run ID starting with &lt;team_id&gt;_&lt;task_id&gt;_&lt;method_used&gt; – e.g.,
UBO_task_2.1_</p>
          <p>TFIDF
manual flag indicating if the run is manual 0,1</p>
          <p>Example of an output file:
[{"run_id":"team1_task_2.1_TFIDF",
"manual":0,
"id":"en_135",
"location":"Pharaohs"},
{"run_id":"team1_task_2.1_TFIDF",
"manual":0,
"id":"en_226",
"location":"bloom"}]</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.3. Participants’ approaches</title>
        <p>
          Eight teams participated in Task 2.1:
1. The AKRaNLU team participants [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] employ the token classification method with a
tagging schema that relies on assigning a tag of 1 to every pun word and 0 to every word
that is not a punning word.
2. The MiCroGerk team [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] chose an large language model (LLM) approach for Task 2.1,
using T5 (SimpleT5), BLOOM, and models from OpenAI and AI21. They also submitted a
baseline that uses last word in the sentence as a prediction, as well as a random baseline.
It is noteworthy that the BLOOM model presented the worst results compared to the
others.
3. The Smroltra team [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] observed that models based on GPT-3, SpaCy, T5, and BLOOM
showed very good performance when it came to Spanish and English, while for French
the results were worse. This was particularly the case for SpaCy, which is believed to be
not as developed for French as for English.
4. TeamCAU [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] used various LLMs. T5 showed good results in comparison to BLOOM and
models from AI21 (albeit for partial runs only).
5. FastText, Ridge, Naive Bayes, SimpleT5, and SimpleTransformersT5 were used by the
participants of ThePunDetectives team [10]. They found the best results to be produced
by the pre-trained models. In particular, T5 achieved good performance, as predicted by
the authors.
6. For the location tasks, the UBO team [11] opted to use T5 (SimpleT5).
7. The Croland team [12] used GPT-3.
8. The Les_miserables team (who did not submit a system description paper) submitted two
baseline runs, one where the system selects the final word of the sentence as the pun
location, and another run that randomly predicts words; they also submitted a run using
the T5 (SimpleT5) model.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.4. Results</title>
        <p>The eight teams submitted 47 runs in total for all three languages. Table 3 reports the
participants’ results for wordplay location tasks in English, French, and Spanish on the test set. This
table provides a comprehensive overview of the participants’ systems’ performances and their
respective scores or metrics achieved in locating wordplay instances in the given languages. As
some participants submitted only partial runs, we provide two sets of accuracy scores: those
labeled A are based on the total number of instances in the test set, while those labeled A* are
calculated only using the actual number of attempted instances (#).</p>
        <p>Accuracy scores for pun location in English and Spanish (A ≈ 80) are roughly twice as good
as those for French (A ≈ 40). For comparison, the predictions made on the last word for test
sets in English and Spanish are around 50% while for French this score goes almost up to 30%.
Thus, the improvement for French over a very simple baseline is rather not high. The significant
improvement over this simple last-word baseline for English and Spanish could be explained by
the fact that participants used large language models (e.g., GPT-3 or BLOOM) that might have
included in their training data some of the same puns found in our corpus. By contrast, the
French wordplay data was largely constructed by us and not previously published online.</p>
        <p>Table 4 shows the results of the participants for wordplay location in English, French, and
Spanish on the training set. We observe a significant diference between the best performance
on the test and training data for French obtained by T5 model trained on our data as well as
the results of the best-performing team AKRaNLU for French. These results suggest overfitting
issues. The performance of a few models is comparable to or even lower than the predictions
made by returning the last word in the text.</p>
        <p>Several teams submitted runs by applying the same methods yet implemented, trained, or
ifne-tuned diferently. Significant disparities can be observed, even in the last-word baseline
outcomes (such as Les_miserables_word and MiCroGerk_lastWord), which could be attributed
to variations in tokenization methods. Distinguishing variations in prediction accuracy are also
noticeable between diferently trained T5 models as well as prompt-based language models.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Task 2.2: Pun interpretation</title>
      <sec id="sec-3-1">
        <title>3.1. Task description</title>
        <p>In this task, systems must describe the semantics –i.e., the two meanings – of the pun. In
JOKER2023, these semantic annotations are in the form of a pair of lemmatized word sets. Following the
practice used in lexical substitution datasets, these word sets contain the synonyms (or if absent,
then hypernyms) of the two words involved in the pun, except for any synonyms/hypernyms
that happen to share the same spelling with the pun as written.</p>
        <p>For example, for the punning joke introduced in Example 1 above, the word sets are {gas,
fuel} and {profane}, and for Example 2, the word sets are {involvement} and {fixed charge , fixed
cost, fixed costs }.</p>
        <p>
          Results for Task 2.2 are scored as the average score for each of punning word senses. Systems
need to guess only one word for each sense of the pun; a guess is considered correct if it matches
any of the words in the gold-standard set. For example, a system guessing {fuel}, {profane} would
receive a score of 1 for Example 1, and a system guessing {fuel}, {prophet} would receive a score
of 1/2.
3.2. Data
For the English pun interpretation data, we manually annotated each pun according to its
senses in WordNet 3.1 and then automatically extracted the synonyms (or if there were none,
the hypernyms) of those words to form the two word sets. In some cases, one or both of the
senses of the pun was not present in WordNet, or WordNet contained neither synonyms nor
hypernyms for the annotated senses. (This was particularly the case with adjectives and adverbs,
which WordNet does not arrange into a hypernymic hierarchy.) In these cases, we sourced the
synonym/hypernym sets from human annotators. For the French data, we used a simplified
FR
A
version of the annotation made in JOKER-2022 [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>Input format The base data for test and training are provided in JSON and CSV formats with
the following fields:
id a unique identifier
text the text of the instance of wordplay
{"id":"en_226",
"text":"At a flower show the first prize is often a bloom ribbon."}
]
Qrels We provide training data in the format of JSON or TSV qrels files with the following
ifelds:
id a unique identifier from the input file
location the portion of the text containing the wordplay
interpretation synonyms or hypernyms of the two meanings of the wordplay</p>
        <p>Example of a qrel file:
[{"id":"en_135",
"location":"Pharaohs",
"interpretation":"Pharaoh;Pharaoh of Egypt / fair"},
{"id":"en_226",
"location":"bloom",
"interpretation":"blossom;flower / bluish;blue;blueish"}
]
Output format Systems were expected to provide their results as a TREC-style JSON or TSV
format with the following fields:
run_id run ID starting with &lt;team_id&gt;_&lt;task_id&gt;_&lt;method_used&gt; – e.g.,
UBO_task_2.2_</p>
        <p>BLOOM
manual flag indicating if the run is manual 0,1
id a unique identifier from the input file
location the portion of the text containing the wordplay
interpretation synonyms or hypernyms of the wordplay meanings
[{"run_id":"team1_task_2.2_manual",
"manual":1,
"id":"en_135",
"location":"Pharaohs",
"interpretation":"Pharaoh;Pharaoh of Egypt / fair"},
{"run_id":"team1_task_2.2_manual",
"manual":1,
"id":"en_226",
"location":"bloom",
"interpretation":"blossom;flower / bluish;blue;blueish"}]</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Participants’ approaches</title>
        <p>
          The teams’ approaches were as follows:
1. For pun interpretation, the AKRaNLU team participants [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] used the results from the pun
location subtask to disambiguate the appropriate senses of the pun word based on the
sentence content and find two synonyms for those senses, sourced from WordNet, that
were most similar to sentence embedding.
2. The MiCroGerk team [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] submitted four runs for the interpretation task based on LLMs,
such as T5 (SimpleT5), BLOOM, and models from OpenAI and AI21.
3. The Smroltra team [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] submitted six runs based on GPT-3 and BLOOM, SpaCy, T5 and
their combinations with WordNet for location prediction.
4. The UBO team [11] applied T5 (SimpleT5) to predict the interpretation of puns in English.
5. The UBO-RT team [13] post-edited output of ChatGPT (C&amp;O in Tables 5 and 6). A
zeroshot strategy was used in their approach and the analysis of the results reveals quite poor
capabilities of ChatGPT in interpreting puns, especially those involving homophonic
components.
6. The Croland team [12] used GPT-3.
7. The Les_miserables team (who did not submit a system description paper) submitted a
run using the T5 (SimpleT5) model to predict pun interpretation in English.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Results</title>
        <p>We show in Table 5 the results of the pun interpretation task for English. We do not report the
results for French as only two teams submitted runs for this language, one of which was heavily
post-processed manually. This means it is impossible to draw any conclusions on the French
dataset.</p>
        <p>For the English pun interpretation task, seven teams submitted 15 runs in total. As we
expected, the majority of participants opted to use LLMs, resulting in the generation of partial
runs due to the eficiency constraints associated with these models. Only five runs out of the
total number of submissions involved the entire testing data (1,192), hence the comparison
run
count
score
is somewhat dificult. Nevertheless, when focussing on the full runs, we observe that the
maximum accuracy was obtained by the team Les_miserables who used the T5 model, achieving
47.4%. A similar result was also obtained by the UBO team who also applied the T5 model. For
comparison, a baseline approach with SpaCy gives 19.76% accuracy. This underscores the utility
of large language models for the interpretation of wordplay. The best partial score was obtained
by the UBO-RT team who used the post-processed results generated by ChatGPT. But even this
heavily manually post-processed run obtained only 70%.</p>
        <p>Additionally, for completeness, Table 5 provides results of participating systems for the
training data. Performance on the training data remains low, which may suggest the overall
dificulty of this task. On the other hand, results of some models such us T5 and SpaCy exhibited
notably superior performance on the training set, which suggests the possibility of overfitting.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>
        In this paper, we have described Task 2 of the JOKER track at CLEF 2023, consisting of pun
location and interpretation challenges. We extended our previously described dataset [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
by introducing semantic annotation for wordplay in English and French. Furthermore, we
constructed a corpus for pun location in Spanish.
      </p>
      <p>Multiple teams submitted runs using similar methods, but with variations in implementation,
training, or fine-tuning approaches for both subtasks. These variations entailed large diferences
in the performance of the systems.</p>
      <p>Our results in general suggest that wordplay location is still a challenge for LLMs despite their
recent significant advances. Interestingly, we found that the results for the French language
when it comes to pun location are quite low (half of the scores of English and Spanish) which
run
count
score
we attribute to diferent data creation procedures. Some of the puns found in our corpus for
English and Spanish might have been “known” by large language models that were used by
participants, as the data were sourced from the web for these languages. On the other hand,
French data was novel and largely constructed by us. This calls for the future to construct
wordplay datasets from scratch rather than sourcing humour from external sources like the
web.</p>
      <p>The results suggest the overfitting problem for models trained on our data (e.g., via the
SimpleT5 library). The diference between training and test data observed for the prompt-based
models is small as they are not actually trained on our data, but might be influenced by examples
from the training set used in the prompts.</p>
      <p>Automatic pun interpretation was the second subtask of Task 2. It is quite a challenging
task due to the inherent ambiguity and creativity of puns, yet it fits into the recent focus on
explainable AI and explainable/interpretable decision systems. Puns often rely on cultural
knowledge, background information, and linguistic subtleties that can be dificult to capture
computationally. Still, researchers continue to explore various approaches, including rule-based
methods, machine learning models, and deep learning techniques to improve automatic pun
interpretation systems. Our results data indicate that the T5 model performs well for this task.
However, we could compare the results only for English language, as for French there were
only two runs that were not fully automatic. We received many partial runs due to token/time
constraints of LLMs and, therefore, apart from efectiveness, the eficiency of the approaches
should be considered in future research. The results suggest the dificulty of pun interpretations
at least in the particular settings that we use in subtask 2.2 (reliance on WordNet synonyms and
hypernyms).</p>
      <p>Additional information on the track is available on the JOKER website: http://www.
joker-project.com/</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This project has received a government grant managed by the National Research Agency under
the program “Investissements d’avenir” integrated into France 2030, with the Reference
ANR-19GURE-0001. JOKER is supported by La Maison des sciences de l’homme en Bretagne. For their
help and support in the first Spanish pun translation contest, we thank Carolina Palma Preciado,
Leopoldo Jesús Gutierrez Galeano, Khatima El Krirh, Nathalie Narváez Bruneau, and Rachel
Kinlay. We also thank all other colleagues and students who participated in data construction,
the translation contests, and the CLEF JOKER track.
[10] F. Ohnesorge, M. Á. Gutiérrez, J. Plichta, CLEF 2023 JOKER Tasks 2 and 3: using NLP
models for pun location, interpretation and translation, in: [14], 2023.
[11] Q. Dubreuil, UBO Team @ CLEF JOKER 2023 track for Task 1, 2 and 3 - applying AI models
in regards to pun translation, in: [14], 2023.
[12] J. Komorowska, I. Čatipović, D. Vujica, CLEF2023’ JOKER Working Notes, in: [14], 2023.
[13] O. Brunelière, C. Germann, K. Salina, CLEF 2023 JOKER Task 2: using Chat GPT for pun
location and interpretation, in: [14], 2023.
[14] Proceedings of the Working Notes of CLEF 2023: Conference and Labs of the Evaluation
Forum, CEUR Workshop Proceedings, CEUR-WS.org, 2023.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Regattin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-G.</given-names>
            <surname>Bosser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Borg</surname>
          </string-name>
          , Élise Mathurin,
          <string-name>
            <given-names>G. L.</given-names>
            <surname>Corre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Araújo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hannachi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Boccou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Digue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Damoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Jeanjean</surname>
          </string-name>
          , Overview of JOKER@CLEF 2022:
          <article-title>Automatic wordplay and humour translation workshop</article-title>
          , in: A. BarrónCedeño, G. D. S. Martino,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Esposti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macdonald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Pasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction: Proceedings of the Thirteenth International Conference of the CLEF Association (CLEF</source>
          <year>2022</year>
          ), volume
          <volume>13390</volume>
          of Lecture Notes in Computer Science, Springer, Cham,
          <year>2022</year>
          , pp.
          <fpage>447</fpage>
          -
          <lpage>469</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -13643-6_
          <fpage>27</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-G.</given-names>
            <surname>Bosser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M. P.</given-names>
            <surname>Preciado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          ,
          <article-title>Overview of JOKER 2023 automatic wordplay analysis task 1 - pun detection</article-title>
          ,
          <source>in: [14]</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-G.</given-names>
            <surname>Bosser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M. P.</given-names>
            <surname>Preciado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          ,
          <article-title>Overview of JOKER 2023 Automatic Wordplay Analysis Task 3 - Pun translation</article-title>
          , in: [14],
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-G.</given-names>
            <surname>Bosser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M. P.</given-names>
            <surname>Preciado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          , Overview of JOKER - CLEF-2023
          <source>Track on Automatic Wordplay Analysis</source>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Giachanou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Aliannejadi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Vlachos</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>CLEF'23: Proceedings of the Fourteenth International Conference of the CLEF Association, Lecture Notes in Computer Science</source>
          , Springer,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Dsilva</surname>
          </string-name>
          ,
          <source>AKRaNLU @ CLEF JOKER</source>
          <year>2023</year>
          :
          <article-title>Using sentence embeddings and multilingual models to detect and interpret wordplay</article-title>
          ,
          <source>in: [14]</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>O.</given-names>
            <surname>Popova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dadić</surname>
          </string-name>
          ,
          <article-title>Does AI have a sense of humor? CLEF 2023 JOKER tasks 1, 2 and 3</article-title>
          :
          <string-name>
            <surname>Using</surname>
            <given-names>BLOOM</given-names>
          </string-name>
          ,
          <article-title>GPT, SimpleT5, and more for pun detection, location, interpretation and translation</article-title>
          , in: [14],
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-G.</given-names>
            <surname>Bosser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          , The JOKER Corpus:
          <article-title>English-French parallel data for multilingual wordplay recognition</article-title>
          ,
          <source>in: SIGIR '23: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , Association for Computing Machinery, New York, NY,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .1145/ 3539618.3591885, to appear.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Prnjak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Davari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Schmitt</surname>
          </string-name>
          ,
          <source>CLEF 2023 JOKER Task 1</source>
          ,
          <issue>2</issue>
          , 3
          <article-title>: pun detection, pun interpretation, and pun translation</article-title>
          ,
          <source>in: [14]</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Anjum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lieberum</surname>
          </string-name>
          ,
          <source>Exploring Humor in Natural Language Processing: A Comprehensive Review of JOKER Tasks at CLEF Symposium</source>
          <year>2023</year>
          , in: [14],
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>