<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Madrid, Spain
* Corresponding author.</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Overview of the CLEF 2025 JOKER Task 2: Wordplay Translation from English into French</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Liana Ermakova</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anne-Gwenn Bosser</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tristan Miller</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ricardo Campos</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Austrian Research Institute for Artificial Intelligence (OFAI)</institution>
          ,
          <addr-line>Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Bretagne INP - ENIB, Lab-STICC CNRS UMR 6285</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Computer Science, University of Manitoba</institution>
          ,
          <addr-line>Winnipeg</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>INESC TEC</institution>
          ,
          <addr-line>Porto</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Université de Bretagne Occidentale</institution>
          ,
          <addr-line>HCTI</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Beira Interior</institution>
          ,
          <addr-line>Covilhã</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>This paper describes Task 2 of the CLEF 2025 JOKER track on the translation of puns from English into French. We outline the overall structure and setup of the shared task, discuss the approaches employed by the participants, and present and analyse the results they achieved. We also describe experiments with a promising new approach for the automatic evaluation of pun translation. Despite the significant improvements observed this year by participating systems, most of which used state-of-the-art large language models, we find wordplay translation to remain a complex and demanding task. Among the manually evaluated translations, 37.5% successfully preserved the meaning and involved wordplay, with success rates per English pun varying widely.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;wordplay</kwd>
        <kwd>puns</kwd>
        <kwd>computational humour</kwd>
        <kwd>machine translation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>of the original wordplay was only 6%. This highlights the need for increased community focus on this
task.</p>
      <p>This year nine teams submitted 52 runs for Task 2, reflecting the community’s stable interest automatic
pun translation. Table 1 shows the number of runs submitted by each participating team.</p>
      <p>One of the innovations made in CLEF 2025 JOKER was infrastructural: for the first time, we ran the
task on Codabench2 [15], the Free Software web-based platform for organising AI benchmarks (see
Figure 1). Codabench greatly facilitated running the track in 2025 and attracted many new participants,
all of whom had full access to the competition, including the submission and leaderboard pages. We
continue to receive new registrations and post-competition submissions. However, in this paper we
present only runs submitted before the oficial results were communicated to the participants.</p>
      <p>
        In the remainder of this paper, we describe the data used in Task 2 (Section 2), the evaluation measures
(Section 3), and the participants’ approaches (Section 4), and then present an analysis of their results
for both the training and test data (Section 5). In addition to traditional machine translation evaluation
measures, such as BLEU [16] and BERTScore [17], we examined the participants’ performance using
the dataset we created to identify words or phrases that have multiple meanings (pun locations) for the
CLEF 2023 JOKER Task 2 [
        <xref ref-type="bibr" rid="ref4">18, 4, 19</xref>
        ]. We show that this approach is promising to evaluate translation of
wordplay based on multiple meanings. Section 6 concludes the paper.
2https://www.codabench.org/competitions/8748/
#
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Data</title>
      <sec id="sec-2-1">
        <title>2.1. Training data</title>
        <p>
          The training data for Task 2, which builds on previous editions [
          <xref ref-type="bibr" rid="ref3 ref4">3, 20, 4, 5</xref>
          ], consists of 1,405 instances of
wordplay in English, with a total of 5,838 French translations sourced from human professionals with
manually annotated words or phrases with multiple meanings (pun locations). Pun location annotations
were collected for previous JOKER evaluation campaigns [
          <xref ref-type="bibr" rid="ref4">19, 4, 18</xref>
          ]. Table 2 shows a histogram of the
number of references and distinct locations per English pun in the training data.
        </p>
        <p>We provide training data in the format of JSON qrels files with the following fields:
• id_en: a unique identifier from the input file. Note that this identifier is not unique in the file, as
the same English pun might have multiple French translations.
• en: the text of the instance of source wordplay in English. Note that the texts in English are not
unique in the file, as the same English pun might have multiple French translations.
• fr: translation of the wordplay into French
Example of a training file:
"id_en":"en_1",
"en":"I used to be a banker but I lost interest",
"fr":"J’ai été banquier mais j’en ai perdu tout l’intérêt."</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Test data</title>
        <p>
          For the 2025 edition, we collected 2,615 new manual translations of 1,682 distinct puns in English with
manually annotated pun locations that we used for the test data. Some of the pun location annotations
were collected for previous JOKER evaluation campaigns [
          <xref ref-type="bibr" rid="ref4">19, 4, 18</xref>
          ]. We expanded this data with
annotations of new references. Table 3 shows a histogram of the number of references and distinct
locations per English pun in the test data. For 25% of English puns from the test data, we have multiple
references and multiple locations. There are 1,382 English puns with a single location, while only 1,252
of them have a single reference. However, we have much more multiple references and distinct locations
on training data.
        </p>
        <p>The test input data is provided in JSON format with the following fields:
• id_en: a unique identifier
• en: the text of the instance of source wordplay in English</p>
        <sec id="sec-2-2-1">
          <title>An input example is as follows:</title>
          <p>"id_en":"en_1",
"en":"I used to be a banker but I lost interest"
The test output was requested to be provided in JSON format with the following fields:
• run_id: Run ID starting with &lt;team_id&gt;_&lt;task_id&gt;_&lt;method_used&gt;, e.g. UBO_task_3_BLOOM
• manual: Whether the run is manual {0,1}
• id_en: a unique identifier from the input file
• en: the text of the instance of source wordplay in English
• fr: translation of the wordplay into French
An ouput example is as follows:
"run_id":"team1_task_3_DeepL",
"manual":0,
"id_en":"en_1",
"en":"I used to be a banker but I lost interest"
"fr":"J’ai été banquier mais j’en ai perdu tout l’intérêt"</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Evaluation</title>
      <sec id="sec-3-1">
        <title>We evaluated the runs with the following metrics:</title>
        <p>BLEU (BiLingual Evaluation Understudy) computes the translation’s overlap in vocabulary overlap
with a reference translation [16]. We used the sacreBLEU implementation [21] with the default
tokeniser 13a. We report the BLEU score (harmonic mean) and the BLEU precisions for -grams
for  = 1, 2, 3, 4.</p>
        <p>BERTScore computes tokenwise similarity scores between the candidate translation and a reference
translation using contextual embeddings [17]. We used the Python implementation from the
bert-score package.3 We report mean values of BERTScore precision, recall, and F1 over all
references.</p>
        <p>Pun location–based evaluation allows for a more fine-grained analysis of generated translations.</p>
        <p>
          We checked for words or phrases with multiple meanings (pun locations) from the reference texts
by combining French reference translations with pun location annotations from the dataset used
for JOKER 2023’s Task 2: Pun Location and Interpretation [
          <xref ref-type="bibr" rid="ref4">4, 18, 19</xref>
          ]. We completed this data
with pun location annotations of the new references.
        </p>
        <p>Manual evaluation consisted of human assessments of 1,297 French translations of 50 distinct source
English puns in terms of meaning preservation and the presence of wordplay. This manual
evaluation was performed by a Master’s student in translation who specialises in wordplay
translation and is a native French speaker.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Participants’ approaches</title>
      <p>Nine teams submitted 52 oficial runs for this task. Statistics on the runs are presented in Table 1. Team
names are reported according to the participant names listed on Codabench. The approaches used were
as follows:
arampageos [9] This team combined neural machine translation systems with a handcrafted
translation dictionary of particularly challenging puns. The machine translation systems included Google
Translate, Argos Translate, the Helsinki-NLP/opus-mt-en-fr models, Facebook’s M2M100 (418M and
1.2B), MBART50, and NLLB (1.3B and distilled 600M). Their two-stage pipelines first checked whether
the input pun matched the curated set, otherwise forwarding the input to the machine translation
system.
verbanex [14] This team relied on extensive data preprocessing, including sentiment classification
and phoneme conversion, to help the trained translation model capture emotional tone and
pronunciation ambiguities. They used two diferent fine-tuning strategies – full parameter optimisation and
parameter-eficient adaptation techniques – with the mBART-50 English-to-French translation model.
rdtaylorjr [13] This participant relied on a three-stage approach. The first stage consisted of training
multiple LLMs (provided by openAI, Google, Mistral, or DeepSeek) using a contrastive learning approach.
In addition to the training set we provided, they used data from the JOKER 2023 shared task on pun
location and interpretation, as well as a contrastive learning dataset constructed by neutralising puns of
their French dataset. The second stage of the approach is based on chain-of-thought prompting making
use of semantic and phonetic embeddings for the French language. Finally, evaluator agents were used
to iterate over various properties of the proposed translations (conserving literal/contextual meaning,
emotion level, and understandability in the target language).
alecs and kamps [8] These participants used a fine-tuned MarianMT sequence-to-sequence model,
T5ForConditionalGeneration, T5-base, Meta AI NLLB-200-1.3B, and mBART-large-cc25.
pjmathematician [12] This team fine-tuned diferent Qwen models, including the Qwen2.5-14B,
experimenting with diferent LoRA parameters on the provided corpus. They then used a simple
prompting approach for requesting translations of puns.
igoranchik [11] This team used supervised fine-tuning with the aim of forcing a model to learn
higher quality responses. They also used an Adaptive Rejection Preference Optimisation (ARPO) [22]
implementation4 in an attempt to enhance humour retention.
cryptix and sarath_kumar [10] These participants used back-translation for data augmentation.
They fine-tuned the MarianMT model and used a loss function combining humour preservation metrics
from a rule-based module evaluating humour preservation with standard BLEU metrics.</p>
      <p>All participants who submitted runs also submitted system description papers to the Working Notes
volume [23]. Two teams from the same university (alecs and kamps) submitted a single joint report, as
did teams cryptix and sarath_kumar, resulting in a total of seven Working Notes from the participants
of Task 2. Despite the requirement to include the team ID in the run name, participants’ submissions
often difered in their run names, registration details, and Codabench IDs. We manually matched the
Working Notes with the submitted runs and report the results using the team names provided in those
submissions.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <sec id="sec-5-1">
        <title>5.1. Test data</title>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Manual evaluation</title>
        <p>Among the translations evaluated manually (1,297 French translations of 50 distinct source English
puns in total), we discovered 50 cases where the output was simply “[unk]”, three empty translations,
72 untranslated texts, and four incomplete translations. Two outputs were a mix of English and French
(“Horloge en forme de the thinker qui annonce l’heure en disant I think it s 20.25 pm” and “I have to keep
count
mean
std
min
25%
50%
75%
max
this fire alight, a crié Tom.”) and seven had useless repetitions (e.g., “Les vendeurs de chips ne peuvent pas
vendre leurs produits, ils ne peuvent pas vendre leurs produits.”)</p>
        <p>There were 572 translations that did not preserve the meaning of the source text, while 464 preserved
it fully and 145 preserved it partially. There were 80 translations containing wordplay that did not
preserve the meaning of the source pun (e.g., “Chips vendors don’t get the dough unless their products
sell.” → “Les agriculteurs ne font pas de blé s’il n’y a pas de blé dans les champs”, or “The scientist had
trouble reducing the liquid, he just couldn’t concentrate.” → “Le brasseur n’arrive pas à maintenir
la mousse, pourtant il se fait mousser.”) Among the translations that preserved the meaning fully or
partially, 487 involved wordplay, while 122 did not, resulting in 37.5% of successful translations over
the total number of manually evaluated translations.</p>
        <p>The number of successful translations per English pun is given in Figure 2. For half of source English
puns the success is less than 25% with mean 10 and median 9 successful translation per source. The
maximal number of successful runs per English source is 26 over 49. For 25% of English puns there
is 3 or less successful translations, highlighting the inherent dificulty of rendering certain wordplays
efectively across languages.</p>
        <p>The only English pun without successful translation was “Having too many axe-like tools to do a
particular job only adze to the confusion.” Two English puns had only a single successful translation:
• “The geneticist taught his students how to mendel defective genes” → “Le généticien a appris à
ses étudiants à repriser leurs jeans. . . et leurs gènes !”
(dsgt_o4_mini_chain_of_thought_phonetic_embeddings)
• “Volts – the dance you perform after an electric shock” → “En anglais, le verbe « voltige » désigne
une danse après une décharge électrique” (duth_xanthi_bloomz3b_local)</p>
        <p>The results of manual evaluation per run are given in Table 7. According to our manual evaluation,
the best runs were dsgt_o4_mini_multi_agent_discriminator and
dsgt_o4_mini_chain_of_thought_phonetic_embeddings [13], achieving 37 and 36 successful translations, respectively. These results
significantly outperformed the third-best run, which achieved only 26 successful translations. These
results are consistent with the evaluation in terms of location-based metric. The runs with lowest
performance according to the location-based metric are also low-scored by the expert.</p>
        <p>The Pearson correlation coeficient between the manually attributed scores and the location-based
metrics is 0.84, indicating a strong positive relationship between the two. The scatter plot in Figure 3
shows this relationship, including a regression line with a 95% confidence interval, which further
illustrates the consistency of the association across the dataset. Such a high correlation suggests that
the location-based evaluation captures key aspects of translation quality that align closely with expert
judgments. Consequently, it can be considered a reliable proxy for assessing the quality of wordplay
translation, ofering a scalable and less resource-intensive alternative to manual evaluation. However,
120
n
ito100
a
c
o
L 80
60
40
20
5
10
15</p>
        <p>20
Success
25
30
35
further analysis is needed, as the current results are based on the oficial scores presented in Tables 6
and 7, and may not fully capture variability across diferent evaluators or contexts. Note that these
results are drastically diferent from the ranking based on BLEU and BERT scores.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Training data</title>
        <p>As in previous years, runs were submitted for both the training and test datasets in order to analyse
potential overfitting and related efects. Tables 8, 9, and 10 report, respectively, the results based on
BLEU, BERTScore, and the pun location–based metric for Task 2 on the training data.</p>
        <p>Three runs (Cryptix_rulebased [10], pjmathematician_Q25-14 [12], yourteam_rulebased [10])
achieved a BLEU score of 100 on the training data, followed by 99.77 for teamX_final [ 14].
For the next four runs, all submitted by the University of Amsterdam
(UvA_finetunedNLLB1.3B, UvA_finetunedNLLB-1.3B&amp;finetunedroBERTa, UvA_finetunedMarianMT,
UvA_finetunedMarianMT&amp;finetunedroBERTa) [ 8], we observe an almost twofold drop in performance.
UvA_finetunedNLLB1.3B remains high on the test data, ranking third according to the BLEU evaluation. Both rule-based
runs (Cryptix_rulebased and yourteam_rulebased) are positioned at the bottom of the table in terms
of test results. The best run in terms of BLEU on test set, Skommarkhos_Lucie_SFT [11] (43.33), has
a BLEU score of 49.98 on the training data. pjmathematician_Q25-14 achieved BLEU scores of 100
on the training data and 39 on the test data. However, the same run ranked third-highest according
to the location-based metric and fell within the second tier of manually ranked runs, highlighting a
discrepancy between traditional BLEU evaluation and alternative assessment methods.</p>
        <p>The BLEU score results on the training data have similar trend as the BERTScore (see Table 9).
However, the top-scored teamX_final [ 14] on the training data has much lower rank on the test data
according to BERTScore.</p>
        <p>Cryptix_rulebased [10], pjmathematician_Q25-14 [12], and yourteam_rulebased [10] have 391
(27.83%) successful locations sharing the first rank on the training data closely followed by
teamX_final [14] with the result 380 (27.05%). Cryptix_rulebased [10], yourteam_rulebased [10], and
teamX_final [14] are in the bottom of the table according to the location-based metric on the test set suggesting
overfitting on the training data. The next runs are
dsgt_o4_mini_chain_of_thought_phonetic_embeddings and dsgt_o4_mini_multi_agent_discriminator [13] with results 184 (13.10%) and 183 (13.02%)
respectively. These runs are the best according to manual evaluation and location-based metric on the
test set, suggesting generalisation capacity. They are followed by the four runs of the University of
Amsterdam [8] with slightly lower scores, and then teamX_aug [14]. Note that teamX_aug is ranked third
according to the location-based metric on the test set. Thus, these two sets of results are comparable on
the test and training data.</p>
        <p>Note that it is problematic to directly compare the absolute scores between the training and test data,
rather than run ranks, as the number of distinct locations per English pun is considerably higher in the
training data than in the test data. (See Tables 3 and 2.)</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>
        In this paper, we have described the wordplay translation task of the JOKER track at CLEF 2025. This
year, we expanded the corpus used in previous editions of the task [
        <xref ref-type="bibr" rid="ref3 ref4">3, 19, 4</xref>
        ] by introducing 1,682
new distinct source texts with 2,615 corresponding reference translations created by professional
French native-speaker translators for the test purpose. We manually annotated pun locations in French
translations in order to provide an automatic evaluation that takes into account pun ambiguity.
      </p>
      <p>Nine teams submitted 52 runs for Task 2, demonstrating stable interest from the community in our
perennial pun translation task. Participants used a variety of methods, including LLMs, commercial
machine translation engines, out-of-the-box translation models, rule-based approaches, and various
ifne-tuning and training techniques to discriminate wordplay from non-wordplay.</p>
      <p>We evaluated the participants’ results using automated measures, specifically BLEU and BERT scores
both on test and training sets. According to these automatic measures on the test data, the best results
were achieved by the fine-tuned approaches of the teams Skommarkhos [ 11] and the University of
Amsterdam [8]. Three runs – Cryptix_rulebased, pjmathematician_Q25-14, and yourteam_rulebased –
achieved perfect BLEU scores of 100 on the training data, with another run scoring 99.77. However,
both rule-based runs (Cryptix_rulebased and yourteam_rulebased) performed poorly on the test data,
ranking near the bottom. The pjmathematician_Q25-14 run also showed signs of overfitting, achieving
a BLEU score of 39 on the test data despite its perfect training score. The BLEU score results on the
training data show a similar trend to those of BERTScore.</p>
      <p>On the pun location–based metric, the best runs were dsgt_o4_mini_multi_agent_discriminator
and dsgt_o4_mini_chain_of_thought_phonetic_embeddings [13]. These were followed by
teamX_aug [14] and pjmathematician_Q25-14 [10], which scored lower on BLEU and BERTScore, and then
by UvA_finetunedMarianMT[ 8], which also ranked fourth by BERTScore. Rule-based runs achieved
top scores on the training set but performed poorly on the test set, suggesting overfitting. In contrast,
dsgt_o4_mini_multi_agent_discriminator and dsgt_o4_mini_chain_of_thought_phonetic_embeddings
generalised well, ranking highest on both manual and location-based evaluations, followed by the
University of Amsterdam runs and teamX_aug.</p>
      <p>Manual evaluation confirmed dsgt_o4_mini_multi_agent_discriminator and
dsgt_o4_mini_chain_of_thought_phonetic_embeddings [13] as the best-performing runs, consistent with the location-based
metric. Conversely, the lowest-ranked runs by this metric were also rated poorly by the expert,
reinforcing the alignment between automated and manual evaluations.</p>
      <p>Overall, we observe significant improvements in participants’ results compared to previous years,
based on both manual (up to 74% of successful translations for one team [13]) and location-based
evaluations, particularly on the training data (up to 28% for rule-based appraoches on the training
data). However, despite these significant improvements, wordplay translation remains a complex and
demanding task. With the exception of runs exhibiting overfitting on the training data, the
locationbased evaluation results are consistent with those of previous years. Among the manually evaluated
translations, 37.5% successfully preserved the meaning and involved wordplay, with success rates per
English pun varying widely –half having under 25% good translations in French, a median of nine
successful translations, and 25% with three or fewer – underscoring the dificulty of efectively rendering
wordplay across languages.</p>
      <p>One of the major obstacles in the development of wordplay machine translation is its evaluation.
Destroying the wordplay may result in the text becoming nonsensical. The existing metrics do not take
into account punning words which can reward translations with completely lost sense. The strong
Pearson correlation (0.84) between manual scores and the location-based metric indicates that the latter
reliably reflects expert judgments. This suggests it can serve as a scalable, less resource-intensive proxy
for evaluating wordplay translation quality. Manual scores and location-based metrics correlate closely
but difer substantially from BLEU and BERTScore rankings, highlighting the limitations of the latter
for evaluating wordplay translation. However, further analyses are needed. In future work, we will
explore new perspectives on evaluating wordplay in machine translation based on the data constructed
within the JOKER track.</p>
      <p>Additional information on the track is available on the JOKER website: https://www.joker-project.
com/</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work has received a government grant managed by the National Research Agency under the
program Investissements d’avenir integrated into France 2030, with the Reference ANR-19-GURE-0001.
It was also financed by National Funds through the Portuguese funding agency FCT through the project
LA/P/0063/2020 (DOI 10.54499/LA/P/0063/2020). Ricardo Campos would also like to acknowledge
project StorySense, with reference 2022.09312.PTDC (DOI 10.54499/2022.09312.PTDC). We thank all
other colleagues and students who participated in data construction, the translation contests, and the
CLEF JOKER track.</p>
    </sec>
    <sec id="sec-8">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used ChatGPT and Grammarly in order to: Grammar
and spelling check and Paraphrase and reword. Further, the authors used Gemini in order to: Generate
images. After using these tools/services, the authors reviewed and edited the content as needed and
take full responsibility for the publication’s content.
(Eds.), Experimental IR Meets Multilinguality, Multimodality, and Interaction, volume 14163,
Springer Nature Switzerland, Cham, 2023, pp. 397–415. doi:10.1007/978-3-031-42448-9_26.
[5] L. Ermakova, T. Miller, F. Regattin, A.-G. Bosser, E. Mathurin, G. L. Corre, S. Araújo, J. Boccou,
A. Digue, A. Damoy, B. Jeanjean, Overview of JOKER@CLEF 2022: Automatic wordplay and
humour translation workshop, in: A. Barrón-Cedeño, G. Da San Martino, M. Degli Esposti, F.
Sebastiani, C. Macdonald, G. Pasi, A. Hanbury, M. Potthast, G. Faggioli, N. Ferro (Eds.), Experimental
IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Thirteenth
International Conference of the CLEF Association (CLEF 2022), volume 13390 of Lecture Notes in Computer
Science, 2022, pp. 447–469.
[6] L. Ermakova, R. Campos, A.-G. Bosser, T. Miller, Overview of the CLEF 2025 JOKER Task 1:</p>
      <p>Humour-aware Information Retrieval, in: [23], 2025.
[7] L. Ermakova, T. Miller, Y. Naud, A.-G. Bosser, R. Campos, Overview of the CLEF 2025 JOKER Task
3: Onomastic Wordplay Translation, in: [23], 2025.
[8] A. Kreeft-Libiu, F. Helms, C. Selçuk, J. Bakker, J. Kamps, University of Amsterdam at the CLEF
2025 JOKER Track, in: [23], 2025.
[9] G. Arampatzis, A. Arampatzis, DUTH at CLEF JOKER 2025 Tasks 2 and 3: Translating Puns and</p>
      <p>Proper Names with Neural Approaches, in: [23], 2025.
[10] S. K. P, B. A, S. M, T. S, REC_Cryptix at JOKER CLEF 2025: Teaching Machines to Laugh:</p>
      <p>Multilingual Humor Detection and Translation, in: [23], 2025.
[11] I. Kuzmin, CLEF 2025 JOKER track: No pun left behind, in: [23], 2025.
[12] P. Vachharajani, pjmathematician at the CLEF 2025 JOKER Lab Tasks 1, 2 &amp; 3: A Unified Approach
to Humour Retrieval and Translation using the Qwen LLM Family, in: [23], 2025.
[13] R. Taylor, B. Herbert, M. Sana, Pun Intended: Multi-Agent Translation of Wordplay with
Contrastive Learning and Phonetic-Semantic Embeddings for CLEF JOKER 2025 Task 2, in: [23],
2025.
[14] D. A. M. Tobon, J. D. Jimenez, J. Serrano, J. C. M. Santos, E. Puertas, UTBNLP at CLEF JOKER 2025
Task 2: mBART-50 Fine-Tuning with Dictionary-Guided Forced Decoding and Phoneme-Based
Techniques for English-French Pun Translation, in: [23], 2025.
[15] Z. Xu, S. Escalera, A. Pavão, M. Richard, W.-W. Tu, Q. Yao, H. Zhao, I. Guyon, Codabench:
Flexible, easy-to-use, and reproducible meta-benchmark platform, Patterns 3 (2022). doi:10.1016/
j.patter.2022.100543.
[16] K. Papineni, S. Roukos, T. Ward, W.-J. Zhu, BLEU: A method for automatic evaluation of machine
translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational
Linguistics, 2002, pp. 311–318. URL: https://www.aclweb.org/anthology/P02-1040. doi:10.3115/
1073083.1073135.
[17] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, Y. Artzi, BERTScore: Evaluating text generation with
BERT, in: International Conference on Learning Representations, 2020. URL: https://openreview.
net/forum?id=SkeHuCVFDr.
[18] L. Ermakova, T. Miller, A.-G. Bosser, V. M. Palma Preciado, G. Sidorov, A. Jatowt, Overview
of JOKER 2023 Automatic Wordplay Analysis Task 2 – pun location and interpretation, in:
M. Aliannejadi, G. Faggioli, N. Ferro, M. Vlachos (Eds.), Working Notes of CLEF 2023 – Conference
and Labs of the Evaluation Forum, volume 3497 of CEUR Workshop Proceedings, 2023, pp. 1804–1817.
[19] L. Ermakova, A.-G. Bosser, A. Jatowt, T. Miller, The JOKER Corpus: English–French parallel
data for multilingual wordplay recognition, in: SIGIR ’23: Proceedings of the 46th International
ACM SIGIR Conference on Research and Development in Information Retrieval, Association for
Computing Machinery, New York, NY, 2023, pp. 2796–2806. doi:10.1145/3539618.3591885.
[20] L. Ermakova, A.-G. Bosser, T. Miller, A. Jatowt, Overview of the CLEF 2024 JOKER Task 3: Translate
puns from English to French, in: G. Faggioli, N. Ferro, P. Galuscakova, A. G. Seco de Herrera (Eds.),
Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2024), volume 3740 of
CEUR Workshop Proceedings, CEUR-WS.org, 2024, pp. 1800–1810.
[21] M. Post, A call for clarity in reporting BLEU scores, in: O. Bojar, R. Chatterjee, C. Federmann,
M. Fishel, Y. Graham, B. Haddow, M. Huck, A. J. Yepes, P. Koehn, C. Monz, M. Negri, A. Névéol,
M. Neves, M. Post, L. Specia, M. Turchi, K. Verspoor (Eds.), Proceedings of the Third Conference
on Machine Translation: Research Papers, 2018, pp. 186–191. doi:10.18653/v1/W18-6319.
[22] H. Xu, K. Murray, P. Koehn, H. Hoang, A. Eriguchi, H. Khayrallah, X-alma: Plug &amp; play modules
and adaptive rejection for quality translation at scale, 2025. URL: https://arxiv.org/abs/2410.03115.
arXiv:2410.03115.
[23] G. Faggioli, N. Ferro, P. Rosso, D. Spina (Eds.), Working Notes of CLEF 2025: Conference and Labs
of the Evaluation Forum, CEUR Workshop Proceedings, CEUR-WS.org, 2025.
[24] L. Ermakova, T. Miller, A.-G. Bosser, V. M. Palma Preciado, G. Sidorov, A. Jatowt, Overview of
JOKER 2023 Automatic Wordplay Analysis Task 3 – pun translation, in: M. Aliannejadi, G. Faggioli,
N. Ferro, M. Vlachos (Eds.), Working Notes of CLEF 2023 – Conference and Labs of the Evaluation
Forum, volume 3497 of CEUR Workshop Proceedings, 2023, pp. 1818–1827.
score
%
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405</p>
      <p>1405
1405</p>
      <p>1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405
1405</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          , A.-G. Bosser,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Campos</surname>
          </string-name>
          ,
          <article-title>Overview of JOKER: Humour in the machine</article-title>
          , in: J.
          <string-name>
            <surname>C. de Albornoz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mothe</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Piroi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Spina</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Sixteenth International Conference of the CLEF Association (CLEF</source>
          <year>2025</year>
          ), Lecture Notes in Computer Science, Springer,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          , A.-G. Bosser,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Campos</surname>
          </string-name>
          ,
          <article-title>CLEF 2025 JOKER lab: Humour in the machine</article-title>
          ,
          <source>in: Advances in Information Retrieval: 47th European Conference on Information Retrieval</source>
          ,
          <string-name>
            <surname>ECIR</surname>
          </string-name>
          <year>2025</year>
          , Lucca, Italy, April 6-
          <issue>10</issue>
          ,
          <year>2025</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>V</given-names>
          </string-name>
          , Springer-Verlag, Berlin, Heidelberg,
          <year>2025</year>
          , p.
          <fpage>389</fpage>
          -
          <lpage>397</lpage>
          . URL: https://doi.org/10.1007/978-3-
          <fpage>031</fpage>
          -88720-8_
          <fpage>59</fpage>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>031</fpage>
          -88720-8_
          <fpage>59</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          , A.-G. Bosser,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M.</given-names>
            <surname>Palma Preciado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          ,
          <article-title>Overview of the CLEF 2024 JOKER track: Automatic humour analysis</article-title>
          , in: L.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Quénot</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Schwab</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Soulier</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. M. D. Nunzio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Galuščáková</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction: Proceedings of the Fifteenth International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ), volume
          <volume>14959</volume>
          of Lecture Notes in Computer Science, Springer, Cham,
          <year>2024</year>
          , pp.
          <fpage>165</fpage>
          -
          <lpage>182</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -71908-
          <issue>0</issue>
          _
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-G.</given-names>
            <surname>Bosser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M.</given-names>
            <surname>Palma Preciado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          , Overview of JOKER - CLEF
          <article-title>-2023 track on automatic wordplay analysis</article-title>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Giachanou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Aliannejadi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Vlachos</surname>
          </string-name>
          , G. Faggioli, N. Ferro
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>