<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Pisa, Italy
* Corresponding author.
$ g.sarti@rug.nl (G. Sarti); t.caselli@rug.nl (T. Caselli);
a.bisazza@rug.nl (A. Bisazza); m.nissim@rug.nl (M. Nissim)
 https://gsarti.com (G. Sarti); https://cs.rug.nl/~bisazza
(A. Bisazza); https://malvinanissim.github.io (M. Nissim)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>EurekaRebus - Verbalized Rebus Solving with LLMs: A CALAMITA Challenge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gabriele Sarti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tommaso Caselli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arianna Bisazza</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Malvina Nissim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Language and Cognition (CLCG), University of Groningen, Oude Kijk in 't Jatstraat 26 Groningen</institution>
          ,
          <addr-line>9712EK</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Language games can be valuable resources for testing the ability of large language models (LLMs) to conduct challenging multi-step, knowledge-intensive inferences while respecting predefined constraints. Our proposed challenge prompts LLMs to reason step-by-step to solve verbalized variants of rebus games recently introduced with the EurekaRebus dataset [1]. Verbalized rebuses replace visual cues with crossword definitions to create an encrypted first pass, making the problem entirely text-based. We introduce a simplified task variant with word length hints and adopt a comprehensive set of metrics to obtain a granular overview of models' performance in knowledge recall, constraints adherence, and re-segmentation abilities across reasoning steps.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Large language models</kwd>
        <kwd>Sequential reasoning</kwd>
        <kwd>Puzzle</kwd>
        <kwd>Rebus</kwd>
        <kwd>Crosswords</kwd>
        <kwd>Enigmistica Italiana</kwd>
        <kwd>CALAMITA</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Challenge: Introduction and</title>
    </sec>
    <sec id="sec-2">
      <title>Motivation</title>
      <p>Language games were adopted as testbeds for measuring
NLP progress in recent years [2, 3, 4], with a particular
focus on (cryptic) crossword solving English [5, 6, 7, 8, 9].
For the Italian language, initial eforts focused on
crossword solving and generation [10, 11] and clue-based word
guessing [12, 13, 9]. Recently, Sarti et al. [1] introduced an
extensive collection of text-adapted Italian rebus puzzles
to evaluate large language models’ (LLMs) knowledge
and sequential reasoning abilities. Rebuses are complex
puzzles combining visual elements and graphic signs to
encode a hidden phrase. Italian can boast a rich and
long-standing rebus tradition dating back to the 19th
century [14], popularized by high-difusion magazines
such as La Settimana Enigmistica1. The structure of
Italian rebuses has, with time, been formalized into beauty
canons [15], and their peculiarities and design principles
were analyzed by several authors [16, 17, 18].</p>
      <p>In Italian rebuses, rebus solving begins by combining
derived by combining graphemes with their underlying
visual elements in a left-to-right fashion, composing a
ifrst pass (prima lettura) representing an intermediate
solution of the puzzle. Then, first pass elements are
re</p>
      <sec id="sec-2-1">
        <title>Verbalized Rebus:</title>
        <p>TES [Dirige la rotta] (Directs the course)
[Le difendono i portieri] (Protected by goalkeepers) CE
N [Calda bevanda rilassante] (Warm relaxing drink)</p>
      </sec>
      <sec id="sec-2-2">
        <title>Solution key (# of chars/word):</title>
        <p>9
9</p>
        <p>Solution: Testimone reticente (reticent witness)
segmented (cesura) according to a solution key
(diagramma), which specifies the length of each word in the
solution (frase risolutiva). The verbalized rebuses
introduced by Sarti et al. [1] are text-only version of real
rebuses published in popular outlets derived by replacing
words corresponding to visual elements with
externallysourced crossword definitions in the transcribed first
passes, using a standardize format. Figure 1 provides a</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Data description</title>
      <p>simple example.</p>
      <p>This work proposes to adopt the EurekaRebus
introduced by Sarti et al. [1] to extend their evaluation of 3.1. Origin of data
LLMs’ multi-step reasoning and linguistic/cultural
awareness to the systems evaluated as part of the CALAMITA The dataset used for this challenge is an extended version
evaluation campaign [19]. We believe the task is par- of EurekaRebus [1], a collection of 222,089 unique
Italticularly relevant since the crossword definitions that ian rebuses extracted from Eureka5 platform3, an open
compose verbalized rebuses rely heavily on idiomatic database of rebuses and other linguistic puzzles
mainexpressions, wordplay, and cultural references specific tained by the Associazione Culturale “Biblioteca
Enigto Italian. Hence, the results of this task could provide mistica Italiana - G. Panini”4. Among these, 83,157 were
valuable insights into the linguistic and cultural compe- converted by the original authors in verbalized form by
tence of LLMs trained on the Italian language. Moreover, leveraging the crossword definitions from the ItaCW
colthe task is especially appealing since it is framed in a lection [10], including 125,202 definition-solution pairs.
templated reasoning format, enabling us to disentangle While Sarti et al. [1] evaluated the performances of
the various components required to successfully solve a prompted and tuned LLMs on rebuses up to June 17th,
verbalized rebus step-by-step. More specifically, several 2024, the current test set include 168 new unseen
exammetrics will be employed to assess LLMs’ factual recall, ples released on Eureka5 after that date.
textual concatenation and re-segmentation capabilities
and, finally, constraint satisfaction given the provided 3.2. Annotation details
cues. We employ the same procedure of Sarti et al. [1] for
ver</p>
      <p>In light of the results reported by [1] for state-of-the- balizing available rebuses. More specifically, only
reart proprietary LLMs, we expect all tested open-source buses having all lowercased or camel-cased words among
systems to perform very poorly, with final solution ac- ItaCW solutions are selected, and every word is replaced
curacies well below 30%. We also note that the high- by sampling one of the available crossword definitions for
est reported overall performance in previous work2 was it at random.5 Moreover, only regular rebuses containing
found by the original authors to be primarily the prod- at least two hidden words are selected, avoiding examples
uct of memorization. We anticipate that this challenge requiring a single definition-solving step and those with
will highlight significant limitations in LLMs’ current more complex templates (e.g., anarebuses using anagrams
factual recall and multi-step reasoning ability and act as of hidden words for the solution).
a catalyst for future improvements in these areas.</p>
      <sec id="sec-3-1">
        <title>3http://www.eureka5.it</title>
        <p>2Namely 58% Solution Exact Match for a LLaMA-3.1 8B model LoRA- 4http://www.enignet.it/home
tuned on 80k EurekaRebus examples [20, 21] 5Words in ItaCW can be associated to multiple definitions.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>2. Challenge: Description</title>
      <p>The proposed challenge aims to evaluate the capabilities
of existing LLMs in solving verbalized Italian rebuses via
prompting at various granularity levels. More specifically,
LLMs will be evaluated in a few-shot prompting setting
with two fixed in-context learning examples pre-selected
at random from the available pool of verbalized rebuses
in EurekaRebus, in two settings:
• Regular, matching the example in table 1 and
the original input format used by Sarti et al. [1].
• Hints, in which the number of characters for
every hidden word is provided alongside definitions
in the verbalized rebus to help the model in
identifying the correct choice. This variant was not
tested by Sarti et al. [1].</p>
      <p>Refer to section 3.3 for the respective example formats.
Models will be evaluated on their performance at each
step required to successfully solve the verbalized rebus
and their overall ability to produce correct final solutions.
3.3. Data format
Each example in the dataset consists of:
• The verbalized rebus (verbalized_rebus)
containing letters from the original rebus and
crossword-style definitions enclosed in square
brackets.
• A variant of the verbalized rebus
containing length hints for definitions
(verbalized_rebus_with_length_hints).
• The solution key, composed by
whitespaceseparated numbers representing the word lengths
in the final solution ( solution_key).
• The first pass words matching definitions in
the verbalized rebus, provided in a
semicolonseparated string in order of occurrence
(word_guesses).
• The first pass obtained by infilling words in
place of their definitions in the verbalized rebus
(first_pass).
3.5. Detailed data statistics</p>
      <sec id="sec-4-1">
        <title>An example is provided in Listing 1. Table 2 from Sarti et al. [1] reports statistics for the full and verbalized subsets of the EurekaRebus dataset.</title>
        <p>3.4. Prompting
Train set contents The training set contains 80,158
Table 1 shows the 2-shot prompting template adopted for examples, which are ignored for the purpose of the
generating a templated solution with the tested LLMs. CALAMITA campaign provided that no adaptation
methThe second in-context example used in the template, ods are evaluated.
omitted for brevity, corresponds to the one shown in
Listing 1. Test set contents The test set contains 3,167 examples</p>
        <p>The task description provided to the model was de- divided as follows, in order of appearance:
rived from a trial-and-error process starting from the
original prompt by Sarti et al. [1]. Notably, compared to
the original authors the task description provides more
detailed descriptions of individual components of the
rebus to provide a clearer overview of the task to the LLM.</p>
        <p>We opted for a 2-shot setting as opposed to the 5-shot
prompting employed by Sarti et al. [1] to accommodate
the limited context length of some of the tested LLMs,
thus ensuring that the total length after model generation
does not exceed 1024 tokens6. The two examples
provided remain the same shown here to simplify evaluation
and ensure consistent results.
• 2000 examples matching the in-domain setting
for models trained by [1], i.e. containing only first
pass words seen by all available trained models.
• 999 examples matching the out-of-distribution
setting for models trained by [1], i.e. containing
at least one first pass word unseen during training
by available trained models.
• 168 new verbalized rebuses added in
EurekaRebus v1.1, added to the Eureka5 platform after
June 17th, 2024. These can be either in-domain
or out-of-distribution for models trained on the
EurekaRebus’s training set.</p>
        <p>Verbalized rebus solving steps Table 1 provide
labels for the steps necessary to solve the verbalized rebus
that are considered in this challenge task. The model
receives a problem input including a verbalized rebus
(possibly with length hints) and a solution key (chiave
di lettura). The first step involves resolving crossword
definitions in order ( Definition resolution ), exploiting
only the model’s parametric knowledge to accomplish
6The LLaMA 3 tokenizer was used to perform this estimate</p>
        <p>While prompted models should obtain similar
performances across all test subsets, the aformentioned division
will enable further comparisons with previously trained
systems.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Metrics</title>
      <p>The challenge employs a comprehensive set of metrics
adapted from the original evaluation of [1]:
Prompt template
Sei un’esperto risolutore di giochi enigmistici. Il seguente
gioco contiene una frase (Rebus) nella quale alcune parole
sono state sostituite da indizi tra parentesi quadre. I numeri in
ogni indizio rappresentano la lunghezza della parola nascosta.</p>
      <p>Il tuo compito è quello di identificare le parole nascoste e
sostituirle agli indizi nel Rebus, producendo una prima lettura
dalla quale poi si deriverà una frase risolutiva. La chiave
di lettura è una sequenza di numeri che rappresentano la
rispettive lunghezze delle parole che compongono la frase
risolutiva. La tua risposta deve essere una frase risolutiva
sensata e che rispetti le lunghezze definite nella chiave di
lettura.</p>
      <p>First
example
Problem</p>
      <p>input
Definition
resolution
First pass
Solution
segmentation
Solution
Second
example
Answer
prefix</p>
      <p># Esempio 1:
⎧Rebus: AC [Un mollusco nell’insalata di
⎪
⎪⎪⎪mare (5)] GLI [Lo è l’operaio che lavora in
⎨cantiere (5)] S TO [Soldati da trincea (5)]
⎪
⎪⎪⎪Chiave di lettura: 11 2 10
⎩</p>
      <p>Procediamo alla risoluzione del rebus passo
per passo:
⎧ - A C = A C
⎪⎪- [Un mollusco nell’insalata di mare] =
⎪
⎪
⎪⎪cozza
⎪
⎪
⎪⎨- G L I = G L I</p>
      <p>- [Lo è l’operaio che lavora in cantiere] =
⎪
⎪⎪⎪edile
⎪
⎪⎪- S T O = S T O
⎪
⎪
⎩- [Soldati da trincea] = fanti
︂{ Prima lettura: AC cozza GLI edile S TO
fanti
Ora componiamo la soluzione seguendo la
chiave risolutiva:
⎧ 11 = Accozzaglie
⎨2 = di
⎩12 = lestofanti
︀{ Soluzione: Accozzaglie di lestofanti
# Esempio 2:
... (same format as the first example)
# Ora tocca a te!
Completa il rebus seguendo il procedimento
descritto, rispondendo esattamente nello
stesso formato utilizzato dagli esempi
precedenti.</p>
      <p>Rebus: {{verbalized_rebus}} or
{{verbalized_rebus_with_length_hints}}
Chiave di lettura: {{solution_key}}</p>
      <p>The Solution Match metric will be used as a primary
metric of correctness, since it captures the model
ability to fully solve the verbalized rebus. While no
baseline evaluation was conducted for the new test set used
in this challenge, we expect the performances of most
capable open-source systems to align with those of
5Table 1 shot prompted LLaMA-3 70B and Qwen-2 72B models
2-shot prompt used for the CALAMITA evaluation. Blue text reported by Sarti et al. [1], which we summarize in
Secrepresent additions for the evaluation in the Hints setting. tion 4. The results show that current models struggle
Template elements are highlighted next to the first in-context
example. Example rebus by Parodi E., Domenica Quiz n. 7
Statistic
# examples
# authors
Year range
# unique words
Avg./SD words/ex.</p>
      <p>Avg./SD word len.</p>
      <p>Avg./SD FP len.
# unique words
Avg./SD words/ex.</p>
      <p>Avg./SD word len.</p>
      <p>Avg./SD Sol. len.</p>
      <p>EurekaRebus</p>
      <p>ItaCW-filtered</p>
      <p>• Word Guess Accuracy: Proportion of correctly
guessed words during definition resolution
(corresponding to the Definition metric in the original
evaluation).
• Word Guess Length Accuracy: Proportion of
word guesses in definition resolution matching
the correct length. This is evaluated only for the
Hints setting, where the length is explicitly
provided (not evaluated in previous works).
• First Pass Accuracy: Proportion of generated
ifrst passes matching the gold reference
(corresponding to the First Pass Exact Match metric in
the original evaluation).
• Solution Word Accuracy: Proportion of correct</p>
      <p>words in the generated solutions.
• Solution Words Lengths Accuracy: Proportion
of generated solution words matching the lengths
specified by the solution key. Lower scores may
indicate dificulty in respecting the given length
constraints (corresponding to the Solution Key</p>
      <p>Match metric in the original evaluation).
• Solution Match: Proportion of generated
solutions matching the gold reference (corresponding
to the Solution Exact Match metric in the original
evaluation).</p>
      <p>Verbalization Simplification The use of verbalized
rebuses, while necessary for text-based LLMs, simplifies
the original visual puzzle. This does not fully capture the
complexity of solving traditional rebuses, which rely on
visual cues and cultural knowledge, making verbalized
rebus solving a much simpler proxy to the multi-step
reasoning required for regular rebuses.</p>
      <p>As reported by the original EurekaRebus dataset license,
the data is redistributed for research purposes only with
the explicit approval of the Associazione Culturale
“Biblioteca Enigmistica Italiana - G. Panini” (here onwards
referred to as the Association), and the rights to each entry
in the EurekaRebus collection are the property of the
reCultural Specificity The selected rebuses and cross- spective copyright holders. The usage and redistribution
word definitions rely heavily on Italian-specific linguistic of these data is allowed only for users providing
approand cultural background. Performance on this task may priate attribution to the original copyright holders and
not generalize to other languages or puzzle types, and the Association, and the creation of derivative works is
it might be unrealistic to expect general-purpose LLMs permitted only for research purposes, using terms no less
to possess the specific lexicon and knowledge used for restrictive than the EurekaRebus license. Researchers are
rebus solving. encouraged to contact the challenge organizers with any
questions or concerns about data usage and licensing.</p>
    </sec>
    <sec id="sec-6">
      <title>7. Data license and copyright issues</title>
      <p>Prompt Sensitivity While the selected prompt
template was observed to perform well for capable
proprietary LLMs in preliminary tests, there are no guarantees
that the instructions provided in the prompt are suficient
for smaller open-source models to perform verbalized
rebus solving proficiently. Moreover, alternative prompt
formulations could lead to potentially better results.
Lack of Human Baseline The challenge currently
lacks a clear human performance baseline, which would
be valuable for contextualizing model performance on
verbalized rebus solving.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Ethical issues</title>
      <p>While this challenge focuses on a relatively benign task
of puzzle-solving, there are some ethical considerations
to keep in mind. First, the dataset captures a very narrow
subset of Italian language and culture. Hence, evaluation</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>We would like to express our gratitude to the following
individuals and organizations:
• The Associazione Culturale "Biblioteca
Enigmistica Italiana - G. Panini" for making their rebus
collection freely accessible on Eureka5.
• The creators of the ItaCW dataset for enabling
the creation of verbalized rebuses.
• The puzzle creators whose work is represented
in this dataset.</p>
      <p>Gabriele Sarti and Arianna Bisazza acknowledge the
support of the Dutch Research Council (NWO) for the
project InDeep (NWA.1292.19.399). Arianna Bisazza
is further supported by the NWO Talent Programme
(VI.Vidi.221C.009). We hope this challenge will contribute
to the difusion of the art of Italian enigmistica among
computational linguistics and artificial intelligence
researchers.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>