<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>MITA: Challenge the Abilities of LAnguage M odels in ITAlian</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giuseppe Attanasio</string-name>
          <email>giuseppe.attanasio@lx.it.pt</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierpaolo Basile</string-name>
          <email>pierpaolo.basile@uniba.it</email>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Federico Borazio</string-name>
          <email>borazio@ing.uniroma2.it</email>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Danilo Croce</string-name>
          <email>croce@info.uniroma2.it</email>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Francis</string-name>
          <email>maria.francis287@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff7">7</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jacopo Gili</string-name>
          <email>jacopo.gili584@edu.unito.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elio Musacchio</string-name>
          <email>elio.musacchio@phd.unipi.it</email>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Malvina Nissim</string-name>
          <email>m.nissim@rug.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Viviana Patti</string-name>
          <email>viviana.patti@unito.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Scalena</string-name>
          <email>d.scalena@campus.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Italian Benchmark, Shared Task, Language Models</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CLCG, University of Groningen</institution>
          ,
          <addr-line>Groningen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Computer Science Department, University of Turin</institution>
          ,
          <addr-line>Turin</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Instituto de Telecomunicações</institution>
          ,
          <addr-line>Lisbon</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Matteo Rinaldi</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Bari “Aldo Moro”</institution>
          ,
          <addr-line>Bari</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Milan Bicocca</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>University of Rome “Tor Vergata”</institution>
          ,
          <addr-line>Rome</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff7">
          <label>7</label>
          <institution>University of Trento</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>The rapid development of Large Language Models (LLMs) has called for robust benchmarks to assess their abilities, track progress, and compare iterations. While existing benchmarks provide extensive evaluations across diverse tasks, they predominantly focus on English, leaving other languages underserved. For Italian, the EVALITA campaigns have provided a long-standing tradition of classification-focused shared tasks. However, their scope does not fully align with the nuanced evaluation required for modern LLMs. To address this gap, we introduce “Challenge the Abilities of LAnguage Models in ITAlian” (CALAMITA), a collaborative efort to create a dynamic and growing benchmark tailored to Italian. CALAMITA emphasizes diversity in task design to test a wide range of LLM capabilities through resources natively developed in Italian by the community. This initiative includes a shared platform, live leaderboard, and centralized evaluation framework. This paper outlines the collaborative process, initial challenges, and evaluation framework of CALAMITA.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR</p>
      <p>ceur-ws.org
https://github.com/crux82 (D. Croce); https://github.com/rosakun
(M. Francis); https://github.com/Jj-source (J. Gili);
https://github.com/m-elio (E. Musacchio);
https://malvinanissim.github.io (M. Nissim);
https://github.com/vivpatti (V. Patti); https://github.com/mrinaldi97
(M. Rinaldi); https://github.com/DanielSc4 (D. Scalena)
0000-0001-6945-3698 (G. Attanasio); 0000-0002-0545-1105
(P. Basile); 0009-0000-0193-2131 (F. Borazio); 0000-0001-9111-1950
(D. Croce); 0009-0007-7638-9963 (M. Francis); 0009-0007-1343-3760
(J. Gili); 0009-0006-9670-9998 (E. Musacchio); 0000-0001-5289-0971
(M. Nissim); 0000-0001-5991-370X (V. Patti); 0009-0004-7488-8855
(M. Rinaldi); 0009-0006-0518-6504 (D. Scalena)
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License
Attribution 4.0 International (CC BY 4.0).</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>In parallel with the ongoing and constant development
of new Large Language Models (LLMs), it has increased
the need for understanding their abilities, how they
differ from one another, and how they improve compared
to previous iterations. To meet this need, the last
couple of years have witnessed multiple eforts to put
together new—or revisiting existing—benchmarks against
which the performance and progress of LLMs can be
monitored. These benchmarks include diferent tasks
to test a variety of characteristics and abilities that are
assumed to be associated with LLMs at diferent degrees.</p>
      <sec id="sec-2-1">
        <title>To mention a few, these span from multiple-choice ques</title>
        <p>
          tions of various sorts, commonsense and mathematical
reasoning, and a variety of linguistic phenomena.
BIGbench [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] is currently the largest and most
comprehensive benchmark, including over 200 tasks, almost all in
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>English, which have been collaboratively contributed by researchers across the globe.</title>
      </sec>
      <sec id="sec-2-3">
        <title>However, benchmarking progress for languages other</title>
        <p>than English has not improved with comparable quality.</p>
      </sec>
      <sec id="sec-2-4">
        <title>In many cases, evaluation datasets are automatic transla</title>
        <p>tions of their English counterparts, yielding not only a
less native and possibly ungrammatical language but also
a cultural picture that is distant from the target language. a strong collaborative nature. The Italian Association for</p>
        <p>
          In the Italian NLP landscape, there is a long tradition Computational Linguistics (AILC, https://www.ai-lc.it)
of evaluation through the contribution of shared tasks. launched a public call, mainly aimed at the Italian NLP
These benchmarks have been collected and run for al- community but spread across the standard international
most 20 years in the context of the EVALITA campaigns communication channels, asking for challenges and
cor(https://www.evalita.it/). The campaigns have fostered responding datasets, that LLMs could be tested on.
the creation of training and evaluation resources and Participants contributing to a challenge were expected
models natively developed for Italian. Based on such to provide an explanation and motivation for a given
chalresources, UINAUIL (Unified Interactive Natural Under- lenge, as well as a dataset that reflects that challenge. It
standing of the Italian Language)[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], an integrated bench- was also asked to provide any information relevant to the
mark for Italian NLU including six tasks has been recently dataset (provenance, annotation, distribution of labels or
proposed, and tested with available Italian and multilin- phenomena, etc.) Evaluation metrics and examples were
gual language models. also expected, along with the task and dataset
submisExcept for CHANGE-IT [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], a generation task focused sion. Existing relevant datasets could also be submitted
on headline transformation and organized within the as long as they made an interesting contribution to the
EVALITA 2020 edition, all EVALITA tasks have focused benchmark and were natively created in Italian. To
stanon classification problems (some have been recast as gen- dardize the contribution to the CALAMITA benchmark,
eration problems as part of a resource release within all proposed tasks with existing or new datasets had to
the “Risorse per la Lingua Italiana” (RiTA) community follow a predefined template created and distributed by
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]). However, to improve upon existing benchmarks, the CALAMITA organizers.
we wanted the core of a dynamic reference benchmark Creating the CALAMITA benchmark and the first
for Italian to include new tasks specifically focused on round of LLM evaluation required several steps. In the
testing LLMs’ abilities. ifrst phase, all prospective participants submitted a
pre
        </p>
        <p>Therefore, in the steps of this solid Italian benchmark- proposal. In case of a positive evaluation, based on
ing tradition, and in line with the most recent devel- compliance with the requirements and balance across
opments regarding the evaluation of LLMs, AILC—the submissions – participants were then asked to submit
Italian Association for Computational Linguistics—has the final and complete challenge, following the provided
launched “Challenge the Abilities of LAnguage Models in CALAMITA template, in phase two. A final report was
ITAlian” (CALAMITA), a large-scale collaborative initia- also requested for each accepted task, providing
informative across the whole Italian NLP community to develop tion on implementing the code for the evaluation.
a dynamic and growing benchmark for evaluating LLMs’ The data and evaluation team set up the final
capabilities in Italian. This strategy would ensure a high CALAMITA benchmark by compiling the data and code
diversity of tasks and, thus, of tested capabilities. It would of all the proposed tasks. We forked the Language Model
distribute the efort of creative resources natively in Ital- Evaluation Harness tool2 to create a custom CALAMITA
ian across many researchers and practitioners. version by including all the accepted tasks. Once the</p>
        <p>In the long term, we aim to establish a continuously benchmark was assembled, the CALAMITA
organizgrowing suite of tasks that can be accessed through a ers ran zero- or few-shot experiments with a selection
shared platform and a live leaderboard so that any newly of LLMs. No tuning materials or experiments are
exdeveloped LLM, either multilingual or Italian monolin- pected at this project stage. Also, while we expect that
gual, can be readily assessed. In the short term, we have CALAMITA, in the longer run, will be further populated
started to build the CALAMITA benchmark through a by additional tasks and will have its own publicly
acseries of challenges collaboratively contributed by the re- cessible leaderboard, allowing for model testing, in this
search community (Section 2). Also, we have established ifrst stage, the choice of LLMs to be evaluated and the
an evaluation framework that enables running the cur- evaluation procedure is centralized.
rent and possibly future challenges in a centralized and
coherent manner. This short paper summarises the
collaborative procedure, the challenges currently included 3. Challenges
in CALAMITA1, and the evaluation procedure.</p>
        <p>The preliminary call for tasks yielded the submission of
over 20 proposals. Almost all of them were retained and
are part of the present CALAMITA challenge, apart from
the proposals that aimed at testing abilities that LLMs
should not be expected to have, such as abilities typical
of information retrieval engines and the proposals that</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Collaborative Methodology</title>
      <sec id="sec-3-1">
        <title>The CALAMITA approach is inspired by standard Natural</title>
        <p>Language Processing shared tasks, giving the benchmark</p>
      </sec>
      <sec id="sec-3-2">
        <title>1The CALAMITA website: https://clic2024.ilc.cnr.it/calamita/.</title>
      </sec>
      <sec id="sec-3-3">
        <title>2https://github.com/EleutherAI/lm-evaluation-harness</title>
        <sec id="sec-3-3-1">
          <title>Ability tested</title>
        </sec>
        <sec id="sec-3-3-2">
          <title>Description</title>
          <p>Commonsense knowledge
Factual knowledge
Linguistic knowledge
Formal reasoning
Fairness and bias
Code generation
Machine translation
Summarization</p>
          <p>General knowledge about the world that is typically taken for granted in everyday
life, e.g., everyday cause-and-efect relationships, situational judgments, physical
properties, and basic social interactions.</p>
          <p>Knowledge of concrete, verifiable facts about the world, e.g., definitions, historical
events, or scientific concepts.</p>
          <p>Linguistically motivated tasks that test specific language skills, e.g., word sense
disambiguation, coreference resolution, or acceptability judgment.
Ability to understand and use formally logical principles to solve problems, e.g.,
mathematical problems.</p>
          <p>Evaluates a model’s capacity to handle sensitive tasks, including exclusive and
stereotyped language understanding and detecting ofensive or biased language
towards social groups.</p>
          <p>Ability to generate fully functioning code for a specific programming language.
Ability to translate a sentence from a source language into another language, with
one of the two being Italian.</p>
          <p>Ability to create relevant summaries of a given excerpt, e.g., news headline
generation or news reduction.</p>
          <p>Count
19
12
22
9
6
tative component as a premise or conclusion. In contrast,
the second and third tasks aim at classifying the type of
premise: legal vs factual, and its corresponding
argumentation scheme. The classes are highly unbalanced, hence
evaluation is based on the macro F1 score.
required manual evaluation. In what follows, we briefly
describe each task included in CALAMITA and refer the
reader to each of the challenges’ reports for further
details. In Table 1, we describe the macro categories under
which the CALAMITA tasks can be grouped, where
categories are broad classes of tested abilities. Table 2 shows
which abilities apply to each challenge.</p>
        </sec>
        <sec id="sec-3-3-3">
          <title>BEEP (BEst DrivEr’s License Performer) [7] is a</title>
          <p>
            benchmark to evaluate large language models in the
conABRICOT (ABstRactness and Inclusiveness in COn- text of a simulated Italian driver’s license exam. This
chaltexT) [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ] is a task designed to evaluate Italian language lenge tests the models’ ability to understand and apply
models on their ability to understand and assess the ab- trafic laws, road safety regulations, and vehicle-related
stractness and inclusiveness of language, two nuanced knowledge through a series of true/false questions. The
features that humans naturally convey in everyday com- dataset is derived from oficial ministerial materials used
munication. Unlike binary categorizations such as ab- in the Italian licensing process, explicitly targeting
Catestract/concrete or inclusive/exclusive, these features exist gory B licenses.
on a continuous spectrum with varying degrees of
intensity. The task is based on a manual collection of sentences
that present the same noun phrase (NP) in diferent
contexts, allowing its interpretation to vary between the
extremes of abstractness and inclusiveness. This
challenge aims to verify how LLMs perceive subtle linguistic
variations and their implications in natural language.
          </p>
          <p>
            BLM-It (Blackbird Language Matrices) [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ] is a task
made of linguistic puzzles (matrices) around
languagerelated problems, focusing on formal and semantic
properties of language. A BLM matrix consists of a context
set and an answer set. The context is a sequence of
sentences that encodes implicitly an underlying generative
linguistic rule. The contrastive multiple-choice answer
AMELIA (Argument Mining Evaluation on Legal set includes negative examples following corrupted
gendocuments in ItAlian) [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] is a challenge consisting erating rules. The models are prompted in a few-shot
of three classification tasks in the context of argument setting. The datasets comprise a few prompts for a
fewmining in the legal domain. The tasks are based on a shot setting.
dataset of 225 Italian decisions on Value Added Tax,
annotated to identify and categorize argumentative text. DIMMI (Drug InforMation Mining in Italian) [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]
The objective of the first task is to classify each argumen- is a task aimed at evaluating the proficiency of Large
Language Models in extracting drug-specific information lingual scenarios. It includes three tasks: (1) the detection
from Patient Information Leaflets. The challenge eval- of gender-marked expressions in Italian sentences, (2)
uates the efectiveness of processing complex medical the rewriting of gendered expressions into gender-fair
information in Italian and is approached as an informa- alternatives, and (3) the generation of gender-fair
lantion extraction task in a zero-shot setting, based on the guage in automatic translation from English to Italian.
model’s pre-existing knowledge or through in-context The challenge relies on three diferent annotated datasets:
learning. Evaluation is performed against a manually the GFL-it corpus, which contains Italian texts extracted
created gold standard. from administrative documents provided by the
University of Brescia; GeNTE, a bilingual test set for
genderECWCA (Educational CrossWord Clues Answering) neutral rewriting and translation built upon a subset of
[
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] is designed to evaluate the knowledge and reasoning the Europarl dataset; Neo-GATE, a bilingual test set
decapabilities of LLMs through crossword clue-answering. signed to assess the use of non-binary neomorphemes in
The challenge consists of two tasks: a standard question- Italian for both fair formulation and translation tasks.
answering format where the LLM is asked to solve
crossword clues and a variation where the model is given hints
about the word lengths of the answers, which is expected
to help models with reasoning abilities.
          </p>
        </sec>
        <sec id="sec-3-3-4">
          <title>GITA (Graded Italian Annotated Dataset) [15] in</title>
          <p>
            vestigates the physical commonsense reasoning
capabilities of large language models, assessing their low-level
understanding of the physical world using a test set in
EurekaRebus [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ] is a task that tests the ability of the Italian language. Three specific tasks are evaluated:
LLMs to conduct multi-step, knowledge-intensive infer- identifying plausible and implausible stories within our
ences while respecting predefined constraints. LLMs dataset, identifying the conflict that generates an
implauare prompted to reason step-by-step to solve verbalized sible story, and identifying the physical states that make
variants of rebus games. Verbalized rebuses replace vi- a story implausible. It is written and annotated by a
sual cues with crossword definitions to create an en- professional linguist.
crypted first pass, making the problem entirely text-based.
          </p>
          <p>Multiple metrics are used to grasp the models’ perfor- INVALSI [16] is a benchmark based on the Invalsi tests
mance in knowledge recall, constraints adherence, and administered to students within the Italian school system.
re-segmentation abilities across reasoning steps. Expert pedagogists prepare these tests with the explicit
goal of testing average students’ performance over time
across Italy. There are two benchmarks: Invalsi MATE
(420 questions), which targets the models’ performance
on mathematical understanding, and Invalsi ITA (1279
questions), which evaluates language understanding in
Italian.</p>
        </sec>
        <sec id="sec-3-3-5">
          <title>GATTINA (GenerAtion of TiTles for Italian News</title>
          <p>
            Articles) [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ] is a task that aims to assess the ability
of LLMs to generate headlines for science news articles.
Aspects such as the appropriateness of the summary,
creativity, and attractiveness are evaluated through a
battery of metrics. The benchmark consists of a large
dataset of science news articles and their corresponding
published headlines from ANSA Scienza and Galileo, two
prominent Italian media outlets.
          </p>
        </sec>
        <sec id="sec-3-3-6">
          <title>ITA-SENSE (ITAlian word SENSE disambiguation)</title>
          <p>
            [17] is a task that assesses LLMs’ abilities in
understanding lexical semantics through Word Sense
Disambiguation. The classical Word Sense Disambiguation task is
GEESE (Generating and Evaluating Explanations cast as a generative problem formalized as two tasks:
for Semantic Entailment) [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ] is focused on evaluat- [T1] Given a target word and a sentence in which the
ing the impact of generated explanations on the predic- word occurs, generate the correct meaning definition;
tive performance of language models for the task of Rec- [T2] Given a target word and a sentence in which the
ognizing Textual Entailment in Italian. Using a dataset word occurs, choose the correct meaning definition from
enriched with human-written explanations, two large a predefined set. For CALAMITA, LLMs are tested in a
language models are employed to generate and utilize ex- zero-shot setting.
planations for semantic relationships between sentence
pairs. GEESE assesses the quality of generated
explanations by measuring changes in prediction accuracy when
explanations are provided.
          </p>
        </sec>
        <sec id="sec-3-3-7">
          <title>MACID (Multimodal ACtion IDentification) [18]</title>
          <p>
            is a task aimed at evaluating LLMs to diferentiate
between closely related action concepts based on textual
descriptions alone. The challenge is inspired by the ”find
GFG (Gender-Fair Generation) [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ] is a task de- the intruder” task, where models must identify an
outsigned to assess and monitor the recognition and gener- lier among a set of 4 sentences that describe similar yet
ation of gender-fair language in both mono- and cross- distinct actions. The dataset highlights action-predicate
mismatches, where the same verb may describe diferent designed to evaluate the ability of LLMs to comprehend a
actions, or diferent verbs may refer to the same action. specific type of complex syntactic construction in Italian:
Although mono-modal (text-only), the task is designed object relative clauses. The challenge is framed as a
for future multimodal integration, linking visual and tex- binary entailment task where, given a complex sentence,
tual representations to enhance action recognition. the model is tasked with determining whether it logically
entails a simpler yes/no implication.
          </p>
          <p>MT (Machine Translation) [19] is a task that aims
at testing the ability of LLMs in automatic translation, Termite [24] focuses on the Text-to-SQL task in
Italfocusing on Italian and English (in both directions). The ian. Natural language queries are written natively in
Italtask proposes a benchmark composed of two datasets ian, and the models are expected to turn them into SQL
covering diferent domains and with varying distribution queries. The dataset is built to be invisible to search
enpolicies. Performances are reported in terms of four eval- gines since it is locked under an encryption key delivered
uation metrics, whose scores allow an overall evaluation along the resource to reduce accidental inclusion in
upof the quality of the automatically generated translations. coming training sets. It contains hand-crafted databases
in diferent domains, each with a balanced set of NL-SQL
Mult-IT [20] is a large-scale Multi-Choice Question An- query pairs. The NL questions are built in such a way
swering (MCQA) dataset for evaluating the factual knowl- that they can be solved by a model relying only on its
edge and reasoning abilities of LLMs in Italian. This linguistic proficiency and an analysis of the schema, with
contribution aims to counteract the disadvantages of us- no external knowledge needed.
ing MCQA benchmarks that are automatically translated
from English and may sound unnatural, contain errors,
or use linguistics constructions that do not align with the
target language. In addition, they may introduce topical
and ideological biases reflecting Anglo-centric
perspectives. Mult-IT comprises over 110,000 manually written
questions sourced directly from preparation quizzes for
Italian university entrance exams or for exams for public
sector employment in Italy.</p>
          <p>VeryfIT [25] is designed to evaluate the in-memory
factual knowledge of language models on data written
by professional fact-checkers, posing it as a true or false
question. Topics of the statements vary, but most are
in specific domains related to the Italian government,
policies, and social issues. The task presents several
challenges: extracting statements from segments of speeches,
determining appropriate contextual relevance both
temporally and factually, and verifying the statements’
accuracy.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>PejorativITy [21] is a task to investigate misogyny</title>
        <p>
          expressed through neutral words that can assume a
negative connotation when functioning as pejorative epithets. ItaEval [26] is a multifaceted evaluation suite
comThis challenge addresses a) the disambiguation of such prising three overarching task categories: (i) natural
ambiguous words in a given context; b) the detection language understanding, (ii) commonsense and factual
of misogyny in instances that contain such polysemic knowledge, and (iii) bias, fairness, and safety [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
ItaEwords. The task is divided into two parts, both framed val is a collection of 18 tasks encompassing existing and
as a binary classification. In Task A, the model is asked new datasets. The so-compiled ItaEval suite provides
to define if, given a tweet, the target word is used in a a standardized, multifaceted framework for evaluating
pejorative or non-pejorative way. In Task B, the model Italian language models, facilitating more rigorous and
is asked whether the whole sentence is misogynous. comparative assessments of model performance.
        </p>
        <sec id="sec-3-4-1">
          <title>TRACE-it (Testing Relative clAuses Comprehension through Entailment in ITalian) [23] is a benchmark</title>
          <p>PERSEID (PERSpEctivist Irony Detection) [22] con- 4. Evaluation Strategy
siders the task of irony detection from short social
media conversations collected from Twitter (X) and Red- Rooted in its very nature, CALAMITA’s biggest challenge
dit. Data is leveraged from MultiPICO, a recent multilin- is standardizing evaluation across many tasks and
scegual dataset with disaggregated annotations and annota- narios. To account for such high variability, we settled
tors’ metadata. The dataset evaluates whether prompting on a few fundamental choices that shape CALAMITA’s
LLMs with additional annotators’ demographic informa- core principles (Design choices) and left broad freedom
tion (gender only, age only, and the combination of the to challenge participants to specify fine-grained aspects
two) improves performance compared to a baseline in of their tasks (Participant choices). Base design choices
which only the input text is provided. shared across all tasks and high task-specific
customization balance standardization and versatility.
ABRICOT
AMELIA
BEEP
BLM-It
DIMMI
ECWCA
EurekaRebus
GATTINA
GEESE
GFG
GITA
INVALSI
ITA-SENSE
MACID
MT
Mult-IT
PejorativITy
PERSEID
Termite
TRACE-it
VeryfIT
ItaEval</p>
          <p>ItaCoLA
Belebele-it
News-Sum
IronITA
SENTIPOLC
SQuAD-it
TruthfulQA-it
ARC-it
XCOPA-it
HellaSwag-it
AMI
HONEST
GeNTE rephrasing
Multilingual HateCheck
HaSpeeDe2
**
**
*
*
*
Design choices. Following recent practices for lan- guage models. Llama’s 3.1 variant introduces
multilinguage model evaluation [e.g., 27, 28], we consider every gual support to the family’s previous iteration. ANITA
received task as a downstream task to be solved via stan- is a fine-tuned version of Llama 3 specializing in English
dard prompting. We support two types of tasks: Multiple- and Italian tasks.</p>
          <p>Choice (MC) and Open-Ended (OE) generation. MC tasks Our choice was driven by three primary reasons. First,
require a model to pick one or more correct answers from both models are open-weight, well-known within the
a finite set. OE tasks require models to generate output Italian NLP community, and explicitly support the Italian
tokens until a stopping criterion is met. For evaluating language. Second, they have been instruction fine-tuned,
multiple-choice tasks, we rank all candidates by their like- a training step that facilitates addressing tasks in
zerolihood conditioned on the prompt and pick the highest shot Third, they are within the 8 billion parameter range,
[29]. We normalize each option probability by the num- which allows for fast iteration and good performance.
ber of tokens. Closed-question question-answering is an
example of an MC task. We do not adopt a single strat- Results. At the time of writing, some of the results
egy for OE tasks, as evaluation depends on the semantics are still being collected. To provide a comprehensive
of the output. Machine translation and summarization and dynamic overview, we refer the reader to the
exare examples of OE tasks. Moreover, we standardize the ternal page where they get regularly updated: https:
decoding strategy across OE tasks. We use beam search //calamita-ailc.github.io/calamita2024/.
( = 5 ) for machine translation and greedy decoding for
all other tasks. See Appendix A for the complete details.</p>
          <p>
            To foster reproducibility, we base CALAMITA’s code- 5. Limitations
base on open-source tools. We forked and built our
evaluation code upon lm-eval [
            <xref ref-type="bibr" rid="ref16">30</xref>
            ]. When possible, we
recommended public and accessible data release to the
participants through the HuggingFace Hub.3 We release our
evaluation code at https://github.com/CALAMITA-AILC/
lm-evaluation-harness.
          </p>
          <p>CALAMITA is not intended to be an exhaustive
benchmark for testing abilities of Italian LLMs, especially at
this first release. Considering the strong collaborative
nature of this benchmark, coherence across tasks might not
be optimal, in spite of the eforts put in by the organisers
to uniform all datasets and the evaluation procedure.
Although we have paid attention to this issue, we cannot be
absolutely certain that none of the datasets, in one form
or another, have ended up in some training set, already.</p>
          <p>Participant choices. In addition to the data
associated with the task and the type (MC or OE), we request
that each participating team provides specifics
regarding compiling an arbitrary prompt and evaluating an
arbitrary model generation. Among prompting details, Acknowledgments
task proposers specified a prompt template and the
number of task demonstrations (0 for zero-shot, N for N- The ItaEval tasks submitted to CALAMITA are the
reshot prompting). In few-shot cases, we requested where sult of a joint efort of members of the “Risorse
to sample the demonstrations and the sampling strat- per la Lingua Italiana” community (rita-nlp.org): we
egy (static, dynamic-random, or dynamic-sequential). thank every member who dedicated their time to the
Among the evaluation details, we requested that par- project. For providing the computational resources we
ticipants specify any post-processing function for model thank CINECA (ISCRA grant: HP10C3RW9F; ISCRA C
raw outputs, one or more evaluation metrics, and relative grant: CALAMITA – HP10CKZDYT), the Center for
information. For reporting purposes, we collected a sin- Information Technology of the University of
Groningle evaluation score (the first metric listed by proposers). gen for their support and for providing access to the</p>
          <p>
            Crucially, we relied upon meta-description and code Hábrók high performance computing cluster and
Univerto streamline the communication between the task pro- sity of Turin for providing access to the HPC4AI cluster
posers and the challenge organizers. Participants were [
            <xref ref-type="bibr" rid="ref19">33</xref>
            ]. Malvina Nissim’s work is also part of the
“Hutasked to provide such information through a single file mane AI” theme of the Dutch Sectorplan for the
Humanfollowing a set of guidelines.4. ities. The work of Viviana Patti was is partially
supported by “HARMONIA” project - M4-C2, I1.3
PartenarModel Selection. We tested Llama 3.1 8B Instruct [
            <xref ref-type="bibr" rid="ref17">31</xref>
            ] iati Estesi - Cascade Call - FAIR - CUP C63C22000770006
and ANITA [
            <xref ref-type="bibr" rid="ref18">32</xref>
            ], two state-of-the-art decoder-only lan- - PE PE0000013 under the NextGenerationEU programme.
The work by Giuseppe Attanasio was supported by the
3Resulting from the efort for CALAMITA, 35 new datasets have Portuguese Recovery and Resilience Plan through project
4bSeeeen trheleeasgeudidweiltihneaspeartmihssttipves:/li/cgeinthsue.b.com/CALAMITA-AILC/ C645008882-00000055 (Center for Responsible AI) and by
calamita2024 and the information file at https://gist.github.com/ Fundação para a Ciência e Tecnologia through contract
g8a9/f5e82d38ce12831323b20dc79b0452c9 UIDB/50008/2020. The work by Pierpaolo Basile and Elio
          </p>
        </sec>
      </sec>
      <sec id="sec-3-5">
        <title>Musacchio was supported by the PNRR project FAIR</title>
        <p>Future AI Research (PE00000013), Spoke 6 - Symbiotic AI
(CUP H97G22000210007) under the NRRP MUR program
funded by the NextGenerationEU. The work of Matteo
Rinaldi and Jacopo Gili has been partly supported by the
Spoke “Future HPC &amp; Big Data” of the ICSC - Centro
Nazionale di Ricerca in “High Performance Computing,
Big Data and Quantum Computing”, funded by European
Union - NextGenerationEU.
I. Gonzalez-Dios, GITA4CALAMITA - Evaluat- Comprehension through Entailment in ITalian: A
ing the Physical Commonsense Understanding CALAMITA Challenge, in: Proceedings of the 10th
of Italian LLMs in a Multi-layered Approach: A Italian Conference on Computational Linguistics
CALAMITA Challenge, in: Proceedings of the 10th (CLiC-it 2024), Pisa, Italy, December 4 - December 6,
Italian Conference on Computational Linguistics 2024, CEUR Workshop Proceedings, CEUR-WS.org,
(CLiC-it 2024), Pisa, Italy, December 4 - December 6, 2024.
2024, CEUR Workshop Proceedings, CEUR-WS.org, [24] F. Ranaldi, E. S. Ruzzetti, D. Onorati, F. M.
Zan2024. zotto, L. Ranaldi, Termite Italian Text-to-SQL: A
[16] G. Puccetti, M. Cassese, A. Esuli, INVALSI - Mathe- CALAMITA Challenge, in: Proceedings of the 10th
matical and Language Understanding in Italian: A Italian Conference on Computational Linguistics
CALAMITA Challenge, in: Proceedings of the 10th (CLiC-it 2024), Pisa, Italy, December 4 - December 6,
Italian Conference on Computational Linguistics 2024, CEUR Workshop Proceedings, CEUR-WS.org,
(CLiC-it 2024), Pisa, Italy, December 4 - December 6, 2024.
2024, CEUR Workshop Proceedings, CEUR-WS.org, [25] J. Gili, V. Patti, L. Passaro, T. Caselli, VeryfIT
2024. Benchmark of Fact-Checked Claims for Italian: A
[17] P. Basile, E. Musacchio, L. Siciliani, ITA-SENSE CALAMITA Challenge, in: Proceedings of the 10th
- Evaluate LLMs’ ability for ITAlian word SENSE Italian Conference on Computational Linguistics
disambiguation: A CALAMITA Challenge, in: Pro- (CLiC-it 2024), Pisa, Italy, December 4 - December 6,
ceedings of the 10th Italian Conference on Com- 2024, CEUR Workshop Proceedings, CEUR-WS.org,
putational Linguistics (CLiC-it 2024), Pisa, Italy, 2024.</p>
        <p>December 4 - December 6, 2024, CEUR Workshop [26] G. Attanasio, M. La Quatra, A. Santilli, B. Savoldi,
Proceedings, CEUR-WS.org, 2024. ItaEval: A CALAMITA Challenge, in: Proceedings
[18] A. A. Ravelli, R. Varvara, L. Gregori, MACID - Mul- of the 10th Italian Conference on Computational
timodal ACtion IDentification: A CALAMITA Chal- Linguistics (CLiC-it 2024), Pisa, Italy, December 4
lenge, in: Proceedings of the 10th Italian Confer- - December 6, 2024, CEUR Workshop Proceedings,
ence on Computational Linguistics (CLiC-it 2024), CEUR-WS.org, 2024.</p>
        <p>Pisa, Italy, December 4 - December 6, 2024, CEUR [27] S. Mehta, M. H. Sekhavat, Q. Cao, M. Horton, Y. Jin,
Workshop Proceedings, CEUR-WS.org, 2024. C. Sun, S. I. Mirzadeh, M. Najibi, D. Belenko, P.
Zat[19] M. Cettolo, A. Piergentili, S. Papi, M. Gaido, M. Ne- loukal, et al., OpenELM: An eficient language
gri, L. Bentivogli, MAGNET - MAchines GeNEr- model family with open training and inference
ating Translations: A CALAMITA Challenge, in: framework, in: Workshop on Eficient Systems
Proceedings of the 10th Italian Conference on Com- for Foundation Models II@ ICML2024, 2024.
putational Linguistics (CLiC-it 2024), Pisa, Italy, De- [28] D. Groeneveld, I. Beltagy, P. Walsh, A. Bhagia,
cember 4 - December 6, 2024, CEUR Workshop Pro- R. Kinney, O. Tafjord, A. H. Jha, H. Ivison, I.
Magceedings, CEUR-WS.org, 2024. nusson, Y. Wang, et al., OLMo: Accelerating the
[20] M. Rinaldi, J. Gili, M. Francis, M. Gofetti, V. Patti, science of language models, in: Proceedings of
M. Nissim, Mult-IT Multiple Choice Questions the 62nd Annual Meeting of the Association for
on Multiple Topics in Italian: A CALAMITA Chal- Computational Linguistics (Volume 1: Long
Palenge, in: Proceedings of the 10th Italian Confer- pers), Association for Computational Linguistics,
ence on Computational Linguistics (CLiC-it 2024), Bangkok, Thailand, 2024, pp. 15789–15809. URL:
Pisa, Italy, December 4 - December 6, 2024, CEUR https://aclanthology.org/2024.acl-long.841. doi:10.</p>
        <p>Workshop Proceedings, CEUR-WS.org, 2024. 18653/v1/2024.acl- long.841.
[21] A. Muti, PejorativITy - In-Context Pejorative Lan- [29] T. B. Brown, B. Mann, N. Ryder, M. Subbiah,
guage Disambiguation: A CALAMITA Challenge, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam,
in: Proceedings of the 10th Italian Conference G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss,
on Computational Linguistics (CLiC-it 2024), Pisa, G. Krueger, T. Henighan, R. Child, A. Ramesh,
Italy, December 4 - December 6, 2024, CEUR Work- D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen,
shop Proceedings, CEUR-WS.org, 2024. E. Sigler, M. teusz Litwin, S. Gray, B. Chess, J. Clark,
[22] V. Basile, S. Casola, S. Frenda, S. M. Lo, PERSEID - C. Berner, S. McCandlish, A. Radford, I. Sutskever,
Perspectivist Irony Detection: A CALAMITA Chal- D. Amodei, Language models are few-shot
lenge, in: Proceedings of the 10th Italian Confer- learners, in: Proceedings of the 34th International
ence on Computational Linguistics (CLiC-it 2024), Conference on Neural Information Processing
Pisa, Italy, December 4 - December 6, 2024, CEUR Systems, NeurIPS ’20, Curran Associates Inc., Red
Workshop Proceedings, CEUR-WS.org, 2024. Hook, NY, USA, 2020, pp. 1877–1901. URL: https:
[23] D. Brunato, TRACE-it: Testing Relative clAuses //proceedings.neurips.cc/paper_files/paper/2020/</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kleyjo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>Beyond the imitation game: Quantifying and extrapolating the capabilities of language models</article-title>
          ,
          <source>Transactions on Machine Learning Research</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bioglio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bosca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <article-title>UINAUIL: A unified benchmark for Italian natural language understanding</article-title>
          , in: D.
          <string-name>
            <surname>Bollegala</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Ritter (Eds.),
          <source>Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>3</volume>
          :
          <string-name>
            <surname>System</surname>
            <given-names>Demonstrations)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Toronto, Canada,
          <year>2023</year>
          , pp.
          <fpage>348</fpage>
          -
          <lpage>356</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .acl-demo.
          <volume>33</volume>
          . doi:
          <volume>10</volume>
          . 18653/v1/
          <year>2023</year>
          .acl- demo.33.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>De Mattei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cafagna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>AI</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nissim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gatt</surname>
          </string-name>
          , Change-it@ evalita 2020:
          <article-title>Change headlines, adapt news, generate, EVALITA Evaluation of NLP and Speech Tools for ItalianDecember 17th</article-title>
          ,
          <year>2020</year>
          (
          <year>2020</year>
          )
          <fpage>235</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Attanasio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Delobelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. La</given-names>
            <surname>Quatra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Santilli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Savoldi</surname>
          </string-name>
          ,
          <article-title>Itaeval and tweetyita: A new extensive benchmark and eficiency-first language model for italian</article-title>
          , in: CLiC-it
          <source>2024: Tenth Italian Conference on Computational Linguistics</source>
          , Date:
          <year>2024</year>
          /12/04- 2024/12/06, Location: Pisa, Italy,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Puccetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Collacciani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Ravelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Esuli</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Bolognesi, ABRICOT - ABstRactness and Inclusiveness in COntexT: A CALAMITA Challenge</article-title>
          ,
          <source>in: Proceedings of the 10th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2024</year>
          ), Pisa, Italy, December 4 - December 6,
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Grundler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Galassi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Santin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fidelangeli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Galli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Palmieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lagioia</surname>
          </string-name>
          , G. Sartor, P. Torroni,
          <article-title>AMELIA - Argument Mining Evaluation on Legal documents in ItAlian: A CALAMITA Challenge</article-title>
          ,
          <source>in: Proceedings of the 10th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2024</year>
          ), Pisa, Italy, December 4 - December 6,
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Mercorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Potertì</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Serino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Seveso</surname>
          </string-name>
          ,
          <string-name>
            <surname>BEEP - BEst DrivEr's License</surname>
          </string-name>
          <article-title>Performer: A CALAMITA Challenge</article-title>
          ,
          <source>in: Proceedings of the 10th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2024</year>
          ), Pisa, Italy, December 4 - December 6,
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Samo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nastase</surname>
          </string-name>
          , P. Merlo, BLMIt
          <article-title>- Blackbird Language Matrices for Italian: A CALAMITA Challenge</article-title>
          ,
          <source>in: Proceedings of the 10th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2024</year>
          ), Pisa, Italy, December 4 - December 6,
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Manna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Di Buono</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. Giordano,</surname>
          </string-name>
          <article-title>DIMMI - Drug InforMation Mining in Italian: A CALAMITA Challenge</article-title>
          ,
          <source>in: Proceedings of the 10th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2024</year>
          ), Pisa, Italy, December 4 - December 6,
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zugarini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zeinalipour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fusco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zanollo</surname>
          </string-name>
          ,
          <string-name>
            <surname>ECWCA - Educational CrossWord Clues Answering A CALAMITA Challenge</surname>
          </string-name>
          ,
          <source>in: Proceedings of the 10th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2024</year>
          ), Pisa, Italy, December 4 - December 6,
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>G.</given-names>
            <surname>Sarti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Caselli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bisazza</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Nissim, EurekaRebus - Verbalized Rebus Solving with LLMs: A CALAMITA Challenge</article-title>
          ,
          <source>in: Proceedings of the 10th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2024</year>
          ), Pisa, Italy, December 4 - December 6,
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Francis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rinaldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gili</surname>
          </string-name>
          , L. De Cosmo,
          <string-name>
            <given-names>S.</given-names>
            <surname>Iannaccone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nissim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          , GATTINA
          <article-title>- GenerAtion of TiTles for Italian News Articles: A CALAMITA Challenge</article-title>
          ,
          <source>in: Proceedings of the 10th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2024</year>
          ), Pisa, Italy, December 4 - December 6,
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zaninello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Magnini</surname>
          </string-name>
          , GEESE - Generating and
          <article-title>Evaluating Explanations for Semantic Entailment: a CALAMITA Challenge</article-title>
          ,
          <source>in: Proceedings of the 10th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2024</year>
          ), Pisa, Italy, December 4 - December 6,
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Frenda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Piergentili</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Savoldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Madeddu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rosola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Casola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ferrando</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Negri</surname>
          </string-name>
          , L. Bentivogli, GFG -
          <string-name>
            <surname>Gender-Fair Generation</surname>
          </string-name>
          :
          <article-title>A CALAMITA Challenge</article-title>
          ,
          <source>in: Proceedings of the 10th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2024</year>
          ), Pisa, Italy, December 4 - December 6,
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>G.</given-names>
            <surname>Pensa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Azurmendi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Etxaniz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Altuna</surname>
          </string-name>
          , file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>S.</given-names>
            <surname>Biderman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schoelkopf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sutawika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Abbasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Aji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Ammanamanchi</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>1</year>
          .
          <string-name>
            <given-names>Technical</given-names>
            <surname>Details S. Black</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clive</surname>
          </string-name>
          , et al.,
          <article-title>Lessons from the trenches on reproducible evaluation of language models, arXiv We run our experiments on the LEONARDO HPC infraspreprint</article-title>
          arXiv:
          <volume>2405</volume>
          .14782 (
          <year>2024</year>
          ).
          <article-title>tructure (Booster partition)</article-title>
          .
          <source>The booster module par-</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dubey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jauhri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pandey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kadian</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Al
          <article-title>- tition is based on BullSequana XH2135 supercomputer Dahle, A</article-title>
          .
          <string-name>
            <surname>Letman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mathur</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Schelten</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Yang, nodes, each with four NVIDIA Tensor Core GPUs (custom</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Fan</surname>
          </string-name>
          , et al.,
          <source>The Llama 3 Herd of Models, arXiv Ampere A100 GPU 64GB HBM2e, NVLink 3</source>
          .0 (
          <issue>200GB</issue>
          /s))
          <source>preprint arXiv:2407.21783</source>
          (
          <year>2024</year>
          ).
          <article-title>and a single Intel CPU.</article-title>
          <year>5</year>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>M.</given-names>
            <surname>Polignano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Basile</surname>
          </string-name>
          , G. Semeraro,
          <article-title>Ad- We forked the lm-eval-harness ofivanced natural-based interaction for the italian cial repository at the commit with hash language: Llamantino-3-anita, arXiv preprint b2bf7bc4a601c643343757c92c1a51eb69caf1d7</article-title>
          .
          <source>arXiv:2405.07101</source>
          (
          <year>2024</year>
          ).
          <source>We report all technical details on our oficial webpage. 6</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>M.</given-names>
            <surname>Aldinucci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rabellino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pironti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Spiga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Viviani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Drocco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guerzoni</surname>
          </string-name>
          , G. Boella,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mellia</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2</year>
          .
          <string-name>
            <surname>Generation Configuration P. Margara</surname>
            ,
            <given-names>I. Drago</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Marturano</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Marchetto, Table 3 reports the generation parameters we used for E</article-title>
          . Piccolo,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bagnasco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lusso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vallero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>At-</surname>
          </string-name>
          Open-Ended tasks. tardi, A.
          <string-name>
            <surname>Barchiesi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Colla</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Galeazzi</surname>
          </string-name>
          ,
          <article-title>Hpc4ai, an ai-on-demand federated platform endeavour</article-title>
          ,
          <source>Parameter Value i2n0:18</source>
          .ACUMRL:Cohmttppus:t/i/nirgis.Furnointoti.ietr/sr,etrIisecvhei/ah, a nItdalley/,
          <source>TBeamtcphesriazteure 01.∗0</source>
          <volume>2318</volume>
          /1765596/689772/2018_hpc4ai_
          <article-title>ACM_CF.pdf</article-title>
          .
          <source>Sampling False doi:10.1145/3203217</source>
          .3205340. Stopping criteria \n\n, &lt;/s&gt;, &lt;|im_end|&gt;, “. ”, &lt;|eot_id|&gt;, &lt;|end_of_text|&gt;
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>