<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Understanding in Italian: A CALAMITA Challenge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giovanni Puccetti</string-name>
          <email>giovanni.puccetti@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Cassese</string-name>
          <email>maria.cassese@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Esuli</string-name>
          <email>andrea.esuli@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Mathematical Understanding, Language Understanding, Invalsi, Large Language Models, Italian Language Models</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CLiC-it 2024: Tenth Italian Conference on Computational Linguistics</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>While Italian is a high resource language, there are few Italian-native benchmarks to evaluate Language Models (LMs) generative abilities in this language. This work presents two new benchmarks: Invalsi MATE to evaluate models performance on mathematical understanding in Italian and Invalsi ITA to evaluate language understanding in Italian.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Challenge: Introduction and</title>
    </sec>
    <sec id="sec-3">
      <title>Motivation</title>
      <p>
        There are several benchmarks to evaluate
mathematical understanding of LLMs based on English tests [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6, 7, 8</xref>
        ]
and there are also several multi-domain benchmarks
inAssessing the quality of Large Language Models is a chal- volving Italian [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], however there aren’t any specifically
lenging task because these models can virtually perform
any task that can be presented through natural language. focus on high-school questions, an English benchmark
focused on mathematical understanding in Italian. We
To address this dificulty, each model needs to be tested
on several tasks at once.
      </p>
      <p>To help provide new benchmarks to evaluate LLMs
in Italian, We propose two benchmarks, Invalsi MATE
and Invalsi ITA the first meant to evaluate LLMs’
mathematical understanding and the second to evaluate their
language understanding, both in Italian.</p>
      <p>
        These benchmarks originate from the Invalsi tests,
which have been used in the past for demographic stud- language.
ies [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ], but, to the best of our knowledge we are the
nEvelop-O
LGOBE
(A. Esuli)
(M. Cassese); https://www.esuli.it/ (A. Esuli)
0000-0002-5725-4322 (A. Esuli)
      </p>
      <p>Attribution 4.0 International (CC BY 4.0).</p>
      <p>
        similar to Invalsi MATE is the GSM8k one, [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which
contains 8,500 high-school questions.
      </p>
      <p>
        Language Models understanding of language in
English is also well studied, there are several benchmarks
meant to measure the ability of language models to
understand language constructs in English, such as [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ],
also arranged into extensive suites [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. On the contrary
there are fewer examples of these tests for the Italian
      </p>
      <p>
        Therefore we propose Invalsi ITA which contains
quesmarks, e.g. MNLI [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], SQuAD [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and others from the
GLUE suite [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The questions in the dataset cover
several aspects of language understanding, ranging from the
ability to extract specific information, such as the date
when something happened to more complex information
such as whether two events implicate each other or not.
      </p>
      <p>
        These two datasets allow us to measure two key
abilities of language models in Italian, to make the comparison
among diferent models more fair we cast all questions as
multiple choice and measure models’ performance by
seifrst to use them to test LLMs performance in Italian [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], tions that are usually split among several diferent
benchfollowed only later by others [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License lecting the answer with the highest likelihood according</p>
      <sec id="sec-3-1">
        <title>Testo</title>
        <p>Elisa è uscita da casa questa mattina alle ore 8:15.
Elisa è rientrata nel pomeriggio alle ore 1:15</p>
      </sec>
      <sec id="sec-3-2">
        <title>Domanda</title>
        <p>Quanto tempo è stata fuori casa Elisa?
A. 5 ore B. 7 ore C. 9 ore D. 11 ore</p>
      </sec>
      <sec id="sec-3-3">
        <title>Testo</title>
        <p>Se moltiplichi per 2 un numero naturale e dal
risultato sottrai 1, ottieni sempre un numero pari.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Domanda</title>
        <p>Vero o Falso?
(a) scelta multipla question from Invalsi MATE.
(b) vero/falso question from Invalsi MATE.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Testo</title>
        <p>Filippo dice: per trovare il numero della mia
maglietta aggiungi una decina e sei unità al numero 4.</p>
      </sec>
      <sec id="sec-3-6">
        <title>Domanda</title>
        <p>Qual è il numero della maglietta di Filippo?</p>
      </sec>
      <sec id="sec-3-7">
        <title>Testo</title>
        <p>Luca lancia due dadi a sei facce non truccati.</p>
      </sec>
      <sec id="sec-3-8">
        <title>Domanda</title>
        <p>Completa la frase inserendo una delle espressioni:
La probabilità che la somma dei punti sia 12 è
maggiore della, minore della, uguale alla probabilità che
la somma sia 2.
(c) numero question from Invalsi MATE.</p>
        <p>(d) completa frase question from Invalsi MATE.
to the model. and the limitations of the data.</p>
        <p>
          We measure the performance of 4 strong large
models, mixtral instruct [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], mistral instruct [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], llama 3 8b 2.1. Task 1: Mathematical Understanding
instruct [18], anita 8b dpo [19], the first three are English- in Italian (Invalsi MATE)
ifrst and the fourth is fine-tuned in Italian, and the current
only Italian-first model minerva 3b. We show that on In- The first task consists in answering mathematical
quesvalsi ITA the best model among those we tested is mixtral tions in Italian. These questions are meant for students
instruct, which reaches an accuracy of 0.8, while on In- from 6 to 18 years of age, therefore the kind of
quesvalsi MATE the highest accuracy is 0.55, also achieved tion can vary significantly, from simpler, example-based,
by mixtral instruct. ones that don’t require any knowledge besides
count
        </p>
        <p>Both language understanding and mathematical un- ing, to more complex ones requiring basic geometry and
derstanding are key abilities for students as well as lan- calculus training and knowledge, never beyond what is
guage models, particularly since these models are often demanded in basic high-school tests.
used in learning environments. By adding these bench- The questions are of 4 kinds, scelta multipla, completa
marks to the CALAMITA suite we hope they will help frase, vero/falso and numero:
the development of LLMs in Italian by providing a more
comprehensive evaluation of their abilities and thus
fostering the research and development of models in this
language.</p>
        <p>The CALAMITA special event [20], which has aims to
establishing a shared benchmark for LLMs in Italian, is a
ifrst step towars a systematic evaluation of LLMs in this
language. We hope that the Invalsi challenge will enrich
the Linguistic and Mathematical understanding branches
of this shared benchmark.
• scelta multipla (multiple choice): the question
requires to pick the right answer among four
possible ones;
• vero/falso (true/false): the question requires to</p>
        <p>pick the right answer between true and false;
• numero (number): the question requires to pick a</p>
        <p>number that is the correct answer to the question;
• completa frase (fill the gap): the question requires
to fill one or more missing words to make the text
coherent.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>2. Challenge: Description</title>
      <p>Of the four question types, scelta multipla and
vero/falso are naturally multiple choice, with scelta multipla
The challenge is composed of two tasks: Invalsi MATE always having 4 possible answers (A, B, C, D) and
vero/and Invalsi ITA. For each task, we provide a detailed falso always 2, (true, false). Questions of the numero type,
description of the data, the metrics used for evaluation, are not naturally multiple choice, since the answer is a
(a) Invalsi MATE
(b) Invalsi ITA
• domanda aperta (open question): the question
requires to pick the passage in the text that answers
the question.
• altro (other): A small share of questions belong
to open-ended questions with varying scope that
are hard to put under a single label.</p>
      <p>Similar to Invalsi MATE, this task involves only
multiple-choice questions, evaluated through a likelihood
approach. Both scelta multipla and binaria are naturally
of this kind, the first with 4 options (A, B, C D) and the
second with 2 options that change for each question.</p>
      <p>Both domanda aperta and altro questions are hard to turn
into multiple-choice ones and therefore we discard them.</p>
      <p>Also for Invalsi ITA, this involves only discarding about
180 questions out of 1297, therefore the task only involves
1117 samples.
number among all possible ones (some of the answers
will be a year, e.g. 1948, while others can be a decimal
number of liters of milk, e.g. 0.2), to address this we add
3 extra answers that are realistic but wrong to make the
questions multiple choice. Finally, completa frase
questions, which are only few (20), are too dificult to turn
into multiple choice without changing their meaning,
and therefore we exclude them.</p>
      <sec id="sec-4-1">
        <title>2.2. Task 2: Language Understanding in</title>
      </sec>
      <sec id="sec-4-2">
        <title>Italian (Invalsi ITA)</title>
        <p>The second task consists in answering Italian language
understanding questions, similarly to task 1, these
questions are also appropriate for students between 6 and 18
years old, and they are overall not too dificult to answer.</p>
        <p>Most of the questions concern a text passage that has to
be included in the model context, making this evaluation
more costly because the context becomes considerably 3. Data description
larger. The text passage is where the dificulty diference
between ages is more evident, since it can be a simple 3.1. Origin of data
and short story for primary school students, while they
are generally longer and more involved texts for older The dataset is built upon the questions from the Invalsi
students. tests of the last 15 years. These tests are administered</p>
        <p>The questions are of 4 diferent types, scelta multipla, to students yearly. There are three diferent Invalsi tests,
binaria, domanda aperta and altro: Language, Mathematics and English. For the scope of
this datasets we don’t look into the English test, but we
• scelta multipla (multiple choice): the question limit ourselves to the Italian language and Mathematics
requires to pick the right answer among four pos- ones. The original questions can be accessed here 1
sible ones; Some of the questions from the original tests contain
• binaria (binary): the question requires to pick visual content as part of the question, we omit these
the right answer about a binary property of a questions since we focus on the language understanding
statement, e.g. True - False, Before - After, etc. abilities of the models.
Question Type
N. Questions
Model
mixtral instruct
mistral instruct
anita 8b dpo
llama 3 8b instruct
minerva 3b
random
0.55
0.44
0.47
0.48
0.20
0.28
scelta multipla vero/falso
244 54</p>
        <p>The data is first collected as is from the webpage and
afterwards it is manually checked for errors and
inconsistencies from two annotators. The annotators have
MSc in Mathematics and Computer Science, which gives
them suficient knowledge to identify issues in the Invalsi
MATE questions. For Invalsi ITA, the annotators don’t
have an appropriate background, however, the questions
are simple enough that they can be easily understood
and checked for errors by anybody who has completed
the mandatory education.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.2. Data format</title>
        <p>The Invalsi MATE dataset has 8 diferent columns:
• testo: this field contains the context needed to
answer the questions, it is often empty for Invalsi
MATE since most of the context is part of the
domanda field itself;
• domanda: this field contains the question itself,
including possible answer options, e.g. for scelta
multipla questions;
• risposta: this field contains the correct answer;
• test_id: this field is just an id to identify each
sample;
• tipo: this field indicates the question type, among
scelta multipla, vero/falso, numero and completa
frase;
• alt1, alt2 and alt3: this three fields indicate the
alternative values for the numero questions since
this we chose ourselves and are not indicated in
the domanda field.</p>
        <p>The Invalsi ITA dataset has the same fields as the
Invalsi MATE one with the exception that the testo field is
often present and generally the longest.</p>
        <p>We evaluate in a zero-shot fashion just providing the
model with question and using a likelihood based method,
we pick as the model’s answer the one with the highest
likelihood among the options available. This is always
possible since we have recast all the questions as multiple
Question Type
N. Questions
Model
mixtral instruct
mistral instruct
anita 8b dpo
llama 3 8b instruct
minerva 3b
random
0.80
0.49
0.71
0.69
0.30
0.27
scelta multipla</p>
        <p>977
Accuracy
0.69
0.51
0.66
0.61
0.54
0.44
choice ones. We also don’t use chain of thought prompts
or similar methods. Since this is the first attempt to
build a dataset on mathematical understanding in Italian,
currently we evaluate with the simplest approach.</p>
      </sec>
      <sec id="sec-4-4">
        <title>3.3. Detailed data statistics</title>
        <p>The data does not have a train and a test split because
we have a limited number of samples. The Invalsi MATE
split is composed of 420 samples, of which 400 are used
in the benchmark, since we exclude the 20 questions
marked as completa frase since they can´t be made into
multiple choice.</p>
        <p>Figure 2a shows the percentage of questions of each
kind in the dataset, scelta multipla has the largest share,
58%, the second most present is numero, 24.7% and then
vero/falso and completa frase are fewer. Table 1 has a
random row that shows the performance if one were to
pick random questions, moreover the correct answers for
each question type are approximately evenly distributed
among labels. Specifically, scelta multipla questions have
answers distributed as follows, 46 are labelled A, 87 B,
71 C and 40 D, showing a moderate balance. Similarly
for vero/falso questions there 24 questions with positive
answer and 30 with negative answer.</p>
        <p>Similarly, Figure 2b shows the percentage of questions
of each of the kinds present in Invalsi ITA, scelta multipla
is by a large margin the most present, composing 76.4%
of all the questions, binaria is second with 10.9% while
domanda aperta and altro are fewer.</p>
        <p>Table 2 shows the performance one can achieve
picking answers at random in each split, and moreover the
correct answers are evenly distributed for each label also
in this dataset. In particular, for Invalsi ITA 254 of the
scelta multipla questions have answer A, 255 B, 263 C
and 205 D which is comparable to the distribution in the
Invalsi MATE dataset and similarly does binaria.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Metrics</title>
      <p>and the second for the evaluation of language
understanding in Italian.</p>
      <p>Since all the datasets and splits we study have balanced For Invalsi MATE we have collected 420 questions
labels, we choose to measure accuracy. In particular, divided into 4 types, scelta multipla, vero/falso, numero
since all the questions in the datasets we propose are and completa frase and we evaluate 4 strong language
multiple choice, it is straightforward to measure accuracy models that are near SOTA in their weight range, mixtral
even if there are questions of diferent kinds, by simply instruct, mistral instruct, llama 3 8b instruct and anita 8b
counting |     |/|   | . dpo. We find that this models are still far from perfect</p>
      <p>To see how challenging our benchmark is , we measure mathematical understanding in Italian with the highest
four powerful language models based on mistral, mixtral accuracy, achieved by mixtral instruct being 55%.
and llama 3, in particular, we measure the performance of For Invalsi ITA we have collected 1297 questions
dimistral instruct, mixtral instruct, anita 8b dpo and llama vided into 4 types, scelta multipla, binaria, domanda
3 8b instruct. aperta and altro, we tested the same models also on this</p>
      <p>These models have between 7 and 54 billion param- benchmark and found that models are stronger at
laneters, three of them, mistral instruct, anita 8b dpo and guage understanding, with the highest accuracy in this
llama 3 8b instruct are purely autoregressive transform- task at 80%, also in this case, achieved by mixtral instruct.
ers, while mixtral instruct is a MoE architecture that has Both mathematical and Language understanding are
54 billion parameters but only uses 14 billion at inference. key abilities for LLMs, we believe that our two
benchWe test them all in the same way, using a likelihood based marks will foster the development of LLMs in Italian and
approach. pave the way for new more challenging benchmarks on</p>
      <p>Table 1 shows the performance of these models on mathematical and language understanding in Italian.
Invalsi MATE on the whole dataset, in the ALL column
and on each split scelta multipla, vero/falso and numero
in the respective columns. mixtral instruct is the clear 6. Limitations
winner among the models we tested, it beats the second
best, llama 3 8b instruct by 7% accuracy on the entire The main limitations of the benchmark we propose lies
dataset. On the separate splits, mixtral instruct is best in Task 2, Invalsi ITA we show that the models we test
overall, however the second best model changes, with achieve very high accuracy, up to 80% on this
benchllama 3 8b instruct being second in scelta multipla with a mark, making it possibly too simple for newer and larger
7% gap, anita 8b dpo second on vero/falso with a smaller models, nevertheless, current Italian first LLMs are not
2% performance gap and mistral instruct being second comparable to larger English-first ones and therefore we
best in numero with a 3% gap. believe it can still be valuable in this transitory phase.</p>
      <p>The total accuracy is bound by 55% showing that the In- On the contrary Invalsi MATE is very challenging and it
valsi MATE task is challenging for models of the sizes we seems that models won’t saturate it soon.
tested, up to 54B parameters, and that the performance We believe that there is a limited risk from
contaminaa model achieves provides valuable insights about how tion from both existing English and Italian tests.
well it can perform mathematical reasoning in Italian. Concerning direct contamination, we were unable to</p>
      <p>Table 2 shows the performance of the same models ifnd any web page that would expose the answers openly
on Invalsi ITA, the model ranking stays the same, with without needing any sort of authentication, making it
difmixtral instruct the strongest and anita 8b dpo second ifcult to crawl these data automatically, therefore, while
best. Performance on Invalsi ITA is higher across all fields the questions might be present in the training set of some
with mixtral instruct achieving 80% accuracy. Unlike of the models, we deem it unlikely that the answers were
what happens for Invalsi MATE, in Invalsi ITA the second there too.
best model is the same across the board, anita 8b dpo is Concerning contamination through translation from
the second in scelta multipla as well as in binaria with a English, the Invalsi questions are carefully crafted to
performance gap around 10% in all question types. match the grade of the students that will undertake them,</p>
      <p>The accuracy of the best model on all the questions at therefore we believe it is unlikely that they are taken
once is 80% showing that the models we tested perform from English questions available in other online sources,
well in the language understanding in Italian. but rather created specifically for each new annual test.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusions</title>
      <p>We propose two Tasks, Invalsi MATE and Invalsi ITA the
ifrst for the evaluation of mathematical understanding</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was also partially supported by: FAIR
(PE00000013) project under the NextGenerationEU
programme, partially by the PNRR project ITSERR (CUP
B53C22001770006) and partially by the Project PRIN
2022EPTPJ9 (WEMB – “Word EMBeddings: From
Cognitive Linguistics to Language Engineering, and Back”),
funded by the Italian Ministry of University and Research
(MUR).
V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V.
Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X.
Martinet, X. Wang, X. E. Tan, X. Xie, X. Jia, X. Wang,
Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song,
Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan,
Z. Chen, Z. Papakipos, A. Singh, A. Grattafiori,
A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A.
Victoria, A. Goldstand, A. Menon, A. Sharma, A.
Boesenberg, A. Vaughan, A. Baevski, A. Feinstein,
A. Kallet, A. Sangani, A. Yunus, A. Lupu, A.
Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan,
A. Ramchandani, A. Franco, A. Saraf, A.
Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A.
Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang,
B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu,
B. Ni, B. Hancock, B. Wasti, B. Spence, B.
Stojkovic, B. Gamido, B. Montalvo, C. Parker, C.
Burton, C. Mejia, C. Wang, C. Kim, C. Zhou, C. Hu,
C.-H. Chu, C. Cai, C. Tindal, C. Feichtenhofer,
D. Civin, D. Beaty, D. Kreymer, D. Li, D. Wyatt,
D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh,
D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland,
E. Dowling, E. Jamil, E. Montgomery, E. Presani,
E. Hahn, E. Wood, E. Brinkman, E. Arcaute, E.
Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Ozgenel,
F. Caggioni, F. Guzmán, F. Kanayet, F. Seide, G. M.
Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern,
G. Thattai, G. Herman, G. Sizov, Guangyi, Zhang,
G. Lakshminarayanan, H. Shojanazeri, H. Zou,
H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk,
H. Aspegren, H. Goldman, I. Damlaj, I. Molybog,
I. Tufanov, I.-E. Veliche, I. Gat, J. Weissman, J.
Geboski, J. Kohli, J. Asher, J.-B. Gaya, J. Marcus, J. Tang,
J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong,
J. Jin, J. Yang, J. Cummings, J. Carvill, J.
Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang,
K. Wu, K. H. U, K. Saxena, K. Prasad, K.
Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan,
K. Michelena, K. Li, K. Huang, K. Chawla, K.
Lakhotia, K. Huang, L. Chen, L. Garg, L. A, L. Silva,
L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich,
L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt,
M. Tsimpoukelli, M. Mankus, M. Hasson, M. Lennie,
M. Reso, M. Groshev, M. Naumov, M. Lathi, M.
Keneally, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel,
M. Vyatskov, M. Samvelyan, M. Clark, M. Macey,
M. Wang, M. J. Hermoso, M. Metanat, M.
Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White,
N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. P.
Laptev, N. Dong, N. Zhang, N. Cheng, O. Chernoguz,
O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh,
P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux,
P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj,
Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy,
R. Nayani, R. Mitra, R. Li, R. Hogan, R. Battey,
R. Wang, R. Maheswari, R. Howes, R. Rinott, S. J.
Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon,
S. Sidorov, S. Pan, S. Verma, S. Yamamoto, S.
Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin,
S. C. Zha, S. Shankar, S. Zhang, S. Zhang, S. Wang,
S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max,
S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad,
S. Gupta, S. Cho, S. Virk, S. Subramanian, S.
Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best,
T. Kohler, T. Robinson, T. Li, T. Zhang, T. Matthews,
T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V.
Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Albiero,
V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov,
W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable,
X. Tang, X. Wang, X. Wu, X. Wang, X. Xia, X. Wu,
X. Gao, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang,
Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Hao,
Y. Qian, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick,
Z. Wen, Z. Yang, Z. Zhao, The llama 3 herd of
models, 2024. URL: https://arxiv.org/abs/2407.21783.
arXiv:2407.21783.
[19] M. Polignano, P. Basile, G. Semeraro, Advanced
natural-based interaction for the italian language:
Llamantino-3-anita, 2024. URL: https://arxiv.org/
abs/2405.07101. arXiv:2405.07101.
[20] G. Attanasio, P. Basile, F. Borazio, D. Croce, M.
Francis, J. Gili, E. Musacchio, M. Nissim, V. Patti, M.
Rinaldi, D. Scalena, CALAMITA: Challenge the
Abilities of LAnguage Models in ITAlian, in:
Proceedings of the 10th Italian Conference on
Computational Linguistics (CLiC-it 2024), Pisa, Italy,
December 4 - December 6, 2024, CEUR Workshop
Proceedings, CEUR-WS.org, 2024.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Bolondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Cascella</surname>
          </string-name>
          ,
          <article-title>Somministrazione delle prove invalsi dal 2009 al 2015: un patrimonio d'informazioni tra evidenze psicometriche e didattiche, in: I dati INVALSI: uno strumento per la ricerca</article-title>
          ,
          <source>Franco Angeli, Milano</source>
          ,
          <year>2017</year>
          , p.
          <fpage>14</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Costanzo</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Desimoni, Beyond the mean estimate: a quantile regression analysis of inequalities in educational outcomes using invalsi survey data, Large-scale Assessments in Education (</article-title>
          <year>2017</year>
          ). URL: https://doi.org/10.1186/s40536-017-0048-4. doi:
          <volume>10</volume>
          . 1186/s40536- 017- 0048- 4.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pietschnig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Oberleiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Tofalini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Giofrè</surname>
          </string-name>
          ,
          <article-title>Reliability of the g factor over time in italian invalsi data (2010-2022): What can achievement-g tell us about the flynn efect?</article-title>
          ,
          <source>Personality and Individual Diferences</source>
          <volume>214</volume>
          (
          <year>2023</year>
          )
          <article-title>112345</article-title>
          . URL: https://www.sciencedirect.com/ science/article/pii/S0191886923002684. doi:https: //doi.org/10.1016/j.paid.
          <year>2023</year>
          .
          <volume>112345</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Esuli</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. Puccetti,</surname>
          </string-name>
          <article-title>The invalsi benchmarks: measuring linguistic and mathematical understanding of large language models in italian, 2024</article-title>
          . URL: https: //arxiv.org/abs/2403.18697. arXiv:
          <volume>2403</volume>
          .
          <fpage>18697</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Mercorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mezzanzanica</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Potertì</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Serino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Seveso</surname>
          </string-name>
          ,
          <article-title>Disce aut deficere: Evaluating llms proficiency on the invalsi italian benchmark</article-title>
          ,
          <year>2024</year>
          . URL: https://arxiv.org/abs/2406.17535. arXiv:
          <volume>2406</volume>
          .
          <fpage>17535</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hendrycks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Burns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kadavath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Basart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Steinhardt</surname>
          </string-name>
          ,
          <article-title>Measuring mathematical problem solving with the math dataset</article-title>
          ,
          <source>NeurIPS</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Fei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>MathBench: Evaluating the theory and application proficiency of LLMs with a hierarchical mathematics benchmark</article-title>
          , in: L.
          <string-name>
            <surname>-W. Ku</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Martins</surname>
          </string-name>
          , V. Srikumar (Eds.),
          <source>Findings of the Association for Computational Linguistics ACL</source>
          <year>2024</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Bangkok, Thailand and virtual meeting,
          <year>2024</year>
          , pp.
          <fpage>6884</fpage>
          -
          <lpage>6915</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .findings-acl.
          <volume>411</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cobbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kosaraju</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bavarian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Plappert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tworek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hilton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Nakano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hesse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schulman</surname>
          </string-name>
          , Training verifiers to solve math word problems,
          <year>2021</year>
          . arXiv:
          <volume>2110</volume>
          .
          <fpage>14168</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>R. J. Das</surname>
            ,
            <given-names>S. E.</given-names>
          </string-name>
          <string-name>
            <surname>Hristov</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>D. I.</given-names>
          </string-name>
          <string-name>
            <surname>Dimitrov</surname>
            ,
            <given-names>I. Koychev</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          , Exams-v:
          <article-title>A multi-discipline multilingual multimodal exam benchmark for evaluating vision language models</article-title>
          ,
          <year>2024</year>
          . arXiv:
          <volume>2403</volume>
          .
          <fpage>10378</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bentivogli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Magnini</surname>
          </string-name>
          , I. Dagan,
          <string-name>
            <given-names>H. T.</given-names>
            <surname>Dang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Giampiccolo</surname>
          </string-name>
          ,
          <article-title>The fith PASCAL recognizing textual entailment challenge</article-title>
          ,
          <source>in: Proceedings of the Second Text Analysis Conference, TAC</source>
          <year>2009</year>
          , Gaithersburg, Maryland, USA, November
          <volume>16</volume>
          -
          <issue>17</issue>
          ,
          <year>2009</year>
          , NIST,
          <year>2009</year>
          . URL: https://tac.nist.gov/publications/2009/additional. papers/RTE5_overview.
          <source>proceedings.pdf.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , X. Liu, J. Liu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Duh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. V.</given-names>
            <surname>Durme</surname>
          </string-name>
          ,
          <article-title>Record: Bridging the gap between human and machine commonsense reading comprehension</article-title>
          ,
          <year>2018</year>
          . URL: https://arxiv.org/abs/
          <year>1810</year>
          .12885. arXiv:
          <year>1810</year>
          .12885.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Pruksachatkun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nangia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Michael</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Bowman</surname>
          </string-name>
          ,
          <article-title>SuperGLUE: a stickier benchmark for general-purpose language understanding systems</article-title>
          , Curran Associates Inc.,
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nangia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bowman</surname>
          </string-name>
          ,
          <article-title>A broadcoverage challenge corpus for sentence understanding through inference</article-title>
          , in: M.
          <string-name>
            <surname>Walker</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Stent (Eds.),
          <source>Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , New Orleans, Louisiana,
          <year>2018</year>
          , pp.
          <fpage>1112</fpage>
          -
          <lpage>1122</lpage>
          . URL: https://aclanthology.org/ N18-1101. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N18</fpage>
          - 1101.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>P.</given-names>
            <surname>Rajpurkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <article-title>Know what you don't know: Unanswerable questions for SQuAD</article-title>
          , in: I. Gurevych, Y. Miyao (Eds.),
          <source>Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Melbourne, Australia,
          <year>2018</year>
          , pp.
          <fpage>784</fpage>
          -
          <lpage>789</lpage>
          . URL: https://aclanthology.org/P18-2124. doi:
          <volume>10</volume>
          .18653/ v1/
          <fpage>P18</fpage>
          - 2124.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Michael</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          , S. Bowman,
          <string-name>
            <surname>GLUE:</surname>
          </string-name>
          <article-title>A multi-task benchmark and analysis platform for natural language understanding</article-title>
          , in: T. Linzen,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chrupała</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Alishahi (Eds.),
          <source>Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, Association for Computational Linguistics</source>
          , Brussels, Belgium,
          <year>2018</year>
          , pp.
          <fpage>353</fpage>
          -
          <lpage>355</lpage>
          . URL: https://aclanthology.org/W18-5446. doi:
          <volume>10</volume>
          .18653/ v1/
          <fpage>W18</fpage>
          - 5446.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A. Q.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sablayrolles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mensch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Savary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bamford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Chaplot</surname>
          </string-name>
          , D. de las Casas,
          <string-name>
            <given-names>E. B.</given-names>
            <surname>Hanna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bressand</surname>
          </string-name>
          , G. Lengyel, G. Bour,
          <string-name>
            <given-names>G.</given-names>
            <surname>Lample</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. R.</given-names>
            <surname>Lavaud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Saulnier</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>A. Lachaux</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Stock</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Subramanian</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Antoniak</surname>
            ,
            <given-names>T. L.</given-names>
          </string-name>
          <string-name>
            <surname>Scao</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Gervet</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lavril</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lacroix</surname>
            ,
            <given-names>W. E.</given-names>
          </string-name>
          <string-name>
            <surname>Sayed</surname>
          </string-name>
          , Mixtral of experts,
          <year>2024</year>
          . arXiv:
          <volume>2401</volume>
          .
          <fpage>04088</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A. Q.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sablayrolles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mensch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bamford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Chaplot</surname>
          </string-name>
          , D. de las Casas,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bressand</surname>
          </string-name>
          , G. Lengyel,
          <string-name>
            <given-names>G.</given-names>
            <surname>Lample</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Saulnier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. R.</given-names>
            <surname>Lavaud</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>A. Lachaux</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Stock</surname>
            ,
            <given-names>T. L.</given-names>
          </string-name>
          <string-name>
            <surname>Scao</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lavril</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lacroix</surname>
            ,
            <given-names>W. E.</given-names>
          </string-name>
          <string-name>
            <surname>Sayed</surname>
          </string-name>
          , Mistral 7b,
          <year>2023</year>
          . arXiv:
          <volume>2310</volume>
          .
          <fpage>06825</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>