<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>B E E P - BEst DrivEr's License Performer:</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fabio Mercorio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniele Potertì</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Serino</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Seveso</string-name>
          <email>andrea.seveso@unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Large Language Models, Benchmarks, CALAMITA, CLiC-it</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CLiC-it 2024: Tenth Italian Conference on Computational Linguistics</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>CRISP Research Centre crispresearch.eu, University of Milano Bicocca</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Dept of Economics, Management and Statistics, University of Milano Bicocca</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Dept of Statistics and Quantitative Methods, University of Milano Bicocca</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>launched by AILC, the Italian Association for Computa-</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present BEEP (BEst DrivEr's License Performer), a benchmark challenge to evaluate large language models in the context of a simulated Italian driver's license exam. This challenge tests the models' ability to understand and apply trafic laws, road safety regulations, and vehicle-related knowledge through a series of true/false questions. The dataset is derived from oficial ministerial materials used in the Italian licensing process, specifically targeting Category B licenses. We evaluate models such as LLaMA and Mixtral across multiple categories. In addition, we simulate a driving license test to assess the models' real-world applicability, where the pass rate is determined based on the number of errors allowed. While scaling up model size improved performance, even larger models struggled to pass the exam consistently. The challenge demonstrates the capabilities and limitations of LLMs in handling real-world, high-stakes scenarios, providing insights into their practical use and areas for further improvement.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR</p>
      <p>ceur-ws.org
A</p>
      <p>CALAMITA</p>
    </sec>
    <sec id="sec-2">
      <title>1. Challenge: Introduction and</title>
    </sec>
    <sec id="sec-3">
      <title>Motivation</title>
      <p>
        come a significant breakthrough in Natural Language
Processing (NLP) and Artificial Intelligence (AI) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Assessing model performance is crucial yet challenging,
involving multiple critical attributes: models must be
precise, resilient, fair, and eficient, among other
characteristics [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Developing efective models in underrepresented
languages such as Italian is a continuing challenge [3]. This
disparity arises from limited and lower-quality data [4]
and a development process often prioritising
Anglocentric perspectives [
        <xref ref-type="bibr" rid="ref3">5</xref>
        ]. Recently, there has been a surge
clusive, moving beyond mere multilingualism to address
deeper cultural contexts [6]. For instance, a structured
benchmark utilising the INVALSI tests—well-established
assessments measuring educational competencies across
Italy—represents one such efort to embed culturally
relevant content in model evaluation [
        <xref ref-type="bibr" rid="ref4">7</xref>
        ].
      </p>
      <p>This work is part of CALAMITA [8] (Challenge the
Abilities of LAnguage Models in ITAlian), an initiative
tional Linguistics. CALAMITA aims to develop a
comprehensive and evolving benchmark for evaluating the
shared platform with a suite of tasks and a live
leaderboard, allowing for ongoing assessments of Italian and
multilingual LLMs. CALAMITA seeks to build this
benchmark through community-driven challenges, inviting
researchers to propose tasks and datasets that evaluate
specific aspects of LLMs’ performance in Italian. This
paper contributes to this collaborative efort by presenting
a benchmark that assesses LLMs’ ability to comprehend
and apply Italian driving regulations, forming one of the
initial tasks in this evolving benchmark.</p>
      <p>This challenge evaluates LLM’s ability to comprehend
While LLMs have shown remarkable capabilities in
understanding and generating human language, their
effectiveness in real-world decision-making scenarios
remains underexplored, especially in languages such as
Italian. This challenge tests whether these models can
perform efectively in a linguistically demanding and
contextually rich domain. Success in this challenge would
demonstrate the model’s ability to generalise language
understanding to practical tasks, a crucial step towards
their broader application in everyday life.</p>
    </sec>
    <sec id="sec-4">
      <title>2. Challenge: Description</title>
      <p>BEst DrivEr’s License Performer (BEEP) is a challenge
benchmark that focuses on assessing LLMs through a
The dataset is formatted with the following columns:
• Categorisation Structure - Each question in
the dataset is organised within a hierarchical
categorisation system consisting of Major
Categories, Minor Categories, and Subcategories
to ensure precise classification. For example, the
Major Category ”Road Signage” includes Minor
Categories like ”Warning Signs” and ”Prohibition
Signs”, which further break down into
Subcategories detailing specific signs such as ”Speed Limit
Signs”;
• Question Text - The actual content of the
question;
• True Answer - Can be either true or false;
• Figure - A reference for the accompanying figure,
if present.</p>
      <sec id="sec-4-1">
        <title>3.3. Example of prompts used</title>
        <sec id="sec-4-1-1">
          <title>The road can be divided into lanes.</title>
          <p>simulated driver’s license exam in Italian. This task re- 3.2. Data format
quires a deep understanding of trafic laws and reasoning
through driving situations.</p>
          <p>In Italy, obtaining a driver’s license is a structured
process involving theoretical and practical assessments to
ensure drivers are well-versed in road safety, trafic
regulations, and practical driving skills. The Italian driver’s
license process is governed by strict rules set forth by
the Ministero delle Infrastrutture e dei Trasporti
(Ministry of Infrastructure and Transport), and the license is
recognised across the European Union.</p>
          <p>Italy ofers several categories of driver’s licenses,
depending on the type of vehicle a person wishes to operate.</p>
          <p>We focus on Category B, which is required for cars (up
to 3.5 tons) and vehicles with up to 8 seats.</p>
          <p>The theoretical exam is crucial to obtaining a driver’s
license in Italy, and it is required, along with the practical
exam. It assesses the applicant’s knowledge of trafic
laws, road signs, and driving regulations. It consists of
multiple-choice questions and is typically administered
electronically. The candidate must understand trafic
regulations, road signs, driving behaviour, and vehicle
maintenance. A Category B license test typically consists
of 30 questions; a candidate can pass up to 3 errors. Question</p>
          <p>The licensing process is not just about learning the
rules; it requires candidates to internalise and apply them
practically. BEEP reflects this focus on real-world
application and safety. The Italian driving system also empha- Options
sises road etiquette and the ability to navigate complex
trafic situations, particularly in high-density urban
areas. Consequently, the challenge aims to mirror this
complexity in evaluating LLMs.</p>
          <p>Options
[ A. True, B. False ]</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>3. Data description</title>
      <sec id="sec-5-1">
        <title>3.1. Origin of data</title>
        <p>Instructions:</p>
        <sec id="sec-5-1-1">
          <title>You must return the letter corresponding to the correct answer in square brackets.</title>
          <p>Answer format: [letter]
BEEP is derived from the publicly accessible PDF ”Listato
A e B”, which includes all quiz questions related to Italian Answer
driver’s license examinations provided by the oficial
ministerial listing1. The quizzes consist of true or false [ A ]
questions for driving license categories A and B, with
data updated as of 01/07/2020. Figure 1: An example question, with instructions and a
cor</p>
          <p>We extracted the data from the oficial PDF file. The rect answer highlighted.
text is segmented by identifying distinct patterns
indicating the start of new questions and sections. These
segments are classified into predefined categories and We exclusively employed the zero-shot setting in our
sub-categories. For each text segment, relevant metadata, evaluation process, where no prior examples were
proquestion types (e.g., true/false) and related image num- vided. An illustrative example of a prompt used in this
bers are extracted and compiled into a structured format. setting is shown in Figure 1, which demonstrates the
The final dataset is exported, ofering a well-organised structure and input format supplied to the model. The
collection of questions for the evaluation. decision to have the language model answer with
’[letter]’ rather than simply ’letter’ or ’True/False’ is due to
1Visit ListatoAB for more information at https://www.neca.it/assets/ our use of pattern matching for response extraction. By
pdf/ListatoAB.pdf. enforcing a consistent answer format with brackets, we</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>VEHICLE EQUIPMENT</title>
        </sec>
        <sec id="sec-5-1-3">
          <title>VISUAL SIGNAL DEVICES AND LIGHTING</title>
        </sec>
        <sec id="sec-5-1-4">
          <title>ACCIDENTS AND INSURANCE</title>
        </sec>
        <sec id="sec-5-1-5">
          <title>CAUSES OF ACCIDENTS</title>
        </sec>
        <sec id="sec-5-1-6">
          <title>ROAD</title>
        </sec>
        <sec id="sec-5-1-7">
          <title>ROAD AND TRAFFIC DEFINITIONS</title>
        </sec>
        <sec id="sec-5-1-8">
          <title>TRAFFIC REGULATIONS</title>
        </sec>
        <sec id="sec-5-1-9">
          <title>STOPPING AND SAFE DISTANCE</title>
        </sec>
        <sec id="sec-5-1-10">
          <title>MANDATORY DOCUMENTS, AGENTS AND LI</title>
        </sec>
        <sec id="sec-5-1-11">
          <title>CENSE PLATES</title>
        </sec>
        <sec id="sec-5-1-12">
          <title>STATIONARY VEHICLE SIGNALS AND ROAD OB</title>
        </sec>
        <sec id="sec-5-1-13">
          <title>STRUCTIONS</title>
        </sec>
        <sec id="sec-5-1-14">
          <title>CLASSIFICATION OF VEHICLES</title>
        </sec>
        <sec id="sec-5-1-15">
          <title>VEHICLE COMPONENTS</title>
        </sec>
        <sec id="sec-5-1-16">
          <title>TIRES, ADHERENCE AND STABILITY</title>
        </sec>
        <sec id="sec-5-1-17">
          <title>WARNING LIGHTS AND SYMBOLS</title>
        </sec>
        <sec id="sec-5-1-18">
          <title>CIVIL AND CRIMINAL LIABILITY AND INSURANCE</title>
        </sec>
        <sec id="sec-5-1-19">
          <title>STOP, STANDING AND PARKING</title>
        </sec>
        <sec id="sec-5-1-20">
          <title>DRIVING ON HIGHWAYS</title>
        </sec>
        <sec id="sec-5-1-21">
          <title>SPEED LIMITS</title>
        </sec>
        <sec id="sec-5-1-22">
          <title>RIGHT-OF-WAY RULES AND PROCESSIONS</title>
        </sec>
        <sec id="sec-5-1-23">
          <title>POSITION ON ROADWAY, DIRECTION CHANGE</title>
        </sec>
        <sec id="sec-5-1-24">
          <title>AND LANE</title>
        </sec>
        <sec id="sec-5-1-25">
          <title>SPEED REGULATION</title>
        </sec>
        <sec id="sec-5-1-26">
          <title>OVERTAKING</title>
        </sec>
        <sec id="sec-5-1-27">
          <title>TRANSPORT OF PEOPLE, LOAD ARRANGEMENT,</title>
        </sec>
        <sec id="sec-5-1-28">
          <title>PANELS AND TOWING</title>
        </sec>
        <sec id="sec-5-1-29">
          <title>FIRST AID TO INJURED PEOPLE</title>
        </sec>
        <sec id="sec-5-1-30">
          <title>SUPPLEMENTARY PANELS</title>
        </sec>
        <sec id="sec-5-1-31">
          <title>TRAFFIC LIGHT SIGNALS AND POLICEMAN</title>
        </sec>
        <sec id="sec-5-1-32">
          <title>PROHIBITION SIGNS</title>
        </sec>
        <sec id="sec-5-1-33">
          <title>INFORMATION SIGNS</title>
        </sec>
        <sec id="sec-5-1-34">
          <title>MANDATORY SIGNS</title>
        </sec>
        <sec id="sec-5-1-35">
          <title>WARNING SIGNS</title>
        </sec>
        <sec id="sec-5-1-36">
          <title>PRIORITY SIGNS</title>
        </sec>
        <sec id="sec-5-1-37">
          <title>ROAD MARKINGS</title>
          <p>Major</p>
        </sec>
        <sec id="sec-5-1-38">
          <title>DOCUMENTS</title>
        </sec>
        <sec id="sec-5-1-39">
          <title>VEHICLES</title>
        </sec>
        <sec id="sec-5-1-40">
          <title>MOTOR VEHICLE</title>
        </sec>
        <sec id="sec-5-1-41">
          <title>FIRST AID</title>
        </sec>
        <sec id="sec-5-1-42">
          <title>TRAFFIC SIGNS</title>
        </sec>
        <sec id="sec-5-1-43">
          <title>SAFETY AND POLLUTION</title>
        </sec>
        <sec id="sec-5-1-44">
          <title>SEAT BELTS, AIRBAG AND PROTECTIVE HELMET</title>
        </sec>
        <sec id="sec-5-1-45">
          <title>TEMPORARY AND SUPPLEMENTARY SIGNS ENVIRONMENTAL AND NOISE POLLUTION</title>
          <p>can reliably parse responses, reducing ambiguity and en- accuracy is commonly used in classification tasks,
particsuring that variations in phrasing or formatting do not ularly in true-false or binary decision evaluations [9]. It
interfere with accurate evaluation. measures the proportion of all correct predictions (true
positives and negatives) out of the total number of
pre3.4. Detailed data statistics dictions made. In other words, it quantifies how well a
binary classification system performs by indicating the
The questions are organised into the categories described fraction of correctly classified instances (both positive
in Tab. 1. This table summarises statistics across various and negative classes) relative to the total number of
inroad safety and vehicle regulation categories, provid- stances evaluated.
ing detailed insight into major and minor classifications.</p>
          <p>Each entry in the table is categorised into broad Major Table 3
Categories such as ”DOCUMENTS,” ”Vehicle Equipment,” Overall accuracy of selected models, ranging from LLaMA to
and ”Road Signage,” which are further subdivided into Mixtral, demonstrating their performance on the dataset.
more specific Minor Categories. For example, the major
category ”DOCUMENTS” includes the minor category Model Overall Accuracy
”Mandatory Documents, Agents, and License Plates,” llama-3-8b-instruct 56.27%
ahnigdhalidgmhtiinnigstdraifetriveentdaestpaielcst.s of document requirements lmmlaiimxxttarraa-3ll---887xx027b2b-ib-ni-nsintsrtsurturcutctct 787737...122993%%%</p>
          <p>We also include figures associated with specific
questions, particularly those addressing trafic signals, road
signs, and right-of-way scenarios. These visual elements Table 3 shows the Overall Accuracy obtained by
provide additional context and enhance the comprehen- LLAMA3 8B - Instruct2 and others State of the Art models.
sion of complex trafic situations. However, for the We evaluate the metrics on the portion of our dataset that
CALAMITA challenge, we opted not to include ques- does not require image processing operations. The
scaltions containing figures, focusing solely on text-based ing laws hold as it is observed that performance increases
questions. This decision ensured that the evaluation of with the number of parameters.</p>
          <p>LLMs remains centred on their language comprehension, Table 2 shows the Overall Accuracy stratified by Major
knowledge and reasoning abilities rather than visual pro- Category for each tested model. Models perform better in
cessing capabilities. Including images would limit partic- the ”SAFETY AND POLLUTION”, ”FIRST AID”, and
”ACipation to multimodal models, excluding many language CIDENTS AND INSURANCE” categories. This may be
models that cannot process visual information. By us- possible given the generality of these major categories, as
ing only text, we maintain a broader, more accessible opposed to more niche categories such as ‘DOCUMENTS’
benchmark. or ‘VEHICLE EQUIPMENT’, where the performance is
worse.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>4. Metrics</title>
      <sec id="sec-6-1">
        <title>4.1. Simulated Driving License Test</title>
        <p>Since the dataset comprises questions that can only be an- We also test the models by simulating a proper driving
swered with true and false, we involved the Overall Accu- licence exam, following the appropriate oficial guidelines
racy to evaluate the models’ answers in our task. Overall
2https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct
and creating a new indicator. We sampled 1000 samples
of 30 questions from the dataset, ensuring each sample
was unique. We then counted the correct and incorrect
answers for each sample and each evaluated model. The
guidelines state that the test is passed if the number of
wrong answers is less than or equal to 3. Therefore, we
built an indicator for each model that considered the
percentage of driving licence exams passed, related to
the number of examinations attempted. The results are
shown in Tab. 4. As expected, smaller models made many
mistakes on average (around 13), which was fatal as it
never passed the test in any of the attempts. Even larger
models like Mixtral-8x22b did not perform well in most
cases. However, we believe more advanced models, such
as GPT-4, might succeed more reliably.
The data are publicly available online and not subject to
copyright restrictions.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We thank Thomas Passera for providing the initial code
for the dataset’s extraction. Evaluation of the
opensource models was conducted on Leonardo
supercomputer with the support of CINECA-Italian Super
Computing Resource Allocation, class C project
IsCb7_LLMEVAL (HP10CIO7T9).
[8] G. Attanasio, P. Basile, F. Borazio, D. Croce, M.
Francis, J. Gili, E. Musacchio, M. Nissim, V. Patti, M.
Rinaldi, D. Scalena, CALAMITA: Challenge the
Abilities of LAnguage Models in ITAlian, in: Proceedings
of the 10th Italian Conference on Computational
Linguistics (CLiC-it 2024), Pisa, Italy, December 4
- December 6, 2024, CEUR Workshop Proceedings,
CEUR-WS.org, 2024.
[9] C. M. Bishop, Pattern recognition and machine
learning, Springer google schola 2 (2006) 1122–1128.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Model Total Tests Passed (%) Avg Errors (Std.) H. Chen</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Yi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Ye</surname>
          </string-name>
          , Y. Zhang, llllaammaa--33--
          <source>780bb-i-ninstsrtuructct 640//11000000 ((60.%4%)) 163.8.187((±±22.6</source>
          .751)
          <string-name>
            <surname>) Y. Chang</surname>
            ,
            <given-names>P. S.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Xie</surname>
          </string-name>
          ,
          <source>A survey on mixtral-8x7b-instruct 61/1000 (6.1%) 6.79 (±2.24) evaluation of large language models</source>
          ,
          <year>2023</year>
          . URL: http:
          <fpage>mixtral</fpage>
          -8x22b
          <source>-instruct 258/1000 (25.8%) 5.01 (±2</source>
          .09) //arxiv.org/abs/2307.03109. arXiv:
          <volume>2307</volume>
          .
          <fpage>03109</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bommasani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tsipras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Soylu</surname>
          </string-name>
          ,
          <article-title>It is important to note that this simulated test is not M. Yasunaga</article-title>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Kuintegral to the CALAMITA benchmark</article-title>
          .
          <article-title>While it provides mar</article-title>
          , et al.,
          <article-title>Holistic evaluation of language models, additional insights into the models' performance in a arXiv preprint</article-title>
          arXiv:
          <volume>2211</volume>
          .09110 (
          <year>2022</year>
          ).
          <article-title>high-stakes, applied setting, the oficial evaluation metric [3</article-title>
          ]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Constant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Botha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Siddhant</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.</surname>
          </string-name>
          <article-title>Fifocuses solely on overall accuracy</article-title>
          . rat, J. Fu, P. Liu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Garrette</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Neubig</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Xtreme-</surname>
          </string-name>
          r:
          <article-title>Towards more challenging and nuanced 5. Limitations multilingual evaluation, in: Proceedings of the 2021 Conference on Empirical Methods in Natural LanConsidering state-of-the-art LLMs, it is possible that guage Processing, Association for Computational one's training sets are contaminated with examples from Linguistics, 2021. the U.S. driving licence test and that these may influence [4</article-title>
          ]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kreutzer</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Caswell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wahab</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>van performance on our benchmark</article-title>
          . Furthermore, although Esch,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ulzii-Orshikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tapo</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Subramani, the benchmark allows the real driving licence test to be A</article-title>
          .
          <string-name>
            <surname>Sokolov</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Sikasote</surname>
          </string-name>
          , et al.,
          <article-title>Quality at a glance: An reproduced, it can only assess true-or-false binary an- audit of web-crawled multilingual datasets, Transacswers and not dialogue or reasoning ability</article-title>
          .
          <source>tions of the Association for Computational Linguistics</source>
          <volume>10</volume>
          (
          <year>2022</year>
          )
          <fpage>50</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Talat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Névéol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Biderman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Clinciu</surname>
          </string-name>
          , M. Dey,
          <volume>6</volume>
          . Ethical issues
          <string-name>
            <given-names>S.</given-names>
            <surname>Longpre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Luccioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Masoud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mitchell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Radev</surname>
          </string-name>
          , et al.,
          <article-title>You reap what you sow: On the chalAlthough the models may demonstrate positive perfor- lenges of bias evaluation under multilingual settings, mance in this benchmark, it is crucial to recognise that in: Proceedings of BigScience Episode# 5-Workshop such results do not equate to an actual ability to drive or on Challenges &amp; Perspectives in Creating Large Lannavigate safely in real-world environments</article-title>
          .
          <source>The bench- guage Models</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>26</fpage>
          -
          <lpage>41</lpage>
          .
          <article-title>mark assesses the models' ability to process</article-title>
          and under- [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Pawar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Myung</surname>
          </string-name>
          , S. Yadav,
          <article-title>stand driving-related questions, a far cry from the com-</article-title>
          F. G. Haznitrama,
          <string-name>
            <surname>I. Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oh</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Augenstein</surname>
          </string-name>
          ,
          <article-title>Surplex task of driving a vehicle, which requires perception, vey of cultural awareness in language models: Text decision-making and real-time motor control. and beyond (</article-title>
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Mercorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mezzanzanica</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Potertì</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Serino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Seveso</surname>
          </string-name>
          ,
          <article-title>Disce aut deficere: Evaluating llms proifciency on the invalsi italian benchmark</article-title>
          ,
          <source>arXiv preprint arXiv:2406.17535</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>