BEEP - BEst DrivEr’s License Performer: A CALAMITA Challenge Fabio Mercorio1,3 , Daniele Potertì2 , Antonio Serino2 and Andrea Seveso1,3,∗ 1 Dept of Statistics and Quantitative Methods, University of Milano Bicocca, Italy 2 Dept of Economics, Management and Statistics, University of Milano Bicocca, Italy 3 CRISP Research Centre crispresearch.eu, University of Milano Bicocca, Italy Abstract We present BEEP (BEst DrivEr’s License Performer), a benchmark challenge to evaluate large language models in the context of a simulated Italian driver’s license exam. This challenge tests the models’ ability to understand and apply traffic laws, road safety regulations, and vehicle-related knowledge through a series of true/false questions. The dataset is derived from official ministerial materials used in the Italian licensing process, specifically targeting Category B licenses. We evaluate models such as LLaMA and Mixtral across multiple categories. In addition, we simulate a driving license test to assess the models’ real-world applicability, where the pass rate is determined based on the number of errors allowed. While scaling up model size improved performance, even larger models struggled to pass the exam consistently. The challenge demonstrates the capabilities and limitations of LLMs in handling real-world, high-stakes scenarios, providing insights into their practical use and areas for further improvement. Keywords Large Language Models, Benchmarks, CALAMITA, CLiC-it 1. Challenge: Introduction and launched by AILC, the Italian Association for Computa- tional Linguistics. CALAMITA aims to develop a com- Motivation prehensive and evolving benchmark for evaluating the In recent years, Large Language Models (LLMs) have be- capabilities of LLMs in Italian. The goal is to establish a come a significant breakthrough in Natural Language shared platform with a suite of tasks and a live leader- Processing (NLP) and Artificial Intelligence (AI) [1]. As- board, allowing for ongoing assessments of Italian and sessing model performance is crucial yet challenging, multilingual LLMs. CALAMITA seeks to build this bench- involving multiple critical attributes: models must be mark through community-driven challenges, inviting precise, resilient, fair, and efficient, among other charac- researchers to propose tasks and datasets that evaluate teristics [2]. specific aspects of LLMs’ performance in Italian. This pa- Developing effective models in underrepresented lan- per contributes to this collaborative effort by presenting guages such as Italian is a continuing challenge [3]. This a benchmark that assesses LLMs’ ability to comprehend disparity arises from limited and lower-quality data [4] and apply Italian driving regulations, forming one of the and a development process often prioritising Anglo- initial tasks in this evolving benchmark. centric perspectives [5]. Recently, there has been a surge This challenge evaluates LLM’s ability to comprehend in research aimed at making LLMs more culturally in- and apply knowledge in a practical, real-world scenario. clusive, moving beyond mere multilingualism to address While LLMs have shown remarkable capabilities in un- deeper cultural contexts [6]. For instance, a structured derstanding and generating human language, their ef- benchmark utilising the INVALSI tests—well-established fectiveness in real-world decision-making scenarios re- assessments measuring educational competencies across mains underexplored, especially in languages such as Italy—represents one such effort to embed culturally rel- Italian. This challenge tests whether these models can evant content in model evaluation [7]. perform effectively in a linguistically demanding and con- This work is part of CALAMITA [8] (Challenge the textually rich domain. Success in this challenge would Abilities of LAnguage Models in ITAlian), an initiative demonstrate the model’s ability to generalise language understanding to practical tasks, a crucial step towards CLiC-it 2024: Tenth Italian Conference on Computational Linguistics, their broader application in everyday life. Dec 04 — 06, 2024, Pisa, Italy ∗ Corresponding author. Envelope-Open andrea.seveso@unimib.it (A. Seveso) 2. Challenge: Description Orcid 0000-0001-6864-2702 (F. Mercorio); 0009-0006-6525-4492 (D. Potertì); 0009-0008-0737-8547 (A. Serino); 0000-0001-7132-7703 BEst DrivEr’s License Performer (BEEP ) is a challenge (A. Seveso) benchmark that focuses on assessing LLMs through a © 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). CEUR ceur-ws.org Workshop ISSN 1613-0073 Proceedings simulated driver’s license exam in Italian. This task re- 3.2. Data format quires a deep understanding of traffic laws and reasoning The dataset is formatted with the following columns: through driving situations. In Italy, obtaining a driver’s license is a structured pro- • Categorisation Structure - Each question in cess involving theoretical and practical assessments to the dataset is organised within a hierarchical cat- ensure drivers are well-versed in road safety, traffic regu- egorisation system consisting of Major Cate- lations, and practical driving skills. The Italian driver’s gories, Minor Categories, and Subcategories license process is governed by strict rules set forth by to ensure precise classification. For example, the the Ministero delle Infrastrutture e dei Trasporti (Min- Major Category ”Road Signage” includes Minor istry of Infrastructure and Transport), and the license is Categories like ”Warning Signs” and ”Prohibition recognised across the European Union. Signs”, which further break down into Subcate- Italy offers several categories of driver’s licenses, de- gories detailing specific signs such as ”Speed Limit pending on the type of vehicle a person wishes to operate. Signs”; We focus on Category B, which is required for cars (up • Question Text - The actual content of the ques- to 3.5 tons) and vehicles with up to 8 seats. tion; The theoretical exam is crucial to obtaining a driver’s license in Italy, and it is required, along with the practical • True Answer - Can be either true or false; exam. It assesses the applicant’s knowledge of traffic • Figure - A reference for the accompanying figure, laws, road signs, and driving regulations. It consists of if present. multiple-choice questions and is typically administered electronically. The candidate must understand traffic 3.3. Example of prompts used regulations, road signs, driving behaviour, and vehicle maintenance. A Category B license test typically consists of 30 questions; a candidate can pass up to 3 errors. Question The licensing process is not just about learning the The road can be divided into lanes. rules; it requires candidates to internalise and apply them practically. BEEP reflects this focus on real-world appli- cation and safety. The Italian driving system also empha- Options sises road etiquette and the ability to navigate complex traffic situations, particularly in high-density urban ar- [ A. True, B. False ] eas. Consequently, the challenge aims to mirror this complexity in evaluating LLMs. Options Instructions: 3. Data description You must return the letter corresponding to the cor- rect answer in square brackets. 3.1. Origin of data Answer format: [ letter] BEEP is derived from the publicly accessible PDF ”Listato A e B”, which includes all quiz questions related to Italian Answer driver’s license examinations provided by the official ministerial listing1 . The quizzes consist of true or false [A] questions for driving license categories A and B, with data updated as of 01/07/2020. Figure 1: An example question, with instructions and a cor- We extracted the data from the official PDF file. The rect answer highlighted. text is segmented by identifying distinct patterns indi- cating the start of new questions and sections. These segments are classified into predefined categories and We exclusively employed the zero-shot setting in our sub-categories. For each text segment, relevant metadata, evaluation process, where no prior examples were pro- question types (e.g., true/false) and related image num- vided. An illustrative example of a prompt used in this bers are extracted and compiled into a structured format. setting is shown in Figure 1, which demonstrates the The final dataset is exported, offering a well-organised structure and input format supplied to the model. The collection of questions for the evaluation. decision to have the language model answer with ’[let- ter]’ rather than simply ’letter’ or ’True/False’ is due to 1 Visit ListatoAB for more information at https://www.neca.it/assets/ our use of pattern matching for response extraction. By pdf/ListatoAB.pdf. enforcing a consistent answer format with brackets, we Table 1 An overview of the dataset categorised by major and minor traffic-related topics. The columns display the number of entries, the percentage of those entries containing figures, and the proportion of correct answers for each category. Category Percent with True Major Minor Rows Figures Answer (%) DOCUMENTS MANDATORY DOCUMENTS, AGENTS AND LI- 261 — 129/261 CENSE PLATES (49.4%) VEHICLE EQUIPMENT VISUAL SIGNAL DEVICES AND LIGHTING 98 — 53/98 (54.1%) STATIONARY VEHICLE SIGNALS AND ROAD OB- 54 — 26/54 STRUCTIONS (48.1%) VEHICLES CLASSIFICATION OF VEHICLES 106 — 48/106 (45.3%) MOTOR VEHICLE VEHICLE COMPONENTS 119 — 63/119 (52.9%) TIRES, ADHERENCE AND STABILITY 134 — 68/134 (50.7%) WARNING LIGHTS AND SYMBOLS 61 100.00 28/61 (45.9%) ACCIDENTS AND INSURANCE CAUSES OF ACCIDENTS 566 — 303/566 (53.5%) CIVIL AND CRIMINAL LIABILITY AND INSURANCE 123 — 53/123 (43.1%) ROAD ROAD AND TRAFFIC DEFINITIONS 203 — 102/203 (50.2%) TRAFFIC REGULATIONS STOPPING AND SAFE DISTANCE 129 — 62/129 (48.1%) STOP, STANDING AND PARKING 208 — 121/208 (58.2%) DRIVING ON HIGHWAYS 59 — 31/59 (52.5%) SPEED LIMITS 81 — 45/81 (55.6%) RIGHT-OF-WAY RULES AND PROCESSIONS 457 86.87 235/457 (51.4%) POSITION ON ROADWAY, DIRECTION CHANGE 27 70.37 13/27 AND LANE (48.1%) SPEED REGULATION 96 — 56/96 (58.3%) OVERTAKING 156 — 82/156 (52.6%) TRANSPORT OF PEOPLE, LOAD ARRANGEMENT, 110 — 55/110 PANELS AND TOWING (50.0%) FIRST AID FIRST AID TO INJURED PEOPLE 96 — 48/96 (50.0%) TRAFFIC SIGNS SUPPLEMENTARY PANELS 59 100.00 27/59 (45.8%) TRAFFIC LIGHT SIGNALS AND POLICEMAN 218 96.33 105/218 (48.2%) PROHIBITION SIGNS 409 100.00 198/409 (48.4%) INFORMATION SIGNS 536 100.00 253/536 (47.2%) MANDATORY SIGNS 402 100.00 190/402 (47.3%) WARNING SIGNS 473 100.00 228/473 (48.2%) PRIORITY SIGNS 201 100.00 99/201 (49.3%) ROAD MARKINGS 147 100.00 73/147 (49.7%) TEMPORARY AND SUPPLEMENTARY SIGNS 189 100.00 89/189 (47.1%) SAFETY AND POLLUTION SEAT BELTS, AIRBAG AND PROTECTIVE HELMET 135 — 70/135 (51.9%) ENVIRONMENTAL AND NOISE POLLUTION 110 — 64/110 (58.2%) Table 2 Overall accuracy of different models across major dataset categories, allowing for comparison of their effectiveness within these distinct areas. Category llama-3-8b llama-3-70b mixtral-8x7b mixtral-8x22b DOCUMENTS 53.26% 66.28% 67.43% 79.69% VEHICLE EQUIPMENT 51.97% 66.45% 71.71% 75.00% VEHICLES 51.89% 77.36% 82.08% 84.91% THE MOTOR VEHICLE 56.13% 82.61% 82.21% 86.56% ACCIDENTS AND INSURANCE 59.22% 85.78% 85.49% 91.15% THE ROAD 51.72% 70.94% 71.92% 81.77% RULES OF CONDUCT 54.36% 71.11% 70.34% 76.85% FIRST AID 61.46% 90.62% 86.46% 88.54% ROAD SIGNAGE 37.50% 75.00% 100.00% 100.00% SAFETY AND POLLUTION 65.31% 88.57% 85.71% 88.57% can reliably parse responses, reducing ambiguity and en- accuracy is commonly used in classification tasks, partic- suring that variations in phrasing or formatting do not ularly in true-false or binary decision evaluations [9]. It interfere with accurate evaluation. measures the proportion of all correct predictions (true positives and negatives) out of the total number of pre- 3.4. Detailed data statistics dictions made. In other words, it quantifies how well a binary classification system performs by indicating the The questions are organised into the categories described fraction of correctly classified instances (both positive in Tab. 1. This table summarises statistics across various and negative classes) relative to the total number of in- road safety and vehicle regulation categories, provid- stances evaluated. ing detailed insight into major and minor classifications. Each entry in the table is categorised into broad Major Table 3 Categories such as ”DOCUMENTS,” ”Vehicle Equipment,” Overall accuracy of selected models, ranging from LLaMA to and ”Road Signage,” which are further subdivided into Mixtral, demonstrating their performance on the dataset. more specific Minor Categories. For example, the major Model Overall Accuracy category ”DOCUMENTS” includes the minor category ”Mandatory Documents, Agents, and License Plates,” llama-3-8b-instruct 56.27% llama-3-70b-instruct 77.23% highlighting different aspects of document requirements mixtral-8x7b-instruct 77.19% and administrative details. mixtral-8x22b-instruct 83.29% We also include figures associated with specific ques- tions, particularly those addressing traffic signals, road signs, and right-of-way scenarios. These visual elements Table 3 shows the Overall Accuracy obtained by 2 provide additional context and enhance the comprehen- LLAMA3 8B - Instruct and others State of the Art models. sion of complex traffic situations. However, for the We evaluate the metrics on the portion of our dataset that CALAMITA challenge, we opted not to include ques- does not require image processing operations. The scal- tions containing figures, focusing solely on text-based ing laws hold as it is observed that performance increases questions. This decision ensured that the evaluation of with the number of parameters. LLMs remains centred on their language comprehension, Table 2 shows the Overall Accuracy stratified by Major knowledge and reasoning abilities rather than visual pro- Category for each tested model. Models perform better in cessing capabilities. Including images would limit partic- the ”SAFETY AND POLLUTION”, ”FIRST AID”, and ”AC- ipation to multimodal models, excluding many language CIDENTS AND INSURANCE” categories. This may be models that cannot process visual information. By us- possible given the generality of these major categories, as ing only text, we maintain a broader, more accessible opposed to more niche categories such as ‘DOCUMENTS’ benchmark. or ‘VEHICLE EQUIPMENT’, where the performance is worse. 4. Metrics 4.1. Simulated Driving License Test Since the dataset comprises questions that can only be an- We also test the models by simulating a proper driving swered with true and false, we involved the Overall Accu- licence exam, following the appropriate official guidelines racy to evaluate the models’ answers in our task. Overall 2 https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct and creating a new indicator. We sampled 1000 samples 7. Data license and copyright of 30 questions from the dataset, ensuring each sample was unique. We then counted the correct and incorrect issues answers for each sample and each evaluated model. The The data are publicly available online and not subject to guidelines state that the test is passed if the number of copyright restrictions. wrong answers is less than or equal to 3. Therefore, we built an indicator for each model that considered the percentage of driving licence exams passed, related to Acknowledgments the number of examinations attempted. The results are shown in Tab. 4. As expected, smaller models made many We thank Thomas Passera for providing the initial code mistakes on average (around 13), which was fatal as it for the dataset’s extraction. Evaluation of the open- never passed the test in any of the attempts. Even larger source models was conducted on Leonardo supercom- models like Mixtral-8x22b did not perform well in most puter with the support of CINECA-Italian Super Com- cases. However, we believe more advanced models, such puting Resource Allocation, class C project IsCb7_LLM- as GPT-4, might succeed more reliably. EVAL (HP10CIO7T9). Table 4 References Driving license Metrics of the Selected Models [1] Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, Model Total Tests Passed (%) Avg Errors (Std.) H. Chen, X. Yi, C. Wang, Y. Wang, W. Ye, Y. Zhang, llama-3-8b-instruct 0/1000 (0%) 13.17 (±2.71) llama-3-70b-instruct 64/1000 (6.4%) 6.88 (±2.65) Y. Chang, P. S. Yu, Q. Yang, X. Xie, A survey on mixtral-8x7b-instruct 61/1000 (6.1%) 6.79 (±2.24) evaluation of large language models, 2023. URL: http: mixtral-8x22b-instruct 258/1000 (25.8%) 5.01 (±2.09) //arxiv.org/abs/2307.03109. arXiv:2307.03109 . [2] P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, It is important to note that this simulated test is not M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Ku- integral to the CALAMITA benchmark. While it provides mar, et al., Holistic evaluation of language models, additional insights into the models’ performance in a arXiv preprint arXiv:2211.09110 (2022). high-stakes, applied setting, the official evaluation metric [3] S. Ruder, N. Constant, J. Botha, A. Siddhant, O. Fi- focuses solely on overall accuracy. rat, J. Fu, P. Liu, J. Hu, D. Garrette, G. Neubig, et al., Xtreme-r: Towards more challenging and nuanced multilingual evaluation, in: Proceedings of the 2021 5. Limitations Conference on Empirical Methods in Natural Lan- Considering state-of-the-art LLMs, it is possible that guage Processing, Association for Computational one’s training sets are contaminated with examples from Linguistics, 2021. the U.S. driving licence test and that these may influence [4] J. Kreutzer, I. Caswell, L. Wang, A. Wahab, D. van performance on our benchmark. Furthermore, although Esch, N. Ulzii-Orshikh, A. Tapo, N. Subramani, the benchmark allows the real driving licence test to be A. Sokolov, C. Sikasote, et al., Quality at a glance: An reproduced, it can only assess true-or-false binary an- audit of web-crawled multilingual datasets, Transac- swers and not dialogue or reasoning ability. tions of the Association for Computational Linguis- tics 10 (2022) 50–72. [5] Z. Talat, A. Névéol, S. Biderman, M. Clinciu, M. Dey, 6. Ethical issues S. Longpre, S. Luccioni, M. Masoud, M. Mitchell, D. Radev, et al., You reap what you sow: On the chal- Although the models may demonstrate positive perfor- lenges of bias evaluation under multilingual settings, mance in this benchmark, it is crucial to recognise that in: Proceedings of BigScience Episode# 5–Workshop such results do not equate to an actual ability to drive or on Challenges & Perspectives in Creating Large Lan- navigate safely in real-world environments. The bench- guage Models, 2022, pp. 26–41. mark assesses the models’ ability to process and under- [6] S. Pawar, J. Park, J. Jin, A. Arora, J. Myung, S. Yadav, stand driving-related questions, a far cry from the com- F. G. Haznitrama, I. Song, A. Oh, I. Augenstein, Sur- plex task of driving a vehicle, which requires perception, vey of cultural awareness in language models: Text decision-making and real-time motor control. and beyond (2024). [7] F. Mercorio, M. Mezzanzanica, D. Potertì, A. Serino, A. Seveso, Disce aut deficere: Evaluating llms pro- ficiency on the invalsi italian benchmark, arXiv preprint arXiv:2406.17535 (2024). [8] G. Attanasio, P. Basile, F. Borazio, D. Croce, M. Fran- cis, J. Gili, E. Musacchio, M. Nissim, V. Patti, M. Ri- naldi, D. Scalena, CALAMITA: Challenge the Abili- ties of LAnguage Models in ITAlian, in: Proceedings of the 10th Italian Conference on Computational Linguistics (CLiC-it 2024), Pisa, Italy, December 4 - December 6, 2024, CEUR Workshop Proceedings, CEUR-WS.org, 2024. [9] C. M. Bishop, Pattern recognition and machine learn- ing, Springer google schola 2 (2006) 1122–1128.