=Paper= {{Paper |id=Vol-3878/127_calamita_long |storemode=property |title=GITA4CALAMITA - Evaluating the Physical Commonsense Understanding of Italian LLMs in a Multi-layered Approach: A CALAMITA Challenge |pdfUrl=https://ceur-ws.org/Vol-3878/127_calamita_long.pdf |volume=Vol-3878 |authors=Giulia Pensa,Ekhi Azurmendi,Julen Etxaniz,Begoña Altuna,Itziar Gonzalez-Dios |dblpUrl=https://dblp.org/rec/conf/clic-it/PensaAEAG24 }} ==GITA4CALAMITA - Evaluating the Physical Commonsense Understanding of Italian LLMs in a Multi-layered Approach: A CALAMITA Challenge== https://ceur-ws.org/Vol-3878/127_calamita_long.pdf
                                GITA4CALAMITA - Evaluating the Physical Commonsense
                                Understanding of Italian LLMs in a Multi-layered Approach:
                                A CALAMITA Challenge
                                Giulia Pensa1 , Ekhi Azurmendi2 , Julen Etxaniz2 , Begoña Altuna2 and Itziar Gonzalez-Dios2
                                1
                                    University of the Basque Country UPV/EHU
                                2
                                    HiTZ Center - Ixa, University of the Basque Country UPV/EHU


                                                  Abstract
                                                  In the context of the CALAMITA Challenge, we investigate the physical commonsense reasoning capabilities of large language
                                                  models (LLMs) and introduce a methodology to assess their understanding of the physical world. To this end, we use a test
                                                  set designed to evaluate physical commonsense reasoning in LLMs for the Italian language. We present a tiered dataset,
                                                  named the Graded Italian Annotated dataset (GITA), which is written and annotated by a professional linguist. This dataset
                                                  enables us to focus on three distinct levels of commonsense understanding. Our benchmark aims to evaluate three specific
                                                  tasks: identifying plausible and implausible stories within our dataset, identifying the conflict that generates an implausible
                                                  story, and identifying the physical states that make a story implausible. We perform these tasks using LLAMA3, Gemma2
                                                  and Mistral. Our findings reveal that, although the models may excel at high-level classification tasks, their reasoning is
                                                  inconsistent and unverifiable, as they fail to capture intermediate evidence.

                                                  Keywords
                                                  Physical commonsense reasoning, large language models, Italian benchmark



                                1. Challenge: Introduction and                                                                             proficiency in reasoning and comprehending language
                                                                                                                                           [5, 6]. As a result, there remains uncertainty regarding
                                   Motivation                                                                                              machines’ ability to truly perform reasoning and whether
                                Physical commonsense understanding refers to the abil-                                                     the existing issues in this regard have been sufficiently
                                ity to comprehend the physical world and the events that                                                   addressed.
                                transpire within it. This capability is a crucial component                                                   In this context, our aim is to contribute to this chal-
                                of human intelligence, enabling us to reason about our                                                     lenge developing an original Italian benchmark that can
                                environment, anticipate future occurrences, and navi-                                                      be used to assess the ability of language models to un-
                                gate our surroundings effortlessly, and recently there has                                                 derstand physical commonsense in a more truthful way,
                                been notable advancement in the development of large                                                       focusing not only on end tasks, but also on intermediate
                                language models (LLMs) that can produce human-like                                                         layer tasks.
                                language and execute a variety of language-related tasks.                                                     In this paper, we present GITA4CALAMITA, the
                                   LLMs have exhibited promising outcomes in grasping                                                      Graded Italian Annotated dataset for the CALAMITA
                                common sense in particular situations [1, 2]. Neverthe-                                                    challenge [7]. GITA4CALAMITA is an adapted version
                                less, it is widely recognized that the most precise evalua-                                                of the GITA dataset proposed in [8]. In particular, we de-
                                tion of their capabilities is attained when assessing their                                                cided to revise the physical states annotation and adapt
                                performance in specific end tasks [3, 4]. The evaluation                                                   it to this challenge. The first version of GITA dataset
                                often emphasizes the capacity of LLMs to replicate rela-                                                   is available in our repository under the license CC BY-
                                tively straightforward tasks, rather than their authentic                                                  NC-SA 4.0.1 . The GITA4CALAMITA dataset is manu-
                                                                                                                                           ally compiled by a professional linguist, which allows
                                CLiC-it 2024: Tenth Italian Conference on Computational Linguistics,
                                                                                                                                           for this multi-layered evaluation of the reasoning pro-
                                Dec 04 - 06, 2024, Pisa, Italy                                                                             cess. With the creation of an Italian dataset we gain
                                ∗
                                     Corresponding author.                                                                                 the linguistic and cultural perspective of Italian, while
                                †
                                    These authors contributed equally.                                                                     commonsense research in Natural Language Processing
                                Envelope-Open giulia.pensa.tr@gmail.com (G. Pensa); ekhi.azurmendi@ehu.eus                                 (NLP) has largely been focused on the English language.
                                (E. Azurmendi); julen.etxaniz@ehu.eus (J. Etxaniz);
                                begona.altuna@ehu.eus (B. Altuna); itziar.gonzalezd@ehu.eus
                                (I. Gonzalez-Dios)
                                GLOBE https://github.com/GiuliaAPensa (G. Pensa)
                                Orcid 0009-0008-4113-890X (E. Azurmendi); 0009–2099-7766
                                (J. Etxaniz); 0000-0002-4027-2014 (B. Altuna); 0000-0003-1048-5403
                                (I. Gonzalez-Dios)
                                            © 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License   1
                                            Attribution 4.0 International (CC BY 4.0).                                                         https://github.com/GiuliaAPensa/GITAdataset




CEUR
                  ceur-ws.org
Workshop      ISSN 1613-0073
Proceedings
Figure 1: Representation of story pair from GITA




2. Challenge: Description                                     Story classification: The end task revolves around
                                                            determining the plausibility of two stories. This deter-
Our aim in this challenge is to assess the understanding of mination is based on the conflicts detected within the
physical commonsense in LLMs for Italian. We configure two stories. By considering the presence of conflicts,
our assessment proposal in the following terms:             the model can assess the viability and coherence of each
    1. given an original dataset of plausible/implausible story, facilitating the classification of the more plausible
       stories related to physical commonsense, systems one.
       must identify the plausible and implausible sto-       By incorporating physical state classification, conflict
       ries;                                                detection, and story classification, we analyze the aspects
    2. systems must recognize the conflicting sentences of coherent reasoning, supported by evidence-driven
       that generate the conflict in implausible stories; analysis.
    3. systems must spot the underlying physical states
       that cause conflict in implausible stories.
                                                               3. Data description
   The recognition of plausible/implausible stories is the
end task envisaged in this benchmark, which must be            The GITA4CALAMITA dataset is composed by plausi-
justified by the second-level and third-level steps. In Fig-   ble and implausible stories. To compose the dataset, we
ure 1 we present a story pair from the GITA4CALAMITA           focused on concrete actions that could be visualized in
dataset and the relation between the layers of annotation.     the physical world, avoiding mental actions such as “to
Story A is a plausible story, Story B is the corresponding     think” or “to like”. We created 5-sentence stories, giving
implausible story where the first and the second sen-          context and requiring reasoning over multiple sentences.
tences are in conflict: Marco closes the refrigerator and      In all the stories, we avoided nonsensical sentences, in
cannot take the milk out of it. In the right part of the       fact, each sentence is plausible alone, but could be im-
figure we can see the reasoning steps that the system          plausible if associated with another specific sentence in
must follow and resolve. This example is presented in          an implausible story. With these characteristics, the task
English for clarity, but our entire dataset is in Italian.     requires reasoning over the entire context.
   We introduce a series of tasks that constitute a human-        An essential part of our evaluation process is consti-
interpretable reasoning process, supported by a chain          tuted by the presence of physical state annotation. Sys-
of evidence, reflecting the assessment methodology out-        tems must identify the underlying physical states that
lined above. To explain this approach, we present the          make a story not plausible in our physical world. Dur-
tasks from the deepest to the shallowest, mirroring hu-        ing the creation of this dataset, we took into account 14
man reasoning:                                                 physical attributes that were included in the annotation
   Physical state classification: Leveraging our phys-         phase, and we composed stories that contained those at-
ical state annotations, systems must recognize the in-         tributes. Following the work of [9] and [10], these are the
volved physical states in the conflicting sentences of im-     14 physical states that we wanted to have in our stories:
plausible stories. If we look at the example in 1, we are
                                                                    • location, conscious, dressed, wet, exist, clean,
able to identify the problematic physical state “open” as
                                                                      power, functional, in pieces, open, temperature,
cause of implausibility.
                                                                      solid, occupied, edible.
   Conflict detection: Next, the task of conflict detec-
tion entails identifying sentence pairs of the form Si → Sj.
Here, Sj represents the breakpoint, indicating the point
                                                               3.1. Dataset creation
at which the story becomes implausible based on the
given context. Si serves as the evidence that explains the     In the first two rows of Table 1 we can see an example
breakpoint, typically causing a conflicting world state.       of plausible story from the GITA4CALAMITA dataset
                sentence 1           sentence 2           sentence 3          sentence 4           sentence 5
        T       Marco ha aperto il   Marco ha preso il    Marco ha preso la   Marco ha versato     Marco ha bevuto
                frigo.               latte dal frigo.     tazza.              il latte nella       il latte.
                                                                              tazza.
                Marco opened         Marco took the       Marco took the      Marco poured         Marco drank the
                the refrigerator.    milk from the re-    cup.                the milk into the    milk.
                                     frigerator.                              cup.
        F       Marco ha preso il    Marco ha aperto il   Marco ha preso la   Marco ha versato     Marco ha bevuto
        (or-    latte dal frigo.     frigo.               tazza.              il latte nella       il latte.
        der)                                                                  tazza.
                Marco took the       Marco opened         Marco took the      Marco poured         Marco drank the
                milk from the re-    the refrigerator.    cup.                the milk into the    milk.
                frigerator.                                                   cup.
        F       Marco ha chiuso il   Marco ha preso il    Marco ha preso la   Marco ha versato     Marco ha bevuto
        (clo-   frigo.               latte dal frigo.     tazza.              il latte nella       il latte.
        ze)                                                                   tazza.
                Marco closed the     Marco took the       Marco took the      Marco poured         Marco drank the
                refrigerator.        milk from the re-    cup.                the milk into the    milk.
                                     frigerator.                              cup.
Table 1
Example of a plausible story, an implausible story from the Order dataset, and an implausible story from the Cloze dataset.



together with the English translation. In this example, the     3.1.1. Order implausible stories
human actor is Marco, and the five sentences are ordered
                                                               The plausible stories only work in the causal sequence
in the required way: the action of opening something,
                                                               that we created. In the first row of Table 1, there is an
picking something up and using it. We can see that some
                                                               example of a plausible story. In the third row, we see the
of the previously listed physical states appear: Marco is
                                                               corresponding implausible story for the order dataset,
conscious because he is doing something, the refrigerator
                                                               in which Marco, first, takes the milk out from the
is open because the actor can take something out of it, the
                                                               refrigerator and then open the refrigerator, generating a
cup is not occupied by anything and can be functional.
                                                               physically impossible situation: it is not possible to take
   We aimed to minimize subjectivity and limit poten-
                                                               something out of a closed refrigerator. By switching the
tial confounding factors from complex language usage.
                                                               first and the second sentences, we created an implausible
By using simple language, we were able to shift our fo-
                                                               story. In the entire dataset, we decided to generate
cus away from linguistic processing and semantic phe-
                                                               implausible stories changing the order of only two
nomena, allowing us to concentrate more on examining
                                                               sentences for story.
machines’ reasoning abilities, particularly their physical
commonsense understanding. Consequently, we created
our simple sentences in a straightforward declarative
structure, typically starting with the agent of the story,      3.1.2. Cloze implausible stories
followed by a verb, a direct object, and optionally, an        The second approach involves the substitution of a sen-
indirect object.                                               tence from the plausible story with a new sentence. Al-
   Implausible stories are built upon the plausible ones,      though the new sentence itself is not inherently implausi-
preserving the same actor and objects; in doing so we en-      ble, its placement within the sequence renders it implau-
sured that implausible variations remained coherent and        sible. In Table 1, the first sentence of the line F (Cloze), in
believable, and we avoided nonsensical information. To         the fifth row, was changed: Marco closes the refrigerator
create implausible stories, we implemented two different       before taking out the milk. Again, the action is physically
methods:                                                       impossible: if the refrigerator is closed, nothing can be
    1. we switched the order of two sentences;                 taken out from it.
    2. we substituted a plausible sentence with an im-
       plausible one.                                  3.2. Origin of data
   These two methods resulted in two different partitions GITA4CALAMITA is a new version of [8], which is based
of our dataset: the Order dataset of implausible stories, on [11]. Our main objective was to create an Italian
and the Cloze dataset of implausible stories respectively. dataset, manually annotated, to assess a pre-trained lan-
                                                           guage model on physical commonsense tiered tasks. To
create the stories, we took inspiration from the Story       overall effort was of approximately 50-55 hours. An ex-
Cloze Test [12] and ROCStories Corpora [13]. The Story       ample of a complete annotation can be found in Appendix
Cloze Test compiles four-sentence stories with a missing     B.
ending so that a system chooses the most appropriate
conclusion; the ROCStories Corpora is composed of five-      3.4. Data format
sentence stories about everyday life for story generation.
                                                             The GITA4CALAMITA dataset was created and anno-
                                                             tated in a JSON format. The following example is story
3.3. Annotation details
                                                             0-C0 of our dataset, the first implausible Cloze story.
GITA4CALAMITA is annotated on three levels. In the
                                                                {
first level, we annotated the plausibility/implausibility of
                                                                     ”0 − C0 ” : {
a story with TRUE or FALSE. In the second level, in im-
                                                                             ” story_id ”: 0 ,
plausible stories we indicated between which sentences
                                                                             ” w o r k e r _ i d ” : ”GAP ” ,
the conflict was, and in the third level we labelled the
                                                                             ” type ” : ” cloze ” ,
involved physical states in each sentence.
                                                                             ” idx ” : 0 ,
   In the dataset, a plausible story is identified using a
                                                                             ” aug ” : f a l s e ,
story number, while implausible stories are identified us-
                                                                             ” a c t o r ” : ” Marco ” ,
ing the same story number as the plausible version, but
                                                                             ” location ” : ” cucina ” ,
with an additional C or O after the story number, where
                                                                             ” objects ” : ” frigo , latte ,
the letter C refers to the Cloze dataset, and the letter O
                                                                                     tazza , cucchiaio ” ,
refers to the Order dataset. Each story has been anno-
                                                                             ” sentences ” : [
tated using these elements: story id, worker id, actor of
                                                                                     ” Marco ha c h i u s o i l f r i g o
the story, objects of the story, physical states, sentences
                                                                                            .” ,
of the story, as well as number of sentences, and conflict-
                                                                                     ” Marco ha p r e s o i l l a t t e
ing sentences, among others. The complete list and the
                                                                                            dal frigo . ” ,
specific meaning of each element are in Appendix A.
                                                                                     ” Marco ha p r e s o l a t a z z a
   In each implausible story, we annotated the physical
                                                                                            .” ,
state that caused a conflict between two sentences. We
                                                                                     ” Marco ha p r e s o i l
annotated both Order and Cloze implausible stories ac-
                                                                                            cucchiaio .” ,
cording to the corresponding physical state involved. If
                                                                                     ” Marco ha messo i l
we consider the stories in Table 1, both implausible stories
                                                                                            cucchiaio nella tazza
(C and O) are annotated using the physical state “open”,
                                                                                            .”
In fact, in both implausible stories the conflict is related to
                                                                             ],
the openness of the refrigerator: in both cases the refrig-
                                                                             ” length ” : 5 ,
erator appears closed when Marco tries to take the milk
                                                                             ” e x a m p l e _ i d ” : ”0 − C0 ” ,
out of it. There are cases where for one plausible story
                                                                             ” plausible ”: false ,
there are two implausible stories that are implausible for
                                                                             ” breakpoint ” : 1 ,
two different reasons, hence the annotated physical state
                                                                             ” confl_sents ”: [0] ,
is different.
                                                                             ” c o n f l _ p a i r s ” : [0 , 1]
   To ensure consistency and reduce human effort, we
                                                                     }
developed a custom environment and a Python script to
                                                                }
streamline the annotation process. This semi-automated
annotation process helped us process sentences from
different story types, extract entities and actors, and or-
                                                                3.5. Example of prompts used for zero
ganize them for manual annotation. The script provided
a user-friendly terminal interface, and it is available in           or/and few shots
our repository. In terms of annotation efficiency, manu- For each of the three proposed tasks we use a different
ally annotating one plausible story and two implausible prompt:
ones typically took around 50 minutes. However, using
our semi-automated annotation interface, we were able               • Task 1: Please read the following story and an-
to complete the same task in approximately 20 minutes.                 swer if the story is plausible taking into account
Consequently, instead of the estimated 100 hours for an-               the order of the events. Please answer with true
notating the entire dataset, we reduced the time to around             or false.
40 hours. Additionally, some annotations required review               Task 2: The following story is implausible. Iden-
and occasional revisions, hence we estimated that the                  tify the breakpoint, and then select the sentence
       responsible for the implausibility. Please iden-        4. Metrics
       tify the breakpoint sentence and the conflicting
       sentence.                                               The metrics involved in our tasks for the
       Task 3: The following story is implausible. Iden-       GITA4CALAMITA benchmark are the following
       tify the physical state that causes the conflict in     ones:
       the story. These are the descriptions of each phys-
       ical state: Power: Indicates whether an object               • Accuracy assesses the traditional measure of end
       is powered or not, relevant for electrical devices.            task accuracy, which quantifies the proportion
       Location: Refers to the spatial position of an                 of testing examples where plausible stories and
       entity, either human or object. Exist: Denotes                 implausible stories are accurately identified.
       whether an object is present or has disappeared.             • Consistency measures the proportion of testing
       Clean: Refers to the cleanliness of an entity, indi-           examples where not only the implausible story is
       cating whether it is clean or dirty. Edible: Identi-           correctly identified, but also the conflicting sen-
       fies whether an object is fit for consumption. Wet:            tence pair for the implausible story is accurately
       Denotes whether an object or person is in a wet                identified. The aim is to demonstrate the model’s
       or dry state. Functional: Refers to whether an                 consistency in recognizing conflicts when reason-
       object is in working condition or broken. Wear-                ing about plausibility.
       ing: Applies to humans, indicating whether they              • Verifiability evaluates the proportion of testing
       are dressed or not. Open: Refers to whether an                 examples where not only the implausible story
       object (e.g., a door or container) is open or closed.          and the conflicting sentence pair for the implau-
       Conscious: Denotes whether a human is con-                     sible story are correctly identified, but also the
       scious or unconscious. Temperature: Refers to                  underlying physical states that contribute to the
       the relative temperature of an entity, e.g., hot or            conflict are accurately identified. This demon-
       cold. Solid: Describes whether an object is in                 strates that the detected conflict can be validated
       a solid state. Occupied: Indicates whether an                  through a correct understanding of the underly-
       object (e.g., a container) is occupied or contains             ing implausible change of physical states.
       something. In pieces: Refers to whether an ob-
       ject is intact or has been broken into pieces. Select  Taking into consideration the three different metrics,
       one of them after reading the story.                in Table 3 we report the results in our test set. We per-
                                                           form experiments using the base and instruct Llama 3.1,
   We select some examples from our GITA4CALAMITA Gemma 2 and Mistral models of various sizes. Each met-
dataset to be used as few-shot examples. For some of the ric is obtained from a different task, where models are
tests we randomly select the examples, for others, we base evaluated in the instances that are only guessed correctly
our choice on their variability. We select stories where in the previous tasks. All tasks are evaluated in a 3-shot
all possible combination of conflicting sentences were setting, using random examples from the test set. For
happening; at the same time, within the selected stories models that support system prompt (Llama3.1 models),
we try to include most of the physical states annotated. the description of each task is included there, for models
                                                           that do not support it (Gemma2 and Mistral models) the
3.6. Detailed data statistics                              task description is included in the first user input. Each
                                                           few-shot instance is formatted as a multiturn conversa-
The GITA4CALAMITA dataset is an Italian test com- tion between user and assistant. Next, we describe the
posed by a total of 356 stories. The statistics of the main findings from these results.
GITA4CALAMITA dataset are in Table 2.

              Measures             GITA4CALAMITA               Model Size and Performance: Generally, larger mod-
    plausible stories                    117                   els (e.g., Llama-3.1 70B) outperform smaller models across
    implausible stories (ORDER)          122                   the metrics. The 70B Llama-3.1 models show improve-
    implausible stories (CLOZE)          117                   ments over their 8B counterparts, particularly in consis-
    total stories                        356                   tency and verifiability. Gemma2 models also show im-
                                                               provements when bigger models are used. There are two
Table 2
                                                               exceptions in the case of the accuracy: Gemma2-Instruct
Statistics of GITA4CALAMITA
                                                               9B and Llama-3.1-Instruct 8B achieve better results than
                                                               their bigger counterparts Gemma2 27B and Llama3 70B.
                                                               They also outperform the base models.
 Model                   Size                Accuracy                        Consistency                 Verifiability
                                Overall   Cloze Order     Plausible    Overall Cloze Order        Overall Cloze Order
 Gemma-2 (base)           9B    72.75     86.96   70.49        61.34    32.35   45.22     20.66    12.18   16.52    8.26
 Gemma-2-Instruct         9B    76.12     85.22   60.66        83.19    38.66   58.26     20.66    17.65   30.43    5.79
 Gemma-2 (base)          27B    75.28     89.57   59.02        78.15    39.07   55.65     23.97    21.85   31.30   13.22
 Gemma-2-Instruct        27B    73.88     80.00   54.10        88.24    39.08   60.87     19.00    24.79   40.87    9.92
 Llama-3.1 (base)         8B    60.96     70.43   60.66        52.10    26.47   33.04     20.66    11.34   13.04    9.92
 Llama-3.1-Instruct       8B    77.25     93.91   90.16        47.90    37.39   53.91     22.31    10.50   16.52    4.96
 Llama-3.1 (base)        70B    82.02     94.78   92.62        58.82    57.14   66.96     47.93    28.99   36.52   21.49
 Llama-3.1-Instruct      70B    74.16     99.13   98.36        25.21    68.07   82.61     54.55    18.07   25.22   11.57
 Mistral-V0.3 (base)      7B    60.39     66.96   54.92        59.66    20.59   27.83     14.05     6.72   11.30    2.48
 Mistral-Instruct-V0.3    7B    59.83     67.82   27.05        85.71    21.00   40.87      2.48     9.24   19.13    0.00

Table 3
Results of the base and instruct Llama 3.1, Gemma 2 and Mistral models of various sizes



Instruction Tuning Effects: Instruction-tuned ver- Acknowledgments
sions (e.g., Gemma-2-Instruct, Llama-3.1-Instruct) typi-
cally outperform their base counterparts. There are ex- This work has been partially funded by:
ceptions such as order accuracy for LLama 3.1 70B and
                                                             • DeepR3 (TED2021-130295B-C31) funded by
Gemma 2 9B. However, Mistral-V0.3-Instruct is very sim-
                                                               MCIN/AEI/10.13039/501100011033 and European
ilar or worse than the base model and generally is more
                                                               Union NextGeneration EU/PRTR.
biased, it tends to classify as plausible the stories and it
performs better in Cloze than in Order.                      • Disargue                (TED2021-130810B-C21)
                                                               MCIN/AEI/10.13039/501100011033 and Eu-
                                                               ropean Union NextGenerationEU/PRTR.
Cloze, Order and Plausible Most models perform
                                                             • DeepKnowledge          (PID2021-127777OB-C21)
generally better on Cloze examples compared to Order
                                                               MCIN/AEI/10.13039/501100011033 and by
examples. This is consistent across models and metrics.
                                                               FEDER, EU.
Models are generally better in Cloze and Order than in
Plausible. This could be explained by the bias of the        • Ixa group A type research group (IT1570-22)
models to answer true or false when they are asked if          Basque Government
the story is plausible. Models also see double implausible   • IKER-GAITU project 11:4711:23:410:23/0808 by
few-shot examples, which could also cause models to            Basque Government
give that answer more frequently.
                                                                References
5. Limitations                                                   [1] J. Huang, K. C.-C. Chang, Towards Reasoning
This study has some limitations that should be acknowl-              in Large Language Models: A Survey, in: Find-
edged. Firstly, only one prompt was tested for each task,            ings of the Association for Computational Linguis-
which may not fully capture the potential variability in             tics: ACL 2023, Association for Computational
performance. Additionally, the models used were mul-                 Linguistics, Toronto, Canada, 2023, pp. 1049–1065.
tilingual but not specifically tailored for the Italian lan-         URL: https://aclanthology.org/2023.findings-acl.67.
guage, potentially affecting the accuracy of the results             doi:10.18653/v1/2023.findings- acl.67 .
for Italian-specific tasks. Furthermore, the dataset used        [2] K. Sakaguchi, R. L. Bras, C. Bhagavatula, Y. Choi,
in this study was limited to stories within the household            WinoGrande: An Adversarial Winograd Schema
domain, which may not generalize well to other contexts.             Challenge at Scale, Commun. ACM 64 (2021)
                                                                     99–106. URL: https://doi.org/10.1145/3474381.
                                                                     doi:10.1145/3474381 .
6. Ethical issues                                                [3] D. Pessach, E. Shmueli, A Review on Fairness in
                                                                     Machine Learning, ACM Comput. Surv. 55 (2022).
The dataset contains stories that may prototypically oc-             URL: https://doi.org/10.1145/3494672. doi:10.1145/
cur in Italian households. While most of these narratives            3494672 .
are likely to be familiar to a broad audience, people from       [4] E. Davis, Benchmarks for Automated Common-
different cultural backgrounds may find some of the sto-             sense Reasoning: A Survey, ACM Comput. Surv.
ries less frequent.
     (2023). URL: https://doi.org/10.1145/3615355. doi:10.          2021, pp. 4902–4918. URL: https://aclanthology.org/
     1145/3615355 , just Accepted.                                  2021.findings-emnlp.422. doi:10.18653/v1/2021.
 [5] T. Linzen, How Can We Accelerate Progress To-                  findings- emnlp.422 .
     wards Human-like Linguistic Generalization?, in:          [12] N. Mostafazadeh, M. Roth, A. Louis, N. Cham-
     Proceedings of the 58th Annual Meeting of the As-              bers, J. Allen, LSDSem 2017 Shared Task: The
     sociation for Computational Linguistics, Associa-              Story Cloze Test, in: Proceedings of the 2nd
     tion for Computational Linguistics, Online, 2020,              Workshop on Linking Models of Lexical, Senten-
     pp. 5210–5217. URL: https://aclanthology.org/2020.             tial and Discourse-level Semantics, Association for
     acl-main.465. doi:10.18653/v1/2020.acl- main.                  Computational Linguistics, Valencia, Spain, 2017,
     465 .                                                          pp. 46–51. URL: https://aclanthology.org/W17-0906.
 [6] E. M. Bender, A. Koller, Climbing towards NLU:                 doi:10.18653/v1/W17- 0906 .
     On Meaning, Form, and Understanding in the                [13] N. Mostafazadeh, N. Chambers, X. He, D. Parikh,
     Age of Data, in: Proceedings of the 58th An-                   D. Batra, L. Vanderwende, P. Kohli, J. Allen, A Cor-
     nual Meeting of the Association for Computa-                   pus and Cloze Evaluation for Deeper Understanding
     tional Linguistics, Association for Computational              of Commonsense Stories, in: Proceedings of the
     Linguistics, Online, 2020, pp. 5185–5198. URL:                 2016 Conference of the North American Chapter
     https://aclanthology.org/2020.acl-main.463. doi:10.            of the Association for Computational Linguistics:
     18653/v1/2020.acl- main.463 .                                  Human Language Technologies, Association for
 [7] G. Attanasio, P. Basile, F. Borazio, D. Croce, M. Fran-        Computational Linguistics, San Diego, California,
     cis, J. Gili, E. Musacchio, M. Nissim, V. Patti, M. Ri-        2016, pp. 839–849. URL: https://aclanthology.org/
     naldi, D. Scalena, CALAMITA: Challenge the Abili-              N16-1098. doi:10.18653/v1/N16- 1098 .
     ties of LAnguage Models in ITAlian, in: Proceed-
     ings of the 10th Italian Conference on Computa-
     tional Linguistics (CLiC-it 2024), Pisa, Italy, Decem-    A. Annotations in the dataset
     ber 4 - December 6, 2024, CEUR Workshop Proceed-
     ings, CEUR-WS.org, 2024.                                  These are the attributes that encode the metadata and
 [8] G. Pensa, B. Altuna, I. Gonzalez-Dios, A Multi-           linguistic information in the GITA dataset:
     layered Approach to Physical Commonsense Un-
     derstanding: Creation and Evaluation of an Ital-               • story_id: refers to the number of the story for
     ian Dataset, in: N. Calzolari, M.-Y. Kan, V. Hoste,              both plausible and implausible stories.
     A. Lenci, S. Sakti, N. Xue (Eds.), Proceedings of the          • worker_id: refers to the name assigned to a spe-
     2024 Joint International Conference on Computa-                  cific worker during the creation of the story.
     tional Linguistics, Language Resources and Evalua-             • type: refers to cloze or order and it is a label used
     tion (LREC-COLING 2024), ELRA and ICCL, Torino,                  only in implausible stories.
     Italia, 2024, pp. 819–831. URL: https://aclanthology.          • idx: refers to the implausible dataset, where there
     org/2024.lrec-main.74.                                           is more than one implausible story for a given
 [9] Q. Gao, M. Doering, S. Yang, J. Chai, Physical                   story number; for example, if we have more than
     Causality of Action Verbs in Grounded Language                   one implausible version of a plausible story (we
     Understanding, in: Proceedings of the 54th An-                   created more than an implausible story chang-
     nual Meeting of the Association for Computational                ing the order of our sentences more than once),
     Linguistics (Volume 1: Long Papers), Association                 the index number indicates to which implausible
     for Computational Linguistics, Berlin, Germany,                  example we are referring.
     2016, pp. 1814–1824. URL: https://aclanthology.org/            • aug: refers to possible automatic data augmenta-
     P16-1171. doi:10.18653/v1/P16- 1171 .                            tion techniques that can be taken into account for
[10] A. Bosselut, O. Levy, A. Holtzman, C. Ennis,                     future works to resolve an overfitting problem.
     D. Fox, Y. Choi, Simulating Action Dynam-                      • actor: refers to the human agent of the story.
     ics with Neural Process Networks,                CoRR          • location: refers to the room where the story
     abs/1711.05313 (2017). URL: http://arxiv.org/abs/                takes place.
     1711.05313. arXiv:1711.05313 .                                 • objects: refers to all the inanimate entities that
[11] S. Storks, Q. Gao, Y. Zhang, J. Chai, Tiered Rea-                we find into each story.
     soning for Intuitive Physics: Toward Verifiable                • sentences: includes the 5 sentences in the story.
     Commonsense Language Understanding, in: Find-                  • length: refers to the number of sentences in each
     ings of the Association for Computational Lin-                   story.
     guistics: EMNLP 2021, Association for Computa-                 • example_id: corresponds to the story number
     tional Linguistics, Punta Cana, Dominican Republic,              and includes letters for implausible stories.
     • plausible: is TRUE when the story is plausible        breakpoint :
       and FALSE when it is implausible.                     −1
     • breakpoint: refers to the sentence where the          c o n f l _ s e n t s ( type o n l y [ ] ) :
       story becomes implausible, where the conflict be-     []
       comes evident; in plausible stories the breakpoint
       is always -1.                                                    Listing 1: Annotation environment.
     • conlict_sents: refers to the other sentence in the
       story that together with the breakpoint sentence
       makes the story implausible; in plausible stories
       this field is blank.
     • conlict_pairs: refers to the conflict pair of sen-
       tences, gathering the two previous labels; in plau-
       sible stories this field is blank.
     • states: includes all the physical states annota-
       tions for all the stories.


B. Annotation environment

actor :
Marco
objects :
frigo l a t t e tazza cucchiaio
s t o r y _ n u m b e r ( same a s s t o r y _ i d in
        quotes ) :
‘0 ’
s t o r y _ i d (NO q u o t e s , NO l e t t e r , o n l y
        number ) :
0
worker_id ( in quotes ) :
‘GAP ’
type ( n u l l f o r p o s i t i v e , o r d e r , o r
        c l o z e , in q u o t e s ) :
null
i d x ( n u l l , o r same a s NUMBER in s t o r y
          number ) :
null
aug ( f a l s e ) :
false
l o c a t i o n ( in q u o t e s ) :
‘ cucina ’
sentences :
Marco ha a p e r t o i l f r i g o .           Marco ha
        preso i l l a t t e .          Marco ha p r e s o
        la tazza .           Marco ha p r e s o i l
        cucchiaio .            Marco ha messo i l
        cucchiaio nella tazza .
length :
5
e x a m p l e _ i d ( same a s s t o r y number , i n
        quotes ) :
‘0 ’
plausible :
true