=Paper=
{{Paper
|id=Vol-3878/127_calamita_long
|storemode=property
|title=GITA4CALAMITA - Evaluating the Physical Commonsense Understanding of Italian LLMs in a Multi-layered Approach: A CALAMITA Challenge
|pdfUrl=https://ceur-ws.org/Vol-3878/127_calamita_long.pdf
|volume=Vol-3878
|authors=Giulia Pensa,Ekhi Azurmendi,Julen Etxaniz,Begoña Altuna,Itziar Gonzalez-Dios
|dblpUrl=https://dblp.org/rec/conf/clic-it/PensaAEAG24
}}
==GITA4CALAMITA - Evaluating the Physical Commonsense Understanding of Italian LLMs in a Multi-layered Approach: A CALAMITA Challenge==
GITA4CALAMITA - Evaluating the Physical Commonsense
Understanding of Italian LLMs in a Multi-layered Approach:
A CALAMITA Challenge
Giulia Pensa1 , Ekhi Azurmendi2 , Julen Etxaniz2 , Begoña Altuna2 and Itziar Gonzalez-Dios2
1
University of the Basque Country UPV/EHU
2
HiTZ Center - Ixa, University of the Basque Country UPV/EHU
Abstract
In the context of the CALAMITA Challenge, we investigate the physical commonsense reasoning capabilities of large language
models (LLMs) and introduce a methodology to assess their understanding of the physical world. To this end, we use a test
set designed to evaluate physical commonsense reasoning in LLMs for the Italian language. We present a tiered dataset,
named the Graded Italian Annotated dataset (GITA), which is written and annotated by a professional linguist. This dataset
enables us to focus on three distinct levels of commonsense understanding. Our benchmark aims to evaluate three specific
tasks: identifying plausible and implausible stories within our dataset, identifying the conflict that generates an implausible
story, and identifying the physical states that make a story implausible. We perform these tasks using LLAMA3, Gemma2
and Mistral. Our findings reveal that, although the models may excel at high-level classification tasks, their reasoning is
inconsistent and unverifiable, as they fail to capture intermediate evidence.
Keywords
Physical commonsense reasoning, large language models, Italian benchmark
1. Challenge: Introduction and proficiency in reasoning and comprehending language
[5, 6]. As a result, there remains uncertainty regarding
Motivation machines’ ability to truly perform reasoning and whether
Physical commonsense understanding refers to the abil- the existing issues in this regard have been sufficiently
ity to comprehend the physical world and the events that addressed.
transpire within it. This capability is a crucial component In this context, our aim is to contribute to this chal-
of human intelligence, enabling us to reason about our lenge developing an original Italian benchmark that can
environment, anticipate future occurrences, and navi- be used to assess the ability of language models to un-
gate our surroundings effortlessly, and recently there has derstand physical commonsense in a more truthful way,
been notable advancement in the development of large focusing not only on end tasks, but also on intermediate
language models (LLMs) that can produce human-like layer tasks.
language and execute a variety of language-related tasks. In this paper, we present GITA4CALAMITA, the
LLMs have exhibited promising outcomes in grasping Graded Italian Annotated dataset for the CALAMITA
common sense in particular situations [1, 2]. Neverthe- challenge [7]. GITA4CALAMITA is an adapted version
less, it is widely recognized that the most precise evalua- of the GITA dataset proposed in [8]. In particular, we de-
tion of their capabilities is attained when assessing their cided to revise the physical states annotation and adapt
performance in specific end tasks [3, 4]. The evaluation it to this challenge. The first version of GITA dataset
often emphasizes the capacity of LLMs to replicate rela- is available in our repository under the license CC BY-
tively straightforward tasks, rather than their authentic NC-SA 4.0.1 . The GITA4CALAMITA dataset is manu-
ally compiled by a professional linguist, which allows
CLiC-it 2024: Tenth Italian Conference on Computational Linguistics,
for this multi-layered evaluation of the reasoning pro-
Dec 04 - 06, 2024, Pisa, Italy cess. With the creation of an Italian dataset we gain
∗
Corresponding author. the linguistic and cultural perspective of Italian, while
†
These authors contributed equally. commonsense research in Natural Language Processing
Envelope-Open giulia.pensa.tr@gmail.com (G. Pensa); ekhi.azurmendi@ehu.eus (NLP) has largely been focused on the English language.
(E. Azurmendi); julen.etxaniz@ehu.eus (J. Etxaniz);
begona.altuna@ehu.eus (B. Altuna); itziar.gonzalezd@ehu.eus
(I. Gonzalez-Dios)
GLOBE https://github.com/GiuliaAPensa (G. Pensa)
Orcid 0009-0008-4113-890X (E. Azurmendi); 0009–2099-7766
(J. Etxaniz); 0000-0002-4027-2014 (B. Altuna); 0000-0003-1048-5403
(I. Gonzalez-Dios)
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License 1
Attribution 4.0 International (CC BY 4.0). https://github.com/GiuliaAPensa/GITAdataset
CEUR
ceur-ws.org
Workshop ISSN 1613-0073
Proceedings
Figure 1: Representation of story pair from GITA
2. Challenge: Description Story classification: The end task revolves around
determining the plausibility of two stories. This deter-
Our aim in this challenge is to assess the understanding of mination is based on the conflicts detected within the
physical commonsense in LLMs for Italian. We configure two stories. By considering the presence of conflicts,
our assessment proposal in the following terms: the model can assess the viability and coherence of each
1. given an original dataset of plausible/implausible story, facilitating the classification of the more plausible
stories related to physical commonsense, systems one.
must identify the plausible and implausible sto- By incorporating physical state classification, conflict
ries; detection, and story classification, we analyze the aspects
2. systems must recognize the conflicting sentences of coherent reasoning, supported by evidence-driven
that generate the conflict in implausible stories; analysis.
3. systems must spot the underlying physical states
that cause conflict in implausible stories.
3. Data description
The recognition of plausible/implausible stories is the
end task envisaged in this benchmark, which must be The GITA4CALAMITA dataset is composed by plausi-
justified by the second-level and third-level steps. In Fig- ble and implausible stories. To compose the dataset, we
ure 1 we present a story pair from the GITA4CALAMITA focused on concrete actions that could be visualized in
dataset and the relation between the layers of annotation. the physical world, avoiding mental actions such as “to
Story A is a plausible story, Story B is the corresponding think” or “to like”. We created 5-sentence stories, giving
implausible story where the first and the second sen- context and requiring reasoning over multiple sentences.
tences are in conflict: Marco closes the refrigerator and In all the stories, we avoided nonsensical sentences, in
cannot take the milk out of it. In the right part of the fact, each sentence is plausible alone, but could be im-
figure we can see the reasoning steps that the system plausible if associated with another specific sentence in
must follow and resolve. This example is presented in an implausible story. With these characteristics, the task
English for clarity, but our entire dataset is in Italian. requires reasoning over the entire context.
We introduce a series of tasks that constitute a human- An essential part of our evaluation process is consti-
interpretable reasoning process, supported by a chain tuted by the presence of physical state annotation. Sys-
of evidence, reflecting the assessment methodology out- tems must identify the underlying physical states that
lined above. To explain this approach, we present the make a story not plausible in our physical world. Dur-
tasks from the deepest to the shallowest, mirroring hu- ing the creation of this dataset, we took into account 14
man reasoning: physical attributes that were included in the annotation
Physical state classification: Leveraging our phys- phase, and we composed stories that contained those at-
ical state annotations, systems must recognize the in- tributes. Following the work of [9] and [10], these are the
volved physical states in the conflicting sentences of im- 14 physical states that we wanted to have in our stories:
plausible stories. If we look at the example in 1, we are
• location, conscious, dressed, wet, exist, clean,
able to identify the problematic physical state “open” as
power, functional, in pieces, open, temperature,
cause of implausibility.
solid, occupied, edible.
Conflict detection: Next, the task of conflict detec-
tion entails identifying sentence pairs of the form Si → Sj.
Here, Sj represents the breakpoint, indicating the point
3.1. Dataset creation
at which the story becomes implausible based on the
given context. Si serves as the evidence that explains the In the first two rows of Table 1 we can see an example
breakpoint, typically causing a conflicting world state. of plausible story from the GITA4CALAMITA dataset
sentence 1 sentence 2 sentence 3 sentence 4 sentence 5
T Marco ha aperto il Marco ha preso il Marco ha preso la Marco ha versato Marco ha bevuto
frigo. latte dal frigo. tazza. il latte nella il latte.
tazza.
Marco opened Marco took the Marco took the Marco poured Marco drank the
the refrigerator. milk from the re- cup. the milk into the milk.
frigerator. cup.
F Marco ha preso il Marco ha aperto il Marco ha preso la Marco ha versato Marco ha bevuto
(or- latte dal frigo. frigo. tazza. il latte nella il latte.
der) tazza.
Marco took the Marco opened Marco took the Marco poured Marco drank the
milk from the re- the refrigerator. cup. the milk into the milk.
frigerator. cup.
F Marco ha chiuso il Marco ha preso il Marco ha preso la Marco ha versato Marco ha bevuto
(clo- frigo. latte dal frigo. tazza. il latte nella il latte.
ze) tazza.
Marco closed the Marco took the Marco took the Marco poured Marco drank the
refrigerator. milk from the re- cup. the milk into the milk.
frigerator. cup.
Table 1
Example of a plausible story, an implausible story from the Order dataset, and an implausible story from the Cloze dataset.
together with the English translation. In this example, the 3.1.1. Order implausible stories
human actor is Marco, and the five sentences are ordered
The plausible stories only work in the causal sequence
in the required way: the action of opening something,
that we created. In the first row of Table 1, there is an
picking something up and using it. We can see that some
example of a plausible story. In the third row, we see the
of the previously listed physical states appear: Marco is
corresponding implausible story for the order dataset,
conscious because he is doing something, the refrigerator
in which Marco, first, takes the milk out from the
is open because the actor can take something out of it, the
refrigerator and then open the refrigerator, generating a
cup is not occupied by anything and can be functional.
physically impossible situation: it is not possible to take
We aimed to minimize subjectivity and limit poten-
something out of a closed refrigerator. By switching the
tial confounding factors from complex language usage.
first and the second sentences, we created an implausible
By using simple language, we were able to shift our fo-
story. In the entire dataset, we decided to generate
cus away from linguistic processing and semantic phe-
implausible stories changing the order of only two
nomena, allowing us to concentrate more on examining
sentences for story.
machines’ reasoning abilities, particularly their physical
commonsense understanding. Consequently, we created
our simple sentences in a straightforward declarative
structure, typically starting with the agent of the story, 3.1.2. Cloze implausible stories
followed by a verb, a direct object, and optionally, an The second approach involves the substitution of a sen-
indirect object. tence from the plausible story with a new sentence. Al-
Implausible stories are built upon the plausible ones, though the new sentence itself is not inherently implausi-
preserving the same actor and objects; in doing so we en- ble, its placement within the sequence renders it implau-
sured that implausible variations remained coherent and sible. In Table 1, the first sentence of the line F (Cloze), in
believable, and we avoided nonsensical information. To the fifth row, was changed: Marco closes the refrigerator
create implausible stories, we implemented two different before taking out the milk. Again, the action is physically
methods: impossible: if the refrigerator is closed, nothing can be
1. we switched the order of two sentences; taken out from it.
2. we substituted a plausible sentence with an im-
plausible one. 3.2. Origin of data
These two methods resulted in two different partitions GITA4CALAMITA is a new version of [8], which is based
of our dataset: the Order dataset of implausible stories, on [11]. Our main objective was to create an Italian
and the Cloze dataset of implausible stories respectively. dataset, manually annotated, to assess a pre-trained lan-
guage model on physical commonsense tiered tasks. To
create the stories, we took inspiration from the Story overall effort was of approximately 50-55 hours. An ex-
Cloze Test [12] and ROCStories Corpora [13]. The Story ample of a complete annotation can be found in Appendix
Cloze Test compiles four-sentence stories with a missing B.
ending so that a system chooses the most appropriate
conclusion; the ROCStories Corpora is composed of five- 3.4. Data format
sentence stories about everyday life for story generation.
The GITA4CALAMITA dataset was created and anno-
tated in a JSON format. The following example is story
3.3. Annotation details
0-C0 of our dataset, the first implausible Cloze story.
GITA4CALAMITA is annotated on three levels. In the
{
first level, we annotated the plausibility/implausibility of
”0 − C0 ” : {
a story with TRUE or FALSE. In the second level, in im-
” story_id ”: 0 ,
plausible stories we indicated between which sentences
” w o r k e r _ i d ” : ”GAP ” ,
the conflict was, and in the third level we labelled the
” type ” : ” cloze ” ,
involved physical states in each sentence.
” idx ” : 0 ,
In the dataset, a plausible story is identified using a
” aug ” : f a l s e ,
story number, while implausible stories are identified us-
” a c t o r ” : ” Marco ” ,
ing the same story number as the plausible version, but
” location ” : ” cucina ” ,
with an additional C or O after the story number, where
” objects ” : ” frigo , latte ,
the letter C refers to the Cloze dataset, and the letter O
tazza , cucchiaio ” ,
refers to the Order dataset. Each story has been anno-
” sentences ” : [
tated using these elements: story id, worker id, actor of
” Marco ha c h i u s o i l f r i g o
the story, objects of the story, physical states, sentences
.” ,
of the story, as well as number of sentences, and conflict-
” Marco ha p r e s o i l l a t t e
ing sentences, among others. The complete list and the
dal frigo . ” ,
specific meaning of each element are in Appendix A.
” Marco ha p r e s o l a t a z z a
In each implausible story, we annotated the physical
.” ,
state that caused a conflict between two sentences. We
” Marco ha p r e s o i l
annotated both Order and Cloze implausible stories ac-
cucchiaio .” ,
cording to the corresponding physical state involved. If
” Marco ha messo i l
we consider the stories in Table 1, both implausible stories
cucchiaio nella tazza
(C and O) are annotated using the physical state “open”,
.”
In fact, in both implausible stories the conflict is related to
],
the openness of the refrigerator: in both cases the refrig-
” length ” : 5 ,
erator appears closed when Marco tries to take the milk
” e x a m p l e _ i d ” : ”0 − C0 ” ,
out of it. There are cases where for one plausible story
” plausible ”: false ,
there are two implausible stories that are implausible for
” breakpoint ” : 1 ,
two different reasons, hence the annotated physical state
” confl_sents ”: [0] ,
is different.
” c o n f l _ p a i r s ” : [0 , 1]
To ensure consistency and reduce human effort, we
}
developed a custom environment and a Python script to
}
streamline the annotation process. This semi-automated
annotation process helped us process sentences from
different story types, extract entities and actors, and or-
3.5. Example of prompts used for zero
ganize them for manual annotation. The script provided
a user-friendly terminal interface, and it is available in or/and few shots
our repository. In terms of annotation efficiency, manu- For each of the three proposed tasks we use a different
ally annotating one plausible story and two implausible prompt:
ones typically took around 50 minutes. However, using
our semi-automated annotation interface, we were able • Task 1: Please read the following story and an-
to complete the same task in approximately 20 minutes. swer if the story is plausible taking into account
Consequently, instead of the estimated 100 hours for an- the order of the events. Please answer with true
notating the entire dataset, we reduced the time to around or false.
40 hours. Additionally, some annotations required review Task 2: The following story is implausible. Iden-
and occasional revisions, hence we estimated that the tify the breakpoint, and then select the sentence
responsible for the implausibility. Please iden- 4. Metrics
tify the breakpoint sentence and the conflicting
sentence. The metrics involved in our tasks for the
Task 3: The following story is implausible. Iden- GITA4CALAMITA benchmark are the following
tify the physical state that causes the conflict in ones:
the story. These are the descriptions of each phys-
ical state: Power: Indicates whether an object • Accuracy assesses the traditional measure of end
is powered or not, relevant for electrical devices. task accuracy, which quantifies the proportion
Location: Refers to the spatial position of an of testing examples where plausible stories and
entity, either human or object. Exist: Denotes implausible stories are accurately identified.
whether an object is present or has disappeared. • Consistency measures the proportion of testing
Clean: Refers to the cleanliness of an entity, indi- examples where not only the implausible story is
cating whether it is clean or dirty. Edible: Identi- correctly identified, but also the conflicting sen-
fies whether an object is fit for consumption. Wet: tence pair for the implausible story is accurately
Denotes whether an object or person is in a wet identified. The aim is to demonstrate the model’s
or dry state. Functional: Refers to whether an consistency in recognizing conflicts when reason-
object is in working condition or broken. Wear- ing about plausibility.
ing: Applies to humans, indicating whether they • Verifiability evaluates the proportion of testing
are dressed or not. Open: Refers to whether an examples where not only the implausible story
object (e.g., a door or container) is open or closed. and the conflicting sentence pair for the implau-
Conscious: Denotes whether a human is con- sible story are correctly identified, but also the
scious or unconscious. Temperature: Refers to underlying physical states that contribute to the
the relative temperature of an entity, e.g., hot or conflict are accurately identified. This demon-
cold. Solid: Describes whether an object is in strates that the detected conflict can be validated
a solid state. Occupied: Indicates whether an through a correct understanding of the underly-
object (e.g., a container) is occupied or contains ing implausible change of physical states.
something. In pieces: Refers to whether an ob-
ject is intact or has been broken into pieces. Select Taking into consideration the three different metrics,
one of them after reading the story. in Table 3 we report the results in our test set. We per-
form experiments using the base and instruct Llama 3.1,
We select some examples from our GITA4CALAMITA Gemma 2 and Mistral models of various sizes. Each met-
dataset to be used as few-shot examples. For some of the ric is obtained from a different task, where models are
tests we randomly select the examples, for others, we base evaluated in the instances that are only guessed correctly
our choice on their variability. We select stories where in the previous tasks. All tasks are evaluated in a 3-shot
all possible combination of conflicting sentences were setting, using random examples from the test set. For
happening; at the same time, within the selected stories models that support system prompt (Llama3.1 models),
we try to include most of the physical states annotated. the description of each task is included there, for models
that do not support it (Gemma2 and Mistral models) the
3.6. Detailed data statistics task description is included in the first user input. Each
few-shot instance is formatted as a multiturn conversa-
The GITA4CALAMITA dataset is an Italian test com- tion between user and assistant. Next, we describe the
posed by a total of 356 stories. The statistics of the main findings from these results.
GITA4CALAMITA dataset are in Table 2.
Measures GITA4CALAMITA Model Size and Performance: Generally, larger mod-
plausible stories 117 els (e.g., Llama-3.1 70B) outperform smaller models across
implausible stories (ORDER) 122 the metrics. The 70B Llama-3.1 models show improve-
implausible stories (CLOZE) 117 ments over their 8B counterparts, particularly in consis-
total stories 356 tency and verifiability. Gemma2 models also show im-
provements when bigger models are used. There are two
Table 2
exceptions in the case of the accuracy: Gemma2-Instruct
Statistics of GITA4CALAMITA
9B and Llama-3.1-Instruct 8B achieve better results than
their bigger counterparts Gemma2 27B and Llama3 70B.
They also outperform the base models.
Model Size Accuracy Consistency Verifiability
Overall Cloze Order Plausible Overall Cloze Order Overall Cloze Order
Gemma-2 (base) 9B 72.75 86.96 70.49 61.34 32.35 45.22 20.66 12.18 16.52 8.26
Gemma-2-Instruct 9B 76.12 85.22 60.66 83.19 38.66 58.26 20.66 17.65 30.43 5.79
Gemma-2 (base) 27B 75.28 89.57 59.02 78.15 39.07 55.65 23.97 21.85 31.30 13.22
Gemma-2-Instruct 27B 73.88 80.00 54.10 88.24 39.08 60.87 19.00 24.79 40.87 9.92
Llama-3.1 (base) 8B 60.96 70.43 60.66 52.10 26.47 33.04 20.66 11.34 13.04 9.92
Llama-3.1-Instruct 8B 77.25 93.91 90.16 47.90 37.39 53.91 22.31 10.50 16.52 4.96
Llama-3.1 (base) 70B 82.02 94.78 92.62 58.82 57.14 66.96 47.93 28.99 36.52 21.49
Llama-3.1-Instruct 70B 74.16 99.13 98.36 25.21 68.07 82.61 54.55 18.07 25.22 11.57
Mistral-V0.3 (base) 7B 60.39 66.96 54.92 59.66 20.59 27.83 14.05 6.72 11.30 2.48
Mistral-Instruct-V0.3 7B 59.83 67.82 27.05 85.71 21.00 40.87 2.48 9.24 19.13 0.00
Table 3
Results of the base and instruct Llama 3.1, Gemma 2 and Mistral models of various sizes
Instruction Tuning Effects: Instruction-tuned ver- Acknowledgments
sions (e.g., Gemma-2-Instruct, Llama-3.1-Instruct) typi-
cally outperform their base counterparts. There are ex- This work has been partially funded by:
ceptions such as order accuracy for LLama 3.1 70B and
• DeepR3 (TED2021-130295B-C31) funded by
Gemma 2 9B. However, Mistral-V0.3-Instruct is very sim-
MCIN/AEI/10.13039/501100011033 and European
ilar or worse than the base model and generally is more
Union NextGeneration EU/PRTR.
biased, it tends to classify as plausible the stories and it
performs better in Cloze than in Order. • Disargue (TED2021-130810B-C21)
MCIN/AEI/10.13039/501100011033 and Eu-
ropean Union NextGenerationEU/PRTR.
Cloze, Order and Plausible Most models perform
• DeepKnowledge (PID2021-127777OB-C21)
generally better on Cloze examples compared to Order
MCIN/AEI/10.13039/501100011033 and by
examples. This is consistent across models and metrics.
FEDER, EU.
Models are generally better in Cloze and Order than in
Plausible. This could be explained by the bias of the • Ixa group A type research group (IT1570-22)
models to answer true or false when they are asked if Basque Government
the story is plausible. Models also see double implausible • IKER-GAITU project 11:4711:23:410:23/0808 by
few-shot examples, which could also cause models to Basque Government
give that answer more frequently.
References
5. Limitations [1] J. Huang, K. C.-C. Chang, Towards Reasoning
This study has some limitations that should be acknowl- in Large Language Models: A Survey, in: Find-
edged. Firstly, only one prompt was tested for each task, ings of the Association for Computational Linguis-
which may not fully capture the potential variability in tics: ACL 2023, Association for Computational
performance. Additionally, the models used were mul- Linguistics, Toronto, Canada, 2023, pp. 1049–1065.
tilingual but not specifically tailored for the Italian lan- URL: https://aclanthology.org/2023.findings-acl.67.
guage, potentially affecting the accuracy of the results doi:10.18653/v1/2023.findings- acl.67 .
for Italian-specific tasks. Furthermore, the dataset used [2] K. Sakaguchi, R. L. Bras, C. Bhagavatula, Y. Choi,
in this study was limited to stories within the household WinoGrande: An Adversarial Winograd Schema
domain, which may not generalize well to other contexts. Challenge at Scale, Commun. ACM 64 (2021)
99–106. URL: https://doi.org/10.1145/3474381.
doi:10.1145/3474381 .
6. Ethical issues [3] D. Pessach, E. Shmueli, A Review on Fairness in
Machine Learning, ACM Comput. Surv. 55 (2022).
The dataset contains stories that may prototypically oc- URL: https://doi.org/10.1145/3494672. doi:10.1145/
cur in Italian households. While most of these narratives 3494672 .
are likely to be familiar to a broad audience, people from [4] E. Davis, Benchmarks for Automated Common-
different cultural backgrounds may find some of the sto- sense Reasoning: A Survey, ACM Comput. Surv.
ries less frequent.
(2023). URL: https://doi.org/10.1145/3615355. doi:10. 2021, pp. 4902–4918. URL: https://aclanthology.org/
1145/3615355 , just Accepted. 2021.findings-emnlp.422. doi:10.18653/v1/2021.
[5] T. Linzen, How Can We Accelerate Progress To- findings- emnlp.422 .
wards Human-like Linguistic Generalization?, in: [12] N. Mostafazadeh, M. Roth, A. Louis, N. Cham-
Proceedings of the 58th Annual Meeting of the As- bers, J. Allen, LSDSem 2017 Shared Task: The
sociation for Computational Linguistics, Associa- Story Cloze Test, in: Proceedings of the 2nd
tion for Computational Linguistics, Online, 2020, Workshop on Linking Models of Lexical, Senten-
pp. 5210–5217. URL: https://aclanthology.org/2020. tial and Discourse-level Semantics, Association for
acl-main.465. doi:10.18653/v1/2020.acl- main. Computational Linguistics, Valencia, Spain, 2017,
465 . pp. 46–51. URL: https://aclanthology.org/W17-0906.
[6] E. M. Bender, A. Koller, Climbing towards NLU: doi:10.18653/v1/W17- 0906 .
On Meaning, Form, and Understanding in the [13] N. Mostafazadeh, N. Chambers, X. He, D. Parikh,
Age of Data, in: Proceedings of the 58th An- D. Batra, L. Vanderwende, P. Kohli, J. Allen, A Cor-
nual Meeting of the Association for Computa- pus and Cloze Evaluation for Deeper Understanding
tional Linguistics, Association for Computational of Commonsense Stories, in: Proceedings of the
Linguistics, Online, 2020, pp. 5185–5198. URL: 2016 Conference of the North American Chapter
https://aclanthology.org/2020.acl-main.463. doi:10. of the Association for Computational Linguistics:
18653/v1/2020.acl- main.463 . Human Language Technologies, Association for
[7] G. Attanasio, P. Basile, F. Borazio, D. Croce, M. Fran- Computational Linguistics, San Diego, California,
cis, J. Gili, E. Musacchio, M. Nissim, V. Patti, M. Ri- 2016, pp. 839–849. URL: https://aclanthology.org/
naldi, D. Scalena, CALAMITA: Challenge the Abili- N16-1098. doi:10.18653/v1/N16- 1098 .
ties of LAnguage Models in ITAlian, in: Proceed-
ings of the 10th Italian Conference on Computa-
tional Linguistics (CLiC-it 2024), Pisa, Italy, Decem- A. Annotations in the dataset
ber 4 - December 6, 2024, CEUR Workshop Proceed-
ings, CEUR-WS.org, 2024. These are the attributes that encode the metadata and
[8] G. Pensa, B. Altuna, I. Gonzalez-Dios, A Multi- linguistic information in the GITA dataset:
layered Approach to Physical Commonsense Un-
derstanding: Creation and Evaluation of an Ital- • story_id: refers to the number of the story for
ian Dataset, in: N. Calzolari, M.-Y. Kan, V. Hoste, both plausible and implausible stories.
A. Lenci, S. Sakti, N. Xue (Eds.), Proceedings of the • worker_id: refers to the name assigned to a spe-
2024 Joint International Conference on Computa- cific worker during the creation of the story.
tional Linguistics, Language Resources and Evalua- • type: refers to cloze or order and it is a label used
tion (LREC-COLING 2024), ELRA and ICCL, Torino, only in implausible stories.
Italia, 2024, pp. 819–831. URL: https://aclanthology. • idx: refers to the implausible dataset, where there
org/2024.lrec-main.74. is more than one implausible story for a given
[9] Q. Gao, M. Doering, S. Yang, J. Chai, Physical story number; for example, if we have more than
Causality of Action Verbs in Grounded Language one implausible version of a plausible story (we
Understanding, in: Proceedings of the 54th An- created more than an implausible story chang-
nual Meeting of the Association for Computational ing the order of our sentences more than once),
Linguistics (Volume 1: Long Papers), Association the index number indicates to which implausible
for Computational Linguistics, Berlin, Germany, example we are referring.
2016, pp. 1814–1824. URL: https://aclanthology.org/ • aug: refers to possible automatic data augmenta-
P16-1171. doi:10.18653/v1/P16- 1171 . tion techniques that can be taken into account for
[10] A. Bosselut, O. Levy, A. Holtzman, C. Ennis, future works to resolve an overfitting problem.
D. Fox, Y. Choi, Simulating Action Dynam- • actor: refers to the human agent of the story.
ics with Neural Process Networks, CoRR • location: refers to the room where the story
abs/1711.05313 (2017). URL: http://arxiv.org/abs/ takes place.
1711.05313. arXiv:1711.05313 . • objects: refers to all the inanimate entities that
[11] S. Storks, Q. Gao, Y. Zhang, J. Chai, Tiered Rea- we find into each story.
soning for Intuitive Physics: Toward Verifiable • sentences: includes the 5 sentences in the story.
Commonsense Language Understanding, in: Find- • length: refers to the number of sentences in each
ings of the Association for Computational Lin- story.
guistics: EMNLP 2021, Association for Computa- • example_id: corresponds to the story number
tional Linguistics, Punta Cana, Dominican Republic, and includes letters for implausible stories.
• plausible: is TRUE when the story is plausible breakpoint :
and FALSE when it is implausible. −1
• breakpoint: refers to the sentence where the c o n f l _ s e n t s ( type o n l y [ ] ) :
story becomes implausible, where the conflict be- []
comes evident; in plausible stories the breakpoint
is always -1. Listing 1: Annotation environment.
• conlict_sents: refers to the other sentence in the
story that together with the breakpoint sentence
makes the story implausible; in plausible stories
this field is blank.
• conlict_pairs: refers to the conflict pair of sen-
tences, gathering the two previous labels; in plau-
sible stories this field is blank.
• states: includes all the physical states annota-
tions for all the stories.
B. Annotation environment
actor :
Marco
objects :
frigo l a t t e tazza cucchiaio
s t o r y _ n u m b e r ( same a s s t o r y _ i d in
quotes ) :
‘0 ’
s t o r y _ i d (NO q u o t e s , NO l e t t e r , o n l y
number ) :
0
worker_id ( in quotes ) :
‘GAP ’
type ( n u l l f o r p o s i t i v e , o r d e r , o r
c l o z e , in q u o t e s ) :
null
i d x ( n u l l , o r same a s NUMBER in s t o r y
number ) :
null
aug ( f a l s e ) :
false
l o c a t i o n ( in q u o t e s ) :
‘ cucina ’
sentences :
Marco ha a p e r t o i l f r i g o . Marco ha
preso i l l a t t e . Marco ha p r e s o
la tazza . Marco ha p r e s o i l
cucchiaio . Marco ha messo i l
cucchiaio nella tazza .
length :
5
e x a m p l e _ i d ( same a s s t o r y number , i n
quotes ) :
‘0 ’
plausible :
true