<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Analysis of Visually Grounded Instructions in Embodied AI Tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Grazioso</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Suglia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AlanaAI</institution>
          ,
          <addr-line>Edinburgh</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Heriot-Watt University</institution>
          ,
          <addr-line>Scotland</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Interdepartmental Research Center Urban/Eco, University of Naples Federico II</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Logogramma s.r.l</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Thanks to Deep Learning models able to learn from Internet-scale corpora, we observed tremendous advances in both text-only and multi-modal tasks such as question answering and image captioning. However, real-world tasks require agents that are embodied in the environment and can collaborate with humans by following language instructions. In this work, we focus on ALFRED, a large-scale instruction-following dataset proposed to develop artificial agents that can execute both navigation and manipulation actions in 3D simulated environments. We present a new Natural Language Understanding component for Embodied Agents as well as an in-depth error analysis of the model failures for this challenge, going beyond the success-rate performance that has been driving progress on this benchmark. Furthermore, we provide the research community with important directions for future work in this field which are essential to develop collaborative embodied agents.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;embodied AI</kwd>
        <kwd>situated interaction</kwd>
        <kwd>visual grounding</kwd>
        <kwd>deep learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1. Introduction
itates the study of both situated language understanding
as well as visual memory, commonsense reasoning, as
In recent years, we experienced tremendous improve- well as long-term action planning.
ments in Natural Language Understanding (NLU) tasks So far, progress on ALFRED has been driven by
thanks to powerful Large Language Models (e.g., [1, 2, 3]). accuracy-based metrics on the oficial leaderboard (e.g.,
These models are trained by leveraging internet-scale tex- [7, 8, 9, 10]). However, considering that the success rate
tual data. However, by having access to text only, they on this benchmark is still below production-level
perleverage only a part of the rich multi-modal training data formance (∼ 40%), this calls for a more in-depth analysis
that can be derived from interaction with the world and of model failures. In this paper, we provide two main
with other agents [4]. Embodied Artificial Intelligence contributions: 1) we train a novel Natural Language
Un(EAI) is the field of AI that aims at developing agents that derstanding component for an EAI agent trained using
can perceive the environment via multi-modal inputs, multi-task learning that has a 0.117 error rate on the
and that can execute actions in the world. validation unseen of ALFRED, an improvement over the</p>
      <p>Many benchmarks have been proposed so far in EAI. one proposed by Min et al.[7]; 2) we provide an in-depth
For instance, Vision+Language Navigation [5] aims at analysis of our model’s failures highlighting lack of
imstudying the capabilities of EAI agents to follow natu- portant situated language understanding capabilities that
ral language instruction in 3D simulated environments. are key for an EAI agent such as referential expression
However, the agent can only output navigation actions resolution, and conversational grounding [11].
limiting the richness of concepts that the agent can learn.</p>
      <p>To simulate a scenario that is closer to the real-world
usage of these systems, Shridhar et al. [6] proposes AL- 2. The ALFRED dataset
FRED, a new instruction-following benchmark that
facil</p>
    </sec>
    <sec id="sec-2">
      <title>In this study, we use ALFRED [6], a benchmark aimed</title>
      <p>CLiC-it 2023: 9th Italian Conference on Computational Linguistics, at assessing the ability of embodied agents to learn from
Nov 30 — Dec 02, 2023, Venice, Italy natural language instructions and egocentric vision to
* Corresponding author. generate sequences of actions for household tasks. The
† These authors contributed equally. ALFRED dataset comprises 25,743 human-annotated
lan($A.mSuargcloia.g)razioso@unina.it (M. Grazioso); a.suglia@hw.ac.uk guage directives corresponding to 8,055 expert
demon0000-0002-4056-544X (M. Grazioso); 0000-0002-3177-5197 stration episodes. Each directive includes a high-level
(A. Suglia) goal and a set of step-by-step instructions. Directives fall
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License under one of the following seven tasks parameterised
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org)
by 84 object classes in 120 scenes: pick and place, stack mentation casts instruction understanding as a
classificaand place, pick two and place, clean and place, heat and tion task and fine-tunes one BERT-based model [ 13] for
place, cool and place, and examine in light. In addition to each of the following tasks: 1) task classification : the
the task category, for each instruction, the dataset also instruction sequence is classified into one of the seven
provides a few relevant annotations: 1) target object: task categories; 2) target classification : the
instructhe principal object involved in the interaction; 2) parent tion sequence is classified into one of the allowed target
object: the final destination of the target object (sink, objects; 3) movable receptacle classification : the
incounter, and similar); 3) movable receptacle: a movable struction sequence is classified into one of the allowed
object containing the target one, e.g. a spoon in a mug; movable receptacle objects; 4) parent classification : the
4) slicing action: true or false respectively if the target instruction sequence is classified into one of the allowed
object must be cut or not; 5) toggle target, indicating an parent objects; 5) slicing classification : the instruction
object to toggle on/of (oven, microwave, etc.). sequence is binary classified to be a slicing/non-slicing</p>
      <p>To better estimate the models’ capability to generalise action.
to new environments, the validation set is composed of However, training five diferent models each one
havtwo subsets called seen and unseen, respectively. In the ing its own BERT encoder can be suboptimal because 1)
former, the agent has to complete tasks in rooms/scenar- it has a high computational cost, and 2) it does not
conios that have already been seen during training, while the sider the semantic relationships between a task and the
latter provides examples in unseen scenarios to assess objects required to solve that task. To take advantage of
the ability of the agent to generalise. these relationships, we implemented a multi-task model
by fine-tuning a single BERT encoder on all the tasks [ 14]
(see Figure 1). As shown in Table 2, thanks to this
multi3. Baseline model and error task setup, our model obtains an improvement in the
analysis overall language understanding performance measured
using the error rate which we define as the proportion of
We implemented the solution proposed in the FILM paper examples for which no mistakes are made (i.e., neither on
[7] which is the base approach for many of the state-of- the high-level task nor on the single slots). Additionally,
the-art models for ALFRED (e.g., [8]). FILM is a modular we report in Table 1, our model performance on specific
architecture that is composed of a trained natural lan- high-level tasks measured using F1-score.
guage component that derives a semantic representation Despite the superior performance of our multitask
represented in terms of intents and slot values akin to con- model, the capabilities of this model were still limited.
ventional NLU systems (e.g., [12]). This representation Therefore, we performed a manual error analysis based
is in turn converted into an action plan by a rule-based on 821 instructions from the validation unseen split of
component. the ALFRED dataset. Particularly, as shown in Table 3,</p>
      <p>In this paper, we focus on improving the language un- the most common errors are about a wrong object
classiderstanding component of FILM, which is essential for ifcation and, even if the object was correctly classified, a
instruction interpretation. Concretely, the FILM
imple</p>
    </sec>
    <sec id="sec-3">
      <title>Thanks to our error analysis, we derive that an embod</title>
      <p>ied agent faces several challenges when fusing multiple
modalities. Moreover, it must take care of the basic
conTable 1 cepts of human-to-human communication [11].
Comparison between our model scores and FILM’s model In this context, the agent’s reasoning can be seen as a
scores in task classification on unseen validation set. sequential process in which it implements a set of
strategies to follow the current language instruction. An
emModel Language processing error rate bodied agent must rely on visual context, commonsense
FILM 0.196 knowledge, and interactive skills. For instance, when
Ours 0.117 the user asks for an object, e.g., soap, the agent must
be able to understand that “soap", “soap bar" and “soap
Table 2 bottle" share enough features to define them as similar
Comparison between our model error rate and FILM’s model objects. Additionally, it should take advantage of the
error rate on all language processing tasks on unseen valida- visual context to resolve ambiguities (if the only soap
tion set. in the agent’s field of view is a soap bar, this should be
the target). Therefore, multi-modal information becomes
Error type Subtype Rate crucial to understanding visually grounded instructions,
Referential ambiguity UMnisdmerastpcehciinfgication 2440//882211 going from spatial language instructions to multi-modal</p>
      <p>Others 32/821 input ones [15]. Finally, if no other strategy resulted in a
Target object search SOpbajteicatl nuontdveirssitbalneding 110761//882211 isnogluctliaornific, aittioshnosutrldateagskiesfo[r
1h6u]m.Faunrtihnetermrvoernet,iionnt,ege.rga.tiunsgOthers interaction errors 218/821 commonsense knowledge can result in better
interpretation (e.g., by leveraging knowledge graphs [17]) as well
Table 3 as better action plans by reasoning over pre-conditions
Error rate for each error type derived from our error analysis. and post-conditions of the actions.</p>
      <p>In collaborative tasks [18, 19, 20], agents have to build
common ground to successfully complete their tasks
failure to find it in the environment. Specifically, we can and adapt to new situations [11]. Therefore,
negotiatcategorise errors in two main classes namely referen- ing meanings becomes a fundamental skill that allows
tial ambiguity, and target object search, which we the agent to learn how the user refers to the
environfurther divide into the following classes: ment, and understand user preferences which will lead
mismatching object reference: the user refers to an to a more efective interaction.
object with a non-conventional name or a particular
linguistic form due to visual ambiguity (brown ball or potato
instead of egg). 5. Conclusion
underspecified object reference: the user refers to an
object using a name which could be ambiguous because In this work, we used the ALFRED dataset as a benchmark
not precise enough (typically “soap" is used to refer to a to investigate the language understanding abilities of
soap bar or a soap bottle). state-of-the-art EAI models. We started by improving the
object not found because not visible: this can happen model originally proposed by Min et al. [7] by training
when the target object is contained in other objects (e.g., using multi-task learning and we showed that even by
spoons are contained in drawers). using the new model several issues remain unsolved.
spatial understanding: the user gives nuanced spatial We categorised these problems into diferent classes to
references for the object but the system does not under- facilitate our analysis. This classification led us to the
stand them (e.g., pick up the salt which is inside the cabinet conclusion that an EAI agent must leverage multi-modal
under the cofee machine ). signals, commonsense knowledge, and interaction with</p>
      <p>Finally, we use a third class (others) which includes the user to solve embodied problems in an efective way.
other interaction errors that do not depend on the lan- According to Schlangen [21], situated interaction is a
guage understanding component. direct, purposeful encounter of free and independent but
similar agents. Following this definition, in the ALFRED
tasks there are two diferent agents: a follower and a
leader. The leader is intended as an oracle that provides
instructions in one go without conversing with the fol- [8] Y. Inoue, H. Ohashi, Prompter: Utilizing large
lower. Moreover, the leader assumes that the follower language model prompting for a data eficient
has perfect capabilities to follow the provided instruc- embodied instruction following, arXiv preprint
tions without considering the notion of uncertainty or arXiv:2211.03267 (2022).
potential mistakes. Finally, there is no concept of conver- [9] A. Suglia, Q. Gao, J. Thomason, G. Thattai,
sational grounding intended as a joint activity in which G. Sukhatme, Embodied bert: A transformer model
the two agents have to negotiate meanings that are re- for embodied, language-guided visual task
complequired to solve the task efectively and eficiently. In this tion, arXiv preprint arXiv:2108.04927 (2021).
sense, even if the ALFRED dataset still represents a chal- [10] A. Pashevich, C. Schmid, C. Sun, Episodic
translenging task, it is far from providing a benchmark that former for vision-and-language navigation, in:
Procan be used to develop artificial agents able to collabora- ceedings of the IEEE/CVF International Conference
tively solve tasks using natural language. on Computer Vision, 2021, pp. 15942–15952.
[11] H. H. Clark, S. E. Brennan, Grounding in
communication., in: Perspectives on socially shared
cogReferences nition., American Psychological Association, 1991,
pp. 127–149.
[1] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Ka- [12] Q. Zhu, Z. Zhang, Y. Fang, X. Li, R. Takanobu, J. Li,
plan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sas- B. Peng, J. Gao, X. Zhu, M. Huang, Convlab-2:
try, A. Askell, et al., Language models are few-shot An open-source toolkit for building, evaluating,
learners, Advances in neural information process- and diagnosing dialogue systems, arXiv preprint
ing systems 33 (2020) 1877–1901. arXiv:2002.04793 (2020).
[2] H. Touvron, L. Martin, K. Stone, P. Albert, A. Alma- [13] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT:
hairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, Pre-training of deep bidirectional transformers for
S. Bhosale, et al., Llama 2: Open foundation and fine- language understanding, in: Proceedings of the
tuned chat models, arXiv preprint arXiv:2307.09288 2019 Conference of the North American
Chap(2023). ter of the Association for Computational
Linguis[3] T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, tics: Human Language Technologies, Volume 1
D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, (Long and Short Papers), Association for
ComM. Gallé, et al., Bloom: A 176b-parameter open- putational Linguistics, Minneapolis, Minnesota,
access multilingual language model, arXiv preprint 2019, pp. 4171–4186. URL: https://aclanthology.org/
arXiv:2211.05100 (2022).
[4] Y. Bisk, A. Holtzman, J. Thomason, J. Andreas, [14] XN.1L9i-u1,4P2.3H.deo,iW:1.0C.h1e8n6,5J.3G/avo1,/MN1u9lt-i-1ta4s2k3d. eep
neuY. Bengio, J. Chai, M. Lapata, A. Lazaridou, J. May, ral networks for natural language understanding,
A. Nisnevich, et al., Experience grounds language, in: Proceedings of the 57th Annual Meeting of the
in: Proceedings of the 2020 Conference on Em- Association for Computational Linguistics,
Associpirical Methods in Natural Language Processing ation for Computational Linguistics, Florence, Italy,
(EMNLP), 2020, pp. 8718–8735. 2019, pp. 4487–4496. URL: https://aclanthology.org/
[5] P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson,</p>
      <p>N. Sünderhauf, I. Reid, S. Gould, A. Van Den Hen- [15] PM1.9G- 1ra4z4i1o.sdoo,Ai: 1.S0..P1o8d6d5a3,S/.vB1a/rPra1,9F-.C1u4t4u1g.no,
Natugel, Vision-and-language navigation: Interpreting ral interaction with trafic control cameras through
visually-grounded navigation instructions in real multimodal interfaces, in: International Conference
environments, in: Proceedings of the IEEE confer- on Human-Computer Interaction, Springer, 2021,
ence on computer vision and pattern recognition, pp. 501–515.</p>
      <p>2018, pp. 3674–3683. [16] V. Russo, A. Mancini, M. Grazioso, M. Di Bratto,
[6] M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, Graph-based representations of clarification
strateW. Han, R. Mottaghi, L. Zettlemoyer, D. Fox, AL- gies supporting automatic dialogue management,
FRED: A Benchmark for Interpreting Grounded In- IJCoL. Italian Journal of Computational Linguistics
structions for Everyday Tasks, in: The IEEE Con- 8 (2022).
ference on Computer Vision and Pattern Recogni- [17] A. Origlia, M. Di Bratto, M. Di Maro, S. Mennella,
tion (CVPR), 2020. URL: https://arxiv.org/abs/1912. A multi-source graph representation of the movie
01734. domain for recommendation dialogues analysis, in:
[7] S. Y. Min, D. S. Chaplot, P. Ravikumar, Y. Bisk, Proceedings of the Thirteenth Language Resources
R. Salakhutdinov, Film: Following instruc- and Evaluation Conference, 2022, pp. 1297–1306.
tions in language with modular methods, 2021. [18] H. De Vries, F. Strub, S. Chandar, O. Pietquin,
arXiv:2110.07342. H. Larochelle, A. Courville, Guesswhat?! visual
object discovery through multi-modal dialogue, in:
Proceedings of the IEEE Conference on Computer
Vision and Pattern Recognition, 2017, pp. 5503–
5512.
[19] N. Ilinykh, S. Zarrieß, D. Schlangen, Meet up! a
corpus of joint activity dialogues in a visual
environment, in: Proceedings of the 23rd Workshop
on the Semantics and Pragmatics of Dialogue-Full</p>
      <p>Papers, 2019.
[20] A. Suhr, C. Yan, J. Schluger, S. Yu, H. Khader,</p>
      <p>M. Mouallem, I. Zhang, Y. Artzi, Executing
instructions in situated collaborative interactions, in:
Proceedings of the 2019 Conference on Empirical
Methods in Natural Language Processing and the
9th International Joint Conference on Natural
Language Processing (EMNLP-IJCNLP), 2019, pp. 2119–
2130.
[21] D. Schlangen, What a situated language-using
agent must be able to do: A top-down analysis,
arXiv preprint arXiv:2302.08590 (2023).</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>