<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of Experi</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1111/cdep.12098</article-id>
      <title-group>
        <article-title>Object Permanence in Embodied Agents using the Animal-AI Environment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Konstantinos Voudouris</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Niall Donnelly</string-name>
          <email>mail.niall.donnelly@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Danaja Rutar</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ryan Burnell</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John Burden</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>José Hernández-Orallo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lucy G. Cheke</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Vienna, Austria</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Psychology, University of Cambridge</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Leverhulme Centre for the Future of Intelligence</institution>
          ,
          <addr-line>Cambridge</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Object Permanence, AI Evaluation, Embodied Agents, Animal-AI Environment</institution>
          ,
          <addr-line>Developmental Psychology, Comparative</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>The College of Engineering</institution>
          ,
          <addr-line>Mathematics</addr-line>
          ,
          <institution>and Physical Sciences, University of Exeter</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>VRAIN, Universitat Politècnica de València</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>34</volume>
      <fpage>164</fpage>
      <lpage>176</lpage>
      <abstract>
        <p>CEUR</p>
      </abstract>
      <kwd-group>
        <kwd>Reinforcement Learning systems perform significantly</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Cognition
Object permanence, the understanding and belief that objects continue to exist even when they are not directly observable, is
important for any agent interacting with the world. Psychologists have been studying object permanence in animals for at
least 50 years, and in humans for almost 50 more. In this paper, we apply the methodologies from psychology and cognitive
science to present a novel testbed for evaluating whether artificial agents have object permanence. Built in the Animal-AI
environment, Object-Permanence In Animal-Ai: GEneralisable Test Suites (O-PIAAGETS) improves on other benchmarks for
assessing object permanence in terms of both size and validity. We discuss the layout of O-PIAAGETS and how it can be used
to robustly evaluate OP in embodied agents.</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <sec id="sec-2-1">
        <title>Object Permanence (OP) is the understanding and belief</title>
        <p>
          that objects continue to exist even when they are not
directly observable. In behavioural terms, an agent has
OP when they behave as though objects continue to exist
when they cannot see them. Human adults use OP to
reason about how objects behave and interact in the external
world. Credited as the first to empirically investigate this
capability, Jean Piaget observed how infants develop the
tendency to search for objects that became occluded [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
Piaget’s insights have been extended considerably by
developmental and comparative psychologists, usually in
the visual modality [
          <xref ref-type="bibr" rid="ref2 ref3 ref4">2, 3, 4</xref>
          ], although OP is an amodal
phenomenon [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
objects continue to exist independently of them, with the
same properties. However, when an object reappears,
fore? Object reidentification has been studied in visual
cognition research with adults [6, 7, 8] and primates [9],
and in developmental psychology with infants [10, 11, 12].
(L. G. Cheke)
        </p>
        <p>0000-0001-8453-3557 (K. Voudouris)</p>
      </sec>
      <sec id="sec-2-2">
        <title>Animal-AI Environment [16] for evaluating whether embodied artificial agents have OP: the</title>
        <p>Object-Permanence
in Animal-Ai: GEneralisable Test Suites. O-PIAAGETS is a
novel attempt to use experiments designed for
investigating whether biological agents have OP for AI evaluation.</p>
      </sec>
      <sec id="sec-2-3">
        <title>First, we examine why OP is a challenge for AI research.</title>
      </sec>
      <sec id="sec-2-4">
        <title>Second, we critically review existing OP testbeds. Third,</title>
        <p>Humans and some animals appear to understand that field of robotics. However, the methods for evaluating
what makes us reidentify this as the same object as be- psychologists have been investigating OP in biological
we outline the structure of the test battery and how it [21], or robust inductive and abductive learning
heuriscan be used to robustly investigate whether agents have tics and biases [20, 19, 8]. It is therefore not as simple as
OP. Finally, we discuss how O-PIAAGETS can be used imputing a Principle of Persistence to build AI systems
for evaluation and how it improves on existing testbeds with OP.
in the field.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Background and Motivations</title>
      <sec id="sec-3-1">
        <title>2.2. Existing Evaluation Methods for OP in AI</title>
        <sec id="sec-3-1-1">
          <title>AI researchers, particularly those working on computer</title>
          <p>2.1. The Logical Problem of OP vision, embodied agents, and robotics, are interested in
OP may appear to be a trivial capacity for an agent to have. building AI systems capable of robustly reasoning about
The agent must simply understand that objects continue visual scenes, in a similar way to how humans and
anito exist when they are not directly observable. Indeed, mals do. Researchers have built several evaluation
frameRenée Baillargeon and colleagues [17] hypothesise that works for assessing whether embodied artificial agents
children are born with a Principle of Persistence, which and computer vision systems have OP.
states exactly this [18, 19]. Why, then, can’t we endow AI Lampinen et al. [23] built OP tasks in a 3D Unity
envisystems with such a principle, bias, or heuristic? Can’t ronment. Here, the agent was fixed as it watched three
we simply tell an agent that objects continue existing boxes. Periodically, objects would leap out of the three
when they are occluded? Fields [20, 19] has discussed boxes, simultaneously or sequentially with or without
how the notion of a Principle of Persistence is untenable, a refractory time lag. The agent would then be turned
due to the Frame Problem (FP). away from the boxes, released, and asked to go to the</p>
          <p>
            The FP implies that endowing an agent, biological or box with a particular object. If it chose the correct box, it
artificial, with a principle of persistence is not trivial. It was rewarded, similar to tasks used with human infants
cannot be overcome with a representation as simple as [24] and non-human primates [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ]. Crosby et al. [16]
objects continue to exist even when they aren’t observable. developed a series of 90 OP tests as part of the Animal-AI
In its raw form, the FP demonstrates that when logically Testbed and Olympics, inspired and directly developed by
describing the efects of particular actions on objects in developmental and comparative psychology. Some work
a domain, we must also describe ad nauseam all the non- has been done comparing embodied deep reinforcement
efects of those actions on those objects. As Fields [ 19] learning agents to humans on these tasks. Children aged
says, it amounts to having to describe everything that 6-10, with limited training, significantly outperformed
doesn’t change in the universe as a result of turning of the Deep Reinforcement Learning systems on the OP tasks
fridge (p. 443). In a domain where objects have certain in the Animal-AI Testbed [13], indicating there is room
properties that can change over time, as in all real-world to improve these systems until they reach human-level
scenarios, the FP implies that we can’t simply say that performance. Leibo et al. [25] developed Psychlab for
the objects stay the same over time, without describing probing psychophysical phenomena in Deep
Reinforcewhich properties remain unchanged and when [21]. ment Learning systems using cognitive science methods
          </p>
          <p>
            When an agent can observe everything in a domain, and qualitatively comparing performance with human
and re-update what has and has not changed at every participants, but they did not investigate OP.
timestep, the FP rarely raises any issues. However, when Having OP is not only applicable to embodied agents,
objects become occluded, it becomes important to track but also to passive computer vision systems engaged in
which properties of those objects do and do not change object tracking. The Localisation Annotations
Compoand when, in order to identify other objects as identi- sitional Actions and TEmporal Reasoning (LA-CATER)
cal or diferent. For example, imagine a lion watching a dataset [26] is prominent in computer vision research.
small antelope pass behind some bushes and then seeing LA-CATER contains 14000 video scenes where objects
a large antelope emerge at the other side. It becomes can move in three dimensions, contain, and carry each
useful to know that antelope don’t change size over such other. Several tasks in this dataset happen to behave
time periods, and therefore the smaller antelope contin- similarly to OP experiments used in psychology. For
ues to exist because of the persistence of its size (and example, one task involves an object being occluded by
other) properties. It also becomes useful to know that one of three identical ‘cups’; once occluded, the cups are
the antelope doesn’t change when the lion changes their moved relative to each other. This bears resemblance to
perspective, or occludes the antelope through its own the cup-tasks used in the Primate Cognition Test Battery
actions, an analogue of the Simultaneous Location and [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ] (see Figure 3) or in the Užgiris and Hunt [24] test
Mapping (SLAM) problem in robotics [22]. Overcoming battery for infants. Other benchmark datasets include
the FP either requires sophisticated deductive techniques ParallelDomain (PD) and KITTI [27]. PD is a synthetic
dataset designed to test occlusions in driving scenarios.
          </p>
          <p>It contains 210 photo-realistic driving scenarios in city [31, 32, 33]. LA-CATER and the procedurally generated
environments, from 3 camera angles, creating a dataset of test sets mentioned earlier were generated according to
630. KITTI [28] has 21 labelled videos of real-world city a series of rules, with training, validation, and test sets
scenes, in which cars, pedestrians, and other objects pass divided arbitrarily. The PD and KITTI datasets were
behind each other and become partially or fully occluded, generated and collected non-procedurally, but again, the
a small fraction of the total KITTI dataset [27]. distinction between training and test sets is often
arbi</p>
          <p>
            Piloto et al. [29] directly applied a measurement frame- trary [27].
work innovated in developmental psychology to probe Moving from i.i.d. to o.o.d. test data promotes
robustphysics knowledge in artificial systems, including OP. Vi- ness in AI systems. Developing a testbed for OP in which
olation of Expectation has been used by the neo-Piagetian training and test data are kept distinct means that we can
school of developmental psychology [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ], investigating in- be more certain that AI systems have OP if they perform
fants’ knowledge about the world by determining when successfully, rather than overfitting to the data
distributhey are surprised to see something, violating their ex- tion. This means we can evaluate whether an AI has an
pectations. For example, infants at about 4.5 months tend ability corresponding to OP, rather than a propensity for
to show surprise (by looking more) if an object appears solving some distribution of tasks that require it1 [35, 36].
to change size whilst occluded [10, 11]. Piloto et al. pro- O.o.d. testing enables researchers to have grounds to
cedurally generated 28 3-second videos that emulated a say they are testing for the presence of abilities. However,
small subset of these studies, and used Kullback-Leibler selecting a test distribution must be guided by some
prindivergence as the AI equivalent of looking time. They ciple that tells us why the training and test distributions
demonstrated the utility of this technique for probing are meaningfully related. This takes us to the second
probphysical knowledge in computer vision systems. lem for OP evaluation in AI: that testing lacks internal
          </p>
          <p>In both computer vision and embodied AI, several validity. Developmental and comparative psychologists
methods for detecting when agents have OP have been have developed numerous experimental designs to test
proposed. However, with the exception of the work of for the presence of cognitive abilities in biological agents,
Piloto et al. [29] and Lampinen et al. [23] at DeepMind introducing numerous controls to eliminate alternative
and the Animal-AI Testbed and Olympics, little attention explanations. As a point of reference, let’s take the classic
has been paid to systematically applying the methodolo- A-not-B paradigm for testing OP. Participants are
pregies of psychology to try to understand and evaluate OP sented with an object of interest that is hidden for several
in artificial agents. trials at location A. In the AI context, this amounts to a
training distribution around location A (with variance
2.3. Problems with Current Evaluation corresponding to minor diferences between trials). To
test the participant to find an object of interest at A, true</p>
          <p>
            Frameworks for OP OP understanding as an explanation is conflated with
Two main problems exist with the methods for evaluating other explanations in terms of memorising spatial
locawhether AI has OP. The first problem is that most of tion or returning to a previously rewarding location as
these benchmarks and testbeds use independent-and- infants under 9 months, and many animals, do [
            <xref ref-type="bibr" rid="ref1">1, 37</xref>
            ].
identically-distributed (i.i.d.) test data, meaning testing To eliminate (some of) these explanations, in the test
data is drawn from the same distribution as training data. condition, participants are faced with an object hidden
This especially applies to LA-CATER, PD, and KITTI. The at location B. The testing distribution now includes
obsecond problem is a lack of internal validity. Suficient jects hidden at B, and the relation between the two is
controls to eliminate alternative explanations for certain meaningful in the context of OP, because an agent needs
behaviours are often lacking. OP to solve the task. The logic is that one would only
          </p>
          <p>The problem with i.i.d. testing data is that it is in prin- perform well on training (A-only) and testing (B-only)
ciple impossible to distinguish between an agent that has if one had OP. Of course, there are further alternative
OP and one using problem-irrelevant shortcuts to max- explanations for correct search at locations A and B, such
imise reward, appearing as if they have OP. This means as simply searching where the experimenter’s hand has
that even if we had an agent that genuinely had OP, our just been [24]. So internal validity tends to increase the
evaluation methods limit how certain we can be of that.</p>
          <p>Geirhos et al. [30] argue that an efective measure against
this is to test AIs on out-of-distribution (o.o.d.) test data,
where training data and test data are drawn from
different (but meaningfully related) distributions. This is
related to the notion of transfer tasks in developmental
more diversity in training and test data there is, as they
become mutually controlling.</p>
          <p>Psychologically-inspired testbeds for evaluating OP
in AI systems, such as Piloto et al. [29], Lampinen et al.
[23] and Crosby et al. [16], remain small and so internal
validity remains relatively low. The confluence of low
internal validity in some testbeds and the lack of o.o.d.
testing means that even if an AI system genuinely has OP,
our evaluation frameworks and metrics are not internally
valid enough to show this. In this paper, we propose a
novel large testbed for conducting o.o.d. testing with
high internal validity.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Introducing O-PIAAGETS</title>
      <sec id="sec-4-1">
        <title>In the previous section, we established three things:</title>
        <p>1. OP poses a challenging logical problem. It is not</p>
        <p>trivial for an agent to have OP.
2. Computer vision and embodied agent results
suggest that trained computer architectures solve
tasks involving OP at a level significantly lower
than that of humans.
3. Current benchmarks and testbeds for evaluating
whether AI systems possess OP have limitations
such that even if an AI system had OP, we might
not be able to tell with reasonable certainty.</p>
        <p>and red lava zones), pink ramps, and transparent and
opaque2 blocks and tunnels (see Figure 1). These objects
can be any size, constrained only by the dimensions of
the arena and the fact that two objects can’t occupy the
same location (apart from lava zones). The lights can also
be switched on or of for preset periods of time, removing
all visual information (see Figure 7 for an example).</p>
        <p>Points are gained and lost through contact with
rewards of difering size and significance, and punishments
of difering severity. Obtaining a yellow sphere increases
Our novel testbed, Object-Permanence in Animal-Ai: points. Obtaining a green sphere also increases points
GEneralisable Test Suites (O-PIAAGETS), overcomes the and is episode-ending. Obtaining a red sphere decreases
limitations of other testbeds by applying an out-of- points and is episode ending, as does touching red lava
distribution testing framework on a large internally valid zones. All spheres can be stationary or in motion through
set of tasks adapted from comparative and developmental all three dimensions. Points start at 0 and decrease
linpsychology and visual cognition research. early with each timestep over an episode, creating time</p>
        <p>O-PIAAGETS uses the Animal-AI Environment to gen- pressure and therefore motivation for fast and decisive
erate individual tasks for training and testing, based on action.
theoretical and empirical findings in the psychology
literature. The testbed has an internal structure in which 3.2. Structure of O-PIAAGETS
certain tasks are designed to test certain aspects of OP
understanding. There is also a tailored training
curriculum to ensure out-of-distribution testing, and more direct
comparison between biological and artificial machines.</p>
        <p>This work complements and extend the work of Piloto
et al. [29], Lampinen et al. [23], Crosby et al. [16], and
Voudouris et al. [13].</p>
      </sec>
      <sec id="sec-4-2">
        <title>O-PIAAGETS adapts some tasks from the open-source</title>
        <p>Animal-AI Testbed, but mostly includes new ones. It
currently contains 5000 tasks, divided into four suites,
although it continues to expand as new features are
released for the Animal-AI Environment. There are three
suites which test diferent aspects of OP and one suite
which contains controls for non-OP based explanations.</p>
        <p>The three suites were motivated a priori by Brian Scholl’s
3.1. The Animal-AI Environment [7] exposition of OP research. Here, Scholl reviews work
The Animal-AI Environment [16] is a 3D world with Eu- on OP from across research in psychology, neuroscience,
clidean geometry and Newtonian physics built in Unity and philosophy, arguing that OP appears to be
under[38]. The environment contains several objects, a single pinned by three key cognitive strategies. Humans appear
agent, and a finite number of actions it can perform (move to reason about objects under occlusion as (a) existing on
and rotate in x-z plane). The agent is situated in a square continuous spatiotemporal trajectories, (b) maintaining
arena. The arena can be populated with appetitive (green certain properties, such as size, but not necessarily
othand yellow spheres) and aversive stimuli (red spheres
2Of any RGB colour combination
ers, such as colour, and (c) existing as unified cohesive
wholes. O-PIAAGETS therefore contains a
Spatiotemporal Continuity suite, a Persistence Through
Property Change suite, and a Cohesion suite. Each suite is
subdivided based on the psychology and AI research into
sub-suites testing diferent aspects of the suites. Those
sub-suites are subdivided into experimental paradigms
from the psychology literature. To maintain high
internal validity, each sub-suite has at least 3 experimental
paradigms. These are further divided into tasks which
are specific instantiations of an experimental paradigm as
used in specific experiments. These tasks are composed
of instances, that are procedurally generated variations
of the global structure of the task, such as right and left
versions or versions with goals of diferent sizes or in
diferent positions. Finally, these instances are composed
of variants, which are procedurally generated variations
of the local structure of instances, with changes to the
colours of walls and the starting orientation of the agent.</p>
        <p>In every test below, the objective is simple: maximise
reward. This involves obtaining yellow and green
rewards while avoiding red rewards and ‘lava’, as quickly
as possible.
3.2.1. Spatiotemporal Continuity</p>
        <p>The Spatiotemporal Continuity suite examines how par- a tunnel and come out of the other side (Burke, 1952).
ticipants reason about objects as persisting in the same However, if the second object appears later than expected
spatiotemporal region, given initial starting velocities or on a diferent trajectory, we do not identify it as the
and other interacting objects. This suite is divided into same object [6] (see Figure 4). The Tunnel Efect tasks
two sub-suites: egocentric OP and allocentric OP. enable us to probe where OP ‘breaks’ in the agent in</p>
        <p>Egocentric OP pertains to reasoning about objects per- question, and how it compares to human performance.
sisting when they pass out of view through the actions of In the Tunnel Efect tasks here and below, the agent
the agent. This allows us to evaluate how well an agent is frozen until they have observed the whole scene, so
can learn about the identity and location of objects in they don’t miss the important occlusion events we are
a region while also moving around that region, a vari- probing, eliminating a potential explanation for why an
ant of the SLAM problem in robotics. An example of agent failed on these tasks.
an egocentric OP task is a detour task where a goal is In line with developments of the Animal-AI
Environobservable but inaccessible behind an obstacle. The way ment, we will introduce allocentric OP tasks involving
to obtain it is to detour around the obstacle such that containment in stationary and moving containers, as
the goal is temporarily left out of sight. The logic here is done in the LA-CATER and Lampinen et al. [23] testbeds
that one would only execute the detouring behaviour if discussed earlier.
one believed that the goal would still exist when one has
ifnished detouring (see Figure 2). 3.2.2. Persistence Through Property Change</p>
        <p>
          Allocentric OP pertains to reasoning about objects that
pass out of view not because of the actions of the agent, The second suite of tests extends the Tunnel Efect tasks,
but because they become occluded by another object. investigating which properties of an object must change
The Cup Task in Figure 3 is an example [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. A goal is under occlusion for the post-occlusion object and
prehidden inside a ‘cup’ for some time. To succeed, the agent occlusion to be classified as diferent. Scholl [ 7] reports
would need to search in the correct ‘cup’. that the Tunnel Efect is not disrupted by colour or
        </p>
        <p>The Tunnel Efect paradigm is a second example. An shape change, only size changes, and the spatiotemporal
object passes behind an occluder, and another emerges changes in the previous sub-suite [39, 9, 6, 40]. Wilcox
some time later. If the second object appears as a human and Baillargeon [10] present evidence that the Tunnel
would expect it to, given the first object’s trajectory, we Efect is disrupted by colour, shape, and texture changes.
perceive it as though the first object has gone through O-PIAAGETS permits more control over the timing and
nature of changes, so can be used for empirical study
with humans to investigate these inconsistent results, as
well as to analyse under what conditions OP breaks in
AI agents.</p>
        <p>Currently, this suite only contains one sub-suite,
testing the Tunnel Efect with apparent size change under
occlusion. However, in line with developments in the
Animal-AI Environment, we are building sub-suites for
apparent shape, colour, and pattern change. An example
of a task in which size appears to change is provided in
Figure 5. The post-occlusion object is smaller than the
pre-occlusion object, so the agent must search for two
distinct objects, not just the visible one.
3.2.3. Object Cohesion</p>
        <p>Scholl [7] argues that OP is not disrupted in human adults
when the contours are partially or completely removed left and seek out the large goal behind the wall, or turn
from a visual object representation, so long as size does right and seek out the entirely visible smaller goal. The
not appear to change. Humans, and many animals, as- smaller goal is larger than the hole in the wall, so agents
sume that an object is of constant size [41], even when that compare the number of green pixels visible at one
contour information is partially occluded or completely time without understanding that size remains constant
removed and replaced with point lights [42, 43]. Cur- under (brief) occlusion will make the wrong choice.
rently, this suite contains only one sub-suite, examining
size constancy under partial occlusion. An example of
this is the aperture task in Figure 6, innovated for
OPIAAGETS based on discussion in Scholl [7]. Agents
watch a large green goal roll behind a wall with a small
hole in it. It is then released and given the choice to turn
they possess OP. If they perform poorly on both, then
there is some issue with understanding the environment
or how to interact with it. If they perform well on the OP
tasks but not the controls, then we have counter-intuitive
evidence that OP can be decoupled from other abilities
required to solve tasks in the environment.
3.3.2. Paradigms, Instances, Variants
Within the test suites themselves, two measures have
been taken to increase internal validity. First, each task
has several instances and variants. We have procedurally
generated many versions of the same task that are mirror
images of each other (left/right versions), have rewards
and goals in diferent positions, or use diferent kinds
of occluders. This counterbalanced design allows us to
Figure 6: The Aperture Task. A and B are before and after par- detect when agents are solving tasks through
problemtial occlusion. Parts 1 and 3 are variants of the same instance, irrelevant shortcuts. For example, in the aperture task
wapitehrtduirfeeritnagskw, athllecomloirurrosr. iPmaartg2e iosfaPdairfetr1en.t instance of the in Figure 6.1, an agent with a bias to turning left might
appear to succeed, but would not succeed at the instance
which is a mirror image of this task, as in 6.2. These
instances have many variants, changing the colour (often
3.3. Increasing Internal Validity randomly) and initial orientation of the agent, as seen in
3.3.1. Control Suite 6.3. This allows us to control for policies such as search
behind the grey obstacle, which may be successful in some
The fourth suite is a set of control tests that serve to de- tasks but does not indicate OP.
termine whether agents can solve tests not measuring OP. Second, the inclusion of several experimental
There are two sub-suites here. The first is an introduction paradigms in each sub-suite means they are mutually
to the environment, introducing basic controls and the controlling. The philosophy of science tells us that
objects present in the environment. These tasks allow an no single experiment would be able to diagnose the
agent, human or artificial, to learn which objects increase presence or absence of OP [44, 45], because there are
reward and which objects decrease reward, and which always alternative explanations that could be appealed
objects are inert. Agents that fail some or all the tasks in to. Using several distinct experimental paradigms means
the above three suites might not be failing because they that they can control for each other and help eliminate
lack OP, but because they do not, for example, navigate these alternative explanations. The cup task in Figure 3
towards green rewards or away from red lava, or under- could be solved by a policy of navigating to where the
stand the utility of ramps for movement in the up-down reward was last seen [46], which is not necessarily the
plane. The second sub-suite contains further control tests same as understanding that the object continues to exist
for the OP tasks in the previous three suites. These are even though the agent can’t see it. An adaptation of
tests that do not require OP to be solved, but introduce Chiandetti and Vallortigara’s [47] paradigm controls for
the kind of landscapes and choices an agent might have this (see Figure 7). Here, the agent watches a reward roll
to make. This means we can determine whether poor away from across lava. Then the lights go out, removing
performance on the OP tasks was a result of a lack of OP, visual information for a short period. When the lights
or a lack of understanding of the landscapes those tests go back on, the goal is not visible. However, there is
took place in. Since every task in the test battery will only one place it can be. Going to where the reward was
require other abilities distinct from OP, these controls last seen would end in failure, by touching lava, and the
allow developers to check whether errors are a result of position of the goal before the lights out provides no cue
a lack of OP or a lack of some other ability. These tasks as to whether the agent should go right or left. The use
can either be used in training or for further testing. An of several experimental paradigms in each sub-suite has
example would be Figure 2 but without a grey wall and the efect of reducing the likelihood of confounds that
with a pink ramp the length of the blue platform. This in- we have not foreseen.
creases internal validity, because if agents performs well
on the control task, but not well on the equivalent OP
task, then we have reason to believe that they lack OP. If
they perform well on both, we have reason to believe that</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Evaluating OP using</title>
    </sec>
    <sec id="sec-6">
      <title>O-PIAAGETS</title>
      <sec id="sec-6-1">
        <title>4.1. Out-of-Distribution Testing</title>
        <p>O-PIAAGETS facilitates out-of-distribution testing by
providing a tailored training set using the control suite,
and a separate test set using the three test suites. The
control suite contains tasks where the positions and
orientations of objects is specified and tasks where those
positions are randomly generated, providing in principle
a very large amount of training data that is on a diferent
distribution to the test data.</p>
      </sec>
      <sec id="sec-6-2">
        <title>4.2. Measurement Layouts</title>
        <p>Each variant in O-PIAAGETS is tagged with its position
in the test battery (i.e., what suite, sub-suite,
experimental paradigm, etc., it is a member of) as well as features
such as goal sizes, the abilities an agent might require
in addition to OP to solve it, and the other variants it
controls for. This leads to an incredibly rich dataset for
evaluating agents beyond merely aggregating their score
or success across the test suites. For example,
developers can explore how relevant and irrelevant features of
the tests, such as goal size, occluder colour, or right/left
variants, correlate with performance [48], and use this
to evaluate whether an agent has OP or is using other
policies to solve OP tasks. For example, assuming any
agent interacting with O-PIAAGETS will make errors,
including humans [13], it is important to evaluate how
those errors are distributed. By hypothesis, an agent
with OP will produce random error, uncorrelated with
experimental paradigms, goal sizes, or the colours of
occluders.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>5. Future Directions and</title>
    </sec>
    <sec id="sec-8">
      <title>Conclusions</title>
      <sec id="sec-8-1">
        <title>Using O-PIAAGETS, developers can robustly evaluate</title>
        <p>whether artificial embodied agents have OP using the
methodologies of cognitive science. It improves on other
benchmarks and testbeds in the field both in terms of
its size, internal validity, and ability to detect the
presence of robust and generalisable OP in artificial systems.
O-PIAAGETS is going through the final stages of
development for general release of Version 1.0, including around
5000 tasks using the current Animal-AI Version 3.0.1.
After validation with human participants and the
development of baseline agents to characterise state-of-the-art
performance in O-PIAAGETS, it will be expanded to
include containment tasks, point lights, and shape, colour,
and pattern changes. In its final form, O-PIAAGETS will
provide a comprehensive and robust evaluation
framework for assessing OP in artificial agents.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <sec id="sec-9-1">
        <title>We thank the anonymous reviewers for their comments.</title>
        <p>This work was funded by the Future of Life Institute,
FLI, under grant RFP2-152, EU’s Horizon 2020 research
and innovation programme under grant agreement No.
952215 (TAILOR), US DARPA HR00112120007
(RECoGAI), and an ESRC DTP scholarship to KV, ES/P000738/1.
M. Sanchez, S. Green, A. Gruslys, S. Legg, D. Hass- ris, G. Lever, A. G. Castañeda, C. Beattie, N. C.
abis, M. M. Botvinick, Psychlab: A Psychology Lab- Rabinowitz, A. S. Morcos, A. Ruderman, N.
Sonoratory for Deep Reinforcement Learning Agents, nerat, T. Green, L. Deason, J. Z. Leibo, D. Silver,
2018. URL: http://arxiv.org/abs/1801.08116, number: D. Hassabis, K. Kavukcuoglu, T. Graepel,
HumanarXiv:1801.08116 arXiv:1801.08116 [cs, q-bio]. level performance in 3D multiplayer games with
[26] A. Shamsian, O. Kleinfeld, A. Globerson, population-based reinforcement learning,
SciG. Chechik, Learning Object Permanence ence 364 (2019) 859–865. URL: https://www.science.
from Video, arXiv:2003.10469 [cs] (2020). URL: org/doi/full/10.1126/science.aau6249. doi:10.1126/
http://arxiv.org/abs/2003.10469, arXiv: 2003.10469. science.aau6249, publisher: American
Associa[27] P. Tokmakov, A. Jabri, J. Li, A. Gaidon, Object tion for the Advancement of Science.</p>
        <p>Permanence Emerges in a Random Walk along [35] J. Hernández-Orallo, Evaluation in
artiMemory, arXiv:2204.01784 [cs] (2022). URL: http: ifcial intelligence: from task-oriented to
//arxiv.org/abs/2204.01784, arXiv: 2204.01784. ability-oriented measurement, Artificial
In[28] A. Geiger, P. Lenz, R. Urtasun, Are we ready for telligence Review 48 (2017) 397–447. URL:
autonomous driving? The KITTI vision bench- https://doi.org/10.1007/s10462-016-9505-7.
mark suite, in: 2012 IEEE Conference on doi:10.1007/s10462-016-9505-7.
Computer Vision and Pattern Recognition, 2012, [36] J. Hernández-Orallo, The Measure of All Minds:
pp. 3354–3361. doi:10.1109/CVPR.2012.6248074, Evaluating Natural and Artificial Intelligence,
CamiSSN: 1063-6919. bridge University Press, 2017.
[29] L. Piloto, A. Weinstein, D. TB, A. Ahuja, M. Mirza, [37] E. Triana, R. Pasnak, Object permanence in cats and
G. Wayne, D. Amos, C.-c. Hung, M. Botvinick, dogs, Animal Learning &amp; Behavior 9 (1981) 135–139.
Probing Physics Knowledge Using Tools from URL: http://link.springer.com/10.3758/BF03212035.
Developmental Psychology, Technical Report doi:10.3758/BF03212035.
arXiv:1804.01128, arXiv, 2018. URL: http://arxiv. [38] A. Juliani, V.-P. Berges, E. Teng, A. Cohen, J. Harper,
org/abs/1804.01128, arXiv:1804.01128 [cs] type: ar- C. Elion, C. Goy, Y. Gao, H. Henry, M. Mattar,
ticle. D. Lange, Unity: A General Platform for
Intelli[30] R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, gent Agents, Technical Report arXiv:1809.02627,
W. Brendel, M. Bethge, F. A. Wichmann, Shortcut arXiv, 2020. URL: http://arxiv.org/abs/1809.02627,
learning in deep neural networks, Nature Ma- arXiv:1809.02627 [cs, stat] type: article.
chine Intelligence 2 (2020) 665–673. URL: https: [39] L. Burke, On the Tunnel Efect,
Quar//www.nature.com/articles/s42256-020-00257-z. terly Journal of Experimental Psychology
doi:10.1038/s42256-020-00257-z, number: 11 4 (1952) 121–138. URL: http://journals.</p>
        <p>Publisher: Nature Publishing Group. sagepub.com/doi/10.1080/17470215208416611.
[31] A. Agrawal, D. Batra, D. Parikh, A. Kembhavi, doi:10.1080/17470215208416611.</p>
        <p>Don’t Just Assume; Look and Answer: Overcom- [40] A. Michotte, G. Thines, G. Crabbé, Les complements
ing Priors for Visual Question Answering, in: amodaux des structures perceptives (Amodal
com2018 IEEE/CVF Conference on Computer Vision pletion of perceptual structures), Studia
Psychoand Pattern Recognition, IEEE, Salt Lake City, UT, logica. Publications Universitaires de Louvain.[GV]
2018, pp. 4971–4980. URL: https://ieeexplore.ieee. (1964).
org/document/8578620/. doi:10.1109/CVPR.2018. [41] C. Fields, Trajectory Recognition as the Basis for
00522. Object Individuation: A Functional Model of
Ob[32] M. Crosby, Building Thinking Machines ject File Instantiation and Object-Token Encoding,
by Solving Animal Cognition Tasks, Minds Frontiers in Psychology 2 (2011). URL: https://www.
and Machines 30 (2020) 589–615. URL: https:// frontiersin.org/article/10.3389/fpsyg.2011.00049.
doi.org/10.1007/s11023-020-09535-6. doi:10.1007/ [42] G. Johansson, Configurations in Event Perception
s11023-020-09535-6. (Uppsala, Sweden, Almqvist and Wiksell.
Johans[33] D. Teney, E. Abbasnejad, K. Kafle, R. Shrestha, son, G.(1973). Visual perception of biological
moC. Kanan, A. van den Hengel, On the Value tion and a model for its analysis. Perception and
of Out-of-Distribution Testing: An Example Psychophysics 14 (1950) 201–211.
of Goodhart’ s Law, in: Advances in Neural [43] G. Johansson, Rigidity, Stability, and
MoInformation Processing Systems, volume 33, tion in Perceptual Space, Nordisk Psykologi
Curran Associates, Inc., 2020, pp. 407–417. URL: 10 (1958) 191–202. URL: https://doi.org/10.1080/
https://proceedings.neurips.cc/paper/2020/hash/ 00291463.1958.10780387. doi:10.1080/00291463.
045117b0e0a11a242b9765e79cbf113f-Abstract.html. 1958.10780387, publisher: Routledge _eprint:
[34] M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Mar- https://doi.org/10.1080/00291463.1958.10780387.
[44] C. Buckner, Understanding associative and
cognitive explanations in comparative psychology, The
Routledge handbook of philosophy of animal minds
(2017) 409–419. Publisher: Routledge.
[45] M. Dacey, Evidence in Default: Rejecting
default models of animal minds, The British
Journal for the Philosophy of Science (2021)
714799. URL: https://www.journals.uchicago.edu/
doi/10.1086/714799. doi:10.1086/714799.
[46] I. M. Pepperberg, M. R. Willner, L. B. Gravitz,
Development of Piagetian Object Permanence in a Grey
Parrot (Psittacus erithacus), Journal of Comparative</p>
        <p>Psychology 111 (1997) 22.
[47] C. Chiandetti, G. Vallortigara, Intuitive
physical reasoning about occluded objects by
inexperienced chicks, Proceedings of the Royal
Society B: Biological Sciences 278 (2011) 2621–2627.</p>
        <p>URL: https://royalsocietypublishing.org/doi/full/
10.1098/rspb.2010.2381. doi:10.1098/rspb.2010.</p>
        <p>2381, publisher: Royal Society.
[48] R. Burnell, J. Burden, D. Rutar, K. Voudouris,</p>
        <p>L. Cheke, J. Hernandez-Orallo, Not a Number:
Identifying Instance Features for
CapabilityOriented Evaluation, forthcoming, p. 9. URL:
https://ryanburnell.com/wp-content/uploads/
Burnell-et-al-2022-Not-a-Number.pdf.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Piaget</surname>
          </string-name>
          ,
          <article-title>The Origins of Intelligence In The Child</article-title>
          , Routledge &amp; Kegan Paul, Ltd.,
          <year>1923</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Baillargeon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. S.</given-names>
            <surname>Spelke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wasserman</surname>
          </string-name>
          ,
          <article-title>Object permanence in five-month-old infants</article-title>
          ,
          <source>Cognition</source>
          <volume>20</volume>
          (
          <year>1985</year>
          )
          <fpage>191</fpage>
          -
          <lpage>208</lpage>
          . URL: https://linkinghub.elsevier. com/retrieve/pii/0010027785900083. doi:
          <volume>10</volume>
          .1016/
          <fpage>0010</fpage>
          -
          <lpage>0277</lpage>
          (
          <issue>85</issue>
          )
          <fpage>90008</fpage>
          -
          <lpage>3</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Baillargeon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gertner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>How Do Infants Reason about Physical Events?</article-title>
          , in: U. Goswami (Ed.),
          <source>The Wiley-Blackwell Handbook of Childhood Cognitive Development</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>48</lpage>
          . Publisher: Wiley Online Library.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E.</given-names>
            <surname>Herrmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Call</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. V.</given-names>
            <surname>Hernàndez-Lloreda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tomasello</surname>
          </string-name>
          ,
          <source>Humans Have Evolved Specialized Skills of Social Cognition: The Cultural Intelligence Hypothesis, Science</source>
          <volume>317</volume>
          (
          <year>2007</year>
          )
          <fpage>1360</fpage>
          -
          <lpage>1366</lpage>
          . URL: https: //www.science.org/doi/10.1126/science.1146282. doi:
          <volume>10</volume>
          .1126/science.1146282.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Bremner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Slater</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , Perception of Object Persistence:
          <article-title>The Origins of Object Permanence in Infancy</article-title>
          ,
          <source>Child Development Perspectives</source>
          <volume>9</volume>
          (
          <year>2015</year>
          )
          <fpage>7</fpage>
          -
          <lpage>13</lpage>
          . URL: https://onlinelibrary.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>