<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Embodied Afordance Grounding using Semantic Simulations and Neural-Symbolic Reasoning: An Overview of the PlayGround Project</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andreas Persson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amy Loutfi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Applied Autonomous Sensor Systems (AASS), Department of Science and Technology, Örebro University</institution>
          ,
          <addr-line>SE-701 82 Örebro</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we present a synopsis of the PlayGround project. Through neural-symbolic learning and reasoning, the PlayGround project assumes that high-level concepts and reasoning processes can be used to advance both symbol grounding and object afordance inference. However, a prerequisite for reasoning about objects and their afordances is integrated object representations that concurrently maintain symbolic values (e.g., high-level concepts), and sub-symbolic features (e.g., spatial aspects of objects). Integrated representations that, preferably, should be based upon neural-symbolic computation such that neural-symbolic models can, subsequently, be used for high-level reasoning processes. Nevertheless, reasoning processes for symbol grounding and afordance inference often require multiple inference steps. Taking inspiration from the cognitive prospects in simulation semantics, the PlayGround project further presumes that these reasoning processes can be simulated by neural rendering complementary to high-level reasoning processes.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Symbol Grounding</kwd>
        <kwd>Semantic World Modeling</kwd>
        <kwd>Afordance Inference</kwd>
        <kwd>Semantic Simulation</kwd>
        <kwd>Neural-Symbolic Reasoning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The meaning of an object goes beyond the symbol used to refer to it. One of the tenets of
intelligence is the ability to solve the symbol grounding problem, also known as the representation
grounding problem. In symbol grounding [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ], it is argued that purely computational symbol
manipulation cannot gain true meaning without reference to an agent’s embodied interaction
with the world. Said diferently, the symbol grounding problem describes the ability to map
words in the language to aspects of the external world. State-of-the-art research can produce
computational models of language and vision that enable us, to some extent, to caption images,
answer questions about images, and generate an image from a natural language description.
Significant progress has been made in each of these tasks. However, one essential task for
robotic systems – and the “holy grail” of language grounding – is to be able to discern impossible
goals from possible ones. For instance, a statement like “the cofee inside the mug” is clearly
discernable from “the mug inside the cofee” . This is what we call afordability inference as what
is discernable is determined by afordances.
      </p>
      <p>1
2
3
????
?</p>
      <p>
        Afordability inference is a task that comes relatively easily to us humans. Our ability to
perform afordability inferences stems from the fact that we are able to index words and phrases
to objects in the world to prototypical symbols of those objects. Once we derive afordances
from those objects, such afordances constrain the way objects can be coherently combined
– they determine what is possible and what is not. Afordability inference is, however, a task
that, in many ways, is challenging, saying the least, for robotic systems, as exemplified in
Figure 1. Cognitive scientists have long advocated that the mechanisms by which afordability
inference is possible are through a process called simulation semantics – the process by which
we understand and reason about utterances by simulating their content, using similar constructs
to both perception and control. Reasoning about afordances can, arguably, proceed at a high
level using symbolic representations. However, sub-symbolic representations are needed to
capture the 3D spatial aspect of objects, as 2D representations are inherently limited in both
capturing adequate dimensional features and features that are invariant to camera motions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>An underlying assumption of the PlayGround project is that high-level concepts and
reasoning processes can be leveraged to improve symbol grounding and afordance inference. This
assumption implies integrated representations, administrated by a symbolic – sub-symbolic
framework, that can cope with both the high-level concepts (symbolic), as well as 3D spatial
aspects of objects (sub-symbolic). Furthermore, this symbolic – sub-symbolic framework should,
intuitively, be based upon neural-symbolic computation such that neural-symbolic models,
subsequently, can be used for high-level reasoning processes. Based on the observation that
reasoning processes for symbol grounding and afordance inference often require multiple
inference steps, a further assumption of the PlayGround project is that these processes can be
simulated (in alignment with the cognitive perspective on simulation semantics). It is, therefore,
natural to develop a simulation framework for reasoning about symbol grounding and
afordances. The simulation framework will use neural rendering to generate possible 3D scenes and
sequences of such scenes, given language statements and instructions. This semantic simulation
renders low-level 3D scenes and is, hence, complementary to the high-level reasoning processes,
supported by the neural-symbolic approach.</p>
      <p>In summary, the overall goal of the PlayGround project is to “[. . . ] contribute novel techniques
for afordance inference and for symbol grounding that are based on 1) an integrated symbolic –
sub-symbolic framework, and 2) a semantic simulation framework.”</p>
    </sec>
    <sec id="sec-2">
      <title>2. Fundamentals and Related Work</title>
      <p>
        Symbol grounding for physically embedded systems has followed several diferent tracks. One
track learns the meaning of words in the sensorimotor space of the robot using neural networks.
Typically, features are extracted from the perceptual data, and the output of the network is
tightly coupled to the exact motor configuration of the robot [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Another approach is to
manually create a symbol system and structures for maintaining the percept-symbol
correspondence [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Semantic perception is a further topic aiming to augment sensor data into semantic
representations. Today, semantic perception is dominated by fundamental topics such as 2D/3D
semantic object recognition and semantic mapping. For example, the task-planning ability of a
robot ultimately presupposes that the symbols that a symbolic planner uses are anchored in the
physical world. Another practical approach with the aim to model semantically meaningful
object representations is semantic world modeling. Initially presented in association with
probabilistic multiple hypothesis anchoring [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], semantic world modeling promotes the modeling
of object structures that captures object properties beyond only numeric properties. Further
explored in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which argued that – unlike multiple hypothesis target tracking – semantic
world modeling should also incorporate specific domain characteristics, e.g., objects can have
features besides location, which makes them distinguishable from each other in general, and
most object states do not change over short periods of time. Symbol grounding and, to some
extent, semantic world modeling are essential for PlayGround. However, PlayGround difers
from this body of literature in robotics by focusing on afordances, neural-symbolic learning
and reasoning, and semantic simulation to determine object feasibility.
      </p>
      <p>
        Learning object afordances has been reported in correlation to both semantic object
recognition [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], as well as computer vision [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], object properties, learned from RGB-D sensory
data, were utilized to identify objects based on natural language queries that contained
appearance and name properties. The work presented in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] promoted, instead, a probabilistic
approach to track the relations between objects and human hand actions to learn the function
of objects. However, in the context of PlayGround, we need to approach the problem
diferently as we approach the problem from the relational setting between all objects (and not just
hand activities) to extract relational afordances. There is also a growing interest in learning
visual concepts from descriptive language in the machine learning community. For example,
network architectures for neural attentions [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], or visually grounded question-answer pairs
[11]. Another interesting architecture is the one for learning disentangled representations
where visual and language features are broken down and learned as separate dimensions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In
PlayGround, we will examine how perceptual systems can learn disentangled representations
between naming objects, observing the actions performed on objects, and generating the efects
of those actions through simulation.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Preliminary and Previous Results</title>
      <p>This section presents a selection of previous results. Results that have paved the way for the
PlayGround project, which we therefore also count as preliminary results of PlayGround.</p>
      <sec id="sec-3-1">
        <title>3.1. Symbolic – Sub-symbolic Framework</title>
        <p>In previous work [12], we have presented ProbAnch – a modular data-driven probabilistic
anchoring framework. The novelty of this framework is the integration of data-driven bottom-up
anchoring [13], together with probabilistic reasoning based on dynamic distributional clauses
(DDC)[14]. This integration allows the ProbAnch framework not only to create and maintain
representations of objects (i.e., anchors) based on perceptual observations (derived from
subsymbolic sensory data), but also to reason about objects in the absence of perceptual inputs
(e.g., in the case of object occlusions), using a combination of logical, probabilistic, and
neuralsymbolic methods. In other words, ProbAnch is a framework for handling semantic world
modeling with an extension for semantic relational object tracking [15, 16], as seen in Figure 2.</p>
        <p>PROBANCH</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Semantic Simulation Framework</title>
        <p>In parallel with the work that resulted in ProbAnch, we have additionally presented initial
work on learning generative image manipulations from language instructions using a semantic
simulation framework [17]. This work has explored whether a perceptual visual system can
simulate human-like cognitive capabilities by training a computational model to predict the
output of actions expressed through language instructions. Using a combination of language
instructions and images pairs of objects before and after state, as the efects of manipulation
actions, the computational model was trained in the settings of a generative adversarial network
(GAN)[18] in order to generate simulated images that visualize the efect of an action on a
given object, i.e., a synthetic generated image that demonstrates the efect of a certain basic
manipulation action (e.g., move, remove, add, and replace). Aiming to bridge the gap between
simulation and the real world, the trained computational model was subsequently tested in
real-world scenarios, as illustrated in Figure 3.</p>
        <p>Language instructions</p>
        <p>Real-world image pairs</p>
        <p>Generated images</p>
        <p>Target images
“replace the green small cube
with a blue big pyramid”
“add a green small cube on
top of red big sphere”</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Future Work and Objectives</title>
      <p>A natural direction of future work within the PlayGround project would be to integrate both
frameworks, as outlined in the previous section. Such integration would allow the symbolic
– sub-symbolic framework to utilize the semantic simulator, in combination with language
instructions, to predict the future whereabouts of objects and thereby inject positional probability
distributions to support the subsequent anchoring of the objects (once the objects are perceived
through sensory observations). However, the data used to train the generative model of the
semantic simulator was neither realistic (image-wise), nor expressive (language-wise). As a
result, the model failed to generate representative images while tested in real-world settings
(as seen by the generated images in Figure 3). Taking inspiration from the CLEVR [19] and
CLEVRER [20] approaches to rendering visual representations together with language questions,
an initial objective of PlayGround is to develop a synthetic generator for generating realistic
synthetic scenarios. This generator should be qualified to generate realistic (and expressive)
scenarios so that the scenarios can be transferred to real-world settings. Essential for the
PlayGround project is that the generator is also incorporating the notion of afordances, both in
terms of afordances given visual representations, as well as afordances in language instructions.
Generated synthetic scenarios can, thereby, be used to advance the development of novel
techniques for both afordance inference and symbol grounding. Given a qualified synthetic
generator, the long-term objectives of PlayGround are, subsequently, to develop integrated
symbolic – sub-symbolic representations for supporting high-level reasoning processes, as well
as a semantic simulator utilizing neural rendering techniques.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>In this paper, we have presented a summary of the PlayGround project. In PlayGround, we
emphasize symbol grounding and afordance inference using neural-symbolic reasoning and
semantic simulation. Inspired by CLEVR [19] and CLEVRER [20], we promote the use of a
synthetic generator for generating scenarios representative of symbol grounding and afordance
inference problems. Based on generated scenarios, the grander ambition of PlayGround is
thereafter to develop both a symbolic – sub-symbolic learning and reasoning framework (i.e.,
a neural-symbolic framework), and a semantic simulator framework. Furthermore, as both
frameworks are tightly connected and likewise intended for symbol grounding and afordance
inference, we expect to additionally be able to exploit synergies between neural-symbolic
reasoning and semantic simulation.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>First of all, we would like to acknowledge the larger consortium of this project that, besides the
authors of this short paper, consists of Prof. Luc De Raedt and two senior researchers, Marjan
Alirezaie and Martin Längkvist. Thank you for all the work on the proposal that realized the
PlayGround project. In addition, this project is funded and supported by the Swedish Research
Council (sv. Vetenskapsrådet), grant number: 2021-05229.
[11] I. Vendrov, R. Kiros, S. Fidler, R. Urtasun, Order-embeddings of images and language, arXiv
preprint arXiv:1511.06361 (2015).
[12] A. Persson, P. M. Zuidberg Dos Martires, L. De Raedt, A. Loutfi, Probanch: a modular
probabilistic anchoring framework, in: Proceedings of the Twenty-Ninth International
Joint Conference on Artificial Intelligence„ International Joint Conferences on Artificial
Intelligence, 2020, pp. 5285–5287.
[13] A. Loutfi, S. Coradeschi, A. Safiotti, Maintaining coherent perceptual information using
anchoring, in: Proc. of the 19th IJCAI Conf., Edinburgh, UK, 2005, pp. 1477–1482.
[14] D. Nitti, T. De Laet, L. De Raedt, Probabilistic logic programming for hybrid relational
domains, Machine Learning 103 (2016) 407–449.
[15] A. Persson, P. Z. Dos Martires, L. De Raedt, A. Loutfi, Semantic relational object tracking,</p>
      <p>IEEE Transactions on Cognitive and Developmental Systems 12 (2020) 84–97.
[16] P. Zuidberg Dos Martires, N. Kumar, A. Persson, A. Loutfi, L. De Raedt, Symbolic learning
and reasoning with noisy data for probabilistic anchoring, Frontiers in Robotics and AI 7
(2020) 100.
[17] M. Längkvist, A. Persson, A. Loutfi, Learning generative image manipulations from
language instructions, in: Concepts in Action: Representation, Learning, and Application
(CARLA 2020), Virtual workshop, September 22-23, 2020, 2020.
[18] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville,
Y. Bengio, Generative adversarial nets, in: Z. Ghahramani, M. Welling, C. Cortes,
N. D. Lawrence, K. Q. Weinberger (Eds.), Advances in Neural Information Processing
Systems 27, Curran Associates, Inc., 2014, pp. 2672–2680. URL: http://papers.nips.cc/paper/
5423-generative-adversarial-nets.pdf.
[19] J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, R. Girshick,
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,
in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017,
pp. 2901–2910.
[20] K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, J. B. Tenenbaum, Clevrer: Collision
events for video representation and reasoning, in: International Conference on Learning
Representations, 2019.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Searle</surname>
          </string-name>
          , Minds, brains, and programs,
          <source>Behavioral and brain sciences 3</source>
          (
          <year>1980</year>
          )
          <fpage>417</fpage>
          -
          <lpage>424</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Harnad</surname>
          </string-name>
          ,
          <article-title>The symbol grounding problem</article-title>
          ,
          <source>Physica D: Nonlinear Phenomena</source>
          <volume>42</volume>
          (
          <year>1990</year>
          )
          <fpage>335</fpage>
          -
          <lpage>346</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kohli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Tenenbaum</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu,</surname>
          </string-name>
          <article-title>The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision</article-title>
          , arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>12584</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Marocco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cangelosi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Belpaeme</surname>
          </string-name>
          ,
          <article-title>Grounding action words in the sensorimotor interaction with the world: experiments with a simulated icub humanoid robot</article-title>
          ,
          <source>Frontiers in neurorobotics 4</source>
          (
          <year>2010</year>
          )
          <article-title>7</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Coradeschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Safiotti</surname>
          </string-name>
          ,
          <article-title>Perceptual anchoring of symbols for action</article-title>
          ,
          <source>IJCAI International Joint Conference on Artificial Intelligence</source>
          (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Elfring</surname>
          </string-name>
          , S. van den Dries, M. van de Molengraft, M. Steinbuch,
          <article-title>Semantic world modeling using probabilistic multiple hypothesis anchoring</article-title>
          ,
          <source>Robotics and Autonomous Systems</source>
          <volume>61</volume>
          (
          <year>2013</year>
          )
          <fpage>95</fpage>
          -
          <lpage>105</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L. L.</given-names>
            <surname>Wong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Kaelbling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lozano-Pérez</surname>
          </string-name>
          ,
          <article-title>Data association for semantic world modeling from partial views</article-title>
          ,
          <source>The International Journal of Robotics Research</source>
          <volume>34</volume>
          (
          <year>2015</year>
          )
          <fpage>1064</fpage>
          -
          <lpage>1082</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fox</surname>
          </string-name>
          ,
          <article-title>Attribute based object identification</article-title>
          ,
          <source>in: Robotics and Automation (ICRA)</source>
          ,
          <source>2013 IEEE International Conference on, IEEE</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>2096</fpage>
          -
          <lpage>2103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kjellström</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kragić</surname>
          </string-name>
          ,
          <article-title>Visual object-action recognition: Inferring object afordances from human demonstration</article-title>
          ,
          <source>Computer Vision and Image Understanding</source>
          <volume>115</volume>
          (
          <year>2011</year>
          )
          <fpage>81</fpage>
          -
          <lpage>90</lpage>
          . doi:http://dx.doi.org/10.1016/j.cviu.
          <year>2010</year>
          .
          <volume>08</volume>
          .002.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-H. Yang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Weakly-supervised disentangling with recurrent transformations for 3d view synthesis</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>28</volume>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>