<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Action Learning and Grounding in Simulated Human-Robot Interactions (Extended Abstract)?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oliver Roesler</string-name>
          <email>oliver@roesler.co.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ann Nowe´</string-name>
          <email>ann.nowe@vub.ac.be</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Artificial Intelligence Lab, Vrije Universiteit Brussel</institution>
          ,
          <addr-line>Brussels</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>3</lpage>
      <abstract>
        <p>Service robots that are employed in human-centered environments with a high degree of complexity, unpredictability, and dynamicity must be able to learn new tasks autonomously, i.e. when only the goal of the task is given. Furthermore, robots must be able to understand natural language instructions to accurately identify the requested tasks, which requires connections between symbols, i.e. words, and their meanings, i.e. percepts. There exist many studies in the literature that investigate action learning or grounding, but few consider both simultaneously. Additionally, action learning studies have been limited to learn a single action while only varying the initial position of the gripper [3,4]. Furthermore, grounding studies were mostly conducted o ine and primarily focused on grounding of object characteristics or spatial concepts [2,1], while conducted action grounding employed simple feature vectors, which cannot be directly translated into motor commands [5]. In this paper, we investigate the possibility of simultaneous action learning and grounding through the combination of reinforcement and cross-situational learning. More specifically, we simulate human-robot interactions during which a human tutor provides instructions and illustrations of the goal states of the corresponding actions. The robot then learns to reach the desired goals taking into account di erent manipulation behaviors for di erent object shapes and grounds the words and detected phrases of the instructions, including synonyms, through obtained percepts.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The employed grounding and action learning system consists of three parts: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
Humanrobot interaction simulation, which generates di erent situations consisting of the initial
gripper and object positions, relative goal positions of the manipulation objects, object
colors, object shapes, and natural language instructions, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) Reinforcement learning
algorithm, which employs Q-learning to learn optimal micro-action patterns for
encountered situations taking into account initial gripper and object positions as well as the
relative goal position of the manipulation object, (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) Cross-situational learning
component, which identifies auxiliary words and phrases, and maps percepts to non-auxiliary
words and phrases in an unsupervised manner by analysing co-occurrences.
? This is an extended abstract of Roesler and Nowe´ [6].
      </p>
      <p>Copyright © 2019 for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0).</p>
    </sec>
    <sec id="sec-2">
      <title>Results</title>
      <p>After about 60,000 situations the reinforcement learner required 1 episode to converge
to the optimal policy, when using a continuously decreasing exploration rate that is
shared across situations. In contrast, when the exploration rate was reset for each
situation, the reinforcement learner required after about 9,000 situations on average 28
episodes. That the agent did not execute the optimal policy immediately in the latter
case, is due to the high exploration rate at the beginning of each situation because it
was reset. Thus, a continuously decreasing exploration rate that is shared across
situations works best for the investigated scenario. The employed CSL algorithm is able to
successfully ground all 39 words used in this study through their corresponding percepts
after about 800 situations. Afterwards the number of correct mappings is constantly 39,
while the number of false mappings oscillates between 0 and 2 because the algorithm
allows a word to be grounded through several percepts to be able to learn homonyms.
The two additional incorrect mappings are for di erent word combinations, depending
on the most recently encountered situations.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>The proposed framework allowed learning of actions through reinforcement learning
as well as identification of auxiliary words and phrases, and grounding of words and
phrases, including synonyms, through cross-situational learning during simulated
humanrobot interactions. In future work, the framework will be extended to handle real shape,
color and preposition percepts obtained with a stereo camera.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taniguchi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taniguchi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>A generative framework for multimodal learning of spatial concepts and object categories: An unsupervised part-of-speech tagging and 3D visual perception based approach</article-title>
          .
          <source>In: IEEE International Conference on Development and Learning and the International Conference on Epigenetic Robotics (ICDL-EpiRob)</source>
          . Lisbon, Portugal (
          <year>September 2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Fontanari</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tikhano</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cangelosi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ilin</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perlovsky</surname>
            ,
            <given-names>L.I.</given-names>
          </string-name>
          :
          <article-title>Cross-situational learning of object-word mapping using neural modeling fields</article-title>
          .
          <source>Neural Networks</source>
          <volume>22</volume>
          (
          <issue>5-6</issue>
          ),
          <fpage>579</fpage>
          -
          <lpage>585</lpage>
          (
          <article-title>July-August</article-title>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Gudimella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Story</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shaker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kong</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brown</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shnayder</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Deep reinforcement learning for dexterous manipulation with concept networks</article-title>
          .
          <source>CoRR</source>
          (
          <year>2017</year>
          ), http://arxiv.org/abs/1709.06977
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Popov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heess</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lillicrap</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hafner</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barth-Maron</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vecerik</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lampe</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tassa</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erez</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riedmiller</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Data-e cient deep reinforcement learning for dexterous manipulation</article-title>
          .
          <source>CoRR</source>
          (
          <year>2017</year>
          ), http://arxiv.org/abs/1704.03073
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Roesler</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taniguchi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayashi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Evaluation of word representations in grounding natural language instructions through computational human-robot interaction</article-title>
          .
          <source>In: Proceedings of the 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI)</source>
          . Daegu, South Korea (
          <year>March 2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Roesler</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nowe</surname>
            <given-names>´</given-names>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Action learning and grounding in simulated human robot interactions. The Knowledge Engineering Review</article-title>
          . (In Press)
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>