<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Manipulation Tasks based on Incremental Demonstrations in a Virtual Environment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giuseppe Rauso</string-name>
          <email>giuseppe.rauso@unina.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Riccardo Caccavale</string-name>
          <email>riccardo.caccavale@unina.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alberto Finzi</string-name>
          <email>alberto.finzi@unina.it</email>
        </contrib>
      </contrib-group>
      <abstract>
        <p>The ability to grasp and manipulate objects is crucial for performing several complex tasks and it is highly desirable to transfer this skill efectively and naturally to robotic systems. Learning by demonstration provides a particularly interesting and promising technique to address this problem, as it allows us to leverage the guidance of the demonstrations provided by an expert to speed up task learning. In this work, we tackle the problem of learning manipulation tasks by demonstration in a virtual environment. Our aim is to develop an incremental, generalizable, and robust method for learning robotic manipulation tasks using a limited number of demonstrations, while assuming minimal information about the objects to be manipulated. The developed method combines imitation learning and reinforcement learning by proposing an incremental approach in which the operator first demonstrates specialized tasks to the robotic system, and subsequently more complex tasks, exploiting the skills learned during the previous phases. The experimental evaluation shows the feasibility and advantage of the proposed method in terms of modularity, low number of demonstrations, and reliability of the trained system.</p>
      </abstract>
      <kwd-group>
        <kwd>Learning from demonstration</kwd>
        <kwd>robot manipulation</kwd>
        <kwd>incremental learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        In this work, we address the problem of learning robotic manipulation tasks by training the
system through demonstrations provided by a human operator in a virtual environment. In
particular, an incremental approach to training is proposed, allowing the operator to first
demonstrate specialized tasks in simple scenarios and then progressively train more complex
and articulated tasks, leveraging the behaviors already trained in previous stages. The objective
is to provide an incremental, generalizable, and robust method for learning robotic manipulation
tasks using a limited number of demonstrations in a virtual environment, assuming restricted
information about the objects to be manipulated. Another relevant aspect of the proposed
approach relies on the use of virtual reality to provide demonstrations to the robotic system in
a safe and simplified manner, in so avoiding the complexity of interacting with a real robotic
system. The proposed method combines reinforcement learning techniques [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] with imitation
learning methods [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to achieve not only accelerated training through demonstration guidance,
but also behavior influenced by the experts’ preferences to leverage their contextual knowledge
for object manipulation. Diferent approaches to learning from demonstration methods in
a virtual environment can be found in the literature. Approaches based on learning from
demonstration have been proposed to guide learning, reducing complexity and improving
eficiency [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] in the presence of more or less complex end-efectors [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. In this context,
experts’ demonstrations are obtained in various manners: through robot teleoperation [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
video demonstrations [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], kinesthetic demonstrations [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ], or through motion capture [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ].
In this work, we focus on demonstrations provided interacting with a virtual manipulator
endowed with a multi-fingered end-efector. A similar problem is addressed in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], where the
authors introduce a technique called Demo Augmented Policy Gradient (DAPG) that incorporates
demonstrations recorded in virtual reality using a motion capture glove, while Behavioral Cloning
(BC) [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ] is employed to support learning; however, here a diferent setup is employed and an
incremental/modular approach is not deployed. An incremental method for training a robotic
manipulator in a virtual environment is presented in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which proposes a combination of
Reinforcement Learning and Generative Adversarial Imitation Learning (GAIL) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. On the
other hand, this approach assumes a two-finger gripper that can only be opened and closed,
therefore the grasp training is diferent from what is considered in the present study where
a multi-fingered end-efector is deployed. Specifically, in this work we consider as a case
study a simulated robotic setup composed of a Barrett WAM Arm manipulator equipped with a
BarrettHand end-efector. As for the virtual environment, the virtual reality headset used for
teleoperating the robots to obtain demonstrations is a Meta Quest 2. The simulated environments
have been developed in Unity 2021.3.16f1, while the training was conducted using its ML-Agents
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] toolkit which seamlessly incorporates the algorithms used in this work and also allows for
a combination of them.
      </p>
    </sec>
    <sec id="sec-3">
      <title>2. Learning manipulation tasks by incremental demonstration</title>
      <p>In the proposed method for learning robotic manipulation tasks, we exploit the expert’s
demonstrations provided in a virtual environment through an incremental approach. Specifically,
we illustrate the approach at work considering two training phases (see Figure 1). In the first
phase, a robotic hand is trained to grasp diferent types of objects from a proximal pose, without
considering the movement of the robotic arm (proximity grasping). In the second phase, the
hand trained during the first phase is exploited to train a more complex task, where the robotic
arm is tasked to reach, grasp, and lift a specific object ( grasp-and-lift ). These two training phases
are therefore sequential, as the result of the first phase is used for the second.</p>
      <sec id="sec-3-1">
        <title>2.1. Proximity Grasping</title>
        <p>
          The environment for this training phase is inspired by the one proposed by [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. In this settings,
the BarrettHand grasper is “suspended” and cannot move in space, i.e., it is not mounted on an
arm or base (see Figure 2). At the beginning of the episode, objects are placed in front of the
hand, and at the end of the episode, a gravity test is applied to verify the grip. In [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], objects
and pre-grasp positions defined in [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] are used. For this work, however, the hand has a fixed
3D position and the objects are cubes, cylinders, and spheres, positioned in front of the agent
(as illustrated in Figure 2) with random position, rotation, scale, and mass defined in reasonable
ranges and suitable for the type of grip. A single episode has a maximum length of 1000 steps.
Gravity for the objects is disabled for 800 steps from the start of the episode. After this interval,
a stability test (or force test) is performed; specifically, similarly to [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], a force with an intensity
VRteleoperation
Environment
        </p>
        <p>Demonstrations
Controler/Learning</p>
        <p>Actions
Reward
Nextstate</p>
        <p>Environment</p>
        <p>
          Trainingand
Testing
equal to gravity acting on the object is applied in four directions: upwards, downwards, to the
right, and to the left. The episode ends with a positive reward if the distance between the hand
and the object never exceeds a certain threshold. Therefore, if the grasp is stable, after the 200
test steps, the object will be at a distance less than the predetermined threshold, and the episode
can be considered successfully completed.
Observations. The agent observes the rotations on the individual rotation axes of the 8
joints of the BarrettHand in reduced space. Additionally, the agent observes a measure of force
applied. The information regarding the object that the agent observes includes the bounding
box (front-top-left and back-bottom-right points, and its rotation), velocity, angular velocity,
and the object type. As hypothesized in [
          <xref ref-type="bibr" rid="ref18 ref20">18, 20</xref>
          ], information about contact with the object
can improve performance and generalize to objects of diferent shapes and sizes. We therefore
define   as a variable with a value of 1 if the sensor detects a touch and 0 otherwise. These
sensors are located only on the fingertips. The recorded value is passed as an observation to
the agent.
        </p>
        <p>Actions. In the proposed setting, the agent can rotate every joint of the robotic hand on
a single axis. The action space is discrete and there are 3 possible actions for each of the 8
articulations: 0 to keep the rotation unchanged, 1 to increase the rotation, 2 to decrease the
rotation. For each joint  ∈ {0, … , 7} , the target angle is increased or decreased starting from the
current target rotation, which is the rotation that the articulation drive aims to reach applying
a force or torque, with a fixed rotation speed. This means that if the agent decides to increase
the target rotation of a joint i, the variation of the angle is fixed and the agent cannot choose
the amount of rotation. The agent can only decide to change the target angle or leave it at the
same value.</p>
        <p>
          Demonstrations. To record demonstrations by directly teleoperating the robotic hand in
simulation, a mapping has been proposed between Oculus hand tracking and the BarrettHand
robotic hand (see Figure 3). In the Unity environment, the operator can directly control the
robotic hand and show the grips for the objects used. Specifically, to record the increments or
decrements of each joint of the robotic hand, the diference between the angle on the rotation
axis of the human hand (tracked by the Oculus headset) and the target angle on the robotic
hand is calculated. The sign of this diference will indicate whether to record an increase, a
decrease, or no variation in the target angle.
Models and Training. A significant aspect of the proposed training method relies on a
specific combination of imitation learning and reinforcement learning techniques. More
precisely, we leverage Proximal Policy Optimization [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] to exploit the environment’s reward,
while simultaneously incorporating the GAIL reward as an intrinsic reward. Additionally, we
perform weight updates of the policy through BC (Behavioral Cloning). This way, our approach
capitalizes on the initial advantage ofered by BC, allowing for the rapid acquisition of a policy
that surpasses random initialization by mimicking the expert’s actions. Subsequently, our goal
is to accumulate significant intrinsic rewards from GAIL in later stages to improve the imitation
of the expert’s behavior. As the training progresses, the influence of BC gradually diminishes
until it converges to a weight of 0 at the conclusion of the training. The environmental reward
in this context serves as a binary signal, indicating either the success or failure of each episode.
For the training described in this section, we recorded only 15 episodes (approximately 3
minutes), specifically 5 episodes per object type. All episodes were successfully completed by the
expert. The model used for the policy is an MLP composed of two fully-connected layers, each
with 128 neurons, and a recurrent layer with LSTM with a hidden state dimensionality of 64
(  _ = 128 , so   _/2 because the initial memory will be divided between the
hidden state and initial cell state and sequence length of 64. As observed in [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ], the use of
memory via LSTM can improve agent performance for object manipulation tasks. The overall
training took about 8 hours on the machine used (i7-9700K and RTX 3070 Ti) for 7 million steps,
but the agent reaches peak performance after approximately 1.3 million steps, which is just one
hour of training.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Grasp-and-lift with manipulator</title>
        <p>In this second stage of learning, our primary objective is to comprehensively train the robotic
system’s full range of motion. Specifically, our aim is to facilitate the acquisition of skillful
hand movements and positioning strategies in close proximity to the target object, enabling
successful grasping. In this context, the robotic hand is attached to the robotic arm, i.e., the
Barrett WAM Arm, positioned deliberately in front of a table designated for object placement.
For this particular experimental scenario, we have opted to exclusively use cylindrical objects.
In this setting, two distinct grasping techniques have been defined, depending on the task: ”top
grasp” for grasping the cylinder from the top and ”side grasp” for grasping it from the side The
environment designed for this second phase, as illustrated in Figure 4, includes a table on which
the cylinder is be placed and the robotic arm positioned in front of it. At the beginning of each
episode, the object is placed on the table in a random position within a defined area, and a scale
for the cylinder is randomly selected within preset intervals. The agent is granted 1000 steps to
position the hand within the designated zone according to the task and activate the grasp Once
the grasp is initiated, the agent becomes ”locked,” meaning that no further progress is made in
the steps for the robotic arm, no decisions are made, and the outcome of the grasp is awaited.
The hand is allotted 800 steps to adjust the finger positions, after which the same 200-step force
test, utilized in the first phase, is activated, this time in the following global directions: right,
left, forward, backward. Furthermore, the cylinder is subjected to gravitational force from the
beginning of the episode and also during the test. In fact, to enhance the test’s robustness, the
hand (while holding the cylinder) is moved upward while the forces are applied. The episode
concludes with a positive reward if the test succeeds, i.e., if the distance between the hand and
the object is less than or equal to  _ throughout or with a zero reward if one of the failure
conditions is met.</p>
        <p>Observations. To move the robotic hand, the agent controls the target position that the
endefector must reach by manipulating the arm joints through inverse kinematics. The agent thus
observes spatial information related to the end-efector, including position, rotation and target
position, wrist rotation (yaw and pitch), hand velocity and hand angular velocity. Furthermore,
the agent observes the type of task (top grasp or side grasp) and the same information about
the object described in the previous section.</p>
        <p>Actions. The displacement of the end-efector, and thus the hand, is achieved through
inverse kinematics. The agent, therefore, controls the target position that the hand must reach.
Additionally, the agent must rotate the wrist about two axes (yaw and pitch) and activate the
grasp to complete tasks. The agent has the following branches of actions at its disposal: the
displacement of the IK target along the x, y, and z axes, the pitch and yaw rotations of the
wrist, and the activation of the grasp. Also in this case, the action space is discrete: actions for
target displacement and yaw and pitch rotations can assume values of 0, 1, and 2 (no change,
increase, decrease). The final action, to activate the grasp, can only take on values of 0 and 1,
representing ”grasp not activated” and ”activate grasp,” respectively.</p>
        <p>Demonstrations. During the demonstration phase, we employ the controller of the Meta
Quest 2 headset to manipulate the target position of the end-efector, activate the grasp using
the controller, and adjust wrist rotation using the stick. Given that we operate within a discrete
action space, during the demonstration, we cannot directly record the controller’s position.
Instead, we determine, for each coordinate, whether to increment, decrement, or maintain the
value unchanged based on the controller’s motion. Thus, we calculate the direction vector
 = poscontroller −postarget between the controller’s position and the target position, and the sign
of each coordinate dictates whether to increase, decrease, or leave unchanged the corresponding
coordinate of the target position. Similarly, for pitch and yaw rotations, the motion is discretized,
and actions are recorded by indicating an increment, a decrement, or no change based on the
analog stick’s movement. In this case as well, the movement and rotation speeds are fixed.
Models and Training. Analogously to the first phase, the algorithm use for training is PPO
in combination with BC and GAIL. In this case, we recorded 30 episodes for demonstrations
(approximately 9 minutes in total), consisting of 15 episodes for top grasp and 15 episodes for
side grasp. Once again, it is worth noting that all of these episodes were successfully completed
by the expert. The neural network model does not utilize LSTM, instead, it has a structure with
3 fully-connected layers, each consisting of 512 neurons. A diferent machine was used for
the training in the second phase (i5-9300H and GTX 1050 Ti). In this setting, the best training
process took approximately 71 hours to complete 50 million steps.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Results and discussion</title>
      <sec id="sec-4-1">
        <title>3.1. Proximity Grasping</title>
        <p>
          In this section, we discuss the performance of proposed method for each training phase.
To evaluate proximity grasping, two tests were conducted. The first test focuses on the same
objects used during the training phase. These objects, characterized by inherent variability and
primitive shapes, can serve as an approximation of more complex objects or object parts. The
second test is carried out using 3 objects from the dataset proposed in [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], i.e., a bottle, a donut,
and a hammer. These objects were chosen for their shapes and the irregularities they exhibit,
which can pose an interesting challenge for the trained model.
        </p>
        <p>As for the first test, we compared the proposed combination of RL, BC, and GAIL with respect
to other ways to combine these algorithms (i.e., RL, RL + BC, RL + GAIL, and RL + BC + GAIL
with diferent parametrizations) showing that the proposed method (and setting) achieves a
significantly higher success rate with respect to the alternatives. In the second test, the agent
achieves a satisfactory success rate (about 75.3% ± 1.5) with the novel dataset. In particular, it
attains a success rate of 66.5% ± 4 for bottles and 71.5% ± 6.9, including the ”awkward” positions,
meaning with the neck of the bottle or the handle of the hammer facing the palm. However,
performance improves (90% ± 2.2) when excluding these unnatural grasp positions for this type
of object.</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Grasp-and-lift with manipulator</title>
        <p>For the second phase, in addition to the success rate, the percentage of successful activations of
the grasp when it is correctly activated has also been calculated. The ratio is then calculated
between the number of episodes completed correctly and the number of grasps activated
correctly. This constitutes an additional quality test for the grasp policy obtained in the first
part and for the hand placement near the object executed by the trained agent in the second
part. Also in this case, the combination of RL+GAIL+BC proves to be the most efective among
those tested, achieving an overall success rate of 84.5% ± 1.3 in the test phase. Specifically, it
achieves an 81% ± 5.5 success rate for the vertical grasp and 87.3% ± 3.4 for the horizontal grasp.
The success rate for correctly activated grasps is at 94.1% ± 0.8. During the experimentation, it
was observed that the top grasp is the most complex. Even when executed successfully, the
applied grasp may not be entirely ”correct” or natural.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion</title>
      <p>In this work, an incremental and modular method for learning object manipulation through
demonstrations in a virtual environment has been presented. Specifically, the proposed
approach addresses the problem of grasping and manipulating objects. The focus was initially
on the problem of proximity grasping before addressing the task of grasp-and-lift. The system
has been designed to allow the operator to interact easily with a simulated robotic system,
thereby avoiding the complexity of interacting with a real robotic system. The defined setup
enables the recording of demonstrations in a simple manner, using cost-efective hardware.
The experimental results illustrate the efectiveness of the proposed approach and how the two
algorithms, BC and GAIL, manage to balance the disadvantages of the two imitation learning
techniques and make the most of their potential when used together.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The research leading to these results has been partially supported by MELODY, grant
P2022XALNS, PRIN 2022 PNRR, European Union - NextGenerationEU; Harmony, grant
101017008, European Union’s Horizon 2020; Inverse, grant 101136067, and euROBIN, grant
101070596, European Union’s Horizon Europe.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Barto</surname>
          </string-name>
          ,
          <article-title>Reinforcement Learning, second edition: An Introduction, Adaptive Computation and Machine Learning series</article-title>
          , MIT Press,
          <year>2018</year>
          . URL: https://books.google.it/ books?id=uWV0DwAAQBAJ.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Verma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Tsang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Imitation learning: Progress, taxonomies</article-title>
          and challenges,
          <year>2022</year>
          . arXiv:
          <volume>2106</volume>
          .
          <fpage>12177</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pathak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <article-title>Third-person visual imitation learning via decoupled hierarchical controller</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1911</year>
          .09676.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sermanet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lynch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chebotar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hsu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Jang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schaal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Levine</surname>
          </string-name>
          ,
          <article-title>Time-contrastive networks: Self-supervised learning from video</article-title>
          ,
          <year>2018</year>
          . arXiv:
          <volume>1704</volume>
          .
          <fpage>06888</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Eppner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Levine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Abbeel</surname>
          </string-name>
          ,
          <article-title>Learning dexterous manipulation for a soft robotic hand from human demonstration</article-title>
          ,
          <year>2017</year>
          . arXiv:
          <volume>1603</volume>
          .
          <fpage>06348</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singhal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rajeswaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Todorov</surname>
          </string-name>
          ,
          <article-title>Learning deep visuomotor policies for dexterous hand manipulation</article-title>
          ,
          <source>in: 2019 International Conference on Robotics and Automation (ICRA)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>3636</fpage>
          -
          <lpage>3643</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICRA.
          <year>2019</year>
          .
          <volume>8794033</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Handa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. V.</given-names>
            <surname>Wyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-W.</given-names>
            <surname>Chao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Birchfield</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ratlif</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fox</surname>
          </string-name>
          ,
          <article-title>Dexpilot: Vision based teleoperation of dexterous robotic hand-arm system</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1910</year>
          .03135.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pinto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <article-title>Multiple interactions made easy (mime): Large scale demonstrations data for imitation</article-title>
          ,
          <year>2018</year>
          . arXiv:
          <year>1810</year>
          .07121.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vecerik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Scholz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Pietquin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Piot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Heess</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rothörl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lampe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Riedmiller</surname>
          </string-name>
          ,
          <article-title>Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards</article-title>
          ,
          <year>2018</year>
          . arXiv:
          <volume>1707</volume>
          .
          <fpage>08817</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Caccavale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Saveriano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Finzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Kinesthetic teaching and attentional supervision of structured tasks in human-robot interaction</article-title>
          ,
          <source>Auton. Robots</source>
          <volume>43</volume>
          (
          <year>2019</year>
          )
          <fpage>1291</fpage>
          -
          <lpage>1307</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rajeswaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          , G. Vezzani,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schulman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Todorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Levine</surname>
          </string-name>
          ,
          <article-title>Learning complex dexterous manipulation with deep reinforcement learning and demonstrations</article-title>
          ,
          <year>2018</year>
          . arXiv:
          <volume>1709</volume>
          .
          <fpage>10087</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Caccavale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Saveriano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Fontanelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ficuciello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Finzi</surname>
          </string-name>
          ,
          <article-title>Imitation learning and attentional supervision of dual-arm structured tasks</article-title>
          ,
          <source>in: 2017 Joint IEEE International Conference on Development and Learning and Epigenetic Robotics</source>
          , ICDL-EpiRob
          <year>2017</year>
          , Lisbon, Portugal,
          <source>September 18-21</source>
          ,
          <year>2017</year>
          , IEEE,
          <year>2017</year>
          , pp.
          <fpage>66</fpage>
          -
          <lpage>71</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Pomerleau</surname>
          </string-name>
          ,
          <string-name>
            <surname>Alvinn:</surname>
          </string-name>
          <article-title>An autonomous land vehicle in a neural network</article-title>
          ,
          <source>in: NIPS</source>
          ,
          <year>1988</year>
          . URL: https://api.semanticscholar.org/CorpusID:18420840.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sammut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hurst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kedzier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Michie</surname>
          </string-name>
          , Learning to fly,
          <source>in: Proceedings of the Ninth International Workshop on Machine Learning</source>
          , ML '
          <fpage>92</fpage>
          , Morgan Kaufmann Publishers Inc., San Francisco, CA, USA,
          <year>1992</year>
          , p.
          <fpage>385</fpage>
          -
          <lpage>393</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kawakami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ishikawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Roxas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Oishi</surname>
          </string-name>
          ,
          <article-title>Learning 6dof grasping using reward-consistent demonstration</article-title>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2103</volume>
          .
          <fpage>12321</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ermon</surname>
          </string-name>
          ,
          <source>Generative adversarial imitation learning</source>
          ,
          <year>2016</year>
          . arXiv:
          <volume>1606</volume>
          .
          <fpage>03476</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Juliani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.-P.</given-names>
            <surname>Berges</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Teng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Harper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Elion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Goy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Henry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mattar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lange</surname>
          </string-name>
          ,
          <article-title>Unity: A general platform for intelligent agents</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <year>1809</year>
          .02627.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>H.</given-names>
            <surname>Merzic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bogdanovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kappler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Righetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bohg</surname>
          </string-name>
          ,
          <article-title>Leveraging contact forces for learning to grasp</article-title>
          ,
          <year>2018</year>
          . arXiv:
          <year>1809</year>
          .07004.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kappler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bohg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schaal</surname>
          </string-name>
          ,
          <article-title>Leveraging big data for grasp planning</article-title>
          ,
          <source>in: 2015 IEEE International Conference on Robotics and Automation (ICRA)</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>4304</fpage>
          -
          <lpage>4311</lpage>
          . doi:
          <volume>10</volume>
          . 1109/ICRA.
          <year>2015</year>
          .
          <volume>7139793</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hermans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fox</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Birchfield</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tremblay</surname>
          </string-name>
          ,
          <article-title>Contextual reinforcement learning of visuo-tactile multi-fingered grasping policies</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1911</year>
          .09233.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Schulman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wolski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Klimov</surname>
          </string-name>
          ,
          <source>Proximal policy optimization algorithms</source>
          ,
          <year>2017</year>
          . arXiv:
          <volume>1707</volume>
          .
          <fpage>06347</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>OpenAI</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Andrychowicz</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chociej</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Jozefowicz</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>McGrew</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Pachocki</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Petron</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Plappert</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Powell</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ray</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Schneider</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sidor</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Tobin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Welinder</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Weng</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Zaremba</surname>
          </string-name>
          ,
          <article-title>Learning dexterous in-hand manipulation</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1808</year>
          .00177.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>