<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Learning Graspability of Unknown Ob jects via Intrinsic Motivation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ercin Temel</string-name>
          <email>ercintemel@itu.edu.tr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Beata J. Grzyb</string-name>
          <email>beata.grzyb@plymouth.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sanem Sariel</string-name>
          <email>sariel@itu.edu.tr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Artificial Intelligence and Robotics Laboratory Computer Engineering Department, Istanbul Technical University</institution>
          ,
          <addr-line>Istanbul</addr-line>
          ,
          <country country="TR">Turkey</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Centre for Robotics and Neural Systems, Plymouth University</institution>
          ,
          <addr-line>Plymouth</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Interacting with unknown objects, and learning and producing e↵ective grasping procedures in particular, are challenging problems for robots. This paper proposes an intrinsically motivated reinforcement learning mechanism for learning to grasp uknown objects. The mechanism uses frustration to determine when grasping of an object is not possible. The critical threshold of frustration is dynamically regulated by impulsiveness of the robot. Here, the artificial emotions regulate the learning rate according to the current task and performance of the robot. The proposed mechanism is tested in a real world scenario where the robot, using the grasp pairs generated in simulation, has to learn which objects are graspable. The results shows that the robot equipped with frustration and impulsiveness learns faster than the robot with standard action selection strategies providing some evidence that the use of artificial emotions can improve the learning time.</p>
      </abstract>
      <kwd-group>
        <kwd>Reinforcement Learning</kwd>
        <kwd>Intrinsic motivation</kwd>
        <kwd>Grasping unknown objects</kwd>
        <kwd>Frustration</kwd>
        <kwd>Impulsiveness</kwd>
        <kwd>Visual scene representation</kwd>
        <kwd>Vision-based grasping</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Robots need e↵ective grasp procedures to interact with and manipulate unknown
objects. In unstructured environments, challenges arise mainly due to
uncertainties in sensing and control, and lack of prior knowledge and model of objects.
E↵ective learning methods are essential to deal with these challenges. One
classic approach here is to use reinforcement learning (RL) where an agent actively
interacts with an environment and learns from the consequences of its actions,
rather than from being explicitly taught. An agent selects its actions on basis
of its past experiences (exploitation) and also by new choices (exploration). The
goal of an agent is to maximize the global reward, therefore the agent needs to
rely on actions that led to high rewards in the past. However, if the agent is
too greedy and neglects exploration, it might never find the optimal strategy for
the task. Hence, to find the best ways to perform an action they need to find a
balance between exploitation of current knowledge and exploration to discover
new knowledge that might lead to better performance in the future.</p>
      <p>We propose a competence-based approach to reinforcement learning where
exploration and exploitation is balanced while learning to grasp novel objects. In
our approach, the dynamics of balancing between exploration and exploitation
is tightly related to the level of frustration. The failures in obtaining a new
goal may significantly increase the robot’s level of frustration, and push it into
searching new solutions in order to achieve its goal. However, a prolonged state
of frustration, when no solution can been found, will lead to a state of learned
helplessness, and the goal will be marked as unachievable at the current state
(i.e., object not graspable). Simply speaking, an optimal level of frustration
favours more explorative behaviour, whereas low or high level of frustration
favours more exploitative behaviour. Additionally, we dynamically change the
robot’s impulsiveness that influences how fast the robot gets frustrated, and
indirectly how much time it devotes to learning a particular task.</p>
      <p>To demonstrate the advantages of our approach, we compare it with three
other action selection methods: "-greedy algorithm, softmax function with
constant temperature parameter, softmax function with variable temperature
depending on agent’s overall frustration level. The results shows that the robot
equipped with frustration and impulsiveness learns faster than the robot with
standard action selection strategies providing some evidence that the use of
artificial emotions can improve the learning time.</p>
      <p>The rest of the paper is organized as follows. We first present related work in
the area. Then, we give the details of the learning system including visual
processing of objects, the RL framework and the proposed action selection strategies.
In the next section, we present the experimental results and then conclude the
paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Our main focus is on learning graspability of objects. Previously, analytical
methods are proposed for grasping objects [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. These methods use contact
point locations on objects and the gripper, and then find the friction coecients
by tactile sensors to compute force [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. With these data, grasp stability values
or promising grasp positions can be determined. Another approach for grasping
is learning by exploration. In a recent work [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], grasp successes are associated
with 3D object models which can lead algorithms to memorize object grasp
coordination. According to their work, grasping unknown objects is a challenging
problem and it varies in accordance with system complexity. This complexity
depends on the chosen sensors, prior knowledge about environment and scene
configuration. In [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], 2D contours are used for approximating the center of mass
of objects for grasping.
      </p>
      <p>
        In our work, we use reinforcement learning (RL) framework for learning and
incorporate competence-based intrinsic motivation for guidance in search. The
complexity of reinforcement learning is high in terms of the number of
stateaction pairs and the computations needed to determine utility values [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
Approximate policy iteration methods can be used to alleviate this problem based
on sampling [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Imitation learning before reinforcement learning [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] is one of
the methods for decreasing the complexity in RL [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Furthermore, it is also used
for robots learn crucial parameters in movement to accomplish the task.
      </p>
      <p>
        In our work, we use a competence-based approach for intrinsic motivation
for balancing exploration in RL. Frustration level of the robot is taken into
account. We further extend this approach by adopting an adaptive frustration
level depending on a task. Intrinsic motivation is investigated in earlier works.
Lenat [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] propose a system considering ”interestingness” and Schmidhuber
introduce curiosity concept for reinforcement learning [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Uchibe and Doya [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
also consider intrinsic motivation as learning objective. Di↵erent from curiosity
and reward functions, Wong [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] point out that ideal level of frustration is
beneficial for exploration and faster learning. In addition, Baranes and Oudeyer [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
propose competence-based intrinsic motivation for learning. In our work, main
di↵erence is that impulsiveness [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] is adapted into the frustration rate in order
to change the learning rate dynamically based on a task in real world
environment for robots.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Learning to Grasp Unknown Objects</title>
      <p>We propose an intrinsically motivated reinforcement learning system for robots
to learn graspability of unknown objects. The system includes two main phases
for determination of grasp points on objects and experimentation of them in the
real world (Fig. 1). The first phase includes the required methods to determine
candidate grasp point pairs in simulation. Note that a robot arm with a
twofingered end e↵ector is selected as the target platform. For this reason, grasp
points are determined as point pairs. In the second phase of the system, the
grasp points determined in the first phase are experimented in the real world
through reinforcement learning. The following subsections explain the details of
these processes.
3.1</p>
      <sec id="sec-3-1">
        <title>Visual Representation of Objects</title>
        <p>
          In our system, objects are detected in the scene by using an ASUS Xtion Pro
Live RGB-D camera mounted on a linear platform for interpreting the scene for
tabletop manipulation scenarios by a robotic arm. We use a scene interpretation
system that can both recognize known objects and detect unknown objects in
the scene [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. For unknown object detection, Organized Point Cloud
Segmentation with Connected Components algorithm [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] from PCL [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] is used. This
algorithm finds and marks connected pixels coming from the RGB-D camera
3D Point Cloud
From Camera
        </p>
        <p>Transfer Point Cloud</p>
        <p>to the
Simulation Environment
The Simulation Environment
Transfer Candidates</p>
        <p>to the</p>
        <p>Real World
Frustration Based
Decision About
Graspability</p>
        <p>Real World
Experimentation via</p>
        <p>Reinforcement</p>
        <p>Learning</p>
        <p>
          The Real World Environment with the Robot and Objects
and finds the outlier 3D edges by RANdom SAmple Consensus (RANSAC)
algorithm [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Hence, the object’s center of mass and its edges are detected to
be used by the grasp point detection algorithm that finds candidate grasp point
pairs for a two-fingered robotic hand.
3.2
        </p>
        <p>Detection of Candidate Grasp Points in the Simulator
Objects are represented by their center of masses (µ) and 3D edges (H). Then
candidate grasp point pairs (⇢ =[p1, p2]) are determined as in Algorithm 1. In
the algorithm, initially the reference points are determined. The center of mass,
the upside and the bottom side center points are chosen as references. Based on
these points, cross section points coplanar with the reference points and parallel
to the table surface are determined. In the next step, the algorithm detects
the closest point to the reference points on the same planar and draw a line
crossing with reference points and closest to it. The second step is determining
the opposite point to the closest one on the same line. This procedure continues
until all points are tested. The algorithm produces the candidate grasp pairs
(two grasp points with x,y,z values) and orientation of each pair according to
(0,0) point in 2D (x,y) plane. These grasp points are tested in the simulator for
finding out only the feasible ones.</p>
        <p>In Fig. 2, the edges and sample grasp points for six di↵erent objects along
with the number of grasp points are presented.
Algorithm 1 Grasp Point Detection (µ, H)</p>
        <p>Input: Object Center Of Mass µ, Edge Point Cloud H
Output: Grasp Pairs P
Detect maxZ, minZ and C as reference point ref .
for each reference point do
cP oints = findPointsOnTheSamePlane()
mP oint = findClosestPointToReferencePoint(cP oints)
slope =findSlope(mP oint,ref )
for each p 2 cP oints do</p>
        <p>P slope =findSlope(mP oint,p)
if onTheSameLine(P slope,slope) then</p>
        <p>P { p, mP oint }
end if
end for
end for</p>
        <p>Edge Detection</p>
        <p>Edge Detection</p>
        <p>Edge Detection
Simulation
Representation</p>
        <p>Simulation
Representation</p>
        <p>Simulation
Representation
84 Grasp points are detected
48 Grasp points are detected
56 Grasp points are detected
Edge Detection</p>
        <p>Edge Detection</p>
        <p>Edge Detection
Simulation
Representation</p>
        <p>Simulation
Representation</p>
        <p>Simulation
Representation
46 Grasp points are detected
92 Grasp points are detected
116 Grasp points are detected
In the system, the output of the simulation environment is fed to the robotic arm
to apply real-world experimentation. Intrinsic motivation with frustration level
and new proposed impulsiveness method are evaluated to increase the learning
process speed for the robot in order to give up quickly for the objects that are
not graspable.</p>
        <p>
          The main task of the robot is to learn which objects are graspable. We use
a Reinforcement Learning (RL) framework with Q-learning [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] algorithm and
softmax action selection [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] strategy. The state space here are all grasp point
pairs generated during the simulation phase. A general state S is defined as:
S = [µ, ⇢, , !,
        </p>
        <p>Ov]
where, µ is the center of mass of the object, ⇢ is the selected set of two
grasp points ⇢ =[p1, p2], is the grasp orientation, ! is the approach direction
of the gripper and Ov is the 3D translation vector for object during grasp trial.
A collision between the robotic arm and the object may occur when there is a
trajectory error that results in a non-zero vector.</p>
        <p>Actions can be represented as follows,
where, ||Rv|| is the slide amount on the x axis and ! represents the approach
vector to the object of interest.</p>
        <p>
          In our framework, the robot receives the reward value of 10 (Rmax) when
the grasp is successful and 0.1 (Rmin) when the grasp is unsuccessful [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The
Q-values are updated according to Eq. 3.
        </p>
        <p>Q0(s, a) = Q(s, a) + ↵ ⇤ [R + ( ⇤ maxQ(s0, a))
Q(s, a)]
(3)
where, Q0(s, a) the next Q-value for state action pair (s, a), Q0(s, a) is the
current Q-value, ↵ is the learning rate, R is the immediate reward after
performing an action a in state s, is the discount factor, maxQ(s0, a) is the maximum
estimate of optimal future value.</p>
        <p>We investigate four action selection strategies. The first (and the simplest)
one is the "-greedy action selection method (M1). This method most of the
time selects the action with the highest estimated action value, but once in
a while (with a small probability "), selects an action at random, uniformly,
independently of the action-value estimates.</p>
        <p>
          The second one is the SoftMax Action selection (Eq. 4) method (M2)with
constant temperature value [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]:
        </p>
        <p>A = [||Rv||, ! ]
P (a)t =</p>
        <p>eQt(a)/⌧
Pn
b=1 eQt(b)/⌧
(1)
(2)
(4)
where, P (a)t is the probability of selecting an action a at the time step t,
Qt(a) is the value function for an action a, and ⌧ is the positive parameter called
the temperature that controls the stochasticity of a decision. A high value of the
temperature will cause the actions to be almost equiprobable and a low value
will cause a greater di↵erence in selection probability for actions that di↵er in
their value estimates.</p>
        <p>
          The third strategy (M3) also uses the Softmax action selection rule. In this
approach, however, the ⌧ parameter is flexible and changes dynamically in
relation to the robot’s level of frustration and sense of control [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. An optimal
L =
expectancy ⇤ value
        </p>
        <p>Z + (T t)
where expectancy represents the probability of getting the highest estimated
action value (as in the greedy action selection method), value refers to the
expected action reward (here value = Rmax), Z is a constant derived from when
rewards are immediate, indicates agent’s sensitivity to delay (impulsiveness)
and (T t) refers to the delay of the reward in terms of “time reward” minus
“time now”.</p>
        <p>The impulsiveness is main focus of ours to develop interaction with
frustration rate competence based motivation. According to triad ”Frustration
Impulse - Temper”, a person who has high impulsiveness is considered as ”short
tempered” and it means quickly get frustrated so that changes on frustration
level for learning behavior. Our proposal with that, di↵erent values on
impulsiveness directly a↵ect rate of leak, L, on frustration formula so frustration rate
of agent also will be dependent on impulsiveness.</p>
        <p>The robot apart from learning how to grasp an object, also needs to learn
whether the target object is graspable or not. The learning of a selected grasp pair
⇢ and action a finishes when overall frustration level becomes equal or greater
than a certain threshold value. This value is determined based on a tolerance
formula:</p>
        <p>T olerance = e|| Ov||⇤ '
where, ||Ov|| denotes the translation of the object on the table because of
the collision with the end e↵ector and ' the number of trials from the beginning
of learning.</p>
        <p>Additionally, the online learning process may also end when the following
criterium has been met:
level of frustration favours more explorative behavior, whereas low or high level
of frustration leads to a more exploitative behavior. For the purpose of our
simulations, frustration was represented as a simple leaky integrator:
df
dt
=</p>
        <p>L ⇤ f + A0
where, f is the current level of frustration, A0 is the outcome of the action
(success or f ailure) and L is the fixed rate of the ’Leak’.</p>
        <p>
          In Eq. 5 the ’leak’ rate (L) was fixed and kept at value 1 for all
simulations [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Higher values of L cause the frustration rate to increase slower
compared to smaller values of L. That means that the robot with a high value of L
spends more time on exploration and possibly learns faster. Hence, we propose
the forth method (M4) that builds on this method and changes the value of L
dynamically using an expected utilization motivation formula [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]:
        </p>
        <p>F rustrationLimit = e1/p n
where, n refers to the number of grasp pairs.
(5)
(6)
(7)
(8)
3.4</p>
      </sec>
      <sec id="sec-3-2">
        <title>Impulsiveness and Learning Rate</title>
        <p>
          The main focus of the presented work is investigating an e↵ect that
impulsiveness has on frustration level and on learning. The learning rate and the speed
of decision making is an important issue in human-robot interaction [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. For
example, when a robot plays a quick game with a human, it has to learn quickly.
However, when the robot is alone, it can spend relatively more time on
exploring di↵erent states. By changing the impulsiveness, the robot may dynamically
control its level of frustration and therefore the time devoted for learning a
particular task. Hence, the robot could behave di↵erently in di↵erent environments
and for di↵erent tasks.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Results</title>
      <p>As mentioned before, the candidate grasping points are first determined in
simulation, and then transferred to a robotic arm for real-word experimentation.
V-REP simulator is used as the simulator and the Cyton-Veta Robotic 7-DOF
robot arm by Robai (shown in Fig 3) is used as the experimental platform. The
reachability of the arm is about 45 cm. Also in the experiments, we used three
objects of di↵erent size and shape (i.e., a small cubic plastic block, a plastic
bowling pin and a spherical plastic ball). We compare the performance of four
(a) Success
lengthwise
on the block.</p>
      <p>of (b) Success of
grasp transverse grasp
on the pin.</p>
      <p>(c) Failure
lengthwise
on the ball.</p>
      <p>of
grasp
di↵erent action selection methods discussed in the previous section. A high value
of impulsiveness results in a faster increase in a frustration level (in other words,
in a “short-tempered” agent). For comparison reasons, we use here two di↵erent
values of impulsiveness: a low value of 0.01 and a high value of 100. The results
of our experiments support our proposed hypothesis. An agent with low
impulsiveness spends more time on exploration, testing more grasp pair possibilities
than an agent with a higher value of impulsiveness. For demonstration purposes,
we chose three di↵erent objects that vary in their graspability properties: a cube
that is relatively easy to grasp, a plastic bowling pin that is easily graspable
but it is liable of toppling down, and finally, a ball that is not graspable at all.
We compare the decision and learning rate of the robot that uses our proposed
strategy (M4) with the one based only on frustration (M3). Fig. 4 shows robot’s
level of frustration for each learning epoch while the robot was learning how to
grasp the block. The 84 possible grasp pairs generated in simulation were used in
a real world scenario. Since the robot can easily grasp the cube, the frustration
level is kept low and the learning process terminates before it reaches its limit
value, 1.115 (i.e., according to Eq. 8.). In case of the pin (see Fig. 5), the
simulation generated 116 possible grasp pair candidates that were subsequently used
by the robotic arm. Since the pin is quite light, the arm pulls it down for some
grasp pairs. When the pin fells down, the frustration threshold is decreased for
the related grasp pairs according to the Eq. 7. Hence, the robot learns that these
grasp pairs should be eliminated from the set and immediately proceeds to test
another grasp pair. While for some grasp pairs grasping of the pin was possible,
the robot was not able to grasp the ball for any of grasp pairs. The ball was made
of a hard plastic material and quite light, so every robot’s attempt to grasp it
resulted in a ball rolling over on the scene Fig. 3(c). After each trial, the robot’s
tolerance for frustration decreased rapidly resulting in that the robot switches
to another grasp pair. With each failure, the overall frustration level was raising
and quickly exceeded the tolerance threshold (that at the same time was being
decreased). Although 92 grasp pairs were transferred to the real world scenario,
only after a few steps the robot learned that the object is not graspable. Fig.
Fig. 5. Frustration Rate Changes For Pin Grasping with Methods M3 and M4.
7 shows the comparison of the results for all four strategies of action selection.
The frustration-based action selection methods require a lower number of trials
to learn the graspability of the objects compared to the standard softmax action
selection with fixed temperature parameter and "-greedy action selection. The
agent with higher value of impulsiveness performs slightly better than the agent
with low value.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We have presented our intrinsically motivated reinforcement learning system
for learning graspability of novel objects. Intrinsic motivation is provided by
frustration-based action selection methods during learning, and tolerance values
are determined based on impulsiveness of the robot. Our claim is that
impulsiveness can be adjusted based on the task that the robot is executing. We have
analyzed this mechanism on a robotic arm to learn graspability of
di↵erentshaped objects. Our results reveal that the intrinsic motivation helps the robot
learn faster. Furthermore, the decision on graspability is made earlier by taking
impulsiveness into account. Our future work includes extending the experiment
set and investigating impulsiveness parameters in detail for di↵erent domains
with varying time constraints.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgment</title>
      <p>This research is funded by a grant from the Scientific and Technological Research
Council of Turkey (TUBITAK), Grant No. 111E-286. TUBITAK’s support is
gratefully acknowledged. We thank Burak Topal for his contribution for robotic
arm movement and also thank Mehmet Biberci for his e↵ort on vision algorithms.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Baranes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oudeyer</surname>
          </string-name>
          , P.Y.:
          <article-title>Maturationally-constrained competence-based intrinsically motivated learning</article-title>
          .
          <source>In: Development and Learning (ICDL)</source>
          ,
          <year>2010</year>
          IEEE 9th International Conference on. pp.
          <fpage>197</fpage>
          -
          <lpage>203</lpage>
          . IEEE (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Barto</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          :
          <article-title>Reinforcement learning: An introduction</article-title>
          . MIT press (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bicchi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>On the closure properties of robotic grasping</article-title>
          .
          <source>The International Journal of Robotics Research</source>
          <volume>14</volume>
          (
          <issue>4</issue>
          ),
          <fpage>319</fpage>
          -
          <lpage>334</lpage>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Buss</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hashimoto</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>J.B.</given-names>
          </string-name>
          :
          <article-title>Dextrous hand grasping force optimization</article-title>
          .
          <source>Robotics and Automation, IEEE Transactions on 12(3)</source>
          ,
          <fpage>406</fpage>
          -
          <lpage>418</lpage>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Chebotar</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kroemer</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Learning robot tactile sensing for object manipulation</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Detry</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baseski</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popovic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Touati</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kruger</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kroemer</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piater</surname>
          </string-name>
          , J.:
          <article-title>Learning object-specific grasp a↵ordance densities</article-title>
          .
          <source>In: Development and Learning</source>
          ,
          <year>2009</year>
          .
          <article-title>ICDL 2009</article-title>
          . IEEE 8th International Conference on. pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          . IEEE (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dimitrakakis</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lagoudakis</surname>
            ,
            <given-names>M.G.</given-names>
          </string-name>
          :
          <article-title>Rollout sampling approximate policy iteration</article-title>
          .
          <source>Machine Learning</source>
          <volume>72</volume>
          (
          <issue>3</issue>
          ),
          <fpage>157</fpage>
          -
          <lpage>171</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>Y.H.</given-names>
            ,
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <article-title>The synthesis of 3-d form-closure grasps</article-title>
          .
          <source>Robotica</source>
          <volume>18</volume>
          (
          <issue>01</issue>
          ),
          <fpage>51</fpage>
          -
          <lpage>58</lpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Ersen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ozturk</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Biberci</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sariel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yalcin</surname>
          </string-name>
          , H.:
          <article-title>Scene interpretation for lifelong robot learning</article-title>
          .
          <source>In: The 9th International Workshop on Cognitive Robotics (CogRob</source>
          <year>2014</year>
          )
          <article-title>held in conjunction with ECAI-2014</article-title>
          . Prague, Czech Republic (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Grzyb</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boedecker</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Asada</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Del Pobil</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>L.B.</given-names>
          </string-name>
          :
          <article-title>Between frustration and elation: Sense of control regulates the lntrinsic motivation for motor learning</article-title>
          .
          <source>In: Lifelong Learning</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Huebner</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruthotto</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kragic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Minimum volume bounding box decomposition for shape approximation in robot grasping</article-title>
          .
          <source>In: Robotics and Automation</source>
          ,
          <year>2008</year>
          .
          <article-title>ICRA 2008</article-title>
          . IEEE International Conference on. pp.
          <fpage>1628</fpage>
          -
          <lpage>1633</lpage>
          . IEEE (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Kober</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Learning motor primitives for robotics</article-title>
          .
          <source>In: Robotics and Automation</source>
          ,
          <year>2009</year>
          . ICRA'09. IEEE International Conference on. pp.
          <fpage>2112</fpage>
          -
          <lpage>2118</lpage>
          . IEEE (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Lenat</surname>
            ,
            <given-names>D.B.</given-names>
          </string-name>
          :
          <article-title>Am: An artificial intelligence approach to discovery in mathematics as heuristic search</article-title>
          .
          <source>Tech. rep., DTIC Document</source>
          (
          <year>1976</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Peters</surname>
          </string-name>
          :
          <article-title>Machine learning of motor skills for robotics (</article-title>
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Platt</surname>
          </string-name>
          , R.:
          <article-title>Learning grasp strategies composed of contact relative motions</article-title>
          .
          <source>In: Humanoid Robots</source>
          ,
          <year>2007</year>
          7th IEEE-RAS International Conference on. pp.
          <fpage>49</fpage>
          -
          <lpage>56</lpage>
          . IEEE (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Rusu</surname>
            ,
            <given-names>R.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cousins</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>3D is here: Point Cloud Library (PCL)</article-title>
          .
          <source>In: IEEE International Conference on Robotics and Automation (ICRA)</source>
          . Shanghai, China (May 9-13
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Rusu</surname>
            ,
            <given-names>R.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cousins</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>3d is here: Point cloud library (pcl)</article-title>
          .
          <source>In: Robotics and Automation (ICRA)</source>
          ,
          <year>2011</year>
          IEEE International Conference on. pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          . IEEE (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Sauser</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Billard</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          :
          <article-title>Biologically inspired multimodal integration: Interferences in a human-robot interaction game</article-title>
          .
          <source>In: Intelligent Robots and Systems</source>
          , 2006 IEEE/RSJ International Conference on. pp.
          <fpage>5619</fpage>
          -
          <lpage>5624</lpage>
          . IEEE (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>iirgen Schmidhuber</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A possibility for lmplementing curiosity and boredom in model-building neural controllers (</article-title>
          <year>1991</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Steel</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Ko¨nig, C.J.:
          <article-title>Integrating theories of motivation</article-title>
          .
          <source>Academy of Management Review</source>
          <volume>31</volume>
          (
          <issue>4</issue>
          ),
          <fpage>889</fpage>
          -
          <lpage>913</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Trevor</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gedikli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rusu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christensen</surname>
          </string-name>
          , H.:
          <article-title>Ecient organized point cloud segmentation with connected components</article-title>
          .
          <source>Semantic Perception Mapping and Exploration (SPME)</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Uchibe</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doya</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Finding intrinsic rewards by embodied evolution and constrained reinforcement learning</article-title>
          .
          <source>Neural Networks</source>
          <volume>21</volume>
          (
          <issue>10</issue>
          ),
          <fpage>1447</fpage>
          -
          <lpage>1455</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Watkins</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dayan</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>Q-learning</article-title>
          .
          <source>Machine learning 8(3-4)</source>
          ,
          <fpage>279</fpage>
          -
          <lpage>292</lpage>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Wong</surname>
          </string-name>
          , P.T.:
          <article-title>Frustration, exploration, and learning</article-title>
          .
          <source>Canadian Psychological Review/Psychologie canadienne 20(3)</source>
          ,
          <volume>133</volume>
          (
          <year>1979</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>