<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>World Modeling for Tabletop Object Construction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arda Inceoglu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Melodi Deniz Ozturk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mustafa Ersen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sanem Sariel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>inceoglua</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ozturkm</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ersenm</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>sarielg@itu.edu.tr</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Artificial Intelligence and Robotics Laboratory Istanbul Technical University</institution>
          ,
          <addr-line>Istanbul</addr-line>
          ,
          <country country="TR">Turkey</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In tabletop construction scenarios, robots work with vertically or horizontally stacked object structures. In order to form such structures, they need to recognize and correctly model closely placed objects in such structures. Depending on the robot's point of view and the objects' positions, it is likely that objects closely located or in contact partially occlude each other, and as a result it is not always possible to model object stacks by relying only on object recognition. However, if the objects are added to the construction consecutively, it becomes possible to sequentially build the model of object stacks. In this work, we propose a scene interpretation system to build and maintain a consistent world model for tabletop construction scenarios. To overcome the challenge of modeling object stacks, we extend our previous scene interpretation system with a semi-closed world assumption and by preserving the models of objects in the formed structures even when they are out of sight. Our extension includes the use of spatial object relations, as well as depth-based segmentation results to model not only single objects, but object combinations. In our system, the LINE-MOD algorithm and an enhanced version with HS histograms are used for recognizing objects along with depth-based segmentation for detecting novel objects. We run numerous construction scenarios using building blocks and show that our system can be successfully used for modeling constructed objects.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>Scene interpretation</kwd>
        <kwd>Tabletop object construction</kwd>
        <kwd>Object manipulation</kwd>
        <kwd>World modeling for tabletop manipulation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        In order to achieve given goals, robots often need to interact with various objects. For
successful interaction, before anything else, they need to collect correct information
about the objects in their environment. The required data includes the accurate
properties of objects, such as their size, shape and color, their locations in the world and
necessary inter-object relations. For this purpose, robots use their sensors to gather
observations from the world, which sometimes do not overlap, are not complete, and sometimes
even contradict with each other. Our previous work presents a scene interpretation
system to cope with these challenges for a ground robot [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In this work, we focus
on tabletop object construction scenarios and extend our previous work for modeling
stacked objects during task execution. This is mainly important for continually
monitoring execution against anomalies (e.g., effects of external interventions) or unexpected
outcomes (e.g., inherent failures). New objects should be correctly localized with their
properties interpreted and information on these objects should be maintained against
any changes (e.g., after disappearing from the scene or displacement in any way).
Besides, when symbol grounding is needed for further cognitive skills such as reasoning
and learning, correct identification of objects is a prerequisite.
      </p>
      <p>
        World modeling is especially challenging when objects are in direct contact with
each other either horizontally or vertically (i.e., when they are on top of each other).
In these scenarios, it is likely that they partially occlude one another from the robot’s
point of view or the vision algorithms may fail in recognizing all objects. For example,
consider the scenario where the robot is tasked to build a rectangular prism structure
from a set of cubical blocks in different colors and sizes as in Figure 1. In this figure,
the LINE-MOD algorithm [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is used to recognize textureless objects in 3D by
considering their surface normals and matching existing measurements with their previously
registered templates. However, since the objects are attached to each other, their
boundaries and some of their surfaces can not be distinguished well which results in errors
in recognition of some of the blocks (e.g., only two blocks on the leftmost column,
indicated with the corresponding markers in the figure, are recognized for this scenario).
This problem can be alleviated by reducing the similarity threshold used for matching
templates in the algorithm. However, this may result in false positives. Humans, on the
other hand, intrinsically use their background and default knowledge when faced with
similar problems, incorparating the recent history of events that have led to the current
situation. In this study, we are inspired by this cognitive skill and propose a system to
reach logical conclusions, similarly to humans, about the robots environment.
      </p>
      <p>Given the requirements in modeling objects in construction scenarios, we propose
new advancements over our previous scene interpretation system. The contributions of
this work are three fold. First, the scene interpretation system is made capable of using
observations taken during execution and building the model of a structure incrementally
using both temporal and spatial relations extracted during runtime and prior semantic
rules for handling occlusions. Second, a truth maintenance mechanism is applied to
store the models of occluded objects even if they can not be recognized but to remove
their models when they are believed to disappear from the scene. Third, for the
identification of objects a semi closed-world assumption is applied for symbol grounding.</p>
      <p>The rest of the paper is organized as follows. First, we mention studies related to
world modeling. Then, we describe our scene interpretation system for consistent world
modeling in tabletop block construction scenarios. We then give empirical results of our
system followed by the conclusions.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Several recent studies address the issue of maintaining a world model from the robot’s
visual observations. Nyga, Balint-Benczedi and Beetz (2014) proposed an ensembles of
experts approach based on Markov Logic Networks [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for fusing different aspects of
information coming from different object recognition methods (e.g., LINE-MOD [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
Google Goggles etc.) enabling robots to answer logical queries about different aspects
of recognized objects [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. WIRE [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is a system based on multiple hypothesis anchoring
for robots to maintain semantically rich world models in unstructured and dynamically
changing environments. It relies on multiple model tracking for incorporating prior
knowledge and multiple hypothesis tracking-based data association for consistently
updating the world model using new observations. Another similar study addresses the
data association problem from a different perspective by using clustering-based
approaches instead of multiple hypothesis tracking [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In our previous work, we
presented a temporal scene interpretation system for maintaining a consistent world model
relying on noisy perception outcomes [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Our system uses segmentation outcomes as
well as object recognition outcomes to be able to detect objects without previously
generated recognition models as unknown object candidates, updates the world model by
evaluating these perceptual outcomes temporally, and takes the robot’s field of view
into account during these updates. In this paper, we enhance our scene interpretation
system in the following directions. First, we replace the 2D model of the robot’s field
of view we used for our ground robot with a 3D model which is necessary for tabletop
object manipulation scenarios. Second, we incorporate a semi-closed world assumption
for keeping track of previously encountered objects. Finally, we present enhancements
for the block construction domain.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Perception Sources</title>
      <p>
        The first step in object manipulation by autonomous robots is maintaining a consistent
and up-to-date world model about their environment. For this task, the robot has to
collect visual recognition and detection data to filter out and to reach conclusions. Our
perception system uses LINE-MOD [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], LINE-MOD&amp;HS [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and 3D segmentation [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
algorithms as means for processing 3D sensory data obtained from an on-board ASUS
Xtion Pro RGB-D camera. LINE-MOD is an object recognition algorithm that uses
surface normals of the objects, calculated from the Point Cloud [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] data regarding the
object, to extract object templates. The algorithm then uses these templates with the
sliding windows approach to detect the modelled objects in new scenes. The
LINEMOD&amp;HS algorithm, in turn, augments LINE-MOD to use the HSV histograms of the
objects in order to integrate the use of color information of the objects in recognition.
Additionally, 3D segmentation is used for detecting objects that are either not previously
modelled, or otherwise cannot be recognized in the current scene.
      </p>
      <p>
        The perception sources are implemented as separate processes, where their
recognition/detection results are asynchronous. The Scene Interpreter system combines these
results to create an accurate representation of the world [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>Scene Interpretation for Tabletop Manipulation</title>
      <p>
        Object recognition is not reliable alone for robotic manipulation tasks, since failures in
recognition or detection occur due to noisy sensor measurements, illumination changes,
dynamic environments or other agents and sensors. In order to automatically build
a consistent and up-to-date model about the environment, visual recognition and
detection outcomes should be filtered out and logical conclusions should be reached in
the face of contradictory outputs. Previous work by the same research team includes a
Scene Interpretation system for ground robots working with objects clearly separated in
the horizontal plane [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which forms the foundation of the proposed system. Necessary
deductions about a robot’s environment include a unique id for each object in the
environment, their type, color, size, shape and location properties, as well as the confidence
of the system about these object’s existence in the environment.
      </p>
      <p>The confidence is represented with a value varying between 0 and 100, proportional
to the degree of belief on the corresponding object’s existence. Confidence values are
updated with every new perception outcome. An object’s confidence value increases as
more consistent recognition or detection results arrive regarding the object.</p>
      <p>The observed facts are kept in the Knowledge base (KB) of the robot, which can be
defined as a collection of reached conclusions about objects, their properties, and
interobject relations. The robot’s KB is initialized as empty. During run time, recognized
objects are inserted into the KB and their corresponding confidence values, as well as
properties, are updated with each newly received recognition message. If an object in
the KB does not receive any corresponding recognition message for a period of time,
even though this object is in the robot’s field of view and should be recognized, the
confidence value regarding the object is gradually decreased. If this value reaches zero,
it is believed that the object is no longer in the robot’s environment, and thus it is
removed from the KB.</p>
      <p>Most humanoid robots have the capability of moving their heads around, making
it possible for them to observe more about their environment. As a result, their visual
field of view (FOV) is bounded by the limitations of their cameras. A robot can receive
reliable information about objects only within its FOV and the field of view constraints
should be taken into account when updating object properties. Extending the 2D
definition in the base system, 3D boundaries are empirically determined for an RGB-D
camera where the objects within are expected to be recognized reliably. An example
scenario regarding FOV calculations can be seen in Figure 2.</p>
      <p>Objects in the environment are considered depending on whether they are inside the
camera’s FOV or not. Objects outside the FOV are not expected to be recognized, and
any data regarding them in the KB are kept static until they re-enter the FOV of the
robot, and new perception data are available.</p>
      <p>
        (a)
(b)
After recognition and localization of the objects in the scene, their spatial relations are
determined as in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. These relations are represented as unary or binary predicates such
as onT able(obj1), near(obj1; obj2) and on(obj2; obj1). Consider a scenario where
three blocks are stacked on top each other, assigned ids 1 to 3 from bottom to top. There
are two on relations expected such that on(3; 2) and on(2; 1). Objects in the bottom are
considered as out of field of view and thus they are not expected to be detected. We
make use of this for modeling objects in block construction as visualized in Figure 3.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Symbolic Models of Tabletop Objects</title>
      <p>The first part of the study focuses on symbol grounding problem for objects ids.
Ambiguities in determination of ids can arise in dynamic scenes. Objects may be displaced,
removed from and put back into the scene, or the robot could be mobile and have
localization problems. As a result, an object might be registered with different ids over
time, which prevents creating and executing plans including object manipulation
successfully.</p>
      <p>
        After successfully detecting objects, the world is assumed to be closed (closed world
assumption) for identity resolution tasks. Closed world assumption [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] can be defined
as having complete knowledge about the world, that is, the numbers and the attributes
of all objects are known apriori. However, robots often have partial information about
the world. Even though object attributes are known, objects’ locations may be dynamic
or unknown which requires obtaining extra information from the environment [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
Whereas, in an open world assumption, no prior knowledge about the world is given,
and every object entering to the scene is assumed to be encountered for the first time.
In contrast, we define a semi-closed world assumption in which the robot builds its KB
itself at runtime and does not use any prior knowledge about the scene contents. At
each object detection, the attributes of the object are compared with that of previously
registered objects in the KB. If an object is believed to have been encountered before,
its previous id is used, otherwise a new id is generated. The corresponding algorithm is
given in Algorithm 1.
      </p>
      <p>Data: Detected object attributes
Result: Object Id
foreach object in KB do
if attributes match and object is not in the scene then</p>
      <p>return object:id;
else
end
end
newId generate new id ;
return newId</p>
      <p>Algorithm 1: Algorithm for Semi-Closed World Assumption</p>
    </sec>
    <sec id="sec-6">
      <title>Symbolic Relations among Tabletop Objects</title>
      <p>The second focus of the study is to correctly model closely located objects. Object
recognition in cluttered scenes is still a challenging problem. Object detection success
is low in such scenes due to their placements. An example scenario with six cubical
blocks is given in Figure 4. The system can not distinguish between objects and only
some of the objects can be added to KB. This is the natural result of assumptions of
vision algorithms. LINE-MOD extracts surface normals on visible surfaces and color
gradients around borders. Furthermore, the 3D segmentation algorithm assumes objects
are clearly separable on a supporting plane.</p>
      <p>The first solution attempt to this problem was decreasing the similarity threshold
of the LINE-MOD algorithm between the object templates and real time detections
(See Figure 5). The threshold is set to 95% by default, and it is decreased gradually.
As a result, the system was able to detect the objects and register them into the KB.
The main drawback of the approach is, as the threshold is lowered, the number of false
positives, i.e. the number of misdetections increase. As the threshold reaches 80% and
below it becomes harder to maintain the number of objects in the KB.</p>
      <p>For a similar scenario where the threshold is set to 95%, even though objects are
placed into the scene one by one but very closely, some of the previously recognized
objects are removed from the KB due to lack of recognition messages after some point.
The second proposed approach utilizes the 3D segmentation algorithm. We can rely
on detections of the 3D segmentation algorithm in terms of the existence of an object
in the scene. If the objects are placed into the scene one by one and clear enough to
be recognized, after successfully being added to the KB, the objects can be marked as
detected, if their centroid lies in one of the last segmented point clouds produced by the
3D segmentation algorithm. Thus, the objects are exempted from being removed from
the KB. The proposed algorithm is given in Algorithm 2.</p>
      <p>The segmentation algorithm is used to maintain object stacks of the same level. In
order to increase the level -the height- of the structure, spatial relations among objects,
namely on relations, are employed. When objects are stacked on top of each other,
corresponding on relations are detected between object pairs. In a pair, the bottommost
object is partially occluded, so it is not expected to be detected, which avoids the update
operation on the object and thus the deletion from the KB. In addition, if the topmost
object of a pair is removed from the scene, the corresponding object model and the on
relation are also removed from the KB, and the bottommost object is expected to be
detected again.</p>
      <p>KB initialize empty knowledge base;
upon receive Objects:;
/* Objects : Recognized objects via LINEMOD and LINEMOD&amp;HS */
foreach object in Objects do
if object in KB then</p>
      <p>update object;
else
end</p>
      <p>KB</p>
      <p>add object ;
end
upon receive Segments:;
/* Segments : Segmented point cloud clusters
foreach object in Objects do
if objectcentroid in Segments then
object mark object as detected
*/
end</p>
      <p>end</p>
    </sec>
    <sec id="sec-7">
      <title>Experiments</title>
      <p>Algorithm 2: Maintaining Closely Located Objects
This section describes the experimental setup and presents the obtained results. First,
we present object recognition and registration to KB during run time. Then, id tracking
capabilities of our system under semi-closed world assumption is demonstrated.</p>
      <sec id="sec-7-1">
        <title>Object Registration to the Knowledge Base</title>
        <p>For the first part, block construction scenario is considered. Red, green and blue colored
blocks are placed into the scene sequentially to form horizontal, vertical and diagonal
structures on the same plane. For comparison purposes, experiments are repeated with
and without employing the proposed segmentation based approach. Each time, after
a block is placed, the number of objects registered to the KB is recorded. Each case
is repeated 10 times, and the mean is calculated. Figure 6 shows the comparison for
horizontal, vertical and diagonal structures. Note that, since the results are recorded in
a sequential manner, errors in the previous steps accumulated to oncoming steps.</p>
        <p>The following conclusions can be drawn from the analysis given in Figure 6.
Employing segmentation based approach fairly increases the number of objects registered
to the KB. The best performance is obtained from vertical placement scenario due to the
fact that the last placed object can be correctly isolated from its surroundings and thus,
it is easier for the vision algorithm to recognize. However, in the diagonal placement
scenario using segmentation does not provide much improvement since objects are in
less contact with each other. Horizontal scenario is the most complicated one in terms
of distinguishing between objects, since objects have more contact with each other.
Improvements become clear when the number of objects in the structure is increased.</p>
        <p>In another experiment, blocks are stacked on top of each other one by one to
measure on relation detection success when new layers are introduced. This time, after each
(b)Vertical
Registered Objects in the KB
(c) Diagonal</p>
        <p>Registered Objects in the KB
6
B5
K
e
h
ijttscenbO43
#
e2
g
a
r1
e
vA
0 2
Without Segmentation</p>
        <p>With Segmentation
3#Objects4in the Sce5ne 6
(e)Vertical</p>
        <p>Without Segmentation</p>
        <p>With Segmentation
3#Objects4in the Sce5ne 6
(f) Diagonal
(a) Horizontal
Registered Objects in the KB</p>
        <p>Without Segmentation</p>
        <p>With Segmentation
3#Objects4in the Sce5ne 6
(d) Horizontal
6
B5
K
e
h
ittscnbO4
je3
#
e2
g
a
r1
e
vA
0 2
stacking operation, the number of on relations is recorded. It is expected to detect one
on relation with a stack of two objects, two on relations with a stack of 3 objects and
so on. Success rates of detecting on relations are 100%, 100% and 93.3% for the
number of layers 2,3 and 4 respectively. Due to object recognition failures, success rate is
decreased in the 4th layer and above.</p>
      </sec>
      <sec id="sec-7-2">
        <title>Semi-Closed World Assumption</title>
        <p>The goal of this experiment is to illustrate id tracking capabilities of the system. The
object set contains red, green and blue colored, medium sized cylinders, and small and
large sized blocks. In the model; type, size and color attributes are taken into account.
An example scenario is visualized in Figure 7.</p>
        <p>The KB is initialized by putting all target objects into the scene, and each object
is assigned a unique id. Then, the objects are removed from the scene. Each object is
put back and id assignments are observed. A confusion matrix based on id assignments
is given in Figure 8. Whenever an object could not be matched with the previously
encountered objects registered to KB, a new id is generated for the object. The reason
of mismatches are originated from errors in recognizing objects due to illumination
conditions. It is observed that if an object is failed to match with an object in the initial
object set, and thus attached a new id, the consecutive recognitions are also matched to
this id. Whereas, some objects could not be recognized at all, which are denoted as not
detected in Figure 8.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Conclusion</title>
      <p>We have presented enhancements for our scene interpretation system in order for it to
be used in tabletop manipulation and construction scenarios for cognitive robots. First,
(a)
(b)
(c)
(d)
the 2D model of a ground robot’s field of view was extended to 3D for a humanoid robot
with a moveable head. Then, we introduced a hybrid model of open and closed world
assumptions for keeping track of object ids in case of dislocations and disappearances
&amp; reappearances. This hybrid model is able to keep track of lost objects, while still
allowing new objects to enter the scene. Deductions about object ids are made based
on physical attributes of the objects and without using any kind of prior knowledge.
Finally, for block construction scenarios, we proposed utilizing 3D segmentation on top
of object recognition to maintain objects in the KB when they are in direct contact with
each other and cannot be recognized. Furthermore, we employed spatial relations to
maintain objects that have other objects on top of them and thus to model higher level
structures. Future work includes improving the system to keep track of more
complicated scenarios that include unknown objects, and modifying the system to operate on
a probabilistic framework.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>This research is partly funded by a grant from the Scientific and Technological Research
Council of Turkey (TUBITAK), Grant No. 111E-286.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>M. D. Ozturk</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ersen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Kapotoglu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Koc</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sariel-Talay</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>H.</given-names>
            <surname>Yalcin</surname>
          </string-name>
          , “
          <article-title>Scene interpretation for self-aware cognitive robots,”</article-title>
          <source>in Proceedings of the 2014 AAAI Spring Symposium: Qualitative Representations for Robots</source>
          , pp.
          <fpage>89</fpage>
          -
          <lpage>96</lpage>
          , AAAI Press,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Ersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Ozturk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Biberci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sariel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Yalcin</surname>
          </string-name>
          , “
          <article-title>Scene interpretation for lifelong robot learning</article-title>
          ,
          <source>” in The 9th International Workshop on Cognitive Robotics (CogRob</source>
          <year>2014</year>
          )
          <article-title>held in conjunction with ECAI-</article-title>
          <year>2014</year>
          , (Prague, Czech Republic),
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>S.</given-names>
            <surname>Hinterstoisser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Cagniart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ilic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Sturm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Navab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fua</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Lepetit</surname>
          </string-name>
          , “
          <article-title>Gradient response maps for real-time detection of textureless objects</article-title>
          ,
          <source>” IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          , vol.
          <volume>34</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>876</fpage>
          -
          <lpage>888</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>M.</given-names>
            <surname>Richardson</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Domingos</surname>
          </string-name>
          , “
          <article-title>Markov logic networks</article-title>
          ,
          <source>” Machine Learning</source>
          , vol.
          <volume>62</volume>
          , no.
          <issue>1-2</issue>
          , pp.
          <fpage>107</fpage>
          -
          <lpage>136</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>D.</given-names>
            <surname>Nyga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Balint-Benczedi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Beetz</surname>
          </string-name>
          , “
          <article-title>PR2 looking at things - Ensemble learning for unstructured information processing with markov logic networks</article-title>
          ,
          <source>” in Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA)</source>
          , pp.
          <fpage>3916</fpage>
          -
          <lpage>3923</lpage>
          , IEEE Press,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>J.</given-names>
            <surname>Elfring</surname>
          </string-name>
          , S. van den Dries,
          <string-name>
            <surname>M. J. G. van de Molengraft</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Steinbuch</surname>
          </string-name>
          , “
          <article-title>Semantic world modeling using probabilistic multiple hypothesis anchoring</article-title>
          ,
          <source>” Robotics and Autonomous Systems</source>
          , vol.
          <volume>61</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>95</fpage>
          -
          <lpage>105</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>L. L. S.</given-names>
            <surname>Wong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Kaelbling</surname>
          </string-name>
          , and
          <string-name>
            <surname>T.</surname>
          </string-name>
          Lozano-Pe´rez, “
          <article-title>Data association for semantic world modeling from partial views</article-title>
          ,”
          <source>International Journal of Robotics Research</source>
          , Accepted for publication.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>M.</given-names>
            <surname>Ersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sariel-Talay</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>H.</given-names>
            <surname>Yalcin</surname>
          </string-name>
          , “
          <article-title>Extracting spatial relations among objects for failure detection</article-title>
          ,”
          <source>in Proceedings of the KI 2013 Workshop on Visual and Spatial Cognition</source>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>20</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>A. J. B. Trevor</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gedikli</surname>
            ,
            <given-names>R. B.</given-names>
          </string-name>
          <string-name>
            <surname>Rusu</surname>
            , and
            <given-names>H. I. Christensen</given-names>
          </string-name>
          , “
          <article-title>Efficient organized point cloud segmentation with connected components,”</article-title>
          <source>in Proceedings of the 3rd Workshop on Semantic Perception, Mapping and Exploration (SPME)</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>A.</given-names>
            <surname>Aldoma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Marton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Tombari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wohlkinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zeisl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Rusu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gedikli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Vincze</surname>
          </string-name>
          , “Tutorial:
          <article-title>Point cloud library: Three-dimensional object recognition and 6 DOF pose estimation</article-title>
          ,
          <source>” IEEE Robotics and Automation Magazine</source>
          , vol.
          <volume>19</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>80</fpage>
          -
          <lpage>91</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. R. Reiter, “
          <article-title>Readings in nonmonotonic reasoning,” ch</article-title>
          .
          <source>On Closed World Data Bases</source>
          , pp.
          <fpage>300</fpage>
          -
          <lpage>310</lpage>
          , Morgan Kaufmann Publishers Inc.,
          <year>1987</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>O.</given-names>
            <surname>Etzioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Golden</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Weld</surname>
          </string-name>
          ,
          <article-title>“Sound and efficient closed-world reasoning for planning</article-title>
          ,
          <source>” Artificial Intelligence</source>
          , vol.
          <volume>89</volume>
          , no.
          <issue>12</issue>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>148</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>