<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>OntoScene: Ontology Guided Indoor Scene Understanding for Cognitive Robotic Tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Snehasis Banerjee ?</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pradip Pramanick</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chayan Sarkar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Balamuralidhar P</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>TCS Research, Tata Consultancy Services</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We have developed a robotics ontology, OntoScene, that extends IEEE CORA [5] and SemNav [1] ontologies. Contrary to the prior work that lacked usage of ontology in scene understanding, the proposed system uses OntoScene to figure out objects and their relations in a scene and create a scene graph for aid in various cognitive robotic tasks where object localization, scene graph generation is important. This work positions semantic web technology as a key enabler in robotic tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>Ontology</kwd>
        <kwd>Cognitive Robotics</kwd>
        <kwd>Semantic Scene Processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Background
Scene understanding and scene graph generation are critical components in most
robotic tasks. It is dificult for a robot to semantically understand the world
(sequence of scenes) based only on sensor inputs, typically an RGB-D camera. To
this extent, approaches presented in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] are enhanced to take the help of
querying the OWL ontology to handle various tasks like navigation, description
based object localization, visual question answering, etc. In contrast to the
endto-end learning approaches, symbolic approaches are more reliable, interpretable,
and safe from the perspective of task actuation in human co-occupied spaces.
      </p>
      <p>Perception
(RGBD)
NavigAation,M
ctuation(Motors)</p>
      <p>anipulation
Find the cup to
the left of bot le</p>
      <p>SReemparenstiecnStacteionne ,f(OediissebCfOthcgir(oiuslopCbr_:ooR0oltoneleTrd:o_P)p0iOnfk) onTopOftbael(seisirdCTetlhoaeOalbfoftnlOre:,B_fle0laftcOkonTfo)p,Of on(iTsbo(CoipstoOCllefooc_rl:u1o{p}r):_W1fOtfelhi,fOetdisebe)</p>
      <p>
        Scene Understanding and Scene Graph Generation
As shown in Fig. 1, suppose a user issues an instruction of finding an object
(cup) with specific criteria (to the left of the bottle). The input to the Cognitive
Module is perception from the robot’s ego view (RGB-D camera image scene
sequences), and output is actuation (navigation or manipulation). We use Yolo [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
and DenseCap [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] as the object detection algorithms to locate the objects within
the current scene. When a ‘cup’ is detected, the Cognitive Module queries the
ontology to get a list of properties corresponding to that object entity. The
corresponding algorithms are selected from the Algorithm Repository to extract the
relevant attributes and relations from the scene. Finally, a semantic scene graph
is generated to aid in complex tasks needing semantic object localization.
      </p>
      <p>
        The standard way to summarize scenes is using deep image captioning [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
and object detection. While machine learning models may infer the presence of
a pair of objects, it is dificult to ground an accurate relationship due to the
large space of possible relationships. For example, it is dificult to disambiguate
between ‘table-beside-cup’ and ‘cup-on-table’. Also, such models can not predict
transitive relationships directly. Thereby utilizing the ontology, commonsense
disambiguation becomes possible, by querying whether such an (un)directed edge
predicate exists or not between entities (objects connected by a property).
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Purushothaman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Semnav: How rich semantic knowledge can guide robot navigation in indoor spaces</article-title>
          .
          <source>In: ISWC (Industry)</source>
          . pp.
          <fpage>398</fpage>
          -
          <lpage>400</lpage>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bruno</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , et. al.:
          <article-title>Knowledge representation for culturally competent personal robots</article-title>
          .
          <source>International Journal of Social Robotics</source>
          , Springer pp.
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hossain</surname>
            ,
            <given-names>M.Z.</given-names>
          </string-name>
          , et. al.:
          <article-title>A comprehensive survey of deep learning for image captioning</article-title>
          .
          <source>ACM Computing Surveys</source>
          <volume>51</volume>
          (
          <issue>6</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>36</lpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Johnson</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Karpathy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fei-Fei</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Densecap: Fully convolutional localization networks for dense captioning</article-title>
          .
          <source>In: IEEE CVPR</source>
          , pages=
          <fpage>4565</fpage>
          -
          <lpage>4574</lpage>
          ,
          <year>year</year>
          =2016
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Prestes</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , et. al.:
          <article-title>Core ontology for robotics and automation</article-title>
          .
          <source>In: Workshop on Knowledge Representation and Ontologies for Robotics and Automation</source>
          . p.
          <volume>7</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Redmon</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farhadi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Yolov3: An incremental improvement</article-title>
          . arXiv preprint arXiv:
          <year>1804</year>
          .
          <volume>02767</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>