<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Robotic Vision: Understanding Improves the Geometric Accuracy</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Javier Civera</string-name>
          <email>fjciverag@unizar.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>I3A, Universidad de Zaragoza</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>38</fpage>
      <lpage>39</lpage>
      <abstract>
        <p>- . Paraphrasing Olivier Faugeras in the foreword of [1], making a robot see is still an unsolved and challenging task after several decades of research. The traditional research has been based on the geometric models of multiple views of a scene, estimating a sparse 3D map of the scene and the camera pose. Recent advances have led to fully dense and real-time 3D reconstructions. Also, there are relevant recent works on the semantic annotation of the 3D maps. This extended abstract summarizes the work of [2], [3], [4], [5], [6] in this direction; in particular using mid and high-level features to improve the accuracy of dense maps.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>II. DENSE MAPPING</title>
      <p>The inverse depth r for each pixel u in a reference image
is estimated by minimizing the following energy E(r)</p>
      <p>Z 3
E(r) = l0C(u; r) + R(u; r) + å P(u; r; rp )¶ u ; (1)
p=1</p>
      <p>C(u; r) is the photometric difference of each pixel u
backprojected at an inverse depth r and projected into
several overlapping images. R(u; r) is a regularization term
–usually the TV-norm. Finally, the three terms in the sum
3
åp=1 P(u; r; rp ) correspond to the three mid and
highlevel scene cues. For more details on each term and the
optimization of the function the reader is referred to [5].
a) SUPERPIXELS (3DS): Superpixels are clusters of
pixels that have been segmented based on their color and 2D
distance. We will assume that such regions of homogeneous
color will be planar. Specifically, we use the superpixel
segmentation of [13].</p>
      <p>We extract the planes P = (p1; : : : ; pk; : : : ; pq) that fit the
superpixels by minimizing a function F of the geometric
error ek of the reprojected contour of the superpixel k in the
rth overlapping frame
m q
Pˆ = arg min å å F(ek) :</p>
      <p>P r=1 k=1
(2)</p>
      <p>
        The inverse depth r1 in equation 1 is the intersection of
the planes P with the backprojected ray from the pixel u.
For details, see [
        <xref ref-type="bibr" rid="ref1">2</xref>
        ], [4].
      </p>
      <p>b) DATA-DRIVEN PRIMITIVES (DDP): A data-driven
3D primitive [14] is a RGB-D pattern learnt from data. The
visual part of the primitive should be discriminative enough
to be detected on another images, and the depth pattern
geometrically consistent.</p>
      <p>The depth pattern is modelled by its normals and the RGB
pattern by a HOG descriptor and a SVM-based classifier. At
detection time, the inverse depth r2 for each pixel is extracted
from the primitive normal and the depth from a multiview
reconstruction. See [5] for more details.</p>
      <p>c) LAYOUT (Lay.): The so-called layout [15] consists
on the estimation of the rough geometry of a room and the
classification of each pixel u into the classes wall, ceiling,
floor and clutter.</p>
      <p>We assume that the room is cuboid, so its model is
composed of six planes. We estimate their normals using
multiview vanishing points and their distances from a
geometric reconstruction. From such layout, the inverse depth r3
is computed as the intersection of each pixel with the room
boundaries if is is classified as that. If the pixel u is classified
as clutter we consider that the depth is not predictable.</p>
    </sec>
    <sec id="sec-2">
      <title>III. EXPERIMENTAL RESULTS</title>
      <p>
        Figure 1 shows an illustrative view of our results in some
selected sequences from the NYU dataset [
        <xref ref-type="bibr" rid="ref2">16</xref>
        ]. Notice how
close our estimation (6th column) is to the ground truth depth
(5th column).
      </p>
      <p>
        Tables I and II show the median depth error of DTAM
[11], the sparse feature-based multiview stereo PMVS [
        <xref ref-type="bibr" rid="ref3">17</xref>
        ]
and our algorithm on low-texture and low-parallax sequences
respectively –typical failures cases for the geometric
estimation. Notice our improvement in every case. Notice also how
it comes from different features depending on the sequence,
      </p>
      <p>DPP</p>
      <p>Layout
sonable results even in the single-view case.</p>
      <p>IV. CONCLUSIONS</p>
      <p>
        In this abstract –and the associated papers [
        <xref ref-type="bibr" rid="ref1">2</xref>
        ], [3], [4],
[5], [6]– we have shown how mid and high-level features
improve the accuracy of a dense point-based reconstruction
from monocular images. The features complement each other
Fig. 1: EstimRateesduldtsepftrhomfrotmhe 3Bseedqrouoemnc1e,s B–eindrrooowms2. 1asntdcoKluitmchneins stheqeureenfecree.nce frame. 2nd column are the extracted
      </p>
      <p>Fig. 6
superpixels, 3rd column the data-driven primitives and 4th column the estimated layout. The 5th column is the ground truth
d5e0pth from a RGB-D camera and the186th one our result. Notice the simi2la5rity between the latest two.</p>
      <p>Median</p>
      <p>Mean 16
40 25%−75% 14 20
showing th)eir complementary natu9r%e−.9F1%or mor)1e2 details on
these and o(cmt3h0er experiments see [5]. (c10</p>
      <p>m
R R</p>
      <p>O O8
SequeERR2n0ce DTAM [11] MeaPnMEVrSro[r1[7c]m(]%) ERR6Ours
nicely, so a fusion o)f all of them improves the accuracy in
a wide array of cas e(cm1s5.</p>
      <p>R</p>
      <p>O
ACRK10NOWLEDGMENTS</p>
      <p>R</p>
      <p>E
This research has been partially funded by projecta</p>
      <p>5
DPI2012-32168 and DGA T04-FSE.</p>
      <p>0</p>
      <sec id="sec-2-1">
        <title>DRTAEMFPEMRVESNLCAEY.S 3DS DDP ALL</title>
        <p>Bedroo1m01 (3DS)
Bedroom1 (DDS)
Bedroom1 (Lay.)</p>
      </sec>
      <sec id="sec-2-2">
        <title>Bedroom0D1T(AAMll)PMVS LAY. 3DS DDP ALL</title>
        <p>Bedroom2 (3DS)
Bedroom2 (DDP)
Bedroom2 (Lay.)(a) Bed7r.1oom1
15.8
7.0 (18%)
415.0
2 4.2</p>
        <p>7.9
0 D5.T9AM PMVS LAY. 3DS DPP ALL
6.7
7.6
7.7
6.8</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Kitchen (All) RGB Image</title>
      <p>[1] R. I. Hartley and A. Zisserman, Multiple View Geometry in Computer
5.7 (22%) (b) Bedroom2Vision. Cambridge University(cP)reKss,itIcShBeNn: 0521540518, 2004.</p>
      <p>
        Bedroom2 (All) [
        <xref ref-type="bibr" rid="ref1">2</xref>
        ] A. Concha and J. Civera, “Using superpixels in monocular SLAM,”
Kitchen (3DS) 5.6 in ICRA, 2014.
      </p>
      <p>KFitcihge.n 7(DDBP)ox and Whiskers plots showing t7h.7e depth error [d3]istAr.ibCuotnicohna, fWor. Hthuessainind,oLo.rMhoingtahn-op, aarnadllJa.xCisveeqrau, e“nMcaensh.attan and
Kitchen (Lay.) 7.2 5.5 (20%) 5.7
ar#e4 t(DhDeP)same t4h2.3an the b28a8s.4e(l9i%n)e DT3A9.M1 and we</p>
      <p>#4 (All) 20.9
only present results for DDP and Layout. As
TABLE pIIr:evMioeuansldyespathide,rtrhorisfiosr aDcTlAeaMr,liPmMiVtaStiaonnd oofu3rsDiSn
low-parallax sequences. (%) is the percentage of pixels
–and in general of multiview geometry– and an</p>
      <p>reconstructed by PMVS.</p>
      <p>advantage of DDP and Layout, that give
rea</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>quennoc.e2s,ppp.r1in67t-e1r81r,o2o0m04. 0001 rect (#1 and #2)</source>
          , [14]
          <string-name>
            <given-names>D. F.</given-names>
            <surname>Fouhey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Hebert</surname>
          </string-name>
          , “
          <article-title>Data-driven 3D primitives bedrfoorosmingl0e1im06agereucntde(rs#tan3d)inag</article-title>
          ,
          <source>n” dinbIeCdCrVo,o2m0130.110 rect [1(5#]4V)</source>
          . Hoefdatuh, eD.dHaotiaemse,ta.ndFDi.gFuorresyt9h,
          <article-title>s“hReocwovseritnhgethBesopxat-ial layout of cluttered rooms</article-title>
          ,” in ICCV,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>N.</given-names>
            <surname>Silberman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hoiem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kohli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          , “
          <article-title>Indoor segmentation and support inference from rgbd images</article-title>
          ,” in ECCV,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Furukawa</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ponce</surname>
          </string-name>
          , “Accurate, dense, and robust multiview stereopsis,
          <source>” IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          , vol.
          <volume>32</volume>
          , no.
          <issue>8</issue>
          , pp.
          <fpage>1362</fpage>
          -
          <lpage>1376</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>