<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the ImageCLEF 2012 Robot Vision Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jesus Martinez-Gomez</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ismael Garcia-Varea</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Barbara Caputo</string-name>
          <email>2bcaputo@idiap.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Idiap Research Institute, Centre Du Parc</institution>
          ,
          <addr-line>Rue Marconi 19 P.O. Box 592, CH-1920 Martigny</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Jesus.Martinez</institution>
          ,
          <addr-line>Ismael.Garcia</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article describes the RobotVision@ImageCLEF 2012 challenge, which addresses the problem of multimodal place classi cation. Participants of the challenge were asked to classify rooms on the basis of image sequences captured by cameras mounted on a mobile robot. The proposals of the participants had to answer the question \where are you?" (I am in the elevator, in the toilet, etc) when presented with a test sequence, acquired within the same building and oor but with di erent lighting conditions than the training sequence. The 2012 edition of the challenge introduced the use of depth images in addition to visual images. Moreover, several techniques for feature extraction and cue integration were also proposed. As in previous editions, two di erent tasks were proposed: task 1 and task 2. In task 1 (mandatory) participants were asked to classify the frames separately, while the temporal continuity of the image sequence could only be exploited in task 2 (optional). Eight di erent groups participated to the 2012 edition of the Robot Vision challenge. The winner in both tasks was the Centro de Investigacion en Informatica para la Ingenier a (CIII), from the Universidad Tecnologica Nacional, Argentina (CIII UTN FRC). This participant obtained an overall score of 2071 (84.70% of the maximum score) in task 1 and 3930 (96.35% of the maximum score) in task 2.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The ImageCLEF 2012 Robot Vision challenge has been the fourth edition of a
competition [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] that started in 2009 within the ImageCLEF 1 as part of the
Cross Lange Evaluation Forum (CLEF) Initiative 2. Since its origin, the Robot
Vision task has been addressing the problem of place classi cation for mobile
robot localization.
      </p>
      <p>
        The 2009@ImageCLEF edition of the task [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], with 7 participating groups,
de ned some details that have been maintained for all the following editions.
Participants were given training data consisting of sequences of frames recorded
in indoor environments. These training frames were labelled with the name of
the rooms they were acquired from. The task consisted on building a system
capable to classify test frames using as class the name of the rooms previously
seen. Moreover, the system could refrain from making a decision in the case of
lack of con dence. Two di erent subtasks were then proposed: obligatory and
optional. The di erence between both subtasks was that the temporal continuity
of the test sequence could only be exploited in the optional task. The score for
each participant submission was computed as the sum of the frames that were
correctly labelled minus a penalty that was applied to the frames that were
misclassi ed. No penalties were applied for frames not classi ed.
      </p>
      <p>
        In 2010, two editions of the challenge took place. The second edition of the
task, 2010@ICPR [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] was held in conjunction with ICPR 2010. 9 groups
participated to this edition, which introduced the use of stereo images and two types
of di erent training sequences (easy and hard) that had to be used separately.
The 2010@ImageCLEF edition [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], with 7 participating groups, was focused on
generalization: several areas could belong to the same semantic category.
      </p>
      <p>Several changes have been proposed for the ImageCLEF 2012 Robot Vision
task. Firstly, stereo images have been replaced by images acquired using two
types of camera: a perspective camera for visual images and a depth camera
(the Microsoft Kinect sensor) for range images. Therefore, each frame consists
of two types of images and the challenge is focused on the problem of
multimodal place classi cation. In addition to the use of depth images, the optional
task contains kidnappings and no unknown rooms appear in the test sequences.
Moreover, several techniques for features extraction and cue integration have
been proposed to the participants.</p>
      <p>We received a total of 23 runs from 8 di erent groups. 18 runs were
submitted to the task 1 (mandatory) and 5 to the task 2 (optional). The best result in
both tasks was obtained by the Centro de Investigacion en Informatica para la
Ingenier a (CIII), from the Universidad Tecnologica Nacional, Argentina (CIII
UTN FRC).</p>
      <p>The rest of the paper details the challenge and is organized as follows:
Section 2 describes the 2012 ImageCLEF edition of the RobotVision task. Section 3
presents all the participants groups, while the results are reported in Section 4.
Finally, in Section 5, conclusions are drawn and future work is outlined.</p>
    </sec>
    <sec id="sec-2">
      <title>The RobotVision Task</title>
      <p>This section describes the details concerning the setup of the ImageCLEF 2012
Robot Vision task. Section 2 gives a description of the task. Section 2.1 describes
all the sequences of frames provided for training and test while the two subtasks
are explained in Section 2.2. The performance evaluation is detailed in Section 2.3
and nally, Section 2.4 describes the information provided by the organizers.
2.1</p>
      <sec id="sec-2-1">
        <title>Description</title>
        <p>The fourth edition of the Robot Vision challenge was focused on the problem of
multi-modal place classi cation. Participants were asked to classify functional
areas on the basis of image sequences, captured by a perspective camera and a
Kinect mounted on a mobile robot (see Fig. 1) within an o ce environment.</p>
        <p>Participants had available visual images and range images that could be used
to generate 3D point cloud les. The di erence between visual images, range
images and 3D point cloud les can be observed in Figure 2. Training and test
sequences were acquired within the same building and oor but with some
variations in the lighting conditions or the acquisition procedure (clockwise and
counter clockwise).</p>
        <p>Two di erent tasks were considered in the Robot Vision challenge: task 1
and task 2. For both tasks, participants should be able to answer the question
\where are you?" when presented with a test sequence imaging a room category
already seen during training. The di erence between both tasks was the presence</p>
        <p>Visual Image</p>
        <p>Range Image</p>
        <p>3D point cloud le
(or lack) of kidnappings in the nal test sequence, and the availability on the
use of the temporal continuity of the sequence.</p>
        <p>The kidnapping (only task 2) is a ected by the robot changing room. Room
changes in sequences without kidnappings were usually represented by a small
number of images showing a smooth transition. On the other side, room changes
with kidnappings were represented by a drastic change for frames, as can be
observed in Figure 3.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>The Data</title>
        <p>Training and validation sequences consisted of a set of the Robot Vision VIDA
dataset. VIDA is a dataset with images acquired within an indoor environment
using a robot platform in the IDIAP research building. This dataset contains
sequences of several rooms belonging to di erent room categories such as
\Corridor" or \Toilet". Sequences were acquired using two cameras: a perspective
visual camera and a 3D range laser sensor Kinect camera. Therefore, there are
two di erent types of images: RGB images and depth images.</p>
        <p>Three di erent sequences of frames were provided for training and two
additional ones for the nal experiment. All training frames were labelled with the
name of the room they were acquired from. There were 9 di erent categories of
rooms, and the distribution for the training sequences can be observed in Table 1</p>
        <p>The di erence between all the room categories can be observed in Figure 4,
where an exemplar visual image for each one of the 9 room categories is shown.</p>
        <p>Elevator Area</p>
        <p>Corridor</p>
        <p>Toilet
Lounge Area</p>
        <p>Technical Room</p>
        <p>Professor O ce
Student O ce</p>
        <p>Visio Conference</p>
        <p>Printer Room</p>
      </sec>
      <sec id="sec-2-3">
        <title>Subtasks</title>
        <p>All the participants of the ImageCLEF 2012 Robot Vision task were allowed to
submit their runs to two di erent subtasks: task1 and task2.</p>
        <p>Task 1 This task was mandatory and the test sequence had to be classi ed
without using the temporal continuity of the sequence. Therefore, the order of the
test frames cannot be taken into account. Moreover, there were not kidnappings
in the nal test sequence.</p>
        <p>Task 2 This task was optional and participants could take advantage of the
temporal continuity of the test sequence. There were kidnapping in the nal test
sequences that allowed participants to obtain additional points when they were
managed correctly
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Performance Evaluation</title>
        <p>The proposals of the participants were compared using the score obtained by
their submissions. These submissions were the classes or room categories assigned
to the frames of the test sequences, and the score was computed using the rules
that are shown in Table 2. Due to wrong classi cations obtaining negative points,
participants were allowed to not classify test frames.</p>
        <p>Each correctly classi ed frame +1 points
Each misclassi ed frame -1 points
Each frame that was not classi ed +0 points
(Task 2) All the 4 frames correctly classi ed after a kidnapping +1 points (additional)
2.5</p>
      </sec>
      <sec id="sec-2-5">
        <title>Additional information provided by the organization</title>
        <p>
          We proposed the use of several techniques for features extraction (PHOG and
NARF) and cue integration (OBSCURE). Thanks to the use of these techniques,
participants could focus on the development of new features while using the
proposed method for cue integration or vice versa. We also provided information
as the point cloud library [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and a basic technique for taking advantage of the
temporal continuity3. In order to evaluate this information, we submitted two
runs (task 1 and task 2) that were obtained using only the provided techniques.
The results obtained with such proposal [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] can be considered as baseline results
that all the groups were expected to improve.
3 http://imageclef.org/2012/robot
Visual Features PHOG features are histogram-based global features that
combine structural and statistical approaches. Other descriptors similar to PHOG
that could also be used are: Sift-based Pyramid Histogram Of visual Words
(PHOW) [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], Pyramid histogram of Local Binary Patterns (PLBP) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ],
SelfSimilarity-based PHOW (SS-PHOW) [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], and Compose Receptive Field
Histogram (CRFH) [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>
          Depth Features NARF features is a novel descriptor technique that has been
included in the point cloud library [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. The number of descriptors that can be
extracted from a range image is not xed, in the same manner as SIFT points.
Cue Integration The algorithm proposed for cue integration was the
OnlineBatch Strongly Convex mUlti keRnel lEarning (OBSCURE) [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. This
SVMbased multiclass learning algorithm obtains state-of-the-art performance in a
considerably lower training time. Other algorithm that could be used was the
Online Independent Support Vector Machines [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] that, in comparison with SVM,
dramatically reduces learning time and space requirements at the price of a
negligible loss in accuracy.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Participation</title>
      <p>In 2012, 43 groups registered to the Robot Vision task but only 8 submitted, at
least, one run, namely:
{ CIII UTN FRC: Universidad Tecnologica Nacional, Cordoba, Argentina.
{ NUDT: National University of Defense Technology, Changsha, China.
{ UAIC2012: Alexandru Ioan Cuza University, Iasi, Romania.
{ USUroom409: Ural Federal University, Yekaterinburg, Russian Federation.
{ SKB Kontur Labs: Kontur Labs, Yekaterinburg, Russian Federation.
{ CBIRITU: Istanbul Technical University, Istanbul, Turkey.
{ SIMD: University of Castilla-La Mancha, Albacete, Spain.
{ Bu aloVision: University at Bu alo, New York, United States.</p>
      <p>A total of 23 runs were submitted, with 18 runs submitted to the task 1
(mandatory) and 5 runs submitted to the task 2 (optional). The limit to the
number of runs that could be submitted was 3.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>This section presents the results of the Robot Vision task of ImageCLEF 2012
for the two subtasks: task1 and task 2.
Eight di erent groups submitted runs for the task 1 as can be observed in Table 3.
The maximum score that could be achieved was 2445 and the winner (CII UTN
FRC) obtained 2071 points. CII UTN FRC and NUDT teams ranked rst and
second respectively and their score was higher than 70% of the maximum score.</p>
      <p>The score obtained by the SIMD/IDIAP team could be considered as a
baseline result that all the groups were expected to improve. Such score was obtained
using the techniques provided by the organizers without new contributions. As
it was expected, 6 out of 7 teams obtained higher scores. The results are
summarized in Figure 5, where only the best run for each team has been considered.
For the optional task, the maximum score was 4079 and only 4 groups submitted
runs. The winner for the task 2 was the CIII UTN FRC group with 3930 points,
only 71 more than the NUDT group, which ranked second. All the results can
be seen in Table 4.</p>
      <p>In view of these results, it should be remarked the high quality of the
participant proposals, due to the score obtained by CIII UTN FRC, NUDT and
CBIRITU groups was higher than 75% of the maximum score. All the results
for task 2 are summarized in Figure 6, where only the best run for each team
has been considered.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>We have described in this article the fourth edition of the Robot Vision task
at ImageCLEF 2012, which attracted a considerable attention with 8 groups
submitting runs. There are 2 main conclusions that can be drawn from the
proposals: (i) despite of two types of images were provided (visual and range),
depth features were not commonly used, and (ii) most of the proposals were
based on the use of SVMs. We plan to continue the task in the next years
with new challenges related to place categorization. Concretely, we have plans
to introduce object categorization in the following editions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>A.</given-names>
            <surname>Bosch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>X.</given-names>
            <surname>Munoz</surname>
          </string-name>
          .
          <article-title>Image classi cation using random forests and ferns</article-title>
          .
          <source>In International Conference on Computer Vision</source>
          , pages
          <fpage>1</fpage>
          <lpage>{</lpage>
          8.
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>O.</given-names>
            <surname>Linde</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Lindeberg</surname>
          </string-name>
          .
          <article-title>Object recognition using composed receptive eld histograms of higher dimensionality</article-title>
          .
          <source>In Proc. ICPR. Citeseer</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Jesus</given-names>
            <surname>Martinez-Gomez</surname>
          </string-name>
          ,
          <article-title>Ismael Garcia-Varea, and Barbara Caputo. Baseline multimodal place classi er for the 2012 robot vision task</article-title>
          .
          <source>In CLEF 2012 working notes</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>T.</given-names>
            <surname>Ojala</surname>
          </string-name>
          , M. Pietikainen, and T. Maenpaa.
          <article-title>Gray scale and rotation invariant texture classi cation with local binary patterns</article-title>
          .
          <source>Computer Vision-ECCV</source>
          <year>2000</year>
          , pages
          <fpage>404</fpage>
          {
          <fpage>420</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>F.</given-names>
            <surname>Orabona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Castellini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Caputo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Luo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Sandini</surname>
          </string-name>
          .
          <article-title>Indoor place recognition using online independent support vector machines</article-title>
          .
          <source>In Proc. BMVC</source>
          , volume
          <volume>7</volume>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>F.</given-names>
            <surname>Orabona</surname>
          </string-name>
          , L. Jie, , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Caputo</surname>
          </string-name>
          .
          <article-title>Online-Batch Strongly Convex Multi Kernel Learning</article-title>
          .
          <source>In Proc. of Computer Vision</source>
          and Pattern Recognition,
          <string-name>
            <surname>CVPR</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>A.</given-names>
            <surname>Pronobis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Christensen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Caputo</surname>
          </string-name>
          .
          <article-title>Overview of the imageclef@ icpr 2010 robot vision track</article-title>
          .
          <source>Recognizing Patterns in Signals, Speech, Images and Videos</source>
          , pages
          <volume>171</volume>
          {
          <fpage>179</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>A.</given-names>
            <surname>Pronobis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fornoni</surname>
          </string-name>
          , HI Christensesn, and
          <string-name>
            <given-names>B.</given-names>
            <surname>Caputo</surname>
          </string-name>
          .
          <article-title>The robot vision track at imageclef 2010</article-title>
          . Working Notes of ImageCLEF,
          <year>2010</year>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>A.</given-names>
            <surname>Pronobis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Xing</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Caputo</surname>
          </string-name>
          .
          <article-title>Overview of the clef 2009 robot vision track</article-title>
          . pages
          <volume>110</volume>
          {
          <fpage>119</fpage>
          . Springer,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>Andrzej</given-names>
            <surname>Pronobis</surname>
          </string-name>
          and
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Caputo</surname>
          </string-name>
          .
          <article-title>The robot vision task</article-title>
          . In Henning Muller, Paul Clough, Thomas Deselaers, and Barbara Caputo, editors,
          <source>ImageCLEF</source>
          , volume
          <volume>32</volume>
          <source>of The Information Retrieval Series</source>
          , pages
          <volume>185</volume>
          {
          <fpage>198</fpage>
          . Springer Berlin Heidelberg,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>R.B. Rusu</surname>
            and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Cousins</surname>
          </string-name>
          .
          <article-title>3d is here: Point cloud library (pcl)</article-title>
          .
          <source>In Robotics and Automation (ICRA)</source>
          ,
          <source>2011 IEEE International Conference on, pages 1{4</source>
          . IEEE,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. E. Shechtman and
          <string-name>
            <given-names>M.</given-names>
            <surname>Irani</surname>
          </string-name>
          .
          <article-title>Matching local self-similarities across images and videos</article-title>
          .
          <source>In IEEE Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2007</year>
          . CVPR'
          <volume>07</volume>
          , pages
          <issue>1{8</issue>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>