<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>REGIMRobvid: Objects and scenes detection for Robot vision 2013</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Amel Ksibi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Boudour Ammar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anis Ben Ammar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chokri Ben Amar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adel M. Alimi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>REGIM (REsearch Group on Intelligent Machines), University of Sfax, National Engineering School of Sfax (ENIS)</institution>
          ,
          <addr-line>BP 1173, Sfax, 3038</addr-line>
          ,
          <country country="TN">Tunisia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the participation of the REGIM team in the ImageCLEF 2013 Robot Vision Challenge. The competition was focused on the problem of objects and scenes classification in indoor environments. Objects and scenes are considered as concepts. During the competition, we aim to classify images according to the room in which they were acquired, using the information provided by the visual images only. Our system is based on PHOW features extraction and PEGASOS SVM algorithm to learn a multi-class classifier that is enable to detect the objects and the adequate scene. For this end, we focus on how to interpret the scores provided by the SVM classifier. Our system was ranked 4th among 6 teams.</p>
      </abstract>
      <kwd-group>
        <kwd>Scene selection</kwd>
        <kwd>object detection</kwd>
        <kwd>PEGASOS SVM</kwd>
        <kwd>PHOW</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        In recent years, robotics has witnessed a large growth and profound change in scope.
Visual detection has becoming one of the most popular research topics and it is playing
an important role in robotics ([
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]). The ImageCLEF 2013 Robot
Vision challenge has been the fifth edition of a competition that started in 2009 within the
ImageCLEF as part of the [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The challenge addresses the problem of semantic place
classification using visual and depth information. This time, the task also addresses
the challenge of object and scene recognition. The rooms/categories that appear in the
      </p>
      <p>
        Fig. 1. Some existing rooms in the database (a) Hall (b) Proffessor office (c) Secretary
database are: Corridor, Hall, ProfessorOffice, StudentOffice, TechnicalRoom, Toilet,
Secretary, VisioConferene, Warehouse and ElevatorArea (see fig. 1). The eight objects
that can appear in any image of the database are: Extinguisher, Computer, Chair, Printer,
Urinal, Screen, Trash, and Fridge (see fig. 2) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>In order to determine the presence or absence of an object, two thresholds are used
(score average and score average plus standard deviation). The object does not exist if
its score is below the average. Else, if its score exceeds the mean standard deviation,
then the object exists. Else if the score is between the two values, this object is
unknown.</p>
      <p>The scene is selected if its score is the maximum. If the score is over than the average,
then the scene is relevant. Else, the scene is unknown.</p>
      <p>The remainder of this paper is organized as follows: the next section explains the
proposed scene and object detection process. Section 3 summarizes the experiments of
the work and discusses the obtained results. Finally, section 4 draws conclusions and
provides suggestions for future work.</p>
    </sec>
    <sec id="sec-2">
      <title>The Proposed system</title>
      <sec id="sec-2-1">
        <title>Architecture of the proposed system</title>
        <p>In this section, we describe the developed system for RobotVision2013 task
participation. Two sets of concepts are used: Object concepts and Scene concepts. In our system,
we aim to learn two appropriate classifiers multi-classes using visual features and
machines learning for objects detection and scene detection.</p>
        <p>Given an image, we hope to describe it using N objects and M locations or scenes. More
specifically, we must detect one appropriate location and some objects.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Concept learning</title>
      </sec>
      <sec id="sec-2-3">
        <title>a)PHOW features extraction</title>
        <p>
          Visual images are represented by a Pyramid Histogram of Visual Words (PHOW) [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ],
which are a variant of dense SIFT descriptors, extracted at multiple scales. In fact, the
PHOW descriptor involves computing visual words on a dense grid [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
To compute these features, a dictionary of visual words was first generated by
quantizing the SIFT descriptors that capture the local spatial distribution of gradients. ELKAN
kmeans clustering is selected to perform quantization. We fixed, here, the dictionary
size to 300 visual words. Then, in order to characterize the joint distribution of
appearance and location of the visual words in an image, each image is divided into regions at
multiple scales (44 subdivisions). So, a spatial histogram is computed for each image
sub-region at each scale.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>b) Object and scene learning</title>
        <p>Object and location detection, in Robot Vision task, requires extremely fast
classification. Or, automatic concept learning relying on computationally heavy kernel-based
classifiers, such as non-linear SVMs, are disabled to accomplish this need resulting in
a computational bottleneck. In fact, we argue that the critical efficiency criterion is the
classifier evaluation cost. There have been numerous approaches to reduce the
computational complexity from the level of standard non-linear SVMs such as the homogeneous
kernel map which is used in our process over the Chi2 kernel SVM.</p>
        <p>
          Firstly, given the obtained spatial histograms, we train the PEGASOS stochastic
gradient descent as a linear SVM classifier in order to build efficiently, concepts models.
While the PEGASOS SVM is very fast to train [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], it cannot typically match the
performance of non-linear, as it is limited to use an inner product to compare descriptors.
Therefore, we perform a step of data pre-transforming through computing the
homogeneous kernel map that provides a linear representation of a Chi2 kernel. This linear
approximation is used, then, to train a Chi2-kernel SVM, by applying the linear SVM
solver PEGASOS.
2.3
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Concept scores estimation</title>
        <p>Given a test image, we classify it using two obtained models, respectively, for objects
and location detection. The outputs of each classifier are the concept having the best
score and two detection scores vectors. Since an image can contain more than one
object, the decision of the object classifier is enough sufficient. So, we need to perform a
process of object selection to deduce other relevant concepts underlying this image.
In addition, the selected concept by the scene classifier can have a low score despite
having the highest one among other concepts. Therefore, the process of scene selection
needs to be improved in a way that for the corresponding case, the system will give the
result ”unknown”.
2.4</p>
      </sec>
      <sec id="sec-2-6">
        <title>Concept Selection</title>
        <p>Concept selection can be either objects selection or scene selection. Objects selection
aims to choose an optimal subset of a predefined concepts list that is able to capture the
semantic content of the corresponding image. In contrast, scene selection aims to find
the most adequate scene.</p>
        <p>
          As the obtained concepts scores are sparse, we need, firstly, a step of normalization
and thresholding to discriminate the most representative objects and the most probably
detected scene [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
a. Scores normalization
The conventional normalization formula is as follows:
        </p>
        <p>Vector[i] =
vector[i]
max</p>
        <p>min
min
(1)
where i=1..N.
i)Scores normalization by each image
For each image, we perform the normalization process using the above formula. We
obtain in each one, obligatory one object having score equal to ”1” and one object having
a score equal to ”0”.</p>
        <p>In case where all objects are present, this normalization will discard the object having
the low scores. In contrast, in case where any object is present, this formula will detect,
always, an object with score equal to ”1”.</p>
        <p>To overcome this problem, we propose to use normalization by each concept.
ii) Scores normalization by each concept
Given a matrix of all obtained scores for all images in the validation collection, we
perform for each concept the above normalization formula.
b. Threshold for concepts selection
After the normalization of the two probability scores, we perform a process of concept
selection which aims to define an adequate threshold that separate relevant concepts
from others. Defining an optimal threshold for all concepts is suboptimal. So, we need
to estimate for each one the corresponding threshold.
i) Object selection
Given an object, the threshold is calculated with respect to the distribution of scores of
this object in all images in the validation dataset.</p>
        <p>Two formulas are tested in experiments:</p>
        <p>&gt;8 1 åN ciq
t = &gt;&gt;&lt; N i=1
&gt; 1 åN ciq + s
&gt;&gt;: N i=1
(2)
Where: N is the number of images, ciq is the score of concept q in image i.
s is the standard deviation of all scores for N images.
ii) Scene selection
For a given image, the system selects the scene having the highest score. Meanwhile,
we need to verify this decision. In fact, we define a threshold to decide if the system is
sure or has an ambiguity. This threshold is equal to the average of all scores of images
according to this scene concept.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments and results</title>
      <p>This section presents the results of the Robot Vision task of ImageCLEF 2013 for the
subtask: task1. Six groups registered to the Robot Vision 2013. A total of 16 runs were
submitted. The limit of the number of runs that could be submitted was 3.
For the competition we submitted two systems: the first is a system based on PHOW
features extraction and PEGASOS SVM classifier with normalization and one threshold,
the second uses the same techniques PHOW + PEGASOS SVM but with normalization
and two thresholds. Our system ranked forth, achieving 4638.250 points on this task.
4</p>
    </sec>
    <sec id="sec-4">
      <title>CONCLUSIONS and future work</title>
      <p>We have described in this article the fifth edition of the Robot Vision task at
ImageCLEF 2013, which attracted an attention of 6 groups submitting runs.
We propose an approach, which detects the location and some objects using PHOW
features extraction and PEGASOS SVM classifier.</p>
      <p>First, an off line module was performed before starting the test. It consists of PHOW
extraction descriptors for object and scene concepts and the training step using
PEGASOS SVM learning method.</p>
      <p>The online process is the concepts scores estimation and the selection of concepts
using the PEGASOS SVM model. Future work aims at exploiting the described methods
to help elderly and disabled persons seated in wheeled chairs. Furthermore, it is also
planned to use different techniques of tracking like Extended Kalman filter or
incremental PCA (Principal Component Analysis) to track the detected objects.</p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGMENT</title>
      <p>The authors would like to acknowledge the financial support of this work by grants from
General Direction of Scientific Research (DGRST), Tunisia, under the ARUB program.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ammar</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rokbani</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alimi</surname>
            <given-names>A. M.</given-names>
          </string-name>
          ,
          <article-title>”Learning System for Standing Human Detection”</article-title>
          ,
          <source>IEEE International Conference on Computer Science and Automation Engineering (CSAE</source>
          <year>2011</year>
          ), Page(s):
          <fpage>300</fpage>
          -
          <lpage>304</lpage>
          , Shanghai 10-12 June 2011
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ammar</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wali</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alimi</surname>
            <given-names>A. M.</given-names>
          </string-name>
          , ”
          <article-title>Incremental Learning Approach for Human Detection and Tracking”</article-title>
          ,
          <source>7th International Conference on Innovations in Information Technology (Innovations'11)</source>
          , Page(s):
          <fpage>128</fpage>
          -
          <lpage>133</lpage>
          ,
          <string-name>
            <given-names>Abu</given-names>
            <surname>Dhabi</surname>
          </string-name>
          25-27 April 2011
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>A.</given-names>
            <surname>Bosch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>X.</given-names>
            <surname>Munoz</surname>
          </string-name>
          .
          <article-title>Image classifcation using random forests and ferns</article-title>
          .
          <source>In Proc. ICCV</source>
          ,
          <year>2007</year>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bousnina</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ammar</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baklouti</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alimi</surname>
            <given-names>M. A.</given-names>
          </string-name>
          ,
          <article-title>”Learning system for mobile robot detection</article-title>
          and tracking” ,
          <source>International Conference on Communications and Information Technology (ICCIT)</source>
          ,
          <source>Page(s)</source>
          :
          <fpage>384</fpage>
          -
          <lpage>389</lpage>
          ,
          <string-name>
            <surname>Hammamet</surname>
            <given-names>Tunisia</given-names>
          </string-name>
          ,
          <year>June 2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>B.</given-names>
            <surname>Caputo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thomee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Paredes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zellhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Goeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bonnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Martinez</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Garcia</given-names>
            <surname>Varea</surname>
          </string-name>
          , M. Cazorla,
          <string-name>
            <surname>ImageCLEF</surname>
          </string-name>
          <year>2013</year>
          :
          <article-title>the vision, the data and the open challenges</article-title>
          .
          <source>Proceedings of CLEF 2013</source>
          ,
          <string-name>
            <surname>Springer</surname>
            <given-names>LNCS</given-names>
          </string-name>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Feki</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ksibi</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ben</surname>
            <given-names>Ammar A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Ben</surname>
            <given-names>Amar C.</given-names>
          </string-name>
          , REGIMvid at ImageCLEF2012:
          <article-title>Improving Diversity in Personal Photo Ranking Using Fuzzy Logic</article-title>
          ,
          <source>Proceedings of CLEF 2011</source>
          ,
          <string-name>
            <surname>Springer</surname>
            <given-names>LNCS</given-names>
          </string-name>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Jesus</given-names>
            <surname>Martinez-Gomez</surname>
          </string-name>
          ,
          <article-title>Ismael Garcia-varea, Miguel Cazorla, Barbara Caputo, Overview of the ImageCLEF 2013 Robot Vision Task</article-title>
          ,
          <source>Working notes of CLEF</source>
          <year>2013</year>
          , Valencia, Spain,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>J.</given-names>
            <surname>Ruiz-del-Solar</surname>
          </string-name>
          and
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Vallejos</surname>
          </string-name>
          , ”
          <article-title>Motion Detection and Object Tracking for an AIBO Robot Soccer Player”</article-title>
          ,
          <string-name>
            <surname>Robotic</surname>
            <given-names>Soccer</given-names>
          </string-name>
          , Pedro Lima (Ed.),
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Shalev-Shwartz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srebro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cotter</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Pegasos: primal estimated sub-gradient solver for SVM</article-title>
          .
          <source>Mathematical Programming</source>
          ,
          <volume>127</volume>
          (
          <issue>1</issue>
          ),
          <fpage>3</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. D. G. Lowe,
          <article-title>Distinctive image features from scale-invariant keypoints</article-title>
          ,
          <source>International Journal of Computer Vision</source>
          , Volume
          <volume>60</volume>
          issue
          <issue>2</issue>
          , pages:
          <fpage>91</fpage>
          -
          <lpage>110</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>