<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MIAR ICT participation at Robot Vision 2013</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ruihan Xu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shuqiang Jiang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xinhang Song</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shuang Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yi Xie</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fang Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiong Lv</string-name>
          <email>lvxiongforyou@foxmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Computing Technology, Chinese Academy of Sciences No.</institution>
          <addr-line>6 Kexueyuan South Road Zhongguancun,Haidian District Beijing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the participation of our team - MIAR ICT in the ImageCLEF 2013 Robot Vision Challenge. The task of the Challenge asked participants to classify imaged indoor scenes and recognize the predefined objects appeared in the imaged scene. Our approach is based on the recently proposed Kernel Descriptors framework, which is an effective representation for images. For the provided visual and depth sequences, we make a simple fusion at feature level. Then we use Linear Support Vector Machine (L-SVM) classifiers for both scene classification and object recognition. At last, the temporal continuity of the given sequences is considered. Our team ranked the first among all the participants, showing the effectiveness of our proposed scheme.</p>
      </abstract>
      <kwd-group>
        <kwd>kernel descriptor</kwd>
        <kwd>scene classification</kwd>
        <kwd>object recognition</kwd>
        <kwd>temporal continuity</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        In the 5th Robot Vision Challenge of the ImageCLEF 2013, image sequences were
captured by a perspective camera and a Kinect[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] mounted on a mobile robot
within an office environment. Visual (RGB) images and depth images generated
from 3D point clouds were available. Training sequences were labelled not only
with semantic labels (corridor, kitchen, office, etc.) but also with the objects that
were represented in them (fridge, chair, computer, etc.). The test sequence were
acquired within the same building and floor, but there could be variations in
the lighting conditions (very bright places or very dark ones) or the acquisition
procedure (clockwise and counter clockwise). Given test sequences, participants
were asked to classify different indoor scenes, and judge the existence of the
given objects within each image.
      </p>
      <p>
        This paper describes the participation of our team in the Robot Vision
Challenge. For the image features extraction part, we used the state-of-the-art Kernel
Descriptors[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] framework, which has proven to be useful for many problems with
RGB-D (visual and depth) information[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. We applied L-SVM[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for
classification, and the temporal continuity is utilized during the classification stage.
      </p>
      <p>The rest part of this paper is organized as follows. In Section 2, we briefly
give a overview of our classification system. In Section 3, we describe in detail
Visual images Depth images
Professor office Professor office</p>
      <p>Corridor</p>
      <p>Corridor
Student office Student office</p>
      <sec id="sec-1-1">
        <title>Training Dataset</title>
        <p>Depth kernel
descriptors
extraction
Dictionary
learning
Depth EMK
feature
extraction
Visual kernel
descriptors
extraction
Dictionary
learning
Visual EMK
feature
extraction
Concatenation
Joint feature
of image</p>
      </sec>
      <sec id="sec-1-2">
        <title>Feature extraction</title>
        <p>Joint feature
of image
Labels
Training
Linear SVM
Model</p>
      </sec>
      <sec id="sec-1-3">
        <title>Training model</title>
        <p>the image feature we used in our scheme. In Section 4 we describe how we apply
classifier for both two tasks, and how we make use of the temporal continuity.
In Section 5, we give some of our experiments and show our final result on test
sequence. In Section 6, we draw some conclusions.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Overview</title>
      <p>In this section, we describe the procedure of our scheme. Both scene classification
and object recognition tasks can be solved using classification framework based
on supervised learning. The training stage for scene classification is shown in
Fig.1 and the test stage is shown in Fig.2. Framework for the recognition of each
object is similar, except that training labels and predicted labels are replaced
by the existence of each predefined object.</p>
      <p>
        During the preprocessing stage of our scheme, the given 3D Point Cloud data
are transformed to depth images,which afterwards will be treated as grayscale
images. Then for both visual images and depth images, we extract Kernel
Descriptors [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] as local descriptors, and use efficient match kernels (EMK) to
transform and aggregate the descriptors to the features of images. We represent each
frame by concatenating the two kinds of features extracted from each visual
image and depth image. Then we choose L-SVM as our classifier for both scene
classification and object recognition. In consideration of temporal continuity, we
assign the averaged L-SVM scores of one frame’s temporal neighbors to its final
score. More details of our scheme will be described in the following sections.
      </p>
      <sec id="sec-2-1">
        <title>Test image sequence</title>
        <p>Visual images
Depth images
Feature
extraction</p>
        <p>Linear SVM
Classifier</p>
        <p>Score for
each category</p>
        <p>Score
smoothing
Prediction by
the max score</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 Image Features</title>
      <p>In this section, we describe the features of image, which have been used in our
work. The feature extraction procedure consists of two steps. The first step is to
design match kernels using pixel attributes, and the second is to learn compact
features.
3.1</p>
      <sec id="sec-3-1">
        <title>Kernel descriptors</title>
        <p>
          Kernel descriptors are able to generate rich patch-level features from different
types of pixel attributes. For visual images, we use gradient, local binary pattern
(LBP)[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and color kernels. For depth images, we use depth gradient, depth LBP,
spin/surface normal kernels.
        </p>
        <p>The gradient match kernel is:</p>
        <p>
          Kgrad(P, Q) = X X m˜ (p) m˜ (q) ko θ˜(p) , θ˜(q) kp (p, q) ,
p∈P q∈Q
(1)
where P and Q are the set of nearby points around the reference point p¯
and p¯, respectively . kp (p, q) = exp −γpkp−qk2 is a Gaussian position kernel
with z denoting the 2D position of a pixel in an image patch (normalized to
[
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ]), and ko (θ (p) , θ (q)) = exp −γokθ (p) −θ (q) k2 is a Gaussian kernel over
orientations.
        </p>
        <p>Kernel view of orientation histograms provides a way to turn pixel attributes
into patch-level features, which can also be extended to LBP match kernel:
KLBP (P, Q) = X X s˜ (p) s˜ (q) kb (b (p) , b (q)) kp (p, q) ,
p∈P q∈Q
(2)
where s˜ = s (p) /qPp∈P s (p)2 + ǫs,s (p) is the standard deviation of values
in the 3 × 3 neighborhood around p ,ǫs is a small constant, and b (p) is a
binary column vector which binarizes the pixel value differences in a local window
around p.</p>
        <p>Similar to gradient and LBP kernels, the color match kernel can be
formulated as:
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Learning Compact Features</title>
        <p>Evaluating kernels is computationally expensive when image patches are large.
For both computational efficiency and representational convenience, the feature
can be extracted as following:</p>
        <p>1. uniformly and densely sample sufficient basis vectors from support region
to guarantee accurate approximation to match kernels.</p>
        <p>2. learn compact basis vectors using kernel principal component analysis.</p>
        <p>Kcol(P, Q) = X X kc (c (p) , c (q)) kp (p, q) ,</p>
        <p>p∈P q∈Q
where c (p) is the pixel color at position z (intensity for gray images and RGB
values for color images). kp (c (p) , c (q)) = exp(−γckc (p)−c (q) k2) measures how
similar two pixel values are.</p>
        <p>
          Since depth images are treated as grayscale images, depth gradient and depth
LBP kernels are constructed in a similar way like the gradient and LBP
kernels for visual images. Here we just describe another one of the depth kernels
spin/surface normal kernel[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>
          In spin images[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], a reference point in a local 3D point cloud is represented
as the pair(p¯, n¯) formed by its 3D coordinate p¯ and surface normal n¯. The spin
image attribute of a point p ∈ P represented by the pair (p¯, n¯) is given by the
triple [ηp, ςp, βp], where the elevation coordinate ηp is the signed perpendicular
distance from the point p to the tangent plane defined by the pair (p¯, n¯), the
radial coordinate ςp is the perpendicular distance from the point p to the line
through the normal n¯, and βp is the angle between the normals n and n¯. The
point attributes [ηp, ςp, βp] can be aggregated into local shape features by the
following kernel:
        </p>
        <p>Kspin (P, Q) = X X ka β¯p, β¯q kspin ([ηp, ςp] , [ηq, ςq]) ,</p>
        <p>p∈P q∈Q
where β¯p = [sin (βp) , cos (βp)], P is the set of nearby points around the
reference point p¯. Gaussian kernels ka and kspin are used to measure the similarities
of attributes β, η and ς, respectively.
(3)
(4)</p>
        <p>
          EMK combines the advantage of both bag-of-words and set kernels. Here we
briefly describe how the EMK transforms kernel descriptors to low dimensional
space to achieve compact features (see [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] for details).
        </p>
        <p>Take feature based on gradient match kernel for example, other kinds of
feature can be extracted in the same way. Rewriting the Eq.1:
the feature over image patches will be:


 ko θ˜ (p) , θ˜(q) = φo θ˜ (p)
⊤</p>
        <p>φo θ˜(q)
kp (p, q) = φp (p)⊤ φp (q)
,
Fgrad (P ) =</p>
        <p>X m (p) φo θ˜(p) ⊗ φp (p) .</p>
        <p>p∈P
where ⊗ is the Kronecker product. A straightforward way to dimension
reduction is to sample sufficient image patches from training images and perform
KPCA for match kernels.</p>
        <p>Sufficient Finite-dimensional Approximation Finite-dimensional features can be
learned by projecting Fgrad (P ) into a set of basis vectors. A key issue in this
projection process is how to choose a set of basis vectors which makes the
finite-dimensional kernel approximate well the original kernel. Given a set of
basis vectors {ϕo (xi)}id=o 1 where xi are sampled normalized gradient vectors, a
infinite-dimensional vector can be approximated by a infinite-dimensional vector
ϕo (θ (p)) by its projection into the space spanned by the set of these do basis
vectors. Such a procedure is equivalent to using a finite-dimensional kernel:
ko θ˜ (p) , θ˜ (q) = ko θ˜(p) , X
˜</p>
        <p>Ko−1 ⊤
ij ko θ˜(p′) , X ,
which can be rewritten as:
⊤
k˜o θ˜(p) , θ˜(q) = hGko θ˜(p) , X i hGko θ˜(q) , X i .</p>
        <p>
          Here ko θ˜(p) , X
⊤
= hko θ˜(p) , x1 , · · · , ko θ˜(p) , xdo i is a do×1 vector,
Ko is a do × do matrix with Koij = ko (xi, xj ), and K = G⊤G. The resulting
feature map ϕo (θ (p)) = Gko (θ (p) , X ) is now only do-dimensional.
Compact Features The size of basis vectors can be further reduced by performing
kernel principal component analysis over joint basis vectors:
ϕo (x1) ⊗ ϕp (y1) , · · · , ϕo (xdo ) ⊗ ϕp ydp
where ϕp (ys) are basis vectors for the position kernel and dp is the number
of basis vectors. The t-th kernel principal component can be written as:
(5)
(6)
(7)
(8)
(9)
do dp
P Ct = X X αitj φo (xi) ⊗ φp (yj) ,
i=1 j=1
where αitj is learned through kernel principal component analysis[
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
        <p>
          Under the framework of kernel principal component analysis, the gradient
kernel descriptor for the patch P has the form:
In our work, we applied the LibLinear[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] as our classifier, since SVM is widely
used for classification task and performs effective especially when the scale of
data is small. For scene classification, we train a multiclass one-vs-all L-SVM
classifier. As there are 10 different concepts of indoor scene, 10 binary L-SVMs
are trained for each concept. For object recognition task, we treat the existence of
each object as a binary classification problem. Frames that contain the predefined
object are taken as positive samples, and the rest are taken as negative samples.
        </p>
        <p>For better comprehension, let us introduce some notation here. Let In be one
image of the test sequence, n ∈ {1, 2, 3, · · · , N }, where N is the number of all
images in the test sequence.</p>
        <p>For scene classification, let Snc be the L-SVM output score for test image
In on concept c, c ∈ {1, 2, 3, · · · , C}, where C is the number of concepts, and
C = 10 in this task. Then the predicted label cpnred of a test image In is decided
following the rule below:
cpred = argmax Snc.</p>
        <p>n c</p>
        <p>For recognition of object objk, k ∈ {1, 2, 3, · · · , K}, where K is the number
of objects to be recognized, and K = 8 in this task. Let Sn,k be the L-SVM
output score of test image In for objk, and cpnr,ked indicates the predicted concept
of In for objk, where cpnr,ked ∈ {−1, 1}, -1 for concept absence and 1 for concept
occurrence. Whether a certain kind of object exists in the test image can be
judged as below:
 1

cpnr,ked =  0
 −1</p>
        <p>Sn,k &gt; 0
Sn,k = 0 ,
Sn,k &lt; 0
(12)
(13)
where the prediction 0 in Eq.(13) means that whether the object exists or
not is ambiguous, and we deal with this situation with not classifying it. This
happens only when Sn,k = 0, which means that for the test image In, object
objk has the same confidence on both concept occurrence and absence.
4.2</p>
      </sec>
      <sec id="sec-3-3">
        <title>Consideration of Temporal Continuity</title>
        <p>Since all the images in training and test sequences are captured continuously, it is
reasonable and feasible to make full use of the temporal continuity. In our work,
we apply a smoothing method for the L-SVM score to improve the classification
performance.</p>
        <p>We empirically think that the concept of an image is quite likely the same
with that of its temporal neighbors, and the L-SVM score should be less changed
compared to its neighbors. Based on this assumption, we smooth the L-SVM
score for both scene classification and object recognition task as bellow:
Snc =</p>
        <p>1 nX+r
2r + 1 k=n−r</p>
        <p>c
Sk,
(14)
where r is the radius of smooth window. Eq.(14) indicates that the final
LSVM score for image In on a certain concept c is determined by all the scores of
neighbors within the smoothing window. With all the L-SVM scores updated,
we do classification and recognition on the basis of these new scores.</p>
        <p>Choosing appropriate r is very important, for it has a relevance with the
specific data and differs from scene classification and object recognition. The
details of choosing r will be discussed in Sec. 5.3.
5
5.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <sec id="sec-4-1">
        <title>Datasets and experimental setup</title>
        <p>For Robot Vision Challenge this year, two training sequences are provided with
1947 and 3316 images respectively. An additional (labelled) validation sequence
with 1869 images is also provided. The final test sequence involves 3315
unlabelled images. For all the sequences, RGB images and Point Cloud Data (PCD)
are available. As mentioned in Sec.1, there can be variations in the lighting
conditions or the acquisition procedure between test sequence and training sequences,
and the validation sequence is similar to the test to some degree.</p>
        <p>For depth features extraction, we transformed all the given PCDs into depth
images, and crop the useless blank border. See the example in Fig.3.</p>
        <p>For the scene classification task, we pick the same size of images for each
class in the training sequences to avoid the imbalance between semantic classes.
In our work, we pick out the concept with minimum training data, count the
number of training images in the concept and set this number as the size of
training data for all the other concepts.</p>
        <p>Point Cloud Data</p>
        <p>Preprocessing for PCD</p>
        <p>RGB image
To choose appropriate kernel descriptors, we evaluate 6 kinds of kernel
descriptors (gradient, LBP, color, depth gradient, depth LBP, spin/surface normal)
through scene classification task on validation sequence. Fig.4 shows the
evaluation of all the 6 kernel descriptors. Depth kernel descriptors perform a little
better than visual kernel descriptors, due to the variations in the lighting
conditions between the training sequences and validation sequence. For the final test,
we select the three descriptors (gradient, depth gradient, depth LBP) which get
the highest classification accuracy on validation sequence, since the test sequence
is similar with validation sequence.</p>
        <p>After 3 optimal kernel descriptors are chosen, we apply spatial pyramid (1, 2×
2, 3×3) and perform the EMK transform with 1000 words. Then we get the image
feature with a total length of 1000 × 1 + 22 + 32 × 3 = 42000.
5.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Smoothing window radius</title>
        <p>As explained in Sec.4.2, the radius of smoothing window is important for the
final performance. We perform experiments on different radius r for validation
se0.760 10 20 radius of smoothing window60 70 80
30 40 50
0.880 10 20radius3o0f smoo4t0hing w5i0ndow 60 70 80
quences, and choose the r which corresponds to the highest accuracy. Fig.5 (left)
shows how the scene classification accuracy varies with the radius of smoothing
window r. We set the step width as 5, and find out that the accuracy reaches
the climax when r is around 25.</p>
        <p>Since the radius is related to the length of continuous scene images with
the same concept, and there is a proportion between the quantity of validation
images and test images, the estimated radius r should be multiplied the
proportion to fit the test sequence. According to Eq.15, we get the estimated radius
rtsecsetne = 44 for test sequence.</p>
        <p>rtsecsetne = rvscaelindeation ×</p>
        <p>Ntest
Nvalidation
.</p>
        <p>(15)</p>
        <p>For object recognition, Fig.5 (right) shows how the object recognition
accuracy varies with the radius of smoothing window r, and it reaches the climax
when r is around 10. Then we use the same method to get the estimated radius
rtoebsjtect = 18 for test sequence.
5.4</p>
      </sec>
      <sec id="sec-4-3">
        <title>Results on validation sequence</title>
        <p>We applied our method on the validation sequence, and computed the
classification accuracy for each scene concept and object category. See Table 1 and
Table 2 for detailed results. We notice that our method has good performance
on each concept except ’StudentOffice’ and ’TechnicalRoom’. This is due to the
large variation of luminance between training and validation sequences on these
two concepts images.
In the 5th edition of the Robot Vision Challenge,Our team ranked the first out
of six participants, results are listed in Table 3.</p>
        <p># Group
1 MIAR ICT
2 NUDT
3 SIMD*
4 REGIM
5 MICA
6 GRAM
In this paper we present our scheme on the 5th Robot Vision Challenge. Our
approach leverages the state-of-the-art methods in the fields of RGB-D image
classification. Among all the participants for the Challenge, our team ranked
the first, showing the effectiveness of our approach. Since the predefined objects
have high dependence on indoor scenes, we achieve high accuracy on object
recognition by using the representation of whole image. This method is useful
for the specific task of this Challenge, and we plan to investigate more effective
methods to better tackle general object recognition problem in future work.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work was supported in part by National Basic Research Program of China
(973 Program): 2012CB316400, in part by National Natural Science
Founda</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Microsoft</given-names>
            <surname>Kinect</surname>
          </string-name>
          . http://www.xbox.com/en-us/kinect.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Liefeng</given-names>
            <surname>Bo</surname>
          </string-name>
          , Xiaofeng Ren, and
          <string-name>
            <given-names>Dieter</given-names>
            <surname>Fox</surname>
          </string-name>
          .
          <article-title>Kernel descriptors for visual recognition</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          ,
          <volume>7</volume>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Liefeng</given-names>
            <surname>Bo</surname>
          </string-name>
          , Xiaofeng Ren, and
          <string-name>
            <given-names>Dieter</given-names>
            <surname>Fox</surname>
          </string-name>
          .
          <article-title>Depth kernel descriptors for object recognition</article-title>
          .
          <source>In Intelligent Robots and Systems (IROS)</source>
          ,
          <year>2011</year>
          IEEE/RSJ International Conference on, pages
          <fpage>821</fpage>
          -
          <lpage>826</lpage>
          . IEEE,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Liefeng</given-names>
            <surname>Bo</surname>
          </string-name>
          and
          <string-name>
            <given-names>Cristian</given-names>
            <surname>Sminchisescu</surname>
          </string-name>
          .
          <article-title>Efficient match kernel between sets of features for visual recognition</article-title>
          .
          <source>Advances in neural information processing systems</source>
          ,
          <volume>2</volume>
          (
          <issue>3</issue>
          ),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Rong-En</surname>
            <given-names>Fan</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kai-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho-Jui</surname>
            <given-names>Hsieh</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiang-Rui</surname>
            <given-names>Wang</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Chih-Jen Lin</surname>
          </string-name>
          .
          <article-title>Liblinear: A library for large linear classification</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          ,
          <volume>9</volume>
          :
          <fpage>1871</fpage>
          -
          <lpage>1874</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Andrew E. Johnson and
          <string-name>
            <given-names>Martial</given-names>
            <surname>Hebert</surname>
          </string-name>
          .
          <article-title>Using spin images for efficient object recognition in cluttered 3d scenes</article-title>
          .
          <source>Pattern Analysis and Machine Intelligence</source>
          , IEEE Transactions on,
          <volume>21</volume>
          (
          <issue>5</issue>
          ):
          <fpage>433</fpage>
          -
          <lpage>449</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Svetlana</given-names>
            <surname>Lazebnik</surname>
          </string-name>
          , Cordelia Schmid, and
          <string-name>
            <given-names>Jean</given-names>
            <surname>Ponce</surname>
          </string-name>
          .
          <article-title>Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories</article-title>
          .
          <source>In Computer Vision and Pattern Recognition</source>
          ,
          <year>2006</year>
          IEEE Computer Society Conference on, volume
          <volume>2</volume>
          , pages
          <fpage>2169</fpage>
          -
          <lpage>2178</lpage>
          . IEEE,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Timo</given-names>
            <surname>Ojala</surname>
          </string-name>
          , Matti Pietikainen, and
          <string-name>
            <given-names>Topi</given-names>
            <surname>Maenpaa</surname>
          </string-name>
          .
          <article-title>Multiresolution gray-scale and rotation invariant texture classification with local binary patterns</article-title>
          .
          <source>Pattern Analysis and Machine Intelligence</source>
          , IEEE Transactions on,
          <volume>24</volume>
          (
          <issue>7</issue>
          ):
          <fpage>971</fpage>
          -
          <lpage>987</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Xiaofeng</given-names>
            <surname>Ren</surname>
          </string-name>
          , Liefeng Bo, and
          <string-name>
            <given-names>Dieter</given-names>
            <surname>Fox</surname>
          </string-name>
          .
          <article-title>Rgb-(d) scene labeling: Features and algorithms</article-title>
          .
          <source>In Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <source>2012 IEEE Conference on</source>
          , pages
          <fpage>2759</fpage>
          -
          <lpage>2766</lpage>
          . IEEE,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Bernhard</surname>
            <given-names>Schölkopf</given-names>
          </string-name>
          , Alexander Smola, and
          <string-name>
            <surname>Klaus-Robert Müller</surname>
          </string-name>
          .
          <article-title>Nonlinear component analysis as a kernel eigenvalue problem</article-title>
          .
          <source>Neural computation</source>
          ,
          <volume>10</volume>
          (
          <issue>5</issue>
          ):
          <fpage>1299</fpage>
          -
          <lpage>1319</lpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Vladimir</given-names>
            <surname>Vapnik</surname>
          </string-name>
          .
          <article-title>The nature of statistical learning theory</article-title>
          . springer,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>