<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantic Contexts and Fisher Vectors for the ImageCLEF 2011 Photo Annotation Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yu Su</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Frederic Jurie</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>GREYC, University of Caen</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the participation of UNICAEN/GREYC to the ImageCLEF 2011 photo annotation task. The proposed approach uses visual image features and binary annotations of concepts only. In this approach, the annotations are predicted by SVM classi ers trained separately for each concept. The classi ers take Bag-of-Words histograms and sher vectors representations as inputs, both being combined at the decision level. Furthermore, contextual information is also embedded into the Bag-of-Words histograms to enhance their performance. The experimental results show that the combination of Bag-of-Words histograms and Fisher vectors brings signi cant performance increase (e.g. 4% for Mean Average Precision). Furthermore, the results of our best-run rank in top 3 for both concept and image level evaluations.</p>
      </abstract>
      <kwd-group>
        <kwd>Image classi cation</kwd>
        <kwd>Photo annotation</kwd>
        <kwd>Bag-of-Words model</kwd>
        <kwd>Semantic context</kwd>
        <kwd>Fisher Vectors</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The aim of the ImageCLEF 2011 photo annotation task is to automatically assign
to each image a set of concepts taken from a list of 99 possible pre-de ned visual
concepts. In this task, the participants are given 8000 training images associated
with the corresponding 99 binary labels, each of which corresponds to a visual
concept, as well as the photo tagging ontology, EXIF data and Flickr user tags. In
the test phase, the participants are requested to give to each test image the labels
of all the visual concepts describing the image. The evaluation of performance
is done at concept and image levels. For the former, Mean Average Precision
(MAP) is computed for each concept. For the latter, F-Measure (F-ex) and the
Semantic R-Precision (SR-Precision) are computed for each image. For more
details on this task, please refer to [7] .</p>
      <p>In our participation, we did not use photo tagging ontology, EXIF data and
Flickr user tags. Our results are only based on visual image features.
Specifically, we extracted di erent types of local features (e.g. SIFT) from images
and then adopted Bag-of-Words (BoW) model to aggregate local features into
a global image descriptor. Our participation is mainly inspired by the work of
Su and Jurie [11], which proposed to embed some contextual information into
the BoW model. In addition, some improvements over [11] are also proposed.
Indeed, Fisher Vectors (FV) have been reported to give good performance on
both object recognition and image retrieval tasks [9]. Thus, we also computed FV
from images and combined them with the context-embedded BoW histograms at
decision level, i.e, training classi ers for Fisher Vectors and context-embedded
BoW histograms separately and combining classi ers by averaging their
outputs. As to photo annotation, the above process is performed for each concept
independently and the averaged classi er outputs are used as the con dences of
concept occurrence.</p>
      <p>The organization of this paper is as follows: In section 2, we describe local
features used in our method. Then, we explain how to combine BoW model with
both semantic contexts (section 3) and FV (section 4). Experimental evaluation
is given in section 5, followed by a conclusion in the last section.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Visual Features</title>
      <p>In our method, 6 kinds of visual features are extracted from each image, which
are introduced in the following paragraph. Before feature computation, the
images are scaled to be at most 300 300 pixels, with their original aspect ratios
maintained. Except for LAB features which encode color information, color
images are rst transformed to grayscale.</p>
      <p>SIFT Vector quantized SIFT descriptors [6] are computed for 5000 image
patches with randomly selected positions and scales (with scales from 16 to
64 pixels), and are quantized to 1024 k-means centers.</p>
      <p>
        HOG HOG descriptors [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] are densely extracted on a regular grid at step of
8 pixels. On each node of the grid a 31 dimensional descriptor is computed and
then 2 2 neighboring descriptors are concatenated to form a descriptor of 124
dimensions. HOG features are nally vector quantized to 256 k-means centers.
Textons Texton descriptors [12] are generated by computing the output of
36 Gabor lters with di erent scales and orientations for each pixel, and then
quantized to 256 k-means centers.
      </p>
      <p>SSIM Self-similarity descriptors [10] are computed on a regular grid at step of
ve pixels. Each feature is obtained by computing the correlation map of a patch
of 5 5 in a window with radius of 40 pixels, then quantizing it in 3 radial bins
and 10 angular bins, obtaining 30 dimensional descriptor vectors. Self-similarity
features are nally quantized to 256 k-means centers.</p>
      <p>LAB LAB descriptors [4] are computed for each pixel and then quantized to
128 k-means centers.</p>
      <p>
        Canny Canny edge descriptors [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] are computed for each pixel and then
quantized to 8 orientation bins.
      </p>
      <p>Finally, concatenating all BoW histograms gives a 1928-dimensional feature
vector which can describe an image or a image region.</p>
    </sec>
    <sec id="sec-3">
      <title>3 Image Representation by Embedding Semantic</title>
    </sec>
    <sec id="sec-4">
      <title>Contexts into BoW Model</title>
      <p>In this section, we rst review how do we de ne the semantic contexts and
embed them into BoW model as introduced in [11]. Then we introduce our
improvements over this method.</p>
      <p>global
scene
(35)
local
scene
(16)
color
(8)
shape
(7)
material
(14)
object
(30)
train station
bedroom
sky
green
triangle
metal
face
motorbike
lab zoo gym jail store casino church harbor
kitchen highway cemetery bathroom
classroom industrial restaurant courtroom
+ tshuepaetremr alaruknetdrliybroafrfyicteutnennenliss_ucbouurbrt
dining_room swimming_pool hospital_room
shopping_mall living_room conference_room
parking_lot indoor/outdoor city/landscape
+ bocueiladninrgivceorarsotaddessneortwfosroeisltsgtrreaests wlaakleltmreoeuntain
+ red black blue gray orange white yellow
+ box cylinder circle cone oval pyramid
+ cpearpaemr ipclsasctlioctrhubfebaetrhsetrognleaswsahtearirwleososdlefautrher
animal_flipper animal_head animal_wing
door window hand hooves screen wheel
+ acihrapilracnoewbidciyncilnegbtairbdlebdooatgbhootrtlseebpuesrscoanr cat
pottedplant sheep sofa train tv/monitor
In [11], 110 semantic contexts are de ned by hand with the intention of providing
abundant semantic information for image description. (see Fig. 1). Two types
of contexts are distinguished: global contexts including global scenes and local
contexts including local scenes, colors, shapes, materials and objects.</p>
      <p>For each semantic context, we learn a SVM classi er with linear kernel
(hereafter called as context classi ers). For the global contexts, the classi ers are
where
and d is a distance function (e.g., the L2 norm).</p>
      <p>Marginalizing p(vj jI) over di erent local contexts gives:
learned on whole images described by BoW histograms. For the local contexts,
the classi ers are learned on some randomly sampled image regions described
again by BoW histograms. The training images are automatically downloaded
from Google image search by using the name of context as query. After the
manual annotation, about 400 relevant images are reserved for each context. They
are used as positive images for the corresponding context while images from the
other contexts are considered as negatives.</p>
      <p>
        In test phase, images (for global contexts) or regions (for local contexts) are
input to context classi ers and a sigmoid function is used to transform the
original decision values to probabilities (refer to [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]).
3.2
      </p>
      <p>Embedding Semantic Contexts into BoW model
Assume that, for an image I, a set of local features fi; i = 1; : : : ; N are extracted
from it, where N is the number of local features. The BoW model consists of
V visual words vj ; j = 1; : : : ; V . The traditional BoW feature for vj measures
the occurrence probability of vj on image I, say p(vj jI). In practice, p(vj jI) is
usually computed by:
p(vj jI) =
where ck is the k-th context, C is the number of local contexts (75 in our case),
p(vj jck; I) is the context-speci c occurrence probability of vj on image I, p(ckjI)
is the occurrence probability of context ck on image I.</p>
      <p>On the other hand, the second term of Eq. 3, which gives the distribution
of di erent contexts on image I, can also provide rich information to describe
the image, as shown by [13]. For example, knowing an image is composed of one
third of sky, one third of sea and one third of beach, brings a lot of information
regarding the content of this image. At the end, images are eventually represented
by multiple context-speci c BoW histograms, i.e., p(vj jck; I) and a vector of
context-occurring probabilities, i.e., p(ckjI).</p>
      <p>In [11], p(vj jck; I) is constructed by modeling the probabilistic distribution
of context ck on image I which is estimated by dividing image I into a set of
regions Ip and predicting the occurrence probabilities of ck for each region (by
using context classi ers). By denoting Ip(fi) the set of image regions which cover
the local feature fi, we de ne:
p(vj jck; I) =
1 XN (fi; vj )p(ckjIp(fi));
where p(ckjIp(fi) can be considered as the weight of local feature fi. In practice,
p(ckjIp(fi)) is computed by averaging the outputs of the context classi er (for
ck) on Ip(fi).</p>
      <p>As to p(ckjI), it can be easily computed by averaging the outputs of the
context classi ers (for ck) on all image regions in Ip. This process is similar to the
computation of p(ckjIp(fi)) in previous subsection. In addition, we also represent
image I by the occurrence probabilities of global contexts. These probabilities are
computed by running the corresponding context classi ers on the whole image.
Finally, an image is represented by concatenating the occurrence probabilities
of both global and local contexts, i.e.,</p>
      <p>(p(c1jI); : : : ; p(cC jI); p(cC+1jI); : : : ; p(cC0 jI));
where C0 is the number of all contexts (110 in our case) and C is the number of
local contexts (75 in our case). We call this image descriptor as semantic features.
3.3</p>
      <sec id="sec-4-1">
        <title>Our improvements over [11]</title>
        <p>The above subsection reviewed the process of constructing context-speci c BoW
histograms introduced in [11]. For our participation to the ImageCLEF 2011
photo annotation task, some improvements over this method are proposed. First,
we learn a speci c vocabulary for each semantic context rather than use a
uniform vocabulary for all contexts as in [11]. Second, instead of selecting a single
context for each visual word as in [11], we train a classi er for each
contextspeci c BoW histogram and then combine all the classi ers. Detailed
implementation of these two improvements are given in the following paragraph.</p>
        <p>In the traditional vocabulary learning process, local features extracted from
a set of images are randomly (or uniformly) sampled and then vector quantized
to get visual words. Di erently, when learning our context-speci c vocabulary,
the sampling of local features is based on the distribution of this context on
images. Speci cally, more local features are sampled at the image regions with
higher context-occurring probabilities (brighter image regions in Fig. 2). In
practice, this process is implemented by assigning each local feature fi a probability
p(ckjIp(fi)) (de ned in section 3.2) and sampling local features based on their
probabilities, which is formulated as follows.</p>
        <p>s(fi) =
1 if p(ckjIp(fi)
0 else
ri
(4)
(5)
where s(fi) indicates whether the local feature fi is selected or not and ri are
random numbers which are uniformly sampled between 0 and 1.</p>
        <p>After sampling local features for each context, k-means is used to build
multiple context-speci c vocabularies. An image is then represented by multiple
context-speci c BoW histograms. The construction of context-speci c BoW
histogram is the same as that in 3.2 (see Eq.4)</p>
        <p>Concatenating all the context-speci c BoW histograms leads to a very high
dimensional feature vector (in our case 1928 75=144,600D). Thus, we train a
classi er for each context-speci c BoW histogram and combine the classi ers by
averaging their outputs.</p>
        <p>semantic contexts</p>
        <p>SPM channels</p>
        <p>Recall that the way we embed contextual information into BoW model is
based on weighting local features (see Eq.4). It is similar to the well-known
spatial pyramid matching (SPM) [5] which divides an image into grids and build
a histogram for each grid. This process can be also considered as weighting local
features: for certain grid, the weights of the local features within it is set to 1
and the weights of other local features are set to 0. Although less exible than
context-based weights, the binary weights in SPM are more stable which is also
favorable. Thus, we also train classi ers for BoW histograms of SPM channels.
In our method, a three level pyramid, 1 1, 2 2, 3 1 (totally 8 channels) is
used. It is worthwhile to point out that, di erent from traditional SPM, we learn
a speci c vocabulary for each SPM grid based on local features within this grid.</p>
        <p>Finally, we train a classi er for the semantic features and combine it with the
classi ers for context-speci c BoW histograms and SPM channels by averaging
their outputs. For both BoW histograms and semantic features, classi ers are
learned by SVM with chi-square kernel. The whole process is illustrated in Fig.2.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Image Representation by Fisher Vectors</title>
      <p>Similar to the BoW model, Fisher Vectors [8] can also be used to aggregate
local features into a global descriptor which is called Fisher Vectors (FV). FV
can be considered as an extension of BoW histograms. They encodes how the
parameters of the model should be changed to represent the image, rather than
only consider the number of occurrences of each visual word as in BoW model.
In our participation, we adopted the improved FV as introduced in [9] which is
shown to outperform BoW histogram on some large-scale image retrieval tasks.</p>
      <p>For an image, we computed the FV for each kind of local features except
Canny for which no actual visual word exist. As in [9], a three level pyramid,
1 1, 2 2, 3 1 (8 channels in total) is used to enhance the performance of
sher vector.</p>
      <p>For SIFT and HOG descriptors, PCA is used to reduce the dimension of
descriptors to 64. For SIFT descriptors, a 64-centroid Gaussian mixture model
(GMM) is computed to construct sher vector whose dimensionality is
therefore 64 64 8 2 = 65; 536. For HOG, Texton, SSIM and LAB descriptors,
64-centroid GMMs are learned therefore the dimensionalities of sher vector for
these descriptors are 16,384, 9,216, 7,680 and 768 respectively. Then we train a
classi er (SVM with linear kernel) for each sher vector and combine all
classiers by averaging their outputs. The whole process is illustrated in Fig.3. Please
note that the semantic contexts are not used for this representation.
5</p>
    </sec>
    <sec id="sec-6">
      <title>ImageCLEF Evaluation</title>
      <p>In our participation, we submitted four runs to the photo annotation task. In
this section, we describe these runs and compare their performances with other
visual-only runs.
5.1</p>
      <sec id="sec-6-1">
        <title>Description of Our Runs</title>
        <p>Run 1: MultiFeat Chi2SVM In this run, context-speci c BoW histograms,
each of which corresponds to a semantic context, as well as the semantic features
fisher vector for HOG (16384D)
fisher vector for Texton (9216D)
fisher vector for SSIM (7680D)
fisher vector for LAB (768D)
Linear SVM
Linear SVM
Linear SVM
Linear SVM
combination
are used to describe images. As illustrated in Fig.2, we trained separated
classi ers (SVMs with chi-square kernel) for both context-speci c BoW histograms
and semantic features and then combine them by averaging their outputs.
Run 2: BoW+FisherKernel In this run, we combined all the classi ers in
run 1 and classi ers for sher vectors of di erent features (refer to Fig.3). The
combination is performed by averaging the outputs of all classi ers.
Run 3: SVMOutput In this run, the con dences of all 99 concepts obtained
from run 2 are used as a new image descriptor. A classi er (SVM with chi-square
kernel) is learned on this descriptor and used to give the con dences of concepts.
By doing so, we hope to bene t from the correlation of di erent concepts.
Run 4: BoW+FisherKernel+SVMOuput In this run, we averaged the
con dences obtained from run 1, 2 and 3.</p>
        <p>
          In our participation, we used the implementation of LIBSVM [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] to learn SVM
classi er. The value of the SVM parameter C and the normalization factor
of chi-square kernel are determined by vefold cross-validation. As to the image
regions used for learning local context classi ers and generating saliency maps,
on each image we sampled 100 regions with random positions and scales (with
scales from 20% to 40% of the image size).
        </p>
        <p>For concept level evaluation, the classi er outputs are used as con dences
directly. For image level evaluation, the real valued con dences are binarized by
a threshold which is determined by vefold cross-validation for each run.
The performances (MAP, F-ex and SR-Precision) of our runs are listed in Table
1. It can be concluded that the performance of context-speci c BoW histograms
is signi cantly enhanced by combining them with sher vectors. It is worthwhile
to point out that, according to our experiments on training data, the performance
of sher vectors alone is comparable to that of context-speci c BoW histograms.
Another conclusion drawn from Table 1 is that using classi er outputs as new
features does not bring any improvement. Thus we need to design more powerful
methods to utilize the correlation of concepts.</p>
        <sec id="sec-6-1-1">
          <title>Runs</title>
        </sec>
        <sec id="sec-6-1-2">
          <title>MultiFeat Chi2SVM</title>
        </sec>
        <sec id="sec-6-1-3">
          <title>BoW+FisherKernel</title>
        </sec>
        <sec id="sec-6-1-4">
          <title>SVMOutput</title>
        </sec>
        <sec id="sec-6-1-5">
          <title>BoW+FisherKernel+SVMOuput MAP 34:2</title>
          <p>38:2
34:5
38:2
F-ex
56:0
60:0
49:1
59:2</p>
        </sec>
        <sec id="sec-6-1-6">
          <title>SR-Precision 69:4</title>
          <p>72:7
65:1
72:5</p>
          <p>For more detailed result, Fig.4 gives the MAPs of 99 concepts obtained from
Run 2. For some concepts, the MAPs are very low, e.g. less than 10%. The
reason is either that the concept is hard to predict (e.g. abstract) or that the
number of training samples is quite small (e.g. only 12 images are annotated
with skateboard).
100%
90%
80%
70%
60%
50%
40%
30%
20%
10%
0% lliittruaannoueNmI lruoNBrssoenoNPtruoodO ykS yaD lsduoC trcseuaeapdnaLNltsnaPitrssenuenuSSlltrrrydueaPBltraaun ittrraoP rseeT traeWtcue iliitsndugghBSIrndoo lirsenoengSPithgN ltduA liiftyeC dooF lcamlisnaAmiilrsydneaFmFitsnaounMiitvcena rcoaM lfeaemaeS lrseoFw liceehV ynunS rrkneadaPGirgupoBGyppha lisycadoaheBHiilfteSLgdo rca ttreeS fynnu lrpuoaSmGitcvae rxsedenopdeU lliccanhoem ittrrccheueAilsuaVtrsAlitfryaeP tunuAmilcycbe keaL iittrcsssehepeonAmI rueSmmiltroonuMBirveR irdb tfcsuouOfo trybodap yoT regaeenT ltsuanpena itrneWiitgannP noSw itscne ybaB ilraenpa liltrvyaeuaOQitrna ycanF trseeD laem rxvseeedopOrcsya irgpnS ildhC tca odhaSw trsopS lrveaT rchhuC rkoWisph iiliftrcaa iredgB rseoh irgbon liIttrcssanuenumMiitffraG irchuepo lrsdoeonpiltcceahn inaR ifsh ttrscaba itryahdB trskaedabo</p>
          <p>Finally, we compare our best run (Run 2: BoW+FisherKernel) with the best
runs (visual-only) of several competitors. It can be seen from Table 2 that no run
gave the best result for both concept and image level evaluation. For concept level
evaluation (MAP as performance measure), TUBFI scores performed best, while
for image level evaluation (F-ex and SR-Precision as performance measures),
ISIS runpa-UvA-coreA performed best. Our best run ranks in the second place
for both MAP and F-ex and the third place for SR-Precision.</p>
        </sec>
        <sec id="sec-6-1-7">
          <title>Runs</title>
        </sec>
        <sec id="sec-6-1-8">
          <title>BPACAD bpacad avg cns</title>
        </sec>
        <sec id="sec-6-1-9">
          <title>ISIS runpa-UvA-coreA</title>
        </sec>
        <sec id="sec-6-1-10">
          <title>LIRIS 4visual model 4</title>
        </sec>
        <sec id="sec-6-1-11">
          <title>TUBFI scores</title>
        </sec>
        <sec id="sec-6-1-12">
          <title>Our best run</title>
          <p>In our participation to the ImageCLEF photo annotation task, multiple visual
features were used for representing the images. We embedded contextual
information into the traditional Bag-of-Words model and further combined it with
sher vector which has been shown to have good performance on image classi
cation and retrieval tasks. The evaluation results showed that the performance
of the Bag-of-Words model can be signi cantly enhanced by combining it with
semantic contexts and sher vector. Our best run gave 38.2%, 59.2% and 72.5%
for MAP and F-ex and SR-Precision respectively, while the best results of
visualonly runs are 38.8%, 61.2% and 73.4% respectively.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgement</title>
      <p>This work was partly realized under the Quaero Programme, funded by OSEO,
French State agency for innovation.
4. Hunter, R.: Photoelectric color di erence meter. JOSA 48(12), 985{993 (1958)
5. Lazebnik, S., Schmid, C., Ponce, J.: Beyond bags of features: Spatial pyramid
matching for recognizing natural scene categories. In: CVPR (2006)
6. Lowe, D.: Distinctive image features from scale-invariant keypoints. International
journal of computer vision 60(2), 91{110 (2004)
7. Nowak, S., Nagel, K., Liebetrau, J.: The clef 2011 photo annotation and
conceptbased retrieval tasks. CLEF 2011 working notes (2011)</p>
      <sec id="sec-7-1">
        <title>8. Perronnin, F., Dance, C.: Fisher kernels on visual vocabularies for image catego</title>
        <p>rization. In: CVPR (2006)</p>
      </sec>
      <sec id="sec-7-2">
        <title>9. Perronnin, F., Sanchez, J., Mensink, T.: Improving the sher kernel for large-scale</title>
        <p>image classi cation (2010)
10. Shechtman, E., Irani, M.: Matching local self-similarities across images and videos.</p>
        <p>In: CVPR (2007)
11. Su, Y., Jurie, F.: Visual word disambiguation by semantic contexts. In: ICCV
(2011)
12. Varma, M., Zisserman, A.: A statistical approach to texture classi cation from
single images. International Journal of Computer Vision 62(1), 61{81 (2005)
13. Vogel, J., Schiele, B.: Semantic modeling of natural scenes for content-based image
retrieval. International Journal on Computer Vision 72(2), 133{157 (2007)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Canny</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A computational approach to edge detection</article-title>
          .
          <source>Pattern Analysis and Machine Intelligence</source>
          ,
          <source>IEEE Transactions on (6)</source>
          ,
          <volume>679</volume>
          {
          <fpage>698</fpage>
          (
          <year>1986</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <issue>2</issue>
          .
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.J.:</given-names>
          </string-name>
          <article-title>LIBSVM: A library for support vector machines</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          <volume>2</volume>
          ,
          <issue>27</issue>
          :1{
          <fpage>27</fpage>
          :
          <fpage>27</fpage>
          (
          <year>2011</year>
          ), software available at http://www.csie.ntu.edu.tw/~cjlin/libsvm
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dalal</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Triggs</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Histograms of oriented gradients for human detection</article-title>
          .
          <source>In: CVPR</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>