<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The University of Amsterdam's Concept Detection System at ImageCLEF 2009</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Koen E. A. van de Sande</string-name>
          <email>ksande@uva.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Theo Gevers</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arnold W. M. Smeulders</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Intelligent Systems Lab Amsterdam (ISLA), University of Amsterdam</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Performance</institution>
          ,
          <addr-line>Measurement, Experimentation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Our group within the University of Amsterdam participated in the large-scale visual concept detection task of ImageCLEF 2009. Our experiments focus on increasing the robustness of the individual concept detectors based on the bag-of-words approach, and less on the hierarchical nature of the concept set used. To increase the robustness of individual concept detectors, our experiments emphasize in particular the role of visual sampling, the value of color invariant features, the in uence of codebook construction, and the e ectiveness of kernel-based learning parameters. The participation in ImageCLEF 2009 has been successful, resulting in the top ranking for the large-scale visual concept detection task in terms of both EER and AUC. For 40 out of 53 individual concepts, we obtain the best performance of all submissions to this task. For the hierarchical evaluation, which considers the whole hierarchy of concepts instead of single detectors, using the concept likelihoods estimated by our detectors directly works better than scaling these likelihoods based on the class priors.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
        <kwd>I</kwd>
        <kwd>4</kwd>
        <kwd>7 [Image Processing and Computer Vision]</kwd>
        <kwd>Feature Measurement</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Robust image retrieval is highly relevant in a world that is adapting swiftly to visual
communication. Online services like Flickr show that the sheer number of photos available online is too
much for any human to grasp. Many people place their entire photo album on the internet. Most
commercial image search engines provide access to photos based on text or other metadata, as
this is still the easiest way for a user to describe an information need. The indices of these search
engines are based on the lename, associated text or (social) tagging. This results in
disappointing retrieval performance when the visual content is not mentioned, or properly re ected in the
associated text. In addition, when the photos originate from non-English speaking countries, such
as China, or the Netherlands, querying the content becomes much harder.</p>
      <p>Sampling
strategy</p>
    </sec>
    <sec id="sec-2">
      <title>Visual feature extraction</title>
    </sec>
    <sec id="sec-3">
      <title>Codebook transform</title>
    </sec>
    <sec id="sec-4">
      <title>Kernelbased learning</title>
      <p>Data flow conventions</p>
      <p>Image
Sampledimageregions
Visualfeatures
Codebook
Codewordfrequencydistribution
Codebooklibrary
Learningparameters
Conceptconfidence</p>
      <p>
        To cater for robust image retrieval, the promising solutions from literature are in majority
concept-based [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], where detectors are related to objects, like a telephone, scenes, like a kitchen,
and people, like big group. Any one of those brings an understanding of the current content. The
elements in such a lexicon o er users a semantic entry by allowing them to query on presence or
absence of visual content elements.
      </p>
      <p>
        The Large-Scale Visual Concept Detection Task [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] evaluates 53 visual concept detectors. The
concepts used are from the personal photo album domain: beach holidays, snow, plants, indoor,
mountains, still-life, small group of people, portrait. For more information on the dataset and
concepts used, see the overview paper [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Based on our previous work on concept detection [
        <xref ref-type="bibr" rid="ref15 ref19">19, 15</xref>
        ], we have focused on improving the
robustness of the visual features used in our concept detectors. Systems with the best performance
in image retrieval [
        <xref ref-type="bibr" rid="ref11 ref19">11, 19</xref>
        ] and video retrieval [
        <xref ref-type="bibr" rid="ref15 ref22">22, 15</xref>
        ] use combinations of multiple features for
concept detection. The basis for these combinations is formed by good color features and multiple
point sampling strategies.
      </p>
      <p>This paper is organized as follows. Section 2 de nes our concept detection system. Section 3
details our experiments and results. Finally, in section 4, conclusions are drawn.
2</p>
      <sec id="sec-4-1">
        <title>Concept Detection System</title>
        <p>We perceive concept detection as a combined computer vision and machine learning problem.
Given an n-dimensional visual feature vector xi, the aim is to obtain a measure, which indicates
whether semantic concept !j is present in photo i. We may choose from various visual feature
extraction methods to obtain xi, and from a variety of supervised machine learning approaches to
learn the relation between !j and xi. The supervised machine learning process is composed of two
phases: training and testing. In the rst phase, the optimal con guration of features is learned
from the training data. In the second phase, the classi er assigns a probability p(!j jxi) to each
input feature vector for each semantic concept.
2.1</p>
        <sec id="sec-4-1-1">
          <title>Sampling Strategy</title>
          <p>
            The visual appearance of a concept has a strong dependency on the viewpoint under which it is
recorded. Salient point methods [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] introduce robustness against viewpoint changes by selecting
points, which can be recovered under di erent perspectives. Another solution is to simply use
many points, which is achieved by dense sampling. We summarize our sampling approach in
Figure 2.
          </p>
          <p>
            Harris-Laplace point detector In order to determine salient points, Harris-Laplace relies on
a Harris corner detector. By applying it on multiple scales, it is possible to select the characteristic
scale of a local corner using the Laplacian operator [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ]. Hence, for each corner the Harris-Laplace
detector selects a scale-invariant point if the local image structure under a Laplacian operator has
a stable maximum.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>HarrisLaplace</title>
    </sec>
    <sec id="sec-6">
      <title>Dense sampling</title>
    </sec>
    <sec id="sec-7">
      <title>Spatial pyramid</title>
      <p>
        Dense point detector For concepts with many homogenous areas, like scenes, corners are
often rare. Hence, for these concepts relying on a Harris-Laplace detector can be suboptimal.
To counter the shortcoming of Harris-Laplace, random and dense sampling strategies have been
proposed [
        <xref ref-type="bibr" rid="ref4 ref6">4, 6</xref>
        ]. We employ dense sampling, which samples an image grid in a uniform fashion
using a xed pixel interval between regions. In our experiments we use an interval distance of 6
pixels and sample at multiple scales.
      </p>
      <p>
        Spatial pyramid weighting Both Harris-Laplace and dense sampling give an equal weight
to all keypoints, irrespective of their spatial location in the image frame. In order to overcome
this limitation, Lazebnik et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] suggest to repeatedly sample xed subregions of an image, e.g.
1x1, 2x2, 4x4, etc., and to aggregate the di erent resolutions into a so called spatial pyramid,
which allows for region-speci c weighting. Since every region is an image in itself, the spatial
pyramid can be used in combination with both the Harris-Laplace point detector and dense point
sampling [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Reported results using concept detection experiments are not yet conclusive in the
ideal spatial pyramid con guration, some claim 2x2 is su cient [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], others suggest to include 1x3
also [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. We use a spatial pyramid of 1x1, 2x2, and 1x3 regions in our experiments.
2.2
      </p>
      <sec id="sec-7-1">
        <title>Visual Feature Extraction</title>
        <p>
          In the previous section, we addressed the dependency of the visual appearance of semantic
concepts on the viewpoint under which they are recorded. However, the lighting conditions during
photography also play an important role. We [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] analyzed the properties of color features under
classes of illumination changes within the diagonal model of illumination change, and speci cally
for data sets consisting of Flickr images. In ImageCLEF, the images used also originate from
Flickr. Here we summarize the main ndings. We present an overview of the visual features used
in Figure 3.
        </p>
        <p>The features are computed around salient points obtained from the Harris-Laplace detector
and dense sampling.</p>
        <p>
          SIFT The SIFT feature proposed by Lowe [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] describes the local shape of a region using edge
orientation histograms. The gradient of an image is shift-invariant: taking the derivative cancels
out o sets [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Under light intensity changes, i.e. a scaling of the intensity channel, the gradient
direction and the relative gradient magnitude remain the same. Because the SIFT feature is
normalized, the gradient magnitude changes have no e ect on the nal feature. To compute SIFT
features, we use the version described by Lowe [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
        <p>OpponentSIFT OpponentSIFT describes all the channels in the opponent color space using
SIFT features. The information in the O3 channel is equal to the intensity information, while the
Invariant visual descriptors</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>SIFT</title>
    </sec>
    <sec id="sec-9">
      <title>OpponentSIFT</title>
    </sec>
    <sec id="sec-10">
      <title>C-SIFT</title>
    </sec>
    <sec id="sec-11">
      <title>RGB-SIFT</title>
      <p>other channels describe the color information in the image. The feature normalization, as e ective
in SIFT, cancels out any local changes in light intensity.</p>
      <p>
        C-SIFT The C-SIFT feature uses the C invariant [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which can be intuitively seen as the
gradient (or derivative) for the normalized opponent color space O1=I and O2=I. The I intensity
channel remains unchanged. C-SIFT is known to be scale-invariant with respect to light intensity.
See [
        <xref ref-type="bibr" rid="ref1 ref19">1, 19</xref>
        ] for detailed evaluation.
      </p>
      <p>
        RGB-SIFT For the RGB-SIFT, the SIFT feature is computed for each RGB channel
independently. Due to the normalizations performed within SIFT, it is equal to transformed color
SIFT [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. The feature is scale-invariant, shift-invariant, and invariant to light color changes and
shift.
2.3
      </p>
      <sec id="sec-11-1">
        <title>Codebook Transform</title>
        <p>
          To avoid using all visual features in an image, while incorporating translation invariance and
a robustness to noise, we follow the well known codebook approach, see e.g. [
          <xref ref-type="bibr" rid="ref19 ref20 ref23 ref6 ref8">8, 6, 23, 20, 19</xref>
          ].
First, we assign visual features to discrete codewords prede ned in a codebook. Then, we use the
frequency distribution of the codewords as a compact feature vector representing an image frame.
Two important variables in the codebook representation are codebook construction and codeword
assignment. An extensive comparison of codebook representation variables is presented by Van
Gemert et al. in [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. Here we detail codebook construction and codeword assignment using hard
and soft variants, following the scheme in Figure 4.
        </p>
        <p>Codebook construction We employ k-means clustering. K-means partitions the visual feature
space by minimizing the variance between a prede ned number of k clusters. The advantage of
the k-means algorithm is its simplicity. A disadvantage of k-means is its emphasis on clusters of
dense areas in feature space. Hence, k-means does not spread clusters evenly throughout feature
space. We x the visual codebook to a maximum of 4000 codewords.</p>
        <p>Hard-assignment Given a codebook of codewords, obtained from clustering, the traditional
codebook approach describes each feature by the single best representative codeword in the
codeCodebook representation</p>
        <p>Softassign
Clustering</p>
        <p>Hardassign</p>
        <p>Codebook
library</p>
        <p>SVM
book, i.e. hard-assignment. Basically, an image is represented by a histogram of codeword
frequencies describing the probability density over codewords.</p>
        <p>
          Soft-assignment In a recent paper [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], it is shown that the traditional codebook approach may
be improved by using soft-assignment through kernel codebooks. A kernel codebook uses a kernel
function to smooth the hard-assignment of image features to codewords. Out of the various forms
of kernel-codebooks, we selected codeword uncertainty based on its empirical performance [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ].
Codebook library Each of the possible sampling methods from Section 2.1 coupled with each
visual feature extraction method from Section 2.2, a clustering method, and an assignment
approach results in a separate visual codebook. An example is a codebook based on dense sampling
of RGB-SIFT features in combination with hard-assignment. We collect all possible codebook
combinations in a visual codebook library. Naturally, the codebooks can be combined using
various con gurations. For simplicity, we employ equal weights in our experiments when combining
codebooks to form a library.
2.4
        </p>
      </sec>
      <sec id="sec-11-2">
        <title>Kernel-based Learning</title>
        <p>Learning robust concept detectors from large-scale visual codebooks is typically achieved by
kernelbased learning methods. From all kernel-based learning approaches on o er, the support vector
machine is commonly regarded as a solid choice. An overview is given together with the codebook
transformations in Figure 4.</p>
        <p>
          Support vector machine We use the support vector machine framework [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] for supervised
learning of concepts. Here we use the LIBSVM implementation [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] with probabilistic output [
          <xref ref-type="bibr" rid="ref13 ref9">13, 9</xref>
          ].
The parameter of the support vector machine we optimize is C. In order to handle imbalance in
the number of positive versus negative training examples, we x the weights of the positive and
negative class by estimation from the class priors on training data. It was shown by Zhang et al.
[
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] that in a codebook-approach to concept detection the earth movers distance and 2 kernel
are to be preferred. We employ the 2 kernel, as it is less expensive in terms of computation.
3.1
        </p>
        <sec id="sec-11-2-1">
          <title>Concept Detection Experiments</title>
        </sec>
      </sec>
      <sec id="sec-11-3">
        <title>Submitted Runs</title>
        <p>
          We have submitted ve di erent runs. All runs use both Harris-Laplace and dense sampling with
the SVM classi er. We do not use the EXIF metadata provided for the photos. Our system has
been developed based on the PASCAL VOC [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] and TRECVID Sound and Vision datasets [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
For ImageCLEF, we have learned new concept models based on the provided annotations. The
only parameter speci cally optimized for this dataset is the slack parameter C of the SVM. All
other parameter settings are the same as in our PASCAL VOC 2008 system [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Extracting
features, training models and applying those models on the test set was nished within 72 hours.
        </p>
        <p>OpponentSIFT: single color descriptor with hard assignment.
2-SIFT: two color descriptors (OpponentSIFT and SIFT) with hard assignment.
4-SIFT: four color descriptors (OpponentSIFT, C-SIFT, RGB-SIFT and SIFT) with hard
assignment.</p>
        <p>Rescaled 4-SIFT: the same ordering of images as 4-SIFT, but with all concept detector
outputs linearly scaled so the number of images with a score &gt; 0:5 is equal to the concept
prior probability in the training set.</p>
        <p>
          Soft 4-SIFT: four color descriptors (OpponentSIFT, C-SIFT, RGB-SIFT and SIFT) with
soft assignment. The soft assignment parameters have been taken from our PASCAL VOC
2008 system [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
3.2
        </p>
      </sec>
      <sec id="sec-11-4">
        <title>Evaluation Per Concept</title>
        <p>In table 1, the overall scores for the evaluation of concept detectors are shown. As for the evaluation
of single detectors only the ranking of the images within a single concept matters, the rescaled
version of 4-SIFT achieves the exact same performance as 4-SIFT. We note that the 4-SIFT run
with hard assignment achieves not only the highest performance amongst our runs, but also over
all other runs submitted to the Large-Scale Visual Concept Detection task.</p>
        <p>In table 2, the Area Under the Curve scores have been split out per concept. We observe
that the three aesthetic concepts have the lowest scores. This comes as no surprise, because these
concepts are highly subjective: even human annotators only agree around 80% of the time with
each other. For virtually all concepts besides the aesthetic ones, either the Soft 4-SIFT or the
Hard 4-SIFT is the best run. This con rms our beliefs that these (color) descriptors are not
redundant when used in combinations. Therefore, we recommend the use of these 4 descriptors
instead of 1 or 2. The di erence in overall performance between the Soft 4-SIFT or the Hard
4SIFT run is quite small. Because the soft codebook assignment smoothing parameter was directly
taken from a di erent dataset, we expect that the soft assignment run could be improved if the
soft assignment parameter was selected with cross-validation on the training set. Together, our
Run name
4-SIFT
Rescaled 4-SIFT
Soft 4-SIFT
2-SIFT
OpponentSIFT</p>
        <p>Codebook
Hard-assignment
Hard-assignment
Soft-assignment
Hard-assignment
Hard-assignment</p>
        <p>Average EER</p>
        <p>Average AUC
0.2345
0.2345
0.2355
0.2435
0.2530
0.8387
0.8387
0.8375
0.8300
0.8217
runs obtain the highest Area Under the Curve scores for 40 out of 53 concepts in the Photo
Annotation task (20 for Soft 4-SIFT, 17 for 4-SIFT and 3 for the other runs). This analysis
has shown us that our system is falling behind for concepts that correspond to conditions we
have included invariance against. Our method is designed to be robust to unsharp images, so for
Out-of-focus, Partly-Blurred and No-Blur there are better approaches possible. For the concepts
Overexposed, Underexposed, Neutral-Illumination, Night and Sunny, recognizing how the scene
is illuminated is very important. Because we are using invariant color descriptors, a lot of the
discriminative lighting information is no longer present in the descriptors. Again, there should
be better approaches possible for these concepts, such as estimating the color temperature and
overall light intensity.</p>
        <p>Our system was developed on other datasets, and only the concept models were speci cally
learned for the Photo Annotation dataset. Its good performance on this dataset, without changing
the parameter settings, shows that it is generic and generalizes to multiple datasets. But, our
system only performs well on this dataset because the train and test set come from the same source
and have been obtained at the same time. Generalization across the boundary of multiple datasets
is still an unsolved problem: for photos downloaded from Flickr in a di erent season or general
web images, the performance will be signi cantly worse. However, all systems participating in the
Photo Annotation task are `overtrained' in this sense, and the models they learned too speci c.
An interesting avenue for future editions is to have a second test set with photos from a di erent
source or moment in time, so this problem can be investigated further.
For the hierarchical evaluation, overall results are shown in table 3. When compared to the
evaluation per concept, the Soft 4-SIFT run is now slightly better than the normal 4-SIFT run. Our
attempt to improve performance for the hierarchical evaluation measure using a linear rescaling
of the concept likelihoods has had the opposite e ect: the normal 4-SIFT run is better than the
Rescaled 4-SIFT run. Therefore, further investigation into building a cascade of concept classi ers
is needed, as simply using the individual concept classi ers with their class priors does not work.
4</p>
        <sec id="sec-11-4-1">
          <title>Conclusion</title>
          <p>Our focus on invariant visual features for concept detection in ImageCLEF 2009 has been
successful. It has resulted in the top ranking for the large-scale visual concept detection task in
terms of both EER and AUC. For 40 individual concepts, we obtain the best performance of all
submissions to the task. For the hierarchical evaluation, using the concept likelihoods estimated
by our detectors directly works better than scaling these likelihoods based on the class priors.</p>
        </sec>
        <sec id="sec-11-4-2">
          <title>Acknowledgements</title>
          <p>This work was supported by the EC-FP6 VIDI-Video project.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Burghouts</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Geusebroek</surname>
          </string-name>
          .
          <article-title>Performance evaluation of local color invariants</article-title>
          .
          <source>Computer Vision</source>
          and Image Understanding,
          <volume>113</volume>
          :
          <fpage>48</fpage>
          {
          <fpage>62</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.-C.</given-names>
            <surname>Chang</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.-J.</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <article-title>LIBSVM: a library for support vector machines</article-title>
          ,
          <year>2001</year>
          . Software available at http://www.csie.ntu.edu.tw/~cjlin/libsvm.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Everingham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Van</given-names>
            <surname>Gool</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. K. I.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Winn</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <source>The PASCAL Visual Object Classes Challenge</source>
          <year>2008</year>
          (
          <article-title>VOC2008) Results</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Fei-Fei</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Perona</surname>
          </string-name>
          .
          <article-title>A bayesian hierarchical model for learning natural scene categories</article-title>
          .
          <source>In IEEE Conference on Computer Vision and Pattern Recognition</source>
          , volume
          <volume>2</volume>
          , pages
          <fpage>524</fpage>
          {
          <fpage>531</fpage>
          , San Diego, USA,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Geusebroek</surname>
          </string-name>
          , R. van den Boomgaard,
          <string-name>
            <given-names>A. W. M.</given-names>
            <surname>Smeulders</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Geerts</surname>
          </string-name>
          .
          <article-title>Color invariance</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          ,
          <volume>23</volume>
          (
          <issue>12</issue>
          ):
          <volume>1338</volume>
          {
          <fpage>1350</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>F.</given-names>
            <surname>Jurie</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Triggs</surname>
          </string-name>
          .
          <article-title>Creating e cient codebooks for visual recognition</article-title>
          .
          <source>In IEEE International Conference on Computer Vision</source>
          , pages
          <volume>604</volume>
          {
          <fpage>610</fpage>
          , Beijing, China,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lazebnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ponce</surname>
          </string-name>
          .
          <article-title>Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories</article-title>
          .
          <source>In IEEE Conference on Computer Vision and Pattern Recognition</source>
          , volume
          <volume>2</volume>
          , pages
          <fpage>2169</fpage>
          {
          <fpage>2178</fpage>
          , New York, USA,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T. K.</given-names>
            <surname>Leung</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Malik</surname>
          </string-name>
          .
          <article-title>Representing and recognizing the visual appearance of materials using three-dimensional textons</article-title>
          .
          <source>International Journal of Computer Vision</source>
          ,
          <volume>43</volume>
          (
          <issue>1</issue>
          ):
          <volume>29</volume>
          {
          <fpage>44</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.-T.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-J. Lin</surname>
            , and
            <given-names>R. C.</given-names>
          </string-name>
          <string-name>
            <surname>Weng</surname>
          </string-name>
          .
          <article-title>A note on Platt's probabilistic outputs for support vector machines</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>68</volume>
          (
          <issue>3</issue>
          ):
          <volume>267</volume>
          {
          <fpage>276</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Lowe</surname>
          </string-name>
          .
          <article-title>Distinctive image features from scale-invariant keypoints</article-title>
          .
          <source>International Journal of Computer Vision</source>
          ,
          <volume>60</volume>
          (
          <issue>2</issue>
          ):
          <volume>91</volume>
          {
          <fpage>110</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Marszalek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Harzallah</surname>
          </string-name>
          , and J. van de Weijer.
          <article-title>Learning object representations for visual object class recognition, 2007. Visual Recognition Challenge workshop</article-title>
          , in conjunction with
          <source>IEEE International Conference on Computer Vision</source>
          , Rio de Janeiro, Brazil.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Nowak</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Dunker</surname>
          </string-name>
          .
          <article-title>Overview of the clef 2009 large scale visual concept detection and annotation task</article-title>
          .
          <source>In CLEF working notes 2009</source>
          , Corfu, Greece,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Platt</surname>
          </string-name>
          .
          <article-title>Probabilities for SV machines</article-title>
          . In A. J.
          <string-name>
            <surname>Smola</surname>
            ,
            <given-names>P. L.</given-names>
          </string-name>
          <string-name>
            <surname>Bartlett</surname>
          </string-name>
          , B. Scholkopf, and D. Schuurmans, editors,
          <source>Advances in Large Margin Classi ers</source>
          , pages
          <volume>61</volume>
          {
          <fpage>74</fpage>
          . MIT Press,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Over</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Kraaij</surname>
          </string-name>
          .
          <article-title>Evaluation campaigns and TRECVid</article-title>
          .
          <source>In ACM International Workshop on Multimedia Information Retrieval</source>
          , pages
          <volume>321</volume>
          {
          <fpage>330</fpage>
          ,
          <string-name>
            <surname>Santa</surname>
            <given-names>Barbara</given-names>
          </string-name>
          , USA,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>C. G. M. Snoek</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. E. A. van de Sande</surname>
            , O. de Rooij,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Huurnink</surname>
            ,
            <given-names>J. C. van Gemert</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>J. R. R.</given-names>
            <surname>Uijlings</surname>
          </string-name>
          , and et al. .
          <article-title>The MediaMill TRECVID 2008 semantic video search engine</article-title>
          .
          <source>In Proceedings of the 6th TRECVID Workshop</source>
          , Gaithersburg, USA,
          <year>November 2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>C. G. M. Snoek</surname>
            and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Worring</surname>
          </string-name>
          .
          <article-title>Concept-based video retrieval</article-title>
          .
          <source>Foundations and Trends in Information Retrieval</source>
          ,
          <volume>4</volume>
          (
          <issue>2</issue>
          ):
          <volume>215</volume>
          {
          <fpage>322</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>T.</given-names>
            <surname>Tuytelaars</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Mikolajczyk</surname>
          </string-name>
          .
          <article-title>Local invariant feature detectors: A survey</article-title>
          .
          <source>Foundations and Trends in Computer Graphics and Vision</source>
          ,
          <volume>3</volume>
          (
          <issue>3</issue>
          ):
          <volume>177</volume>
          {
          <fpage>280</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>K. E. A. van de Sande</surname>
            , T. Gevers, and
            <given-names>C. G. M.</given-names>
          </string-name>
          <string-name>
            <surname>Snoek</surname>
          </string-name>
          .
          <article-title>A comparison of color features for visual concept classi cation</article-title>
          .
          <source>In ACM International Conference on Image and Video Retrieval</source>
          , pages
          <volume>141</volume>
          {
          <fpage>150</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>K. E. A. van de Sande</surname>
            , T. Gevers, and
            <given-names>C. G. M.</given-names>
          </string-name>
          <string-name>
            <surname>Snoek</surname>
          </string-name>
          .
          <article-title>Evaluating color descriptors for object and scene recognition</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          , (in press),
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>J. C. van Gemert</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          <string-name>
            <surname>Veenman</surname>
            ,
            <given-names>A. W. M.</given-names>
          </string-name>
          <string-name>
            <surname>Smeulders</surname>
            , and
            <given-names>J. M.</given-names>
          </string-name>
          <string-name>
            <surname>Geusebroek</surname>
          </string-name>
          .
          <article-title>Visual word ambiguity</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          , (in press),
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>V. N.</given-names>
            <surname>Vapnik</surname>
          </string-name>
          .
          <source>The Nature of Statistical Learning Theory</source>
          . Springer-Verlag, New York, USA, 2nd edition,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <surname>B. Zhang.</surname>
          </string-name>
          <article-title>Video diver: generic video indexing with diverse features</article-title>
          .
          <source>In ACM International Workshop on Multimedia Information Retrieval</source>
          , pages
          <volume>61</volume>
          {
          <fpage>70</fpage>
          ,
          <string-name>
            <surname>Augsburg</surname>
          </string-name>
          , Germany,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Marszalek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lazebnik</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <article-title>Local features and kernels for classication of texture and object categories: A comprehensive study</article-title>
          .
          <source>International Journal of Computer Vision</source>
          ,
          <volume>73</volume>
          (
          <issue>2</issue>
          ):
          <volume>213</volume>
          {
          <fpage>238</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>