<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>TELECOM ParisTech at ImageCLEF 2010 Photo Annotation Task: Combining Tags and Visual Features for Learning-Based Image Annotation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hichem Sahbi ?</string-name>
          <email>hichem.sahbi@telecom-paristech.fr</email>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2010</year>
      </pub-date>
      <abstract>
        <p>In this paper, we describe the participation of TELECOM ParisTech in the ImageCLEF 2010 Photo Annotation challenge. This edition focuses on promoting combination between visual and tag features in order to enhance photo annotation. An image collection is supplied with tags which are used both for training and testing. Our training approach consists of building SVM classi ers and kernels which take into account the similarity between visual features as well as tags. The results clearly corroborate (i) the complementarity of tags and visual descriptors and (ii) the e ectiveness of SVM classi ers in photo annotation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Xi Li ?;??
Recent years have witnessed a rapid increase of image sharing spaces, such as
Flickr, due to the spread of digital cameras and mobile devices. An urgent need
is how to e ectively search these huge amounts of data and how to exploit the
structure of these sharing spaces. A possible solution is CBIR (Content-Based
Image Retrieval); where images are represented using low-level visual features
(color, texture, shape, etc.) and searched by analyzing and comparing those
features. However, low-level visual features are usually unable to deliver satisfactory
semantics, resulting in a gap between them and the high-level human
interpretations. To address this problem, a variety of machine learning techniques were
introduced in order to discover the intrinsic correspondence between visual
features and semantics of images and allow to predict keywords for images.</p>
    </sec>
    <sec id="sec-2">
      <title>Work</title>
      <p>
        Conventionally, image annotation is converted into a classi cation problem.
Existing state of the art methods (for instance [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]) treat each keyword or concept
as an independent class, and then train the corresponding concept-speci c
classi er to identify images belonging to that class, using a variety of machine
learning techniques such as hidden Markov models [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], latent Dirichlet allocation [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
probabilistic latent semantic analysis [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and support vector machines [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The
aforementioned annotation methods may also be categorized into two branches;
region-based requiring a preliminary step of image segmentation [
        <xref ref-type="bibr" rid="ref12 ref2">2, 12</xref>
        ], and
holistic [
        <xref ref-type="bibr" rid="ref25 ref6">6, 25</xref>
        ] operating directly on the whole image space. In both cases,
training is achieved in order to learn how to attach keywords with the corresponding
visual features.
      </p>
      <p>
        The above annotation methods heavily rely on their visual features for image
annotation. Due to the semantic gap, they are unable to fully explore the
semantic information inside images. Another class of annotation methods has emerged
that takes advantage of extra information (tags, context, users' feedback,
ontologies, etc.) in order to capture the correlations between images and concepts.
A representative work is the cross-media relevance model (CMRM) [
        <xref ref-type="bibr" rid="ref6 ref9">6, 9</xref>
        ], which
learns joint statistics of visual and concepts and its variants [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. The model
uses the keywords shared by similar images to annotate new ones. In [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], the
similarity measure between images integrates contextual information for
concept propagation. Semi-supervised annotation techniques were also studied and
usually rely on graph inference [10{13]. The original work, in [
        <xref ref-type="bibr" rid="ref26 ref3">3, 26</xref>
        ], is inspired
from machine translation and considers images and keywords as two di erent
languages; in that case, image annotation is achieved by translating visual words
into keywords.
      </p>
      <p>
        Other existing annotation methods focus on how to de ne an e ective distance
measure for exploring the semantic relationships between concepts in large scale
databases. In [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], the Normalized Google similarity Distance (NGD) is
proposed by exploring the textual information available on the web. It is a measure
of semantic correlations derived from counts returned by Google's search engine
for a given set of keywords. Following the idea of [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], the Flickr distance [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
is proposed to precisely characterize the visual relationships between concepts.
Each one is represented by a visual language model in order to capture its
underlying visual characteristics. Then, a Flickr distance is de ned, between two
concepts, as the square root of Jensen-Shannon (JS) divergence between the
corresponding visual language models. Other techniques consider extra knowledge
derived from ontologies (such as the popular WordNet [14{16]) in order to enrich
annotations [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The method in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] introduces a visual vocabulary in order to
improve translation model in the preprocessing stage of visual feature extraction.
A directed acyclic graph is used to model the causal strength between concepts,
and image annotation is performed by inference on this graph [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref17 ref18">17, 18</xref>
        ],
the semantic ontology information is integrated in the post processing stage in
order to further re ne initial annotations.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Motivation and The Proposed Method at a Glance</title>
      <p>
        Among the most successful annotation methods, those based on machine learning
and mainly support vector machines; show a particular interest as they are
performant and theoretically well grounded [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Support vector machines [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ],
basically require the design of similarity measures, also referred to as kernels, which
should provide high values when two images share similar structures/appearances
and should be invariant, as much as possible, to the linear and non-linear
transformations. They also satisfy positive de niteness which ensures, according to
Vapnik's SVM theory [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], optimal generalization performance and also the
uniqueness of the SVM solution. In practice, kernels should not depend only
on intrinsic aspects of images (as images with the same semantic may have
different visual and textual features), but also on di erent sources of knowledge
including context.
      </p>
      <p>
        In this work, we introduce an image annotation framework based on a new
similarity measure which takes high values not only when images share the same
visual content but also the same context. The context of an image is de ned
as the set of images, with the same tags, and exhibiting better semantic
descriptions, compared to both pure visual and tag based descriptions. The issue
of combining context and visual content for image retrieval is not new (see for
instance [28{30]) but the novel part of this work aims to (i) integrate context,
in similarity design useful for classi cation and annotation, and (ii) plug this
similarity in support vector machines in order to take bene t from their well
established generalization power [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. This type of similarity will be referred to as
context-based while those relying only on the intrinsic visual or textual content
will be referred to as context-free. Again, our proposed method goes beyond the
naive use of low level features and context-free similarities (established as the
standard baseline in image retrieval) in order to design a similarity applicable
to annotation and suitable to integrate the \contextual" information taken from
tagged datasets. In the proposed method, two images (even with di erent visual
content and even sharing di erent tags) will be declared as similar if they share
the same visual context1. This is usually useful as tags in data may be noisy and
misspelled. Furthermore, the intrinsic visual content of images might not always
be relevant especially for categories exhibiting large variation of the underlying
visual aspects.
      </p>
      <p>Through this work, an image database is modeled as a graph where nodes are
pictures and edges correspond to shared tags (links) between images. We
design our similarity as the solution of a constrained energy function containing
a delity term which measures visual similarity between images and a context
criterion that captures the similarity between the underlying links.
1 Visual context is de ned as the set of images sharing the same tags.</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>4.1</p>
      <p>
        MIR Flickr/ImageCLEF Collection
We evaluated our annotation method on the MIR Flickr dataset containing
18; 000 images belonging to 93 categories (for instance \sky, clouds, water, sea,
river,..."), among them 8; 000 are used for training and 10; 000 for testing. The
whole dataset is annotated but ground truth is provided only for the training
set. The MIR Flickr collection contains 1; 386 tags (provided by the Flickr users)
which occur in at least 20 images, with an average total number of 8:94 tags per
image (see Fig. 1 and [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]).
Recent years have witnessed a great success of the bag-of-features
representation in a wide range of application, such as image retrieval, image classi cation,
image segmentation, object recognition, etc. Inspired by text classi cation,
visual feature spaces are conventionally partitioned by vector quantization (e.g.
kmeans) into several subspaces, each of which corresponds to a visual word. As a
consequence, the bag-of-feature representation is converted to the bag-of-words
(BoW). Since using a basic histogram of orderless visual words, the BoW
representation only re ects the global statistical properties of visual words, and
ignores their spatial layout. Therefore, the orderless BoW representation has a
low descriptive capability of capturing the geometric relationships among visual
words. Motivated by this, we use, in this evaluation campaign, the same approach
as in [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] in order to better capture the spatial layout of images. The algorithm
is based on a spatial pyramid representation, which constructs a multi-level
spatial pyramid by block division. For any block at each level, a traditional BoW
representation in the SIFT feature space is used. In this way, we have a set of
block-speci c BoW histograms at multiple levels. As a result, the geometric
relationships among visual words can be e ectively captured.
      </p>
      <p>Given a test picture, the goal is to predict which categories (object classes)
are present into that picture. This task is commonly known as concept detection.
For this purpose, we trained "one-versus-all" SVM classi ers for each category;
we repeat this training process through di erent folds (20 times), for each
category, and we take the average score of the underlying SVM classi ers on the
test picture. This makes classi cation results less sensitive to sampling and
unbalanced classes. Performances are reported using the Mean Average Precision
(MAP), the Equal Error Rate (EER) and the Area Under Curve (AUC). Higher
MAP, AUC and lower EER imply better performance. Figs. (2, 3, 4) show the
annotation results of our best ImageCLEF run through di erent classes.</p>
      <p>1
We introduced in this work our participation in the ImageCLEF 2010 Photo
Annotation Task. Our annotation method takes into account image features as
well as their context links (taken from tags in the MIR Flickr collection) in order
to achieve SVM learning and classi cation. Future extensions of this work include
extra processing of these tags prior to SVM learning and further evaluations in
the next campaigns.</p>
      <p>Classes</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgement</title>
      <p>This work is supported by the French National Research Agency (ANR) under
the AVEIR project.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>G.</given-names>
            <surname>Carneiro</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Vasconcelos</surname>
          </string-name>
          , \
          <article-title>Formulating semantic image annotation as a supervised learning problem"</article-title>
          ,
          <source>in Proc. of CVPR</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          , \
          <article-title>Automatic linguistic indexing of pictures by a statistical modeling approach,"</article-title>
          <source>IEEE Trans. on PAMI.</source>
          ,
          <volume>25</volume>
          (
          <issue>9</issue>
          ):
          <fpage>1075</fpage>
          -
          <lpage>1088</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>K.</given-names>
            <surname>Barnard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Duygululu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Forsyth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Blei</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Jordan</surname>
          </string-name>
          , \
          <article-title>Matching words and pictures,"</article-title>
          <source>The Journal of Machine Learning Research</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>F.</given-names>
            <surname>Monay</surname>
          </string-name>
          and D. GaticaPerez, \
          <article-title>PLSA-based Image AutoAnnotation: Constraining the Latent Space,"</article-title>
          <source>in Proc. of ACM International Conference on Multimedia</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xue</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Jain</surname>
          </string-name>
          , \
          <article-title>Automatic Image Annotation by Incorporating Feature Hierarchy and Boosting to Scale up SVM Classi ers,"</article-title>
          <source>in Proc. of ACM MULTIMEDIA</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>J.</given-names>
            <surname>Jeon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lavrenko</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Manmatha</surname>
          </string-name>
          , \
          <article-title>Automatic image annotation and retrieval using cross-media relevance models,"</article-title>
          <source>in Proc. of ACM SIGIR</source>
          , pp.
          <fpage>119</fpage>
          -
          <lpage>126</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>V.</given-names>
            <surname>Lavrenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Manmatha</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Jeon</surname>
          </string-name>
          , \
          <article-title>A model for learning the semantics of pictures,"</article-title>
          <source>in Proc. of NIPS</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>S.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Manmatha</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Lavrenko</surname>
          </string-name>
          , \
          <article-title>Multiple Bernoulli relevance models for image and video annotation,"</article-title>
          <source>in Proc. of ICCV</source>
          , pp.
          <fpage>1002</fpage>
          -
          <lpage>1009</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ma</surname>
          </string-name>
          , H.Lu, and S.Ma, \
          <article-title>Dual cross-media relevance model for image annotation,"</article-title>
          <source>in Proc. of ACM MULTIMEDIA</source>
          , pp.
          <fpage>605</fpage>
          -
          <lpage>614</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>X.</given-names>
            <surname>Wan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Xiao</surname>
          </string-name>
          , \
          <article-title>Manifold-ranking based topic-focused multidocument summarization,"</article-title>
          <source>in Proc. of IJCAI</source>
          , pp.
          <fpage>2903</fpage>
          -
          <lpage>2908</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>D. Zhou</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Gretton</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Bousquet</surname>
            , and
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Scho</surname>
          </string-name>
          <article-title>lkopf. Ranking on data manifolds</article-title>
          ,
          <source>in Proc. of NIPS</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>J. Liu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Lu</surname>
            , and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Ma</surname>
          </string-name>
          ,
          <article-title>"Image annotation via graph learning,"</article-title>
          <source>Pattern Recognition</source>
          ,
          <volume>42</volume>
          (
          <issue>2</issue>
          ):218C228,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>J. Liu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Ma</surname>
            , Q. Liu, and
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Lu</surname>
          </string-name>
          , \
          <article-title>An adaptive graph model for automatic image annotation,"</article-title>
          <source>in Proc. of ACM International Workshop on Multimedia Information Retrieval</source>
          , pp.
          <fpage>61C70</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>M. Srikanth</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Varner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Bowden</surname>
            , and
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Moldovan</surname>
          </string-name>
          , \
          <article-title>Exploiting ontologies for automatic image annotation,"</article-title>
          <source>in Proc. of SIGIR</source>
          , pp.
          <fpage>552</fpage>
          -
          <lpage>558</lpage>
          ,
          <year>2005</year>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. Y.</given-names>
            <surname>Chang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B. L.</given-names>
            <surname>Tseng</surname>
          </string-name>
          . \
          <article-title>Multimodal metadata fusion using causal strength,"</article-title>
          <source>in Proc. of ACM MULTIMEDIA</source>
          , pp.
          <fpage>872</fpage>
          -
          <lpage>881</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Miller</surname>
          </string-name>
          , \
          <article-title>Wordnet: a lexical database for English,"</article-title>
          <source>Commun. ACM</source>
          ,
          <volume>38</volume>
          (
          <issue>11</issue>
          ):
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>C. Wang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Jing</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            , and
            <given-names>H. J.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , \
          <article-title>Image annotation re nement using random walk with restarts,"</article-title>
          <source>in Proc. of ACM MULTIMEDIA</source>
          , pp.
          <fpage>647</fpage>
          -
          <lpage>650</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Awad</surname>
          </string-name>
          , \
          <article-title>Image annotations by combining multiple evidence &amp; wordNet,"</article-title>
          <source>in Proc. of ACM MULTIMEDIA</source>
          , pp.
          <fpage>706</fpage>
          -
          <lpage>715</lpage>
          ,
          <year>2005</year>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>R.</given-names>
            <surname>Cilibrasi and P. M. B. Vitanyi</surname>
          </string-name>
          , \
          <article-title>The google similarity distance,"</article-title>
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. L.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Hua</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
            , W. Ma, and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          , \
          <article-title>Flickr distance,"</article-title>
          ,
          <source>in Proc. of ACM MULTIMEDIA</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Gong</surname>
          </string-name>
          , \
          <article-title>Translating Topics to Words for Image Annotation,"</article-title>
          <source>in Proc. of ACM CIKM</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Zhiwu</surname>
            <given-names>Lu</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horace H.S. Ip</surname>
            , and
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>He</surname>
          </string-name>
          , \
          <article-title>Context-Based Multi-Label Image Annotation,"</article-title>
          <source>in Proc. of ACM CIVR</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Boser</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guyon</surname>
            <given-names>I.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Vapnik</surname>
            <given-names>V</given-names>
          </string-name>
          ,
          <article-title>" An training algorithm for optimal margin classi ers"</article-title>
          in
          <source>In Fifth Annual ACM Workshop on Computational Learning Theory, Pittsburgh</source>
          ,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>V.</given-names>
            <surname>Vapnik</surname>
          </string-name>
          ,
          <article-title>"</article-title>
          <source>Statistical Learning Theory."</source>
          , in A Wiley-Interscience
          <source>Publication"</source>
          ,
          <year>1998</year>
          ".
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>C. Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , H. Zhang, \
          <article-title>Multi-Label Sparse Coding for Automatic Image Annotation,"</article-title>
          ,
          <source>in Proc. of CVPR</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <given-names>P.</given-names>
            <surname>Duygulu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Barnard</surname>
          </string-name>
          , J. de Freitas, and
          <string-name>
            <given-names>D.</given-names>
            <surname>Forsyth</surname>
          </string-name>
          , \
          <article-title>Object Recognition as Machine Translation: Learning a Lexicon for a Fixed Image Vocabulary,"</article-title>
          <source>in Proc. of ECCV</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <given-names>S.</given-names>
            <surname>Lazebnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ponce</surname>
          </string-name>
          , \
          <article-title>Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories"</article-title>
          ,
          <source>in Proc. of CVPR</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>A.C. Gallagher</surname>
            ,
            <given-names>C.G.</given-names>
          </string-name>
          <string-name>
            <surname>Neustaedter</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Luo</surname>
            , and
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>"Image Annotation Using Personal Calendars as Context"</article-title>
          ,
          <source>in Proc. of ACM Multimedia"</source>
          ,,
          <year>2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Cao</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Huang</surname>
            <given-names>T.S.</given-names>
          </string-name>
          ,
          <article-title>"Annotating Photo Collection by Label Propagation According to Multiple Similarity Cues"</article-title>
          ,
          <source>in Proc. of ACM Multimedia"</source>
          ,
          <year>2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <given-names>Y.H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.T.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.W.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.H</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.H.</given-names>
            <surname>Hsu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          , \
          <article-title>ContextSeer: Context Search and Recommendation at Query Time for Shared Consumer Photos,"</article-title>
          <source>in Proc. of ACM Multimedia"</source>
          ,
          <year>2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31. D. Haussler, \
          <article-title>Convolution Kernels on Discrete Structures,"</article-title>
          <source>in Technical Report UCSC-CRL-99-10</source>
          , University of California in Santa Cruz, Computer Science Department, July,
          <year>1999</year>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32. S. Nowak and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Huiskes</surname>
          </string-name>
          , \
          <article-title>New Strategies for Image Annotation: Overview of the Photo Annotation Task at ImageCLEF</article-title>
          <year>2010</year>
          ,
          <article-title>"</article-title>
          <source>in The Working Notes of CLEF</source>
          <year>2010</year>
          , Padova, Italy,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>