<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Renmin University of China at ImageCLEF 2013 Scalable Concept Image Annotation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Xirong Li</string-name>
          <email>xirong@ruc.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shuai Liao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Binbin Liu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gang Yang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Qin Jin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jieping Xu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiaoyong Du</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Key Lab of Data Engineering and Knowledge Engineering</institution>
          ,
          <addr-line>MOE</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Multimedia Computing Lab, School of Information, Renmin University of China Zhongguancun</institution>
          <addr-line>Street 59, Beijing 100872</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <abstract>
        <p>In this paper we describe our image annotation system participated in the ImageCLEF 2013 scalable concept image annotation task. The system leverages multiple base classifiers, including singlefeature and multi-feature k NN classifiers and histogram intersection kernel SVMs, all of which are learned from the provided 250K web images and provided features with no extra manual verification. These base classifiers are combined into a stacked model, with the combination weights optimized to maximize the geometric mean of F-samples, F-concepts, and AP-samples metrics on the provided development set. By varying the configuration of the system, we submitted five runs. Evaluation results show that for all of our runs, model stacking with optimized weights performs best. Our system can annotate diverse Internet images purely based on the visual content, at the following accuracy level: F-samples of 0.290, F-concepts of 0.304, and AP-samples of 0.380. What is more, a system-to-system comparison reveals that our system and the best submission this year are complementary with respect to the best annotated concepts, suggesting the potential for future improvement.</p>
      </abstract>
      <kwd-group>
        <kwd>Image annotation</kwd>
        <kwd>learning from web</kwd>
        <kwd>stacked model</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Annotating unlabeled images by computers is crucial for organizing and
retrieving the ever-growing amounts of images on personal devices and the Internet.
Due to the semantic gap, i.e., the lack of coincidence between visual features
extracted from the visual data and a user’s interpretation on the same data,
image auto-annotation is challenging.</p>
      <p>
        In the context of annotating Internet images, the semantic gap becomes even
bigger, as a specific concept exhibits significant diversity in its visual appearance.
The imagery of a concept does not limit to realistic photographs, but can also
be artificial correspondences such as posters, drawings, and cartoons, as
demonstrated in Fig. 1. To annotate the uncontrolled visual content with a large set of
concepts, a promising line of research is to learn from web data which contains
airplane cloud daytime
poster sky vehicle
bird cartoon
bird flower painting
plant
many images but with unreliable annotations [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6">1–6</xref>
        ]. In these works, k Nearest
Neighbors (k NN) and Support Vector Machines (SVMs) are two popular
classifiers, as have been separately used in [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1–3</xref>
        ] and [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4–6</xref>
        ]. Given the difficulty of
the Scalable Concept Image Annotation 2013 task [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], we believe that a single
classifier is inadequate. In that regard, we develop an image annotation system
that combines a number of base classifiers into a stacked model. In our 2013
experiments we submitted five runs, with the purpose of verifying the effectiveness
of model stacking.
      </p>
      <p>The remainder of the paper is organized as follows. We first describe our
image annotation system in Section 2. Then we detail our experiments in Section
3, with conclusions given in Section 4.
2</p>
      <p>The RUC Image Annotation System
Our image annotation system consists of multiple base classifiers, which are then
combined into a stacked model to make final predictions. The base classifiers are
learned from the 250K Internet images provided by the 2013 task, while the
stacked model is optimized using the development set of 1,000 ground-truthed
images. A conceptual diagram of the system is given in Fig. 2.</p>
      <p>Next, we describe the stacked model in Section 2.1, followed by the base
classifiers in Section 2.2.
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Stacked Image Annotation Model</title>
      <p>For the ease of consistent description, let x be a specific image. For a given
concept ω, let g(x, ω) be an image annotation model which produces a relevance
score of ω with respect to x. A set of t models are denoted by {g1(x, ω), . . . , gt(x, ω)}.</p>
      <p>We define our stacked model as a convex combination of the t models:
t
gΛ(x, ω) = X λi · gi(x, ω),
i=1
(1)</p>
      <sec id="sec-2-1">
        <title>Internet images</title>
      </sec>
      <sec id="sec-2-2">
        <title>Construct base classifiers</title>
      </sec>
      <sec id="sec-2-3">
        <title>Select relevant examples</title>
      </sec>
      <sec id="sec-2-4">
        <title>Single/Multi feature kNN</title>
      </sec>
      <sec id="sec-2-5">
        <title>HikSVMs</title>
        <p>development set
t
where λi is the nonnegative weight of the i-th classifier, with Pi=1 λi = 1. Notice
that gΛ(x, ω) is to indicate that the stacked model is parameterized by Λ = {λi}.</p>
        <p>
          We look for the setting of {λi} that maximizes the image annotation
performance. The 2013 task specifies three performance metrics [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]: F-measure for
images, F-measure for concepts, and Average Precison for images, as given in
the Appendix. To jointly maximize the three metrics, we take their geometric
mean as a combined metric. Weights of Eq. (1) optimized with respect to the
combined metric are found by a coordinate ascent algorithm.
2.2
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Base Classifiers</title>
      <p>We choose k Nearest Neighbors (k NN) and Support Vector Machines (SVMs)
as two types of base classifiers, for their good performance.</p>
      <p>
        Base classifier I: k NN. Given a test image x, we define the k NN classifier as
k
gknn(x, ω) = X rel(xi, ω) · wi,
i=1
(2)
where rel(xi, ω) denotes the relevance score between the i-th neighbor and the
concept, and wi is the neighbor weight. In this work, we instantiate rel(xi, ω)
using the provided textual feature, which is derived from the web page of xi [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
For wi, we choose a Bayesian implementation [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], computing wi as k−1 Pj=i j−1
k
to give more importance to closer neighbors.
      </p>
      <p>
        Base Classifier II: SVMs. We choose the histogram intersection kernel SVMs
(hikSVMs) which is known to effective for bag of visual words features of medium
size [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], such as the 5,000-dim features provided by the task. More importantly,
the decision function of a hikSVMs can be efficiently computed by a few linear
interpolations on a set of precomputed points.
      </p>
      <p>To obtain relevant positive training examples for a given concept from the
250K Internet images, we utilize both the textual feature and the query log
generated when collecting the images from web image search engines. We describe
the connection of an Internet image x to a web image search engine s by a triplet
&lt; q, r, s &gt;, where q represents a query keyword, r is the rank of x in the search
results of q returned by s. An image can be associated with multiple triplets. To
determine the positiveness of the image x with respect to the given concept ω,
we propose to compute a search engine based score as
l</p>
      <p>
        w(si) ,
positiveness(x, ω) = X sim(qi, ω) √ri
i=1
(3)
where l is the number of triplets associated with the image, and sim(qi, ω) is a
tag-wise similarity measure, computed on the base of tag co-occurrence in 1.2
million Flickr images [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The variable w(si) indicates the weight of a specific
search engine, which we empirically set to be 1, 0.5, and 0.5 for Google, Yahoo,
and Bing, respectively. The positiveness score in Eq. 3 is further combined with
the given textual feature. For each concept, we sort images labeled with the
concept in descending order by their positiveness scores, and select the top 500
ranked images as positive training examples.
      </p>
      <p>
        For the acquisition of negative training examples, we consider the following
two approaches. The first approach is to sample negative examples at random.
Given a specific concept and the 500 selected positive examples, we randomly
sample a set of 500 negative examples, to make the training data perfectly
balanced. We repeat the random sampling 10 times, yielding a set of 10 hikSVMs.
The second approach is Negative Bootstrap [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Different from random sampling,
this approach iteratively selects negative examples which are most misclassified
by present classifiers, and thus most relevant to improve classification. Per
iteration, the approach randomly samples 5,000 examples to form a candidate
set. An ensemble of classifiers obtained in the previous iterations are used to
classify each candidate example. The top 500 most misclassified examples are
selected and used together with the 500 positives to train a new hikSVMs. We
conduct Negative Bootstrap with 10 iterations, producing 10 hikSVMs for each
concept. For efficient classification, we leverage a model compression technique
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] to compress the ensemble into a single classifier such that the prediction
time complexity is independent of the ensemble size.
3
3.1
      </p>
      <sec id="sec-3-1">
        <title>Experiments</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Submitted Runs</title>
      <p>This year we submitted five runs, which are listed as follows. For all runs, we
empirically preserve for each image the top six ranked concepts as its final
annotation.</p>
      <p>Run: RUC Crane. This run uses a k NN classifier with early fused multiple
features (Multi-Feat-k NN). We use all the 7 provided features, i.e., getlf (256-D),
colorhist (576-D), gist (480-D), sift (5000-D), csift (5000-D), rgbsift (5000-D),
and opponentsift (5000-D). For each feature, we compute the l1 distance.
Distance values of individual features are zero-score normalized, and then averaged
to obtain k = 256 nearest neighbors.</p>
      <p>Run: RUC Snake. This run combines multiple single feature k NN
(Single-Featk NN) and SVMs classifiers with uniform weights. In particular, five features
(colorhist, gist, csift, rgbsift, and opponentsift) are used separately for
SingleFeat-k NN, again with the l1 distance and k = 256. For SVMs, we use the three
variants of sift, i.e., csift, rgbsift, and opponentsift. In total, this run employs
5+3+3=11 base classifiers.</p>
      <p>Run: RUC Monkey. This run uses the same 11 base classifiers as have been used
in the Snake run, but with optimized weights. Consequently, the effectiveness of
the stacked model can be verified by comparing the Monkey and the Snake.
Run: RUC Mantis. This run combines the Multi-Feat-k NN, multiple
SingleFeat-k NN, and 6 variants of SVMs with uniform weights. So the run employs 12
base classifiers in total.</p>
      <p>Run: RUC Tiger. This run uses the same 12 base classifiers as have been used
in the Mantis run, but with optimized weights. The Tiger is our primary run.
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>The performance scores of the five runs are summarized in Table 1. For all
performance metrics, the Tiger run, which combines all the base classifiers using
the coordinate ascent optimized weights, is the best. On the test set, this run
reaches an F-samples score of 0.290, an F-concepts score of 0.304, and an
APsamples score of 0.380. The Monkey run using optimized weights is better than
the Snake run using uniform weights. Moreover, by comparing the performance
on the development set and on the test set, we find that the system generalizes
well to unseen images and concepts. These results justify the importance of
model stacking and weights optimization for image annotation.
sunset
firework
galaxy
space
underwater
helicopter
lightning
aerial
fog
cartoon
pool
portrait
desert
bus
water
rainbow
protest
mountain
sand
tricycle
boat
tree
logo
snow
castle
nebula
plant
newspaper
but erfly
person
motorcycle
sea
drink
sky
flower
drum
road
forest
furniture
moon
wagon
building
church
violin
bot le
beach
food
lake
fire
bicycle
silhouet e
horse
cloud
sun
s traf ic
t sign
p vehicle
e phone
c teenager
n reptile
o toy</p>
      <p>footwear
C fish
instrument
night ime
chair
highway
arthropod
poster
reflection
countryside
painting
truck
car
harbor
guitar
train
sculpture
diagram</p>
      <p>river
overcast
airplane
grass
rain
bridge
cityscape
soil
hat
indoor
sport
submarine
cat
dog
coast
spider
baby
embroidery
garden
elder
female
rodent
child
book
table
park
bird
smoke
outdoor
shadow
male
unpaved
daytime
monument
spectacle
cloudless
closeup
Median score
Best submission
RUC Tiger
0
0.1
0.2
0.3
0.4
0.6
0.7
0.8
0.9</p>
      <p>1
0.5</p>
      <p>F−concepts
average level. Moreover, when
compared
with
the
best
submission
of
this
year,
we
find
that
the
two
systems
are
complementary,
as
their
best
annotated
concepts
differ.</p>
      <p>Discussions and Conclusions
This paper documents our experiments in the ImageCLEF 2013 Scalable
Concept Image Annotation, a testbed for developing image annotation systems using
generic web data. We have built such a system.</p>
      <p>Our system annotates images purely based on the visual content. It combines
multiple base classifiers, i.e., variants of k NN and SVMs, into a stacked model.
In all of our five submitted runs, the stacked model with optimized weights
performs best.</p>
      <p>Through a system-to-system comparison, we find that our system and the
best submission this year is complementary in the sense of the best annotated
concepts. Given that we use relatively simple base classifiers, we consider this
finding interesting, as it suggests the potential of future improvement.
Acknowledgments. This research was supported by the Basic Research funds
in Renmin University of China from the central government (13XNLF05). The
authors are grateful to the ImageCLEF coordinators for the benchmark
organization efforts.</p>
      <sec id="sec-5-1">
        <title>Appendix</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>A1. Performance Measures</title>
      <p>(4)
(5)
(6)
F-samples. Given a test image x with its relevant tag set Rx and a predicted
tag set Px, its F-samples score is computed as</p>
      <p>F-samples(x) =
2 ∗ precision(x) ∗ recall(x)
precision(x) + recall(x)
where precision(x) is |Rx ∩ Px|/|Px|, and recall(x) is |Rx ∩ Px|/|Rx|.
F-concepts. Given a test concept ω with its relevant image set Rω and a set
of images Pω labeled with ω by the annotation system, its F-concepts score is
computed as</p>
      <p>F-concepts(ω) =
2 ∗ precision(ω) ∗ recall(ω)
precision(ω) + recall(ω)
where precision(ω) is |Rω ∩ Pω|/|Pω|, and recall(ω) is |Rω ∩ Pω|/|Rω|.
AP-samples. Given a test image x with m tags sorted in descending order by
relevance scores, its AP-samples score is computed as</p>
      <p>AP-samples(x) =</p>
      <p>1 Xm ri δ(i),
|Rx| i=1 i
where ri is the number of relevant tags among the top i tags, and δ(i) is 1 if the
i-th tag is in Rx, 0 otherwise.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.J.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            ,
            <surname>Ma</surname>
          </string-name>
          , W.Y.:
          <article-title>Annotating images by mining image search results</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>30</volume>
          (
          <issue>11</issue>
          ) (Nov.
          <year>2008</year>
          )
          <fpage>1919</fpage>
          -
          <lpage>1932</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snoek</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Worring</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Annotating images by harnessing worldwide usertagged photos</article-title>
          .
          <source>In: ICASSP</source>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paredes</surname>
          </string-name>
          , R.:
          <article-title>A k-nn approach for scalable image annotation using general web data</article-title>
          .
          <source>In: BigVision</source>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snoek</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Worring</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smeulders</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Harvesting social images for biconcept search</article-title>
          .
          <source>IEEE Transactions on Multimedia</source>
          <volume>14</volume>
          (
          <issue>4</issue>
          ) (Aug.
          <year>2012</year>
          )
          <fpage>1091</fpage>
          -
          <lpage>1104</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>Y.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngo</surname>
            ,
            <given-names>C.W.</given-names>
          </string-name>
          :
          <article-title>Sampling and ontologically pooling web images for visual concept learning</article-title>
          .
          <source>IEEE Transactions on Multimedia</source>
          <volume>14</volume>
          (
          <issue>4</issue>
          ) (Aug.
          <year>2012</year>
          )
          <fpage>1068</fpage>
          -
          <lpage>1078</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kordumova</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snoek</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Evaluating sources and strategies for learning video concepts from social media</article-title>
          .
          <source>In: CBMI</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paredes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomee</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the imageclef 2013 scalable concept image annotation subtask</article-title>
          .
          <source>In: CLEF 2013 working notes</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Maji</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berg</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malik</surname>
          </string-name>
          , J.:
          <article-title>Classification using intersection kernel support vector machines is efficient</article-title>
          .
          <source>In: CVPR</source>
          . (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snoek</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Worring</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smeulders</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Social negative bootstrapping for visual categorization</article-title>
          .
          <source>In: ICMR</source>
          . (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snoek</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Worring</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koelma</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smeulders</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Bootstrapping visual categorization with relevant negatives</article-title>
          .
          <source>IEEE Transactions on Multimedia</source>
          <volume>15</volume>
          (
          <issue>4</issue>
          ) (Jun.
          <year>2013</year>
          )
          <fpage>933</fpage>
          -
          <lpage>945</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>