<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CEA LIST@imageCLEF 2013: Scalable Concept Image Annotation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Herve Le Borgne</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adrian Popescu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amel Znaidia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CEA, LIST, Vision &amp; Content Engineering Gif-sur-Yvettes</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We report the participation of the CEA LIST to the Scalable Concept Image Annotation Subtask of ImageCLEF 2013. The full system is based on both textual and visual similarity to each concept, that are merged by late fusion. Each image is visually represented with a bag of visterm, computed from a dense grid of SIFT every 3 pixels, that a locally soft coded and max pooled on a codebook of size 1024 and spatially extended with a pyramid 1 1 + 3 1 + 2 2, resulting into a vector of size 8192. The visual neighbors of a query are given by the L1 distance to the images of the learning database. The similarity of a query to one of the 95/116 concepts to identify is the sum of the similarity of each neighbor to the concept. The decision is set at 1 for all concepts above one standard deviation from the average similarity to all concepts for the considered query. The similarity between a training image and a concept is computed from an intermediate vectorial representation of its tags. We tested several spaces of representation for the tags, including wikipedia concepts sorted according to their popularity or characterized according FlickR data. As well, the size of space was pruned at several values, from 5000 to 200; 000. The tag representation are max-pooled to make the training image vector, such that the resulting vector contain the maximal similarity to each concept of the intermediate space. The 96/116 concepts to identify are represented in the same space and their similarity to the training image is the cosine between the intermediate space representation. As well, we ranked all training images to each concept to identify in order to learn visual classi ers (linear SVM). We tested several strategies to select positive and negative examples, including the visual coherency, but the simplest strategies was nally the most e cient. It consisted in setting the 100 most similar images as positive and the 500 least similar as negative. Finally, a simple weighted late fusion of the visual and textual similarity scores appeared to be more e cient than more sophisticated strategies, resulting to 0:4 MAP on the development query and 0:34 on the testing ones.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The Scalable Concept Image Annotation Subtask of ImageCLEF 2013[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is
described in detail in [7]. The system we propose rely on both visual and textual
cues. We conducted many preliminary experiments in order to iteratively
improve the provided baseline system (see section 1.1). These experiments dealt
with visual features to nd image neighbors of queries (section 2), several models
of tag (section 3 to 5), the way we learnt visual models (section 6) and nally
the decision process (section 7).
1.1
      </p>
      <p>baseline system
A baseline system based on the co-occurence of concepts and tags of the visual
neighbors of each query is provided[7]. Each image I of the development set
has to be annotated according to concepts Cp; p = 1:::95. Such an image has
Kv neighbors in the train test, according to a visual descriptor (csift BoV
provided). Each of these neighbors (k = 1 : : : Kv) has Tk tags with a given score
(tk;1; sk;1 : : : tk;i; sk;i : : : tk;Tk ; sk;Tk ). Each of these tags is described with Nk;i
weighted concepts (Ck;i;1; Wk;i;1 : : : Ck;i;j ; Wk;i;j : : : Ck;i;Nk;i ; Wk;i;Nk;i ). Thus, the
score of concept Cp for image I is:</p>
      <p>ScoreI (Cp) =
1 XKv PiT=k1 sk;iWk;i;Cp
Kv k=1</p>
      <p>PiT=k1 sk;i
In practice, each tag is described by the same number Kconcepts of concepts
(default: 6).
2</p>
    </sec>
    <sec id="sec-2">
      <title>Searching visual neighbors</title>
      <p>In the original system, visual neighbors are provided and said to be found with
a C-SIFT based descriptor. We tested several alternative methods.</p>
      <p>
        Descriptors are bag of visterms. SIFT local descriptors are densely extracted
every 3 pixels. The bag are coded using local soft coding [3] and max pooling.
Then two di erent spatial pyramid matching scheme [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] are used: 1 1 + 3 + 2 2
for BoV1 and 1 1 + 2 24 4 for BoV2. Further details on bag-of-visterm design
can be found in [6]. Several distances were tested on these vectors to nd the
neighbours. The histogram intersection (HI) distance implemented as:
and the classical L1 distance:
      </p>
      <sec id="sec-2-1">
        <title>DistHI (x</title>
        <p>y) = 1</p>
      </sec>
      <sec id="sec-2-2">
        <title>DistL1(x</title>
        <p>y) =
1 XD min(xi; yi)
D i=1 max(xi; yi)
1 XD
D
i=1
jxi
yij
Results are shown in table 1, showing a non-signi cant improvement with the
BoV1 signature and the L1 distance.
(1)
(2)
(3)</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Using a FlickR-based tag model</title>
      <p>We used a FlickR-based tag model built from the selection of the 95 concepts
(F lickr95) and another one built from 30; 000 wikipedia concept (F lickr30k). See
[5, 8] for details about the way similarities are computed for these models.Note
that both models were built from the FlickR tags. The F lickr95 tag model
was injected into the system provided, in conjunction with two di erent visual
models. The mAP is reported in table 2. Note this performance measure should
be independant from the parameter Kconcepts that was xed to 6. The F-measure
only uses the annotation decisions and is computed in two ways, one by analyzing
each of the testing samples and the other by analyzing each of the concepts.
Results are reported on table 3 and 4.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Window-restricted FlickR-based tag models</title>
      <p>The tag model of each training document is built with a restriction of the
wordimage distance. In the original web page a training image has been found, the
method consists in taking into account words that are less than a given distance</p>
      <p>Tag
11.92 12.49 11.54 10.81 10.32
15.53 16.48 16.60 16.37 15.77
17.90 18.96 18.60 17.88 17.54
19.27 19.71 19.61 19.03 18.65
20.09 20.14 19.86 19.59 19.23
from the considered image. Moreover, we considered a lemmatized and
nonlemmatized version of the models.</p>
      <p>Results are comparable to those obtained with F lickr30k (around 0:31) but
no improvement is actually observed.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Wikipedia-based tag models</title>
      <p>Similarly to the FlickR-based tag models, tags are projected onto 1187980 wikipedia
concepts. The concepts being ranked to the numbers of their incoming links, the
representation can be pruned to a lower dimension.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Learning visual models</title>
      <p>For a given tag-model, training images are sorted according to their score for each
concept. Then we select positive and negative examples according to di erent
strategies to learn linear SVM models from the BoV1 signatures. Text model used
was F lickr3w0k0 i.e the FlickR tags projected on the 30k wikipedia concepts, with
restriction a window of size 0.</p>
      <p>Tag
csift
7.25 5.49 4.33 3.60 2.91
11.33 9.13 8.47 7.81 6.11
13.98 12.48 10.67 9.41 8.00
15.34 13.17 11.87 10.46 9.39
16.18 14.15 12.67 11.72 10.43</p>
      <p>The simplest strategy consisted in setting the 100 most similar images as
positive and the 500 least similar as negative. It leaded to a mAP of 0:219 on
the devel queries.</p>
      <p>A second strategy consisted to select as positive all images above a given
score (0.8) and as negative all those below a smaller threshold (0.1). Given that
classes were then strongly ubalanced, we limited negative image to nine times
the positive on for each SVM model, leading to a mAP of 0:207. When the
number of negative samples are forced to be equal to the positive ones, the mAP
is 0:212.</p>
      <p>A last strategy was tested, for wich the choice of the images was based on
the visual coherency [4]. The 1000 most similar images to each concept are
resorted according to their VCscore and the 100 best are thus selective as positive
examples. Negative examples are chosen as the 1000 least similar images to the
concept. Although promising, this approach leaded to a mAP of 0:209 only.</p>
      <p>We thus nally decided to keep the rst and simplest strategy.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Decision</title>
      <p>The decision value (0/1) is taken independently on each query, according to
its similarity to the 95 or 116 concepts. We compute the average ( ) and the
standard deviation ( ) of the scores and set the decision to 1 for all concepts
having a score above + .</p>
    </sec>
    <sec id="sec-8">
      <title>Participation to the campaign</title>
      <sec id="sec-8-1">
        <title>Submitted runs</title>
        <p>We submitted ve runs to the campaign, based on the obsvation made during
preliminary experiments:</p>
        <p>Run 1: we computed the FlickR-based tag model and merged it with the
visual similarities. The weights were respectively 0.8 and 0.2.</p>
        <p>Run 2: we added to Run 1 a Wikipedia-based tag model with a
representation pruned to 5; 000 dimensions.</p>
        <p>Run 3: we added to Run 2 a FlickR-based tag model with a representation
pruned to 50; 000 dimensions.</p>
        <p>Run 4: similar to Run 3 with a visual model selected on scores
Run 5: similar to Run 4 with a FlickR-based tag model with a representation
pruned to 200; 000 dimensions.
8.2</p>
      </sec>
      <sec id="sec-8-2">
        <title>Results</title>
        <p>Results are quite close from each other and around 10 points (in term of
mAP) below the best run of the campaign.
3. Lingqiao Liu, Lei Wang, and Xinwang Liu. In Defense of Soft-assignment Coding.</p>
        <p>In IEEE International Conference on Computer Vision, 2011.
4. Debora Myoupo, Adrian Popescu, Herve Le Borgne, and Pierre-Alain Moellic.
Multimodal image retrieval over a large database. In Proceedings of the 10th
international conference on Cross-language evaluation forum: multimedia experiments,
CLEF'09, pages 177{184, Berlin, Heidelberg, 2010. Springer-Verlag.
5. Adrian Popescu and Gregory Grefenstette. Social media driven image retrieval. In</p>
        <p>ACM International Conference on Multimedia Retrieval, pages 33:1{33:8, 2011.
6. Aymen Shabou and Herve Le Borgne. Locality-constrained and spatially
regularized coding for scene categorization. In IEEE Conference on Computer Vision and
Pattern Recognition, pages 3618{3625, 2012.
7. Mauricio Villegas, Roberto Paredes, and Bart Thomee. Overview of the imageclef
2013 scalable concept image annotation subtask. In CLEF 2013 working notes,
2013.
8. Amel Znaidia, Aymen Shabou, Adrian Popescu, Herve Le Borgne, and Celine
Hudelot. Multimodal feature generation framework for semantic image classi cation. In
ICMR, International Conference on Multimedia Retrieval, ICMR '12, Hong Kong,
China, June 5-8, 2012, page 38, 2012.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>B.</given-names>
            <surname>Caputo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thomee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Paredes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zellhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Goeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bonnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Martinez</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Garcia</given-names>
            <surname>Varea</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Cazorla</surname>
          </string-name>
          .
          <source>Imageclef</source>
          <year>2013</year>
          :
          <article-title>the vision, the data and the open challenges</article-title>
          .
          <source>In Proc CLEF</source>
          <year>2013</year>
          , LNCS,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Svetlana</given-names>
            <surname>Lazebnik</surname>
          </string-name>
          , Cordelia Schmid, and
          <string-name>
            <given-names>Jean</given-names>
            <surname>Ponce</surname>
          </string-name>
          .
          <article-title>Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories</article-title>
          .
          <source>In IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <volume>2169</volume>
          {
          <fpage>2178</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>