<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IPL at ImageCLEF 2014: Scalable Concept Image Annotation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Spyridon Stathopoulos</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Theodore Kalamboukis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Processing Laboratory, Department of Informatics, Athens University of Economics and Business</institution>
          ,
          <addr-line>76 Patission Str, 104.34, Athens</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <fpage>398</fpage>
      <lpage>403</lpage>
      <abstract>
        <p>In this article we report on the experiments conducted by the IPL team within the context of the ImageCLEF 2014 challenge on Scalable Concept Image Annotation. Our approach encompasses, a CBIR phase following with a concept extraction procedure. The content based retrieval utilizes Latent Semantic Analysis on a set of multiple Compact Composite Features to retrieve the most similar images and in the sequel a number of concepts are extracted from the associated textual information, based on their posterior probabilities.</p>
      </abstract>
      <kwd-group>
        <kwd>LSA</kwd>
        <kwd>Image Annotation</kwd>
        <kwd>Data Fusion</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The continuous increase of digital images over the web has led to the need for
an e cient method of indexing and retrieving content, based on the semantic
information presented in an image. A most common approach to achieve this,
is by assigning metadata in the form of keywords to the image. Most methods
rely on the use of a manually labeled training set of images. However, manually
annotating images is a costly process and has obvious scalability problems.</p>
      <p>
        In the Scalable Concept Image Annotation task of the ImageCLEF 2014 [
        <xref ref-type="bibr" rid="ref1 ref2">1,
2</xref>
        ], the goal is to develop a fully automatic procedure, that is able to annotate
an image with a prede ned set of labels without the use of any hand labeled
training data. Our baseline algorithm is divided into two phases. Visual retrieval
step: Given a test image, a sample of the K most visually similar images is
retrieved. Annotation step: From the texts associated to the K retrieved images
a set of candidate keywords is selected as labels. The nal assigned keywords for
the test image are determined by a probability score based on the co-occurrence
of labels in the selected sample.
      </p>
      <p>Section 4 presents the results obtained from our experiments using the
development and test set made available by the task. We compare our results with
those from the task's baseline system.</p>
      <p>
        Image Representation and Retrieval
In order to retrieve a sample of the visually nearest images, several low-level
visual descriptors were extracted form each image. Those descriptors are then
combined using LSA to provide a latent semantic vector representation for each
image. The low level features (CEDD,FCTH) were selected based on our
previous research [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and experience from our participation in CLEF, medical Image
retrieval task [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Furthermore, the following descriptors were selected for our
nal submitted runs since their combination gave the best results on the
development set:
1. Color and Edge Directivity Descriptor (CEDD)[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
2. Fuzzy Color and Texture Histogram (FCTH)[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
3. Opponent-SIFT [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], provided by the organizers with a codebook of 1,000
features used to create the histograms for the images.
      </p>
      <p>
        The descriptors CEDD and FCTH were locally extracted from a 3x3 grid,
resulting in a vector size of 1,296 and 1,728 features respectively. Content based
retrieval was based on applying LSA to the feature matrix X=[CEDD; FCTH;
OppSIFT] using the Matlab's routine eigs for matrix XXT with k=50. It has
been shown [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] that, this is an e ective and e cient approach that overcomes
the de ciencies of using the singular value decomposition analysis.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Image Annotation</title>
      <p>Each image in the training data is associated with a set of keywords with a
score assigned to each keyword. The keywords were extracted from the text
surrounding the image within its webpage and the scores were calculated using:
{ The term frequency (TF).
{ The document object model (DOM) attributes.</p>
      <p>{ The word distance to the image.</p>
      <p>
        More information on the data is provided by the organizers in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Additionally,
we have removed the stopwords from the keywords sets.
      </p>
      <p>
        A concept is a construct, an idea, of something formed by mentally combining
all its characteristics. In our perception, a concept c, corresponding to a keyword,
w, is de ned by the set C = fw; s1(w); ::::; sn(w)g where si(w) are the synonyms
of w extracted from WordNet [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. For a test image, g, labels were selected based
on the posterior probabilities p(cjg), (probability to select concept c, given a test
image g), [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], de ned by:
p(cjg) =
      </p>
      <p>
        K
X p(cjj)p(jjg)
j=1
(1)
Probabilities p(cjg) are approximated from the K nearest neighbors visually
retrieved images from the training set when the test image g is submitted as
query. To obtain an estimate of p(jjg) we use the method proposed by Platt [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
to extract probabilistic outputs from SVM. The basic idea is that the retrieval
problem can be considered equivalent to classi cation. The hyperlane in the
case of retrieval is de ned from the query vector Q for and an appropriate
constant (h(x) = Qx + ). The matching function (cos) for a test example
is inversely proportional to the distance from the hyperplane de ned by the
classi er (d(j; g) = 2 2cos(j; c)). Thus following Platt's method we approximate
the probability p(jjg) by :
      </p>
      <p>The conditional probability p(cjj) is calculated by:
1
p(jjg) = 1 + ea d(j;g)
p(cjj) = X
w2Cj</p>
      <p>ps(w; j)
P
w02Wj ps(w0; j)
(2)
(3)
where Cj is the set of concepts of the ith retrieved image and Wj the set of
keywords assigned to image j. Moreover, s(w; j) are the scores provided by the
organizers as part of the training set. Finally, the l concepts with the highest
posterior probability p(cjg) calculated from Equation (1), are selected as the
concepts being present in the test image g. Di erent values of l were tested with
the development set, with l = 8 giving the best results and thus this value was
selected for our submitted runs.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Submitted Runs</title>
      <p>Several initial experiments were performed using the development set. The most
of notable ones were those which study the impact of the rate parameter a and
the number of K for the top retrieved images. The corresponding results in
Tables 2 and 3, show that these parameters can have an important impact on
annotation performance. For the test set, a total of 10 runs were submitted:
{ Run 1: K-NN with K = 1000 neighbors, retrieved by Early fusion and LSA
on Opponent SIFT, CEDD, FCTH. Parameter a = 6. No synonym usage.
{ Run 2: Same as Run 1, a = 10.
{ Run 3: Same as Run 1, a = 16.
{ Run 4: Same as Run 1, but with K = 800 and a = 16.
{ Run 5: Same as Run 1, but with K = 450 and a = 16.
{ Run 6: K-NN with K = 1000 neighbors, retrieved by Early fusion and LSA
on Opponent SIFT, CEDD, FCTH. Parameter a = 6. Concepts include
WordNet synonyms.
{ Run 7: Same as Run 6, a = 10.
{ Run 8: Same as Run 6, a = 16.
{ Run 9: Same as Run 6, but with K = 800 and a = 16.
{ Run 10: Same as Run 6, but with K = 450 and a = 16.
Run</p>
      <p>MAP-Samples
%
34.7
37.2
39.1
40.1
40.7
41.0
41.0
41.0
Run
Run 1
Run 2
Run 3
Run 4
Run 5
Run 6
Run 7
Run 8
Run 9
Run 10
We have presented a baseline algorithm for image annotation based on the visual
retrieval from the train set and extracted the labels of concepts from the top
K-NN retrieved images. Our approach enhances the representation of an image,
by fusing di erent low-level features using LSA. Results show that image
representation play an important role in the annotation problem. Furthermore, the
number of the retrieved images K, and the way these are compared with a test
image g (p(jjg)), have an important impact on annotation.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Caputo</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Martinez-Gomez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Acar</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patricia</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marvasti</surname>
            , N., Uskudarl ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paredes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cazorla</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Varea</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morell</surname>
          </string-name>
          , V.:
          <article-title>ImageCLEF 2014: Overview and analysis of the results</article-title>
          .
          <source>In: CLEF proceedings. Lecture Notes in Computer Science</source>
          . Springer Berlin Heidelberg (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paredes</surname>
          </string-name>
          , R.:
          <article-title>Overview of the ImageCLEF 2014 Scalable Concept Image Annotation Task</article-title>
          . In:
          <article-title>CLEF 2014 Evaluation Labs</article-title>
          and Workshop, Online Working Notes. (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Stathopoulos</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalamboukis</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>An svd-bypass latent semantic analysis for image retrieval</article-title>
          . In Greenspan, H.,
          <string-name>
            <surname>Muller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Syeda-Mahmood</surname>
          </string-name>
          , T., eds.:
          <article-title>Medical Content-Based Retrieval for Clinical Decision Support</article-title>
          . Volume
          <volume>7723</volume>
          of Lecture Notes in Computer Science. Springer Berlin Heidelberg (
          <year>2013</year>
          )
          <volume>122</volume>
          {
          <fpage>132</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Muller, H.,
          <string-name>
            <surname>de Herrera</surname>
            ,
            <given-names>A.G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalpathy-Cramer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antani</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eggel</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Overview of the imageclef 2012 medical image retrieval and classi cation tasks</article-title>
          . In Forner, P.,
          <string-name>
            <surname>Karlgren</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Womser-Hacker</surname>
          </string-name>
          , C., eds.: CLEF (Online Working Notes/Labs/Workshop). (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Chatzichristo s,
          <string-name>
            <given-names>S.A.</given-names>
            ,
            <surname>Boutalis</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.S.:</surname>
          </string-name>
          <article-title>Cedd: Color and edge directivity descriptor: A compact descriptor for image indexing and retrieval</article-title>
          . In: ICVS. (
          <year>2008</year>
          )
          <volume>312</volume>
          {
          <fpage>322</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Chatzichristo s,
          <string-name>
            <given-names>S.A.</given-names>
            ,
            <surname>Boutalis</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.S.</surname>
          </string-name>
          : Fcth:
          <article-title>Fuzzy color and texture histogram - a low level feature for accurate image retrieval</article-title>
          .
          <source>In: WIAMIS</source>
          . (
          <year>2008</year>
          )
          <volume>191</volume>
          {
          <fpage>196</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. van de Sande,
          <string-name>
            <given-names>K.E.A.</given-names>
            ,
            <surname>Gevers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Snoek</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.G.M.:</surname>
          </string-name>
          <article-title>Evaluating color descriptors for object and scene recognition</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>32</volume>
          (
          <issue>9</issue>
          ) (
          <year>2010</year>
          )
          <volume>1582</volume>
          {
          <fpage>1596</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>WordNet: An Electronic Lexical Database</article-title>
          . Bradford
          <string-name>
            <surname>Books</surname>
          </string-name>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paredes</surname>
          </string-name>
          , R.:
          <article-title>A k-nn approach for scalable image annotation using general web data</article-title>
          .
          <source>In: Big Data Meets Computer Vision: First International Workshop on Large Scale Visual Recognition and Retrieval, held in conjunction with NIPS</source>
          <year>2012</year>
          .
          <article-title>(</article-title>
          <year>2012</year>
          ) 1{
          <fpage>5</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Platt</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          :
          <article-title>Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods</article-title>
          .
          <source>In: ADVANCES IN LARGE MARGIN CLASSIFIERS</source>
          , MIT Press (
          <year>1999</year>
          )
          <volume>61</volume>
          {
          <fpage>74</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>