<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Run Title
IAM Southampton</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>IAM@ImageCLEFPhotoAnnotation 2009: Nave application of a linear-algebraic semantic space</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jonathon S. Hare</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul H. Lewis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Intelligence Agents Multimedia Group</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Electronics and Computer Science, University of Southampton</institution>
          ,
          <addr-line>Southampton</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <volume>30</volume>
      <issue>2</issue>
      <abstract>
        <p>This paper describes Southampton's submissions to the 2009 ImageCLEF photo annotation task. For the task we used an annotation system based on the idea of constructing semantic spaces, which was developed previously at Southampton. To represent the image content, we used a combination of di erent SIFT and Colour-SIFT features detected using the di erence-of-Gaussian and MSER techniques. These features were converted into a visual term representation by applying vector quantisation using a codebook learnt from a hierarchical k-means clustering. In terms of EER and AUC, the annotator performs reasonably well, however, it struggles when evaluated using the hierarchical measure proposed for the task, due to the way the annotation con dences are thresholded.</p>
      </abstract>
      <kwd-group>
        <kwd>Image Content Analysis</kwd>
        <kwd>Data Fusion</kwd>
        <kwd>Semantic Space</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Methodology</title>
      <p>As with many automatic annotation approaches, the methodology applied to this task involved
extracting feature vectors for each of the images, and then feeding the features of a training set,
together with annotations to a machine learning system. The machine learning system attempts
to learn low-level relationships between all of the features and annotations. Once the training
phase is complete, features from un-annotated images can be fed into the system to use the learnt
relations to get predictions of annotations.
2.1</p>
      <sec id="sec-1-1">
        <title>Visual Features</title>
        <p>
          The images were represented by vectors of visual-term occurrences [
          <xref ref-type="bibr" rid="ref10">11</xref>
          ]. The visual-terms were
created by nding interest points and extracting local feature descriptors, and then quantising to
a pre-determined codebook. For the experiments in this task we used a combination of
multiscale di erence-of-Gaussian interest regions with SIFT features [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], MSER regions [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] with SIFT
features, and MSER regions with colour-SIFT features [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Each of the three region/feature
combinations had its own 3125 term codebook created by applying hierarchical k-means [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] (5 levels
with 5 clusters per node). The codebook size was not optimised in any way, and was chosen based
on a best guess basis from previous experience with these feature morphologies and the machine
learning technique described in the next subsection. The nal image representation was created by
appending the term-occurrence vectors from each of the region/feature representations to create
a vector with 9375 dimensions.
2.2
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>Machine Learning</title>
        <p>
          The machine-learning component is based on a linear-algebraic semantic space [
          <xref ref-type="bibr" rid="ref2 ref3">3, 2</xref>
          ], which is a
development and generalisation of a text indexing technique called Cross-Language Latent
Semantic Indexing [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. This technique produces a vector space into which both visual-terms and keyword
terms are mapped along with the images. Un-annotated images can then be projected into this
space. Annotation was performed by projecting the test images into the space, and ranking the
possible annotations based on their cosine similarity.
        </p>
        <p>The use of the cosine similarity measure gives each possible annotation a score between -1
and 1; however, these scores are not themselves all that informative. A higher score does mean
more con dence in an annotation, but only when considered against all the other annotations.
For the purposes of evaluating the annotator using the EER and AUC measures this is not too
much of a problem | we can just scale the scores to the 0..1 range (i.e add one and divide
by two). However, as will be discussed in the next section, the hierarchical scoring measure
[10] thresholds the annotation con dences at 0.5 to produce a binary indication of
present/notpresent. Unfortunately, for the semantic space this poses a big problem as the position of the
threshold should ideally be set di erently for each image, based on con dences of all the predicted
annotations.
3</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Experiments, Results and Discussion</title>
      <p>
        We submitted three di erent runs to the task organisers. The rst run was trained on the raw
annotations. The second included a partial expansion of the annotation hierarchy [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] provided
by the organisers, based on the non-abstract nodes (i.e. if Lake=true then water=true). The
third included a full expansion of the hierarchy. The hierarchical expansion just means that a few
extra annotation terms are fed into the machine learning component together with the leaf-node
annotations already present for the image in question. The run titles and settings are shown in
Table 1.
      </p>
      <sec id="sec-2-1">
        <title>Description</title>
        <p>Raw annotations
Partial hierarchical expansion</p>
        <p>Full hierarchical expansion
The two runs that included the hierarchical information did not perform as well (based on average
EER and AUC) as the one based on the raw annotations. Looking at the EER scores for each
annotation term, the hierarchical methods were consistently worse performing. For these reasons
the results for these runs will not be discussed further.</p>
        <p>The EER and AUC scores are summarised in Table 2. Our scores are better than the averages
of the other participants, however, they are still a fair way o of the top scores. Using a semantic
space approach for annotation, from past experience, we expect that there will be a large diversity
in the performance for di erent terms. This is due to di erences in the amount of training data
for each term, and also the amount of visual diversity that might be associated with a term. For
example, visually speci c annotation terms require less training data than diverse ones. Table 3
shows the top- and bottom-most 5 annotation terms. It is interesting that the worst performing
terms are those that are rather general and unspeci c. The top performing terms all have very
speci c visual representations.</p>
        <p>Table 4 shows the results of our annotator using the hierarchical measure [10]. Unfortunately,
because this scoring measure performs a binary thresholding operation on the con dence scores,</p>
      </sec>
      <sec id="sec-2-2">
        <title>Annotation Term</title>
        <p>Sunset-Sunrise
Landscape-Nature</p>
        <p>Night</p>
        <p>Sea</p>
        <p>Mountains
Aesthetic-Impression</p>
        <p>Overall-Quality
Neutral-Illumination</p>
        <p>Sports
Fancy</p>
        <p>EER
0.232
0.234
0.237
0.243
0.249
0.416
0.436
0.466
0.470
0.478
the performance of our technique measures at the lower end of the spectrum of results from the
di erent participants. As previously discussed, the semantic space annotation approach doesn't
really permit the global setting of such a threshold.
The feature extraction phase was performed in parallel (4 images being processed at once) on a
quad core machine (Intel Core 2 Quad @ 2.66Ghz, 8G ram, Redhat Enterprise 5.3). The time for
image processing varied based on both the size of the image, and the image content. Timings for
a typical image from the training set are shown in Table 5.</p>
        <p>Training the semantic space took approximately 1 hour on a dual quad core 2.8GHz Xeon
workstation running Mac OS X (the semantic space code is single threaded, so only uses a single
core). We would estimate that no more than 1G of ram was used during the semantic space
training phase. Projecting all the test image in bulk took under 2 minutes, and it took about 5
minutes to generate annotations for all the 13000 images; so, in general, it took less than .05s to
get from a list of visual terms to the suggested annotations for a single image.
Implementation. The semantic-space software is written in C and makes use of Doug Rohde's
SVDLIBC 1 for e ciently performing the large sparse SVD. The feature detector and descriptor
software is written in C and C++. The image processing components were driven by a standard
UNIX make le, which enabled easy parallelisation.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>For this task we applied an older technique for automatically annotating images using a semantic
space. The performance of the technique in terms of EER and AUC is fairly competitive, however,
the technique does not mesh well with the hierarchical scoring measure proposed for this task. The
semantic space technique is reasonably computationally e cient; the most time is spent processing</p>
      <sec id="sec-3-1">
        <title>Feature</title>
        <p>Di erence-of-Gaussian detection + SIFT extraction</p>
        <p>MSER detection</p>
        <p>SIFT extraction on MSER
Colour-SIFT extraction on MSER</p>
        <p>Vector quantisation</p>
        <p>Estimated total</p>
        <p>Time
1.8s/image
0.1s/image
2.7s/image
1.0s/image
&lt;0.1s per set of extracted features
5.9s/image
the images to extract features. In our experiments, the use of the hierarchy did not lead to any
improvement in the annotation quality.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>The authors wish to thank the European Union, which supported this work under the Seventh
Framework project LivingKnowledge (IST-FP7-231126) and the LiveMemories project, graciously
funded by the Autonomous Province of Trento (Italy).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Gertjan</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Burghouts</surname>
          </string-name>
          and
          <string-name>
            <surname>Jan-Mark Geusebroek</surname>
          </string-name>
          .
          <article-title>Performance evaluation of local colour invariants</article-title>
          .
          <source>Computer Vision</source>
          and Image Understanding,
          <volume>113</volume>
          (
          <issue>1</issue>
          ):
          <volume>48</volume>
          {
          <fpage>62</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Jonathan</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Hare</surname>
            , Sina Samangooei,
            <given-names>Paul H.</given-names>
          </string-name>
          <string-name>
            <surname>Lewis</surname>
            , and
            <given-names>Mark S.</given-names>
          </string-name>
          <string-name>
            <surname>Nixon</surname>
          </string-name>
          .
          <article-title>Semantic spaces revisited: investigating the performance of auto-annotation and semantic retrieval using semantic spaces</article-title>
          .
          <source>In ACM CIVR '08</source>
          , pages
          <fpage>359</fpage>
          {
          <fpage>368</fpage>
          . ACM,
          <year>July 2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Jonathon</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Hare</surname>
            ,
            <given-names>Paul H.</given-names>
          </string-name>
          <string-name>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <surname>Peter G. B. Enser</surname>
            , and
            <given-names>Christine J.</given-names>
          </string-name>
          <string-name>
            <surname>Sandom</surname>
          </string-name>
          .
          <article-title>A LinearAlgebraic Technique with an Application in Semantic Image Retrieval</article-title>
          . In Hari Sundaram, Milind Naphade,
          <string-name>
            <surname>John R. Smith</surname>
          </string-name>
          , and Yong Rui, editors,
          <source>CIVR</source>
          <year>2006</year>
          , volume
          <volume>4071</volume>
          <source>of LNCS</source>
          , pages
          <volume>31</volume>
          {
          <fpage>40</fpage>
          . Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Mark J.</given-names>
            <surname>Huiskes</surname>
          </string-name>
          and
          <string-name>
            <given-names>Michael S.</given-names>
            <surname>Lew</surname>
          </string-name>
          .
          <article-title>The mir ickr retrieval evaluation</article-title>
          .
          <source>In MIR '08: Proceedings of the 2008 ACM International Conference on Multimedia Information Retrieval</source>
          , New York, NY, USA,
          <year>2008</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T K</given-names>
            <surname>Landauer and M L</surname>
          </string-name>
          <article-title>Littman</article-title>
          .
          <article-title>Fully automatic cross-language document retrieval using latent semantic indexing</article-title>
          .
          <source>In Proceedings of the Sixth Annual Conference of the UW Centre for the New Oxford English Dictionary and Text Research</source>
          , pages
          <volume>31</volume>
          {
          <fpage>38</fpage>
          ,
          <string-name>
            <surname>UW</surname>
          </string-name>
          <article-title>Centre for the New OED</article-title>
          and Text Research, Waterloo, Ontario, Canada,
          <year>October 1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>David</given-names>
            <surname>Lowe</surname>
          </string-name>
          .
          <article-title>Distinctive image features from scale-invariant keypoints</article-title>
          .
          <source>IJCV</source>
          ,
          <volume>60</volume>
          (
          <issue>2</issue>
          ):
          <volume>91</volume>
          {
          <fpage>110</fpage>
          ,
          <year>January 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Jiri</given-names>
            <surname>Matas</surname>
          </string-name>
          , Ondrej Chum,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Urban</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Pajdla</surname>
          </string-name>
          .
          <article-title>Robust wide baseline stereo from maximally stable extremal regions</article-title>
          . In Paul L.
          <article-title>Rosin and A</article-title>
          . David Marshall, editors,
          <source>BMVC. British Machine Vision Association</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>David</given-names>
            <surname>Nister</surname>
          </string-name>
          and
          <string-name>
            <given-names>Henrik</given-names>
            <surname>Stewenius</surname>
          </string-name>
          .
          <article-title>Scalable recognition with a vocabulary tree</article-title>
          . In In CVPR, pages
          <volume>2161</volume>
          {
          <fpage>2168</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Stefanie</given-names>
            <surname>Nowak</surname>
          </string-name>
          and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Dunker</surname>
          </string-name>
          .
          <article-title>Overview of the CLEF 2009 Large Scale - Visual Concept Detection and Annotation Task</article-title>
          .
          <source>In CLEF working notes 2009</source>
          , Corfu, Greece,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J</given-names>
            <surname>Sivic</surname>
          </string-name>
          and
          <string-name>
            <given-names>A</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <article-title>Video google: A text retrieval approach to object matching in videos</article-title>
          . In ICCV, pages
          <volume>1470</volume>
          {
          <fpage>1477</fpage>
          ,
          <year>October 2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>