<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>I2R ImageCLEF Photo Annotation 2009 Working Notes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Appendix: Selected Features</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Jiquan Ngiam and Hanlin Goh Institute for Infocomm Research, Singapore</institution>
          ,
          <addr-line>1 Fusionopolis Way</addr-line>
          ,
          <country country="SG">Singapore</country>
          <addr-line>138632</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the method that was used for our two submission runs for the ImageCLEF Photo Annotation task. Image Features Global Features The following global features were each computed over the entire image. Each feature essential provides a histogram for the image. In features where a quantized HSV space was used, the following quantization parameters were employed: 12 Hue Bins, 3 Saturation Bins, 3 Value Bins. Bins were of equal width in each dimension. This results in a total of 108 bins. The choice of these parameters was motivated by [6] HSV Histogram - 108 dim</p>
      </abstract>
      <kwd-group>
        <kwd>Photo annotation</kwd>
        <kwd>Global and local features</kwd>
        <kwd>Support Vector Machines (SVM)</kwd>
        <kwd>feature selection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The ImageCLEF Photo Annotation 2009 task involved 53 concepts spanning from abstract
concepts (aesthetics, blur) to visual elements (trees, people). Although an ontology was provided, our
method did not rely on the ontology heavily, except for handling disjoint cases.</p>
      <p>
        Our method follows the framework in [
        <xref ref-type="bibr" rid="ref8">9</xref>
        ] , involving Support Vector Machines and extended
Gaussian Kernels over the 2 distance. We use a variety of global features, novel local region
selectors and simple greedy feature selection.
2.1
2.1.1
The quantized HSV histogram (as above) was used as a feature vector.
      </p>
      <p>Color Auto Correlogram (CAC) - 432 dim
We computed the CAC over a quantized HSV space with 4 distances 1,3,5,7. This gives us a
feature vector of 108*4 = 432 dimensions.</p>
      <p>
        For each color &amp; distance pair (c, d), we computed the probability of nding the same color
at exactly distance d away. Refer to [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] for details.
2.1.3
      </p>
      <p>
        Color Coherence Vector (CCV) - 216 dim
We computed the CCV over a quantized HSV space. Since there are two states - coherent and
incoherent, this gives us a feature vector of 108*2 = 216 dimensions. We set the tau parameter to
1% of the image size. Refer to [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] for details.
2.1.4
      </p>
      <p>Census Transform (CT) - 256 dim
The CT histogram is a simple transformation of each pixel into a 8-bit value based on its 8
surrounding neighbors (two states, either &gt;= or &lt; its neighbor). This provides a feature vector of
256 dimensions (histogram of the CT values). Refer to [8] for details.
2.1.5</p>
      <p>Edge Orientation Histogram - 37 dim
We used the LTI-Lib's Canny Edge detector to compute the edge orientation histogram. Each
pixel is assigned to either an edge (with orientation) or non-edge. Orientations are quantized into
5 degree angle bins, giving a total of 36 bins for 180 degrees. An additional bin is concatenated
for non-edges. This gives a nal vector of 37 dimensions.
2.1.6</p>
      <p>
        Interest Point Based SIFT - 500 dim
We used the SIFT binary provided by David Lowe [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to compute SIFT descriptors. The descriptors
were quantized into 500 visual words. A visual words dictionary was computed using k-Means
clustering.
2.1.7
      </p>
      <p>
        Densely Sampled SIFT - 1500 dim
We densely sampled SIFT points at 10 pixel spacings, 4 scales (4, 8, 12, 16 px radius) and 1
orientation. The points are similarly quantized into 1500 visual words by k-Means. This scheme
follows that of [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
2.2
      </p>
      <sec id="sec-1-1">
        <title>Local Region Features</title>
        <p>For a number of concepts, the classi cation problem can be framed in the Multiple Instance
Learning (MIL) framework. In essence, a concept (e.g. Mountain) exists if and only if a region
within the image demonstrates the concept. Hence, one is motivated to consider whether it is
possible to improve performance by nding appropriate region(s) to consider in an image.</p>
        <p>We de ne a local region to be a bounding box. Given a bounding box in an image, one can
compute image features similar to the global features. Hence, we de ne a local region feature to be
a feature vector that is extracted based only on region in a bounding box. Therefore, the problem
of nding local region features is reduced to one of nding good bounding boxes for each image.
2.2.1</p>
        <sec id="sec-1-1-1">
          <title>Local Region Selectors</title>
          <p>To nd good local region selectors (bounding boxes), we frame the problem in a MIL setting. In
this setting, each image is a bag-of-regions and regions are considered to be true i they contain
the target concept. Furthermore, a bag is true i it contains a true region.</p>
          <p>
            We used EM-Diverse Density [
            <xref ref-type="bibr" rid="ref9">10</xref>
            ] together with E cient Subwindow Search (ESS) [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ] to search
for a target concept with good diverse density. We note that since ESS is able to consider all
possible rectangular subwindows, the algorithm essentially considers all possible bounding boxes.
However, multiple restarts are required since the algorithm is susceptible to local minimas.
          </p>
          <p>For each concept, we learned a local region selector based on the densely sample SIFT features.
From each local region, we extract only the HSV, interest point SIFT and densely sampled SIFT
histograms. These three histograms form the (concept-speci c) local features for each image.
3
3.1</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Learning and Feature Selection</title>
      <sec id="sec-2-1">
        <title>Support Vector Machines</title>
        <p>
          For the nal classi cation, we used LIBSVM [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] with probability estimates (provided with the
software). Each concept was treated separately as a individual classi cation task. Following [
          <xref ref-type="bibr" rid="ref8">9</xref>
          ],
we used extended Gaussian kernels with the 2 distance.
        </p>
        <p>K(Si; Sj ) =</p>
        <p>X</p>
        <p>1
f2features f
2(f (Si); f (Sj ))</p>
        <p>Both local and global features were treated in the same manner. f is the average 2 distance
for a particular feature; we used it to normalize the distances across di erent features.</p>
        <p>We also performed experiments on cost-sensitive SVMs but the results did not vary that much.
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Feature Selection</title>
        <p>
          Unsurprisingly, di erent features work well with di erent concepts. To combine di erent features,
one could incorporate weighting into the kernel function. This method is adopted by the INRIA
group in their VOC2007 submission [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. However, learning these weights is non-trivial and one
often resorts to ad-hoc methods such as genetic algorithms.
        </p>
        <p>We chose a simpler method for feature selection in which a greedy algorithm is used.
Furthermore, we do employ any partial weighting scheme for the features. Using equal error rate (ERR)
as out performance measure, the algorithm is described as follows:</p>
        <sec id="sec-2-2-1">
          <title>Greedy Algorithm</title>
          <p>1. F = all global features
2. For each feature f 2 F : Compute error rate if f is removed
3. Remove the feature which results in best improvement
4. Repeat (2-3) until removing any feature results in worse performance
5. Consider each feature f 2 All F eatures</p>
          <p>F : Compute error rate if f is added
6. Consider each feature f 2 F : Compute error rate if f is removed
7. Add or remove the feature which gives best improvement
8. Repeat (5-7) until local optima is reached
9. Return F
The appendix contains a list of concepts and the features that were selected for each of them.
While there was a new hierarchical measure introduced for the task, we did not speci cally optimize
for it. The lack of the annotator agreement values also made it more di cult to optimize for this
new measure.</p>
          <p>However, for the classes that were speci ed as disjoint, we did simple post processing on the
probability estimates to ensure that exactly 1 of the concepts is 0:5. This was achieved by
simply moving the probability estimates to 0:5 .</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Submissions and Results</title>
      <p>Two nal submissions were made for this task. One used all the features available while the other
used only the global features. The o cial results of the two runs in terms of Average Equal Error
Rate (Avg. EER), Average Area Under Curve (Avg. AUC), Average Annotation Score with
Annotator Agreement (Avg. AS with AA) and Average Annotation Score without Annotator
Agreement (Avg. AS without AA) are reported in Table 1.</p>
      <p>Based on the Avg. EER measure, our run with all features was ranked 6 out of 74 submitted
runs, while the run using only global features was ranked 11. Comparing the best runs from
each group, were reported to be third in the list of twenty participating groups. Evaluating our
performance based on the Average Annotation Score with and without the use of Annotator
Agreement, our runs were ranked second (global features only) and third (all features).</p>
      <p>Submission Run</p>
      <p>Avg. EER</p>
      <p>Avg. AUC</p>
      <p>Avg. AS
with AA</p>
      <p>Avg. AS
without AA
All Features
(CVIUI2R 22 2 1244628714641.txt)
Global Features Only
(CVIUI2R 22 2 1244629050173.txt)
[8] Jianxin Wu and James M. Rehg. Where am i: Place instance and category recognition using
spatial pact. Computer Vision and Pattern Recognition, IEEE Computer Society Conference
on, 0:1{8, 2008.</p>
      <p>Concept
0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52</p>
      <p>Features1
gdsift1500 gsift500 gcac ldsift1500
gdsift1500 gsift500 gcac gccv gcthist
gsift500 gcac gccv gcthist
gdsift1500 gsift500 gcac gcthist
gdsift1500 gsift500 gcac ldsift1500
gdsift1500 gsift500 gcac ghsv
gdsift1500 gsift500 gcac gcthist
gdsift1500 gsift500 gcac
gdsift1500 gcac gccv ghsv gcthist
gdsift1500 gsift500 gcac ghsv gcthist ldsift1500
gsift500 gcac gccv ghsv gcthist
gdsift1500 gsift500 gcac gccv ghsv lsift500 ldsift1500 gedgehist
gdsift1500 gsift500 ghsv gcthist
gdsift1500 gsift500 ghsv ldsift1500
gdsift1500 gsift500 ghsv gcthist
gdsift1500 gsift500 gccv
gdsift1500 gsift500 gccv gcthist
gdsift1500 gsift500 gcac
gdsift1500 gsift500 gcac ghsv gedgehist
gdsift1500 gsift500 ldsift1500
gdsift1500 gcac
gdsift1500 gsift500 gccv
gdsift1500 gsift500 gccv gcthist gedgehist
gdsift1500 gsift500 gcac gccv
gsift500 gcac ghsv gcthist
gdsift1500 gcac gccv ghsv gcthist
gdsift1500 lsift500 ldsift1500 lhsv
gdsift1500 gsift500 gcac
gsift500 gcac gccv ghsv gcthist
gdsift1500 gsift500 gcac
gdsift1500 gsift500 gcac gccv gcthist
gdsift1500 gsift500 gcac gccv
gdsift1500 gsift500 gcthist
gdsift1500 gsift500 gcac gcthist
gdsift1500 gccv gcthist lsift500
gdsift1500 gsift500 gcac lhsv gedgehist
gdsift1500 gsift500 gccv ghsv
gdsift1500 gsift500 gccv ghsv
gdsift1500 gsift500 gccv
gdsift1500 gsift500 gcac
gdsift1500 gsift500 gccv ghsv gcthist
gdsift1500 gsift500 gcthist
gdsift1500 gsift500 gcthist
gdsift1500 gsift500 gcac gccv
gdsift1500 gsift500 gcthist gedgehist lsift500 lhsv
gdsift1500 gsift500 gccv gcthist
gdsift1500 gsift500 gcac
gdsift1500 gsift500 gcac ldsift1500 gedgehist
gdsift1500 gsift500 gcac gccv gcthist gedgehist
gdsift1500 gsift500 gcthist ldsift1500
gdsift1500 gsift500 gcac gcthist
gdsift1500 gsift500 gcac gccv ldsift1500 lsift500
gdsift1500 gsift500 gcac gccv ghsv ldsift1500 gedgehist</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Anna</given-names>
            <surname>Bosch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Muoz</surname>
          </string-name>
          .
          <article-title>Scene classi cation using a hybrid generative/discriminative approach</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          ,
          <volume>30</volume>
          (
          <issue>4</issue>
          ):
          <volume>712</volume>
          {
          <fpage>727</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Chih-Chung Chang</surname>
          </string-name>
          and
          <string-name>
            <surname>Chih-Jen Lin</surname>
          </string-name>
          .
          <article-title>LIBSVM: a library for support vector machines</article-title>
          ,
          <year>2001</year>
          . Software available at http://www.csie.ntu.edu.tw/~cjlin/libsvm.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Everingham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Van</given-names>
            <surname>Gool</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. K. I.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Winn</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <source>The PASCAL Visual Object Classes Challenge</source>
          <year>2007</year>
          (
          <article-title>VOC2007) Results</article-title>
          . http://www.pascal-network. org/challenges/VOC/voc2007/workshop/index.html.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C. H.</given-names>
            <surname>Lampert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Blaschko</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          .
          <article-title>Beyond sliding windows: Object localization by e cient subwindow search</article-title>
          .
          <source>In Computer Vision and Pattern Recognition</source>
          ,
          <year>2008</year>
          .
          <article-title>CVPR 2008</article-title>
          . IEEE Conference on, pages
          <volume>1</volume>
          {
          <issue>8</issue>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>David G.</given-names>
            <surname>Lowe</surname>
          </string-name>
          .
          <article-title>Distinctive image features from scale-invariant keypoints</article-title>
          .
          <source>International Journal of Computer Vision</source>
          ,
          <volume>60</volume>
          :
          <fpage>91</fpage>
          {
          <fpage>110</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T</given-names>
            <surname>Ojala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Rautiainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E</given-names>
            <surname>Matinmikko</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M</given-names>
            <surname>Aittola</surname>
          </string-name>
          .
          <article-title>Semantic image retrieval with hsv correlograms</article-title>
          .
          <source>In 12th Scandinavian Conference on Image Analysis</source>
          , pages
          <volume>621</volume>
          {
          <fpage>627</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Greg</given-names>
            <surname>Pass</surname>
          </string-name>
          , Ramin Zabih, and
          <string-name>
            <given-names>Justin</given-names>
            <surname>Miller</surname>
          </string-name>
          .
          <article-title>Comparing images using color coherence vectors</article-title>
          .
          <source>In MULTIMEDIA '96: Proceedings of the fourth ACM international conference on Multimedia</source>
          , pages
          <volume>65</volume>
          {
          <fpage>73</fpage>
          , New York, NY, USA,
          <year>1996</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Marszalek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lazebnik</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <article-title>Local features and kernels for classication of texture and object categories: A comprehensive study</article-title>
          .
          <source>Int. J. Comput. Vision</source>
          ,
          <volume>73</volume>
          (
          <issue>2</issue>
          ):
          <volume>213</volume>
          {
          <fpage>238</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Qi</given-names>
            <surname>Zhang and Sally</surname>
          </string-name>
          <string-name>
            <given-names>A.</given-names>
            <surname>Goldman</surname>
          </string-name>
          .
          <article-title>Em-dd: An improved multiple-instance learning technique</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , pages
          <fpage>1073</fpage>
          {
          <fpage>1080</fpage>
          . MIT Press,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>