<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CLaC at ImageCLEFPhoto 2008</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Osama El Demerdash</string-name>
          <email>el@cse.concordia.ca</email>
          <email>osama@cse.concordia.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leila Kosseim</string-name>
          <email>kosseim@cse.concordia.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sabine Bergler</string-name>
          <email>bergler@cse.concordia.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Concordia University</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2008</year>
      </pub-date>
      <abstract>
        <p>This paper presents our participation at the ImageCLEFPhoto 2008 task. We submitted six runs, experimenting with our own block-based visual retrieval as well as with query expansion. The results we obtained show that despite the poor performance of the visual and text retrieval components, good results can be obtained through Pseudo-relevance feedback and the fusion of the results.</p>
      </abstract>
      <kwd-group>
        <kwd>Image Retrieval</kwd>
        <kwd>Pseudo-relevance feedback</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>This paper presents our participation at the ImageCLEFPhoto 2008 task. We submitted six runs
experimenting with our own block-based visual retrieval as well as with query expansion. While
the 2008 task introduced a focused clustering theme for the first time, we did not attempt to use
the cluster target information. Our resources comprise a text search engine and a content-based
search system that we developed. The results we obtained show that despite the poor performance
of the visual and text retrieval components, good results can be obtained through Pseudo-relevance
feedback and the inter-media fusion of the results.</p>
      <p>Visual query</p>
      <p>Image</p>
      <p>Database
Text Query</p>
      <p>Text Search</p>
      <p>Query Expansion
Results</p>
      <p>Combined</p>
      <p>
        Results
For text retrieval, we used the Apache Lucene engine [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] , which implements a TF-IDF paradigm.
Stemming was done using the Snow-ball stemmer based on the porter algorithm and available
separately under the BSD license at http://snowball.tartarus.org/. For image analysis, we utilized
the Java Advanced Imaging (JAI) API v. 1.1.3. JAI is available from http://www.java.net.
4
      </p>
    </sec>
    <sec id="sec-2">
      <title>Text Retrieval</title>
      <p>As mentioned earlier, for text retrieval, we used the Apache Lucene engine, which implements
a TF-IDF paradigm. Stop-words were removed, the rest of the terms were stemmed using the
porter stemming algorithm. The documents were then indexed as field data retaining only the
title, notes and location fields, all of which were concatenated into one field. All text query terms
were joined using the OR operator.</p>
      <p>When searching the text, the query is also stemmed and stop words are removed. we found
that using the first sentence of the narrative field in addition to the title field improves the result.
By contrast, the rest of the narrative needs semantic processing to avoid introducing noise.
For visual retrieval, we implemented a system based on unsupervised analysis of the image. We
sought to capture basic global and local color, texture and shape information. In order to achieve
this goal, we divided the image into 2X2, 3X3, 4X4 and 5X5 blocks yielding 4, 9, 16 and 25 equal
partitions respectively. We also used the image as a whole as well as a center block occupying half
the image dimensions. Figure 2 shows the different regional divisions used to analyze the image.
The image was first converted to the Intensity/Hue/Saturation (IHS) color space, a perceptual
color space which more intuitive and reflective of human color perception than the RGB color
Space, then the following features were extracted:
• A three-band color histogram for each of the image divisions
• A histogram of the grey level image
• A histogram of the gradient magnitude image for each of the divisions of the grey-level image
• A three-band color histogram of the thumbnail of the image
The first feature captures the color characteristics of the image, while the grey-level histogram
conveys some texture information. The gradient magnitude adds the outline of the shapes in the
image.</p>
      <p>For retrieval, the different partitions are compared to their counter parts in the query image.
We did not experiment with assigning different weights to features. After experimenting with
several measures including the Euclidean distance, the Manhattan distance was chosen as the
distance measure. The images in the database were ranked according to their highest proximity
to any of the three query images.</p>
    </sec>
    <sec id="sec-3">
      <title>Query Expansion and Fusion of the Results</title>
      <p>The next step in processing the query is the text query expansion involving a pseudo-relevance
feedback mechanism and the fusion of the text and visual search results.
6.1</p>
      <sec id="sec-3-1">
        <title>Query Expansion</title>
        <p>We attempted several ways of query expansion. The highest ranked results from each of the
text and visual search engines were passed to the other engine. We also attempted adding noun
synonyms from WordNet to the query.</p>
        <p>Our best run uses pseudo-relevance feedback for query expansion. While the MAP of the visual
only run is only 0.055, its precision at five retrieved documents (0.328) is significantly higher than
that of the text only run (0.236). For this reason, we use the highest ranked document for expansion
of the text query. This is only done if the document meets a confidence level that we determined
empirically. The confidence score is assigned based on the proximity score to the query image.
6.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Fusion of Image and Text Search Results</title>
        <p>To combine the results from the different media searches, we again took into consideration the
confidence level in the visual results (i.e. the level of proximity from the query images). A
maximum of three highest ranked images is taken from the visual query result depending on the
confidence score, followed by a majority of the text results after query expansion.
7</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>We submitted six runs at ImageCLEFPhoto 2008. Following is their description:
• clacTX: Uses text search with title field only
• clacTXNR: uses Text Search with Title field and the first sentence of Narrative field with
query expansion
• clacIR: uses visual search only
• clacIRTX: combines the results from clacTXNR and clacIR
• clacNoQE: uses Text Search on Title and Narrative Field without query expansion
• clacNoQEMX: same as clacNoQE combined with clacIR</p>
      <p>Run ID
clacTX
clacTXNR
clacIR
clacIRTX
clacNoQE
clacNoQEMX
Average run
Median run
Best run</p>
      <p>Modality
Text
Text
Visual
Mixed
Mixed
Mixed
N/A
N/A
N/A</p>
      <p>MAP
0.1201
0.2577
0.0552
0.2622
0.2034
0.218
0.2187
0.2096
0.4288</p>
      <p>P10
0.1872
0.4103
0.2282
0.4359
0.3205
0.4026</p>
      <p>P20
0.1487
0.3449
0.1615
0.3744
0.2705
0.3269
0.3203
0.3203
0.6962</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Otis</given-names>
            <surname>Gospodnetic</surname>
          </string-name>
          and
          <string-name>
            <given-names>Erik</given-names>
            <surname>Hatcher</surname>
          </string-name>
          . Lucene in Action.
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Jun-Hua Han</surname>
          </string-name>
          and
          <string-name>
            <surname>De-Shuang Huang</surname>
          </string-name>
          .
          <article-title>A novel BP-Based Image Retrieval System</article-title>
          .
          <source>In International Symposium on Circuits and Systems (ISCAS</source>
          <year>2005</year>
          ),
          <fpage>23</fpage>
          -
          <lpage>26</lpage>
          May
          <year>2005</year>
          , Kobe, Japan, pages
          <fpage>1557</fpage>
          -
          <lpage>1560</lpage>
          . IEEE,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Nicolas</given-names>
            <surname>Maillot</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jean-Pierre Chevallet</surname>
          </string-name>
          , and
          <string-name>
            <surname>Joo-Hwee Lim</surname>
          </string-name>
          .
          <article-title>Inter-media pseudo-relevance feedback application to imageclef 2006 photo retrieval</article-title>
          . In Carol Peters, Paul Clough, Fredric C. Gey, Jussi Karlgren, Bernardo Magnini, Douglas W. Oard, Maarten de Rijke, and Maximilian Stempfhuber, editors,
          <source>CLEF</source>
          , volume
          <volume>4730</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>735</fpage>
          -
          <lpage>738</lpage>
          . Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Takala</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ahonen</surname>
            <given-names>T.</given-names>
          </string-name>
          , and Pietik¨ainen M.
          <article-title>Block-based methods for image retrieval using local binary patterns</article-title>
          .
          <year>2005</year>
          . In:
          <article-title>Image Analysis</article-title>
          ,
          <source>SCIA 2005 Proceedings, Lecture Notes in Computer Science 3540</source>
          , Springer,
          <fpage>882</fpage>
          -
          <lpage>891</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>