<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IIIT-H at ImageCLEF Wikipedia MM 2009</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Srinivasarao Vundavalli International Institute of Information Technology</institution>
          ,
          <addr-line>Hyderabad-500032, Andhra Pradesh</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>TF-IDF, Information Retrieval</institution>
          ,
          <addr-line>Image Retrieval, Vector Space Model, Boolean Model</addr-line>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>4</lpage>
      <abstract>
        <p>In this paper, we describe the IIITH retrieval system used for the ImageCLEF Wikipedia MM task. The system automatically ranks the most similar images to a given textual query. The system preprocesses the data set in order to remove the non-informative terms. For each query, the system finds a ranked list of its most similar images using the textual information only.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>7 Digital Libraries</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Our system uses a combination of the Vector Space Model (VSM) of Information Retrieval and
the Boolean model to determine how relevant a given Document is to a User’s query. In general,
the idea behind the VSM is the more times a query term appears in a document relative to the
number of times the term appears in all the documents in the collection, the more relevant that
document is to the query. It uses the Boolean model to first narrow down the documents that
need to be scored based on the use of boolean logic in the Query specification.</p>
      <p>
        The score is calculated based on TF-IDF model[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
2.1
      </p>
      <sec id="sec-1-1">
        <title>Term Frequency-Inverse Document Frequency Model</title>
        <p>
          In the TF-IDF model [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], each document in the collection and a query are represented by their
associated vector of the length of the vocabulary.
        </p>
        <p>tf (t in d) correlates to the term’s frequency, defined as the number of times term t appears in
the currently scored document d. Documents that have more occurrences of a given term receive
a higher score. The default computation of tf (t in d) in our system is sq.rt(f requency).</p>
        <p>idf (t) stands for Inverse Document Frequency. This value correlates to the inverse of docF req
(the number of documents in which the term t appears). This means rarer terms give higher
contribution to the total score. The default computation of idf (t) in our system is 1+log(numDocs/(docF req+
1))
2.2</p>
      </sec>
      <sec id="sec-1-2">
        <title>Vector Space Model(VSM)</title>
        <p>
          In a Vector Space Model[
          <xref ref-type="bibr" rid="ref3 ref4">3,4</xref>
          ], a document is represented as a vector. Each dimension corresponds
to a separate term. If a term occurs in the document, its value in the vector is non-zero. Several
different ways of computing these values, also known as (term) weights, have been developed. One
of the best known schemes is tf-idf weighting[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>The score of query q for document d correlates to the cosine-distance or dot-product between
document and query vectors. A document whose vector is closer to the query vector in that model
is scored higher. The scoring we used was the lucene[5] scoring.
3</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Runs and Results</title>
      <p>By using the image’s filename and the text associated with the image, we submitted one run based
on the models described in the previous section.</p>
      <p>The results for the wikipediaMM task have been computed with the trec eval tool (version
8.1). The submitted runs have been corrected (where necessary) so as to correpond to valid runs
in the correct TREC format. The following corrections have been made:
• The runs comply with the TREC fomat as specified in the submission guidelines for the task
• When a topic contains an image example that is part of the wikipediaMM collection, this
image is removed from the retrieval results, i.e., we are seeking relevant images that the
users are not familiar with (as they are with the images they provided as examples).
• When an image is retrieved more than once for a given topic, only its highest ranking for that
topic is kept and the rest are removed (and the ranks in the retrieval results are appropriately
fixed).
• Ensure that each of the submitted runs has a unique name.</p>
      <p>The interpolated recall precision averages are shown in Figure 1, and the summary statistics
for the runs sorted by MAP are shown in Table 1.</p>
      <p>Run</p>
      <sec id="sec-2-1">
        <title>Modality Topic Fields FB/QE MAP</title>
        <p>TXT</p>
      </sec>
      <sec id="sec-2-2">
        <title>TITLE</title>
        <p>NOFB</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>In this working note, we presented a retrieval system which automatically ranks the most similar
images to a given textual query. The system preprocesses the data set in order to remove the
noninformative terms. For each query, the system finds a ranked list of its most similar images using
the textual information only. Even though our system did not perform well, we are happy that we
participated in the ImageCLEF WikipediaMM task and are looking forward to ImageCLEF2009.
5</p>
      <sec id="sec-3-1">
        <title>5. http://lucene.apache.org</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. http://inex.is.informatik.uni-duisburg.de/2007/mmtrack.html</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Salton</surname>
            , Gerard and Buckley,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>1988</year>
          ).
          <article-title>Term-weighting approaches in automatic text retrieval</article-title>
          .
          <source>Information Processing and Management</source>
          <volume>24</volume>
          (
          <issue>5</issue>
          ):
          <fpage>513</fpage>
          -
          <lpage>523</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>G.</given-names>
            <surname>Salton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Yang</surname>
          </string-name>
          (
          <year>1975</year>
          ),
          <article-title>”A Vector Space Model for Automatic Indexing”</article-title>
          .
          <source>Communications of the ACM</source>
          , vol.
          <volume>18</volume>
          , nr.
          <volume>11</volume>
          , pages
          <fpage>613620</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>F.</given-names>
            <surname>Song</surname>
          </string-name>
          and W.B.
          <string-name>
            <surname>Croft</surname>
          </string-name>
          (
          <year>1999</year>
          ).
          <article-title>A General Language Model for Information Retrieval</article-title>
          . Research and Development in Information Retrieval:
          <fpage>279</fpage>
          -
          <lpage>280</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>