<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>I ma geC L EF  2006 E xp er imen t s a t  th e C h emn it z T ech n ica l  Un iver sity </article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Chemnitz University of Technology Faculty of Computer Science</institution>
          ,
          <addr-line>Media Informatics Strasse der Nationen 62, 09107 Chemnitz</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>5</lpage>
      <abstract>
        <p>Abstr act.  We  present  a  report  about  our  participation  in  the  ImageCLEF  photo  task  2006  and  a  short  description  of  our  new  framework  for  future  use  in  further  CLEF  participations.  We  described and analysed  our participation in the monolingual English task. Special Lucene­aligned  query expansion and histogram comparisons are helping to improve the baseline results.  ACM Categor ies and Subject Descr iptor s  1 Intr oduction  Our  goal  was  to  develop  a  new  information  retrieval  framework  which  can  handle  different  custom  retrieval  mechanisms  (i.e.  Lucene,  the  GNU  Image­Finding  Tool  and  others).  The  framework  should  be  so  general  that  one can use either one of the custom retrieval mechanisms or combine different ones. Therefore our efforts were  not focused on better results but to get the framework up and running.  2 System descr iption  For our participation at CLEF 2006 we constructed a GUI (graphical user interface) quite easy to handle. It bases  on Lucene version 1.4 and includes several user­configurable components. For example we wrote an Analyzer  in  which TokenFilter s can be added and rearranged.  For the participation in ImageCLEF 2006 we needed to extend this system with image retrieval capabilities. But  instead of extending the old system we rewrote it and developed a general IR framework. This framework does  not necessarily require Lucene and makes it possible to build an abstract layer for other search engines like GIFT  (GNU Image­Finding Tool) and others.  The main component of our framework is the class Run, which includes all information required to perform and  repeat searches. A Run contains the Topic s, which can be preprocessed by TopicFilters a search method called Searcher , which can read and search the Index and HitSet s, which contain the results.  To use another search engine one will have to extend the abstract classes Indexer  and Searcher . To use another  data collection type (e.g. GIRT4, IAPR, ...) one will have to extend the abstract class DataCollection.  </p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>an Index, which was created by an Indexer  from a DataCollection
In  order  to  merge  results  there  is  a  special  extension  of  the  class  Searcher :  the  MergeSearcher .  It  combines 
several Searcher s and uses an abstract class Merger to merge their results.</p>
    </sec>
    <sec id="sec-2">
      <title>2.1 Lucene </title>
      <p>The  text  search  engine  used  is  Lucene  version  1.9  with  an  adapted  Analyzer .  As  described  above  we  create  a 
LuceneIndexer   and  an  abstract  LuceneSearcher   that  provides  basic  abilities  to  underlying  Searcher s  like  a 
selectable  Lucene Analyzer .  The Analyzer   is  based  on  the StandardTokenizer  and LowerCaseFilter   by  Lucene 
and is enhanced by a custom Snowball filter for stemming. We also added a positional stop word filter and used 
the  stop  word  list  suggested  on  the  CLEF  website.  The Analyzer   details  (stemmer  and  stop  word  list)  can  be 
configured  through  a  GUI.  The  LuceneStandardSearcher   works  with  the  Lucene  MultiFieldQueryParser   and 
searches in specified fields of the index. </p>
    </sec>
    <sec id="sec-3">
      <title>2.2 GIFT (GNU Image­Finding Tool) </title>
      <p>The  GIFT  was  not  used  for  a  submitted  run  but  it  should  be  mentioned  here  to  show  the  abilities  of  the 
framework.  We  only  implemented  a  Searcher ,  which  can  access  a  GIFT  server  and  takes  images  to  generate 
queries by example. The underlying index must be created with the command line utility provided by the GIFT. </p>
    </sec>
    <sec id="sec-4">
      <title>2.3 IAPR Data Collection </title>
      <p>To access the IAPR data collection of ImageCLEF we implemented the class IaprDataCollection that provides 
two special options for the DataDocument s to contain special data:
· 
· 
option 1: the DataDocument s contain the annotations
option 2: the DataDocument s contain the histograms 
In both cases the associated images can be retrieved with the help of the DataDocument . </p>
    </sec>
    <sec id="sec-5">
      <title>3 Descr iption of the r uns submitted </title>
      <p>We focused on the monolingual task using the English topics and annotations. 
The  Lucene  index  of  the  annotations  contains  all  fields.  All  fields  are  indexed  and  tokenized  except  DATE, 
which is not tokenized. The DOCNO is the only field  which is stored, resulting in a quite small index (only 2 
MB). 
We used the default BooleanQuery from Lucene with no further weighting of any field pairs in order to avoid 
any corpus specific weighting. 
The default Lucene search classes (default similarity, etc.) were used to search the topic­annotation field pairs as 
shown in table 1. </p>
      <sec id="sec-5-1">
        <title>Topic field </title>
      </sec>
      <sec id="sec-5-2">
        <title>TITLE </title>
      </sec>
      <sec id="sec-5-3">
        <title>NARR </title>
      </sec>
      <sec id="sec-5-4">
        <title>Annotations searched </title>
      </sec>
      <sec id="sec-5-5">
        <title>TITLE, LOCATION, DESCRIPTION, NOTES </title>
      </sec>
      <sec id="sec-5-6">
        <title>DESCRIPTION, NOTES </title>
        <p>For all runs, except “tucEEANT”, the query expansion was applied once, using the first 20 results of the original 
query. 
In  the  run  “tucEEAFT2”  the  query  expansion  is  applied  twice.  The  first  time  using  the  annotations  of  the 
example images and the second time using the first 20 results of the original query. 
To speed up the colour histogram comparison, we also created an index for the histogram data. For each image</p>
      </sec>
      <sec id="sec-5-7">
        <title>Rank  Rank  EN­EN  ALL  1  2 </title>
        <p>As  one  can  see,  run  “tucEEAFT2”  is  our  best  run  which  is  no  surprise  if  one  takes  into  account  that  three 
definitely relevant documents are used for query expansion. Our second best run is “tucEEAFTI” which shows 
that the additional histogram comparison produces better results than without it. 
The  good  results  in  ImageCLEF  2006  show  that  our  framework  works  and  that  our  first  implemented  search 
methods provide good results. 
The  normal  query  expansion  improves  the  mean  average  precision  by  more  than  1%.  Considering  the  poor 
overall  mean  average  precision  this  seems  to  be  a  good  improvement  of  the  results.  Although  the  histogram 
comparison  is a  very  simple  and  presumably  poor  measurement  to  compare  images, it improves  our  text­only 
retrieval slightly. 
The  next  steps  to  improve  the  results  is  to  implement  a  better  image  retrieval  algorithm  due  to  the  weak 
performance of the histogram comparison. </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Refer ences </title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.  Stevens, J. S., Husted T., Cutting D., &amp; 
          <string-name>
            <surname>Carlson</surname>
          </string-name>
           P. (
          <year>2005</year>
          ). Apache Lucene ­ Overview ­ 
          <string-name>
            <surname>Apache</surname>
          </string-name>
           Lucene.  Retrieved August 
          <fpage>10</fpage>
          , 
          <year>2006</year>
          , from http://lucene.apache.org/java/. 
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.  Grubinger M., Leung C., &amp; 
          <string-name>
            <surname>Clough</surname>
          </string-name>
           P. (
          <year>2005</year>
          ). 
          <article-title>The IAPR Benchmark for Assessing Image Retrieval  Performance in Cross Language Evaluation Tasks</article-title>
          . Retrieved August 
          <fpage>10</fpage>
          , 
          <year>2006</year>
          , from  http://ir.shef.ac.uk/cloughie/papers/muscle­imageclef2005.
          <fpage>pdf</fpage>
          . 
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.  Roos, M. (
          <year>2005</year>
          ).
          <article-title>The GNU Image­Finding Tool ­ GNU Project ­ Free Software Foundation (FSF)</article-title>
          .
          <source>  Retrieved August 10</source>
          ,
          <year>2006</year>
          , from http://www.gnu.org/software/gift/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>