<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SINAI at ImageCLEF 2007</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>M.C. D´ıaz-Galiano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M.A. Garc´ıa-Cumbreras</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M.T. Mart´ın-Valdivia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Montejo-Raez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>L.A. Uren˜a-L´opez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Campus Las Lagunillas</institution>
          ,
          <addr-line>Ed. A3, E-23071, Ja ́en</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Ja ́en. Departamento de Informa ́tica</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2075</year>
      </pub-date>
      <abstract>
        <p>This paper describes the SINAI team participation in the ImageCLEF campaign. The SINAI research group has participated in both the ad hoc task and the medical task. The experiments accomplished in both cases result from very different approaches. For the ad hoc task the main Information Retrieval (IR) system used combines the document lists retrieved by two IR systems, and uses online translators for the bilingual experiments. For the medical task, we have used the MeSH ontology to expand the queries. The expansion consists in searching terms of the query in the MeSH ontology in order to add similar terms. We have processed the set of collections using Information Gain (IG) in the same way as in ImageCLEFmed 2006.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        This is the third participation of the SINAI research group at the ImageCLEF campaign. We have
participated in the ad hoc task [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and the medical task [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>The ad hoc task involves retrieving relevant images using the text associated to each image
query. As a cross-language retrieval task, multilingual image retrieval based on query translation
can achieve a higher performance than monolingual retrieval.</p>
      <p>This year, a new IR module has been tested. This module works with two different IR systems
and the final relevant list is the result of the combination of both IR lists. The Machine Translation
Module developed last year has been updated and used for the bilingual task. English, Spanish,
French, Italian and Portuguese are the languages used this year.</p>
      <p>
        The goal of the medical task is to retrieve relevant images based on an image query. This year,
two new collections have been introduced. We have filtered all the collections using IG to select
the best tags of each one [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Moreover, we have expanded the queries using MeSH ontology: we
have selected similar terms to the query in the ontology and have added them to the query itself.
      </p>
      <p>The following section describes the ad hoc experiments. In Section 3, we explain the
experiments for the medical task. Finally, conclusions and futher work are presented in Section 4.</p>
    </sec>
    <sec id="sec-2">
      <title>The Ad Hoc Task</title>
      <p>Given a multilingual query, the goal of the ad hoc task is to find as many relevant images as
possible from an image collection.</p>
      <p>The proposal of the ad hoc task is to compare results with and without pseudo-relevant feedback
(PRF), with or without query expansion, using different methods of query translation or using
different retrieval models and weighing functions.
2.1</p>
      <sec id="sec-2-1">
        <title>Experiments Description</title>
        <p>In our experiments we have used the five following languages: English, French, Italian, Portuguese
and Spanish.</p>
        <p>
          This year we have combined lists of relevant documents returned by two different IR systems:
Lemur1 and Jirs [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>
          As translation module we have used SINTRAM (SINai TRAnslation Module), our Meta
Machine Translation system that uses some online Machine Translators for each language pair and
implements some heuristics to combine the different translations [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. After a complete research we
have found that the best translators were:
• Systran for French, Italian and Portuguese
• Prompt for Spanish
        </p>
        <p>The dataset is the collection IAPR. The IAPR TC-12 image collection consists of 20,000
images taken from different locations around the world and comprises a varying cross-section of
still natural images. It includes pictures of a range of sports and actions, photographs of people,
animals, cities, landscapes and many others of contemporary life.</p>
        <p>The collections have been preprocessed using stopwords removal and the Porter’s stemmer.</p>
        <p>The dataset of the collection has been indexed using both IR systems, namely, Lemur and Jirs.
One parameter for each experiment is the weighing function, such as Okapi or TFIDF. Another
is the use or not of PRF.</p>
        <p>A simple fusion method has been implemented to obtain a simple list of relevant documents. In
the first step, both lists are normalized between 0 and 1. Then, some heuristics are implemented:
• Weighing each list. Some experiments are based on a weighing function that gives a
percentage of importance to the Lemur list and a different one to the Jirs list. The final score
of each relevant document is calculated by the sum of each score multiplied by its weight.</p>
        <p>Finally, the documents are sorted by their final fusion score.
• Using a threshold. Filtering relevant documents by a threshold value is another heuristic.</p>
        <p>If the score of a document is worse than this parameter, it is not included in the final list.</p>
        <p>Finally, the resultant list is sorted by the score of the included documents.</p>
        <p>
          With the ad hoc 2006 framework (using the same collection, queries and relevance judgements)
[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] and the heuristics already described we have evaluated several configurations in order to obtain
the best ones.
        </p>
        <p>1. The basic case, with Lemur, uses English queries, Lemur as IR system and Okapi with PRF
as weighing function. It obtains a MAP value of 0.1672
2. The basic case, with Jirs, uses English queries, Jirs as IR system and Okapi with PRF as
weighing function. It obtains a MAP value of 0.1513
3. For the first heuristic the weight for Lemur and Jirs changes between 1 and 0.1. For instance,
the experiment that weighs both lists in the same way uses 0.5 as the weight value for both
Lemur and Jirs. The best result was 0.1678, using a weight of 0.6 for Lemur and 0.4 for Jirs.</p>
        <sec id="sec-2-1-1">
          <title>Experiment</title>
          <p>EN-EN-Exp2
EN-EN-Exp1
ES-EN-Exp9
ES-EN-Exp10
PO-EN-Exp8
PO-EN-Exp7
FR-EN-Exp4
FR-EN-Exp3
IT-EN-Exp5
IT-EN-Exp6</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Language</title>
          <p>English
Spanish
Portuguese
French
Italian</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>Experiment</title>
          <p>EN-EN-Exp11
ES-EN-Exp15
PO-EN-Exp14
FR-EN-Exp12
IT-EN-Exp13
IR
Fusion
Fusion
Fusion
Fusion
Fusion
Expansion
without
without
without
without
without
without
without
without
without
without</p>
        </sec>
        <sec id="sec-2-1-4">
          <title>Expansion</title>
          <p>without
without
without
without
without</p>
        </sec>
        <sec id="sec-2-1-5">
          <title>Weight</title>
          <p>Okapi
Okapi
Okapi
Okapi
Okapi
4. In the case of the second heuristic, different values, from 0.1 to 0.9, are tested as threshold.</p>
          <p>The best result was 0.1524, obtained with 0.1 value as threshold.</p>
          <p>Finally, all the Lemur weighing functions have been tested and the best one was again Okapi
with feedback, with a result of 0.1672.
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Results and Discussion</title>
        <p>With the results obtained with the 2006 framework, the new 2007 queries were run. We sent 15
runs: five runs using Lemur, five using Jirs and five with the fusion of both lists.</p>
        <p>The results obtained with each IR system (using only text) and the best MAP for each language
is shown in Table 1.</p>
        <p>Good results have been obtained this year with both IR systems. Only the English runs have
obtained a loss of MAP of around 25%. Our best Spanish result is similar to the best one obtained.
For Portuguese we have obtained the best one, and for French and Italian our results are a bit
worse: only a loss of MAP of around 8%. The MAP values are low because only the title is used
as query.</p>
        <p>From these results we can conclude that Lemur IR system works better than Jirs, but the
difference is not very significant.</p>
        <p>The results obtained by applying the fusion method and the best MAP for each language is
shown in Table 2.</p>
        <p>The main conclusion for these fusion results is that there may be a problem, because it is not
logical to obtain so poor results.</p>
        <p>Experiments with the 2006 framework gave us similar results between the ones obtained with
each simple IR system and the fusion ones.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>The Medical Task</title>
      <p>
        The main goal of the medical ImageCLEF task is to improve the retrieval of medical images from
heterogeneous and multilingual document collections containing images and text. Queries are
formulated with sample images and some textual description explaining the research goal. For the
medical task we have used the list of retrieved images by FIRE2 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which was supplied by the
organizers of this track.
      </p>
      <p>
        Last year, our efforts concentrated on manipulating the text descriptions associated with these
images and mixing the results partial lists with the GIFT lists [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We also focused on preprocessing
the collection using Information Gain (IG) in order to improve the quality of results and to
automatize the tag selection process. However, this year we have concentrated on improving the
queries using MeSH ontology.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Preprocessing the Collection</title>
        <p>In order to generate the textual collection we have used the ImageCLEFmed.xml file that links
collections with their images and annotations. It has external links to the images and to the
associated annotations in XML files. It contains relative paths from the root directory to all the
related files.</p>
        <p>The entire collection consists of six datasets (CASImage, Pathopic, Peir, MIR, endoscopic and
MyPACS) containing about 66,600 images (16,600 more than the previous year). Each
subcollection is organized into cases that represent a group of related images and annotations. In every
case a group of images and an optional annotation is given. Each image is part of a case and
has optional associated annotations, which enclose metadata and/or a textual annotation. All
the images and annotations are stored into separated files. ImageCLEFmed.xml only contains the
connections between collections, cases, images and annotations.</p>
        <p>
          The collection annotations are in XML format and most of them are in English. We have
preprocessed the collections to generate a textual document per image [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>We have used the IG measure to select the best XML tags in the collection. Once the document
collection was generated, experiments were conducted with the Lemur retrieval information system
by applying the KL-divergence weighing scheme.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Expanding Queries with MeSH Ontology</title>
        <p>The Medical Subject Headings (MeSH) is a thesaurus developed by the National Library of
Medicine3. MeSH contains two organization files, an alphabetic list with bags of synonymous
and related terms, and a hierarchical organization of descriptors associated to the terms. A term
is composed by one o more words.</p>
        <p>
          We have used the bags of terms to expand the queries. If all the words of a term are in the
query, we generate a new expanded query by adding all its bag of terms. To compare the words
of a particular term and those of the query, we first put all the words in lowercase and we do not
remove stopwords. In order to reduce the number of terms that could expand the query, we have
only used those that are in A, C or E categories of MeSH (A: Anatomy, C: Diseases, E: Analytical,
Diagnostic and Therapeutic Techniques and Equipment) [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Experiment Description</title>
        <p>Our main objective is to investigate the effectiveness of the query expansion together with filtering
tags using IG in the text collection. We have carried out these experiments using a corpus with
20%, 30%, 40%, 50% and 60% of tags for the 2007 collection, because these settings led to the
best results on the 2006 corpus.</p>
        <p>Finally, the expanded textual list and the FIRE list are merged in order to obtain one final list
(FL) with relevant images ranked by relevance. The merging process was done by giving different
importance to the visual (VL) and textual lists (TL):</p>
        <p>F L = T L ∗ α + V L ∗ (1 − α)
(1)
2http://www-i6.informatik.rwth-aachen.de/˜deselaers/fire.html
3http://www.nlm.nih.gov/mesh/</p>
        <sec id="sec-3-3-1">
          <title>Experiment LIG-MRIM-LIG MU A.eval</title>
          <p>SINAI-SinaiC100T80.eval
miracleTxtENN.txt.eval
OHSU-oshu as is 1000.eval
UB-NLM-UBTI 1.eval
IPAL-IPAL1 TXT BAY ISA0.1.eval
RWTH-FIRE-ME-tr0506.eval
GE EN.treceval.eval
iclefmed2007 text2 out.txt.eval
UNALCO-nni FeatComb.eval
DEU CS-DEU R2.eval
The total runs submitted to ImageCLEFmed2007 for textual retrieval and mixed retrieval were
more than 100.</p>
          <p>Table 3 shows the best results for the groups participating in ImageCLEFmed2007. In this
table we can observe that the best result obtained by our system is the experiment with 100% of
the tags and α = 0.8 (80% of textual information).</p>
          <p>In Table 4 we can see the results of only textual experiments. The best result obtained is 0.3668
of precision value, using 100% of tags in the collection. Surprisingly, the IG selection experiments
do not improve the results. However, by using 40% of tags the loss of precision is lower than 5%.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Further Work</title>
      <p>This year the ad hoc task obtained good results, but always with a low MAP, because only the title
was used. It could be interesting to develop a new fusion module and to apply a query expansion
module based on Google, tasks on which we are already working.</p>
      <p>This year two new collections have been included in the ImageCLEFmed2007. Adding this
new information in the collection improves the results obtained with the Lemur IR system. We
want to investigate what type of query (textual, visual or mixed) is influenced the most by these
new collections. Moreover, our next step will focus on using the UMLS ontology4, in order to
include the multiligual features of the collections.</p>
      <p>4http://www.nlm.nih.gov/pubs/factsheets/umls.html</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This project has been partially supported by a grant from the Spanish Government, project
TIMOM (TIN2006-15265-C06-03).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grubinger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deselaers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and Mu¨ller, H.:
          <article-title>Overview of the ImageCLEF 2006 Photographic Retrieval and Object Annotation Tasks</article-title>
          .
          <source>In Proceedings of the Cross Language Evaluation Forum (CLEF</source>
          <year>2006</year>
          ),
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Deselaers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weyand</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keysers</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Macherey</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ney</surname>
          </string-name>
          , H.:
          <article-title>FIRE in ImageCLEF 2005: Combining Content-based Image Retrieval with Textual Information Retrieval</article-title>
          .
          <source>Working Notes of the CLEF Workshop</source>
          , Vienna, Austria,
          <year>September 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D</given-names>
            <surname>´</surname>
          </string-name>
          ıaz-Galiano,
          <string-name>
            <given-names>M.C.</given-names>
            ,
            <surname>Garc</surname>
          </string-name>
          <article-title>´ıa-</article-title>
          <string-name>
            <surname>Cumbreras</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <article-title>Mart´ın-</article-title>
          <string-name>
            <surname>Valdivia</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montejo-Raez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <article-title>and Uren˜a-L´opez</article-title>
          , L.A.:
          <article-title>SINAI at ImageCLEF 2006</article-title>
          .
          <source>In Proceedings of the Cross Language Evaluation Forum (CLEF</source>
          <year>2006</year>
          ),
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G</given-names>
            <surname>´omez-Soriano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.M.</given-names>
            ,
            <surname>Montes-</surname>
          </string-name>
          y-G´omez,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Sanchis-Arnal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            , and
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>A Passage Retrieval System for Multilingual Question Answering</article-title>
          .
          <source>8th International Conference of Text, Speech and Dialogue 2005 (TSD'05). Lecture Notes in Artificial Intelligence (LNCS/LNAI 3658)</source>
          . pp.
          <fpage>443</fpage>
          -
          <lpage>450</lpage>
          . Karlovy Vary,
          <string-name>
            <given-names>Czech</given-names>
            <surname>Republic</surname>
          </string-name>
          .
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Garc</surname>
          </string-name>
          <article-title>´ıa-</article-title>
          <string-name>
            <surname>Cumbreras</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uren˜</surname>
            a-L´opez,
            <given-names>L.A.</given-names>
          </string-name>
          ,
          <article-title>Mart´ınez-</article-title>
          <string-name>
            <surname>Santiago</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Perea-Ortega</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          :
          <string-name>
            <given-names>BRUJA</given-names>
            <surname>System</surname>
          </string-name>
          . The University of Ja´
          <article-title>en at the Spanish task of QA@CLEF 2006</article-title>
          .
          <article-title>In Proceedings of the Cross Language Evaluation Forum (CLEF</article-title>
          <year>2006</year>
          ),
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Grubinger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and Mu¨ller, H.:
          <article-title>Overview of the ImageCLEF 2007 Photographic Retrieval Task</article-title>
          .
          <source>Working Notes of the 2007 CLEF Workshop</source>
          . Sep,
          <year>2007</year>
          . Budapest, Hungary.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7] Mu¨ller, H.,
          <string-name>
            <surname>Deselaers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalpathy-Cramer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deserno</surname>
            ,
            <given-names>T.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Hersh</surname>
          </string-name>
          , W.:
          <article-title>Overview of the ImageCLEFmed 2007 Medical Retrieval</article-title>
          and
          <string-name>
            <given-names>Annotation</given-names>
            <surname>Tasks</surname>
          </string-name>
          .
          <source>Working Notes of the 2007 CLEF Workshop</source>
          . Sep,
          <year>2007</year>
          . Budapest, Hungary.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Chevallet</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lim</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Radhouani</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Using Ontology Dimensions and Negative Expansion to solve Precise Queries in CLEF Medical Task</article-title>
          .
          <source>Working Notes of the 2005 CLEF Workshop</source>
          . Sep,
          <year>2005</year>
          . Vienna, Austria.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>