<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>J.L. Martínez-Fernández</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ana García Serrano</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julio Villena</string-name>
          <email>jvillena@daedalus.es</email>
          <email>jvillena@it.uc3m.es</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Víctor David Méndez Sáenz</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Santiago González Tortosa</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michelangelo Castagnone</string-name>
          <email>mcastagnone@isys.dia.fi.upm.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Javier Alonso</string-name>
          <email>jalonso@daedalus.es</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Advaced Databases Group, Computer Science Department, Universidad Carlos III de Madrid</institution>
          ,
          <addr-line>Avda. Universidad 30, 28911 Leganés, Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Artificial Intelligence Department, Universidad Politécnica de Madrid.</institution>
          <addr-line>Campus de Montegancedo s/n, Boadilla del Monte 28660</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>DAEDALUS - Data, Decisiond and Language, S.A. Centro de Empresas “La Arboleda”</institution>
          ,
          <addr-line>Ctra. N-III km. 7,300 Madrid 28031</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Department of Telematic Engineering, Universidad Carlos III de Madrid</institution>
          ,
          <addr-line>Avda. Universidad 30, 28911 Leganés, Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2004</year>
      </pub-date>
      <abstract>
        <p>The second participation of the MIRACLE (Multilingual Information RetrievAl for the CLEf campaign) research group in the ImageCLEF task is described in this paper. New techniques, devoted to the combination of linguistic and statistical language processing methods, have been tested, continuing with the experiments carried out in last year This is the second time for the MIRACLE (Multilingual Information RetrievAl for the CLEf campaign) research group as a participant in the Image CLEF task. The work presented in this paper is the continuation of the experiments carried out in CLEF 2003. Some new techniques, like the inclusion of linguistic information for monolingual English tasks or the application of EuroWordnet as a translation and query expansion tool, have been developed and tested.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Adhoc Retrieval Task</title>
      <p>As can be seen in the figure, different tools have been used to process english queries:



</p>
      <p>
        BRILL: A tagger, based on Brill's work [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], can be used to attach a morphosyntactic tag to each word.
Proper Names: A Proper Name detection module can be applied at the output of the Brill tagger,
although plain text can also be used as input to this module.
      </p>
      <p>SETA Module: In a final step, the text can be divided in sentences and their constituent phrases can be
extracted using this module. Thus, this module is a linguistic parser for english, implemented using
prolog.</p>
      <p>EWN synonym: This box represents a subsystem used to extract the corresponding WordNet synonyms
for a given word. So, a query expansion is implemented based on semantic information contained in
WordNet database. Optionally, the linguistic category of a given word is used when the semantic
expansion is performed. For example, if a word acting as a name is going to be expanded, only
synonyms of the given word that can act as a name are considered.</p>
      <p>
        The rest of languages have been treated using EuroWordNet, where available, or translation tools, like Systran
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] or Translation Experts [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. These experiments where devoted to test the quality of EuroWordNet when used
in translation tasks and as a synonym expansion tool. For translation purposes, the inter-lingual index (ILI)
supplied with EuroWordNet has been applied. Again, it is possible to consider the linguistic category of the
word when asking for its translations. So, if a name is going to be translated, only words that can act as a name
in the target language are taken into account.
      </p>
      <sec id="sec-2-1">
        <title>Monolingual English experiments</title>
        <p>For the rest of languages, only the baseline index database has been used to search translated queries.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Index/Tasks</title>
        <p>Tokenize</p>
      </sec>
      <sec id="sec-2-3">
        <title>Filter</title>
        <p>Stopwords</p>
      </sec>
      <sec id="sec-2-4">
        <title>Brill</title>
        <p>Tagger</p>
      </sec>
      <sec id="sec-2-5">
        <title>Filter</title>
        <p>Nouns</p>
      </sec>
      <sec id="sec-2-6">
        <title>Stem</title>
        <p>ming</p>
      </sec>
      <sec id="sec-2-7">
        <title>Filter Proper Nouns</title>
        <sec id="sec-2-7-1">
          <title>DB1 - Baseline</title>
        </sec>
        <sec id="sec-2-7-2">
          <title>DB2 - Only Nouns</title>
        </sec>
        <sec id="sec-2-7-3">
          <title>DB3 - Proper Names + Baseline DB4 - Proper Names + Nouns</title>
          <p>√
×
√
×
√
×
√
×
×
√
×
√
×
√
×
√
×
×
√
√
×
×
√
√
Query Process</p>
        </sec>
        <sec id="sec-2-7-4">
          <title>Topic Words</title>
        </sec>
        <sec id="sec-2-7-5">
          <title>Topic Words + Synonyms</title>
        </sec>
        <sec id="sec-2-7-6">
          <title>Nouns</title>
        </sec>
        <sec id="sec-2-7-7">
          <title>Nouns + Synonyms without category</title>
        </sec>
        <sec id="sec-2-7-8">
          <title>Nouns + Synonyms with category</title>
        </sec>
        <sec id="sec-2-7-9">
          <title>Topic Words + Proper Names</title>
        </sec>
        <sec id="sec-2-7-10">
          <title>Topic Words +Synonyms + Proper Names</title>
        </sec>
        <sec id="sec-2-7-11">
          <title>Nouns + Proper Names</title>
        </sec>
        <sec id="sec-2-7-12">
          <title>Nouns + Synonyms without category + Proper</title>
        </sec>
        <sec id="sec-2-7-13">
          <title>Names</title>
        </sec>
        <sec id="sec-2-7-14">
          <title>Nouns + Synonyms with category + Proper Names</title>
        </sec>
        <sec id="sec-2-7-15">
          <title>Topic and Narration Words</title>
        </sec>
        <sec id="sec-2-7-16">
          <title>Topic and Narration Words + Synonyms with category</title>
        </sec>
      </sec>
      <sec id="sec-2-8">
        <title>Database Searched DB 1 DB 1</title>
        <p>DB 2
DB 2
DB 2
DB 3
DB 3
DB 4
DB 4
DB 4
DB 3
DB 3</p>
        <p>Run Name
mirobaseen
mirosbaseen
mironounen
mirosnounen
miroscnounen
miroppbaseen
mirosppbaseen
miroppnounen
mirosppnounen
miroscppnounen
mirorppbaseen
mirorscppbaseen
In this table, 'Topic Words' means that all simple words (excluding stopwords) are used to search the
corresponding index database. 'Synonyms' means that all synonyms for a word found in WordNet are used to
expand the query, without any refinement. 'Nouns' stands for the situation where the query text is tagged and
only words acting as nouns are selected as part of the final query. 'Proper Names' is used to mark that only
recognized proper names in the text are used as part of the query. 'Synonyms with category' is used to distinguish
the process in which not all the synonyms of a words are taken into account, but only those synonyms that can
act with the same category than the initial word are included in the query. Finally, in the last two experiments
included in the table, the narrative of the query (only available for the english queries) is used as the input to the
SETA module, in charge of parsing the text and getting a more precise category for the word.</p>
      </sec>
      <sec id="sec-2-9">
        <title>Monolingual English Results</title>
        <p>Regarding these results, it is important to highlight some points: first of all, the basic experiment (taken as the
baseline) produces the best results. Very different results are obtained from the fourth result in advance and
again a gap in the average precision can be seen from the eighth result in advance. These differences in precision
show that, when all words are used in the characterization of the textual captions results are better and the
inclusion of more linguistic information (like proper nouns or synonyms) does not lead to an improvement. On
the other side, if only common or proper nouns are used to represent the documents there is a loss in precision,
perhaps due to the fewer number of words used for document characterization. Also, its worth mentioning that
the experiment using all linguistic information that available tools can extract is among the worse ones</p>
      </sec>
      <sec id="sec-2-10">
        <title>Bilingual Experiments</title>
        <p>For the bilingual experiments two different approaches, depending on available information, have been applied.
These two approaches are:

</p>
        <p>A EuroWordNet based approach, where information contained in the ILI index provided by
EuroWordNet is used to translate the original query. This approach has been applied for the following
languages: Spanish, German, French and Italian.</p>
        <p>A translator based approach, where online translation tools, in particular Systran and Translation
Experts tools, have been applied to translate the queries from the initial language to the target language
(English in Image CLEF tasks).</p>
        <p>In both approaches, the index database used corresponds to the baseline, i.e., the one where all words (excluding
stopwords) are considered as indexes. Table 4 summarizes the features of the experiments defined for this
multilingual task.</p>
        <p>Tokenize</p>
        <p>Remove</p>
        <p>Stopwords</p>
        <p>EWN
Languages</p>
        <sec id="sec-2-10-1">
          <title>Automatic translation using</title>
          <p>web translator: BabelFish
(http://babelfish.altavista.com)</p>
        </sec>
        <sec id="sec-2-10-2">
          <title>Automatic translation using</title>
          <p>web translator: TransExp
(http://www.tranexp.com)</p>
        </sec>
      </sec>
      <sec id="sec-2-11">
        <title>Index Database</title>
        <p>DB 1
DB 1
DB 1
Run Names
mirowbaseit
mirowbasefr
mirowbasees
mirowbaseesc
mirowbasege
mirobaseru
mirobaseja
mirobasezh
mirobasedu
mirobasesw
mirobaseda</p>
      </sec>
      <sec id="sec-2-12">
        <title>Bilingual Results</title>
        <p>
          According to these results, one important fact to mention is the loss of precision, taking into account the best
monolingual experiment. As can be seen, a decrease of 34,07% in precision, marking again the importance of the
quality of the translators used in multilingual environments. Cases where EuroWordNet has been used as a
translation tool can be compared with CLEF 2003 obtained results [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and an important decrease in precision can
be noticed. Last year bilingual experiments with French, German, Italian and Spanish where around 40%
average precision, while this year average precision for these languages is around 30%. It is also worth
mentioning that other participants, according to official results, have obtained only a decrease of 10% in
precision for some bilingual tasks, so, in our situation, there is room for improvement.
        </p>
        <p>Mixing text based retrieval with content based image retrieval (CBIR) for the Adhoc task
This year, the MIRACLE team has made a first step in image content retrieval. This first step has led to the
definition of experiments where content based image retrieval (CBIR) is applied. This is the case of the adhoc
retrieval task, where some runs mixing results obtained using textual search and CBIR search have been
submitted. The CBIR subsystem used for this experiment is based on GIFT 0.1.9 and will be described in the
next section. The text retrieval subsystem is the one used in text based experiments, although for initial test and
tuning of the overall system, last year data and text search systems have been used.</p>
        <p>The process of mixing textual and image results begins taking the list with the images returned by the text search
subsystem and their relevance figures and building a query for the CBIR subsystem. The content search is
performed and a new search is performed considering the 5 first elements returned. Finally, results obtained with
this last relevance feedback approach are combined with the original results list returned by the textual search
subsystem. The expression used to combine these partial lists is:
k REL _ VIS weight _ vis × REL _ TXT weight _ txt ,
factor_vis,
factor_txt,
In this expression, REL_VIS and REL_TXT are the relevance value returned by the CBIR subsystem and the text
search subsystem respectively. factor_vis, factor_txt, weight_vis and weight_txt are parameters to be defined and
can be used to adjust the overall system according to obtained results, for example, giving more importance to
textual results or CBIR results.</p>
        <p>Results of the text and CBIR mixing experiments
Two sets of experiments have been done. Results for the first set are included in Table 6, where the initial set of
text search results have been some of the experiments defined in Table 2.
Comparing to Table 3, these results are very close (and always below) of the ones where only textual search is
applied. This can be due to the chosen configuration of the combination algorithm. More tests should be made to
extract a valid conclusion.</p>
        <p>Some other experiments where executed using a different textual search subsystem. Obtained results for these
experiments (Table 7 and Table 8) have been always worse than the previously mentioned ones. One of these
sets of experiments, Table 8, is a bilingual one with English as a target language and Spanish as the initial
language.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Medical Retrieval Task</title>
      <p>This year ImageCLEF organizers have defined a new task where the main focus is image content based retrieval.
For this purpose a set of medical images, including scans, x-ray images and photographs of different illness has
been made available to ImageCLEF participants.</p>
      <p>
        The CBIR system used has been GIFT 0.1.9 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] developed under GNU licence which allows query by example,
using an image as a starting point for the search process, and implements relevance feedback methods. This
software has been developed by the Vision Group at the CUI of the University of Geneva.
      </p>
      <p>Although the first step in the search process for this task must involve an image, textual descriptions of the
medical cases have been used to try to improve retrieval results.</p>
      <p>The search process can be divided in the following steps:
1. The initial query, formed by one image, is introduced in the CBIR system to obtain a set of images to define
the query.
2. The CBIR system returns a list of images along with the corresponding relevance values. The number of
images used in the search process is called relevance threshold and constitutes a system configuration
parameter.
3. Previous steps have produced a valid query which is introduced in the overall system. The complete system
is formed by a textual subsystem and a CBIR subsystem. In a first step both subsystemas are used to
perform the search process.
4. Partial results lists are combined using an intersection operator: images not appearing in both partial lists are
dropped. Two special parameters make it possible to consider textual results more important than CBIR
ones or vice versa.
5. The previous step produces a unique results list that is again introduced in the CBIR subsystem. The new
results list obtained is again combined, applying the intersection operator, with the output of the textual
subsystem.</p>
      <p>The overall process is depicted in Figure 2. The expression used to obtain a unique relevance value according to
the partial results lists produced by textual and CBIR subsystems is:</p>
      <p>k REL _ VIS weight _ vis × REL _ TXT weight _ txt ,
Rel =
for elements appearing only in one
partial list
Some other models, based on different ways to combine the output of the textual and CBIR subsystems, have
been tested but the one described here has produced the best results. Four different runs have been defined
according to different values for the configuration parameters defined for the overall system. These parameters
include: the minimum threshold to build the initial query, the number of results used for relevance feedback and
the weights given to textual and image results.</p>
      <sec id="sec-3-1">
        <title>Medical Retrieval Task Results</title>
        <p>Average precision figures obtained for the submitted experiments are included in Table 9. As can be seen, the
difference in precision among the first and the last run is around 2%, not enough to extract some conclusions
about which method (or configuration parameters set) is the best. On the other hand, according to the obtained
rank for these runs, there is still room for improvement in this task, perhaps testing new configurations or new
values for defined parameters or taking the most of textual descriptions related to each medical case.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>6. Conclusions</title>
      <p>This is the second year for the MIRACLE team taking part in the CLEF campaign and in the ImageCLEF track
in particular. The main goal pursued this year was to continue with the research in finding a right combination of
linguistic and statistical methods to improve the Information Retrieval process. MIRACLE group is also very
interested in the field of multimedia retrieval so, the content based image retrieval task defined this year as part
of the ImageCLEF track was a great opportunity to take a first step in the field. From out point of view, obtained
results for the adhoc retrieval task are very good. Average precision values for the monolingual english task are a
little bit better than the ones obtained last year, pointing that is difficult to improve results for this task. Perhaps
the best performance figures that can be obtained with actual technology have been reached. On the contrary,
bilingual tasks, in the way we have developed them, can be improved.</p>
      <p>A mention apart must be made for the content based image retrieval task, where obtained results are not as good
as for the textual task. This fact drives us to increase efforts devoted to this kind of retrieval for the following
campaigns.</p>
    </sec>
    <sec id="sec-5">
      <title>7. Acknowledgements</title>
      <p>This work has been partially supported by the projects OmniPaper (European Union, 5th Framework Programme
for Research and Technological Development, IST-2001-32174) and MIRACLE (Regional Government of</p>
      <sec id="sec-5-1">
        <title>Madrid, Regional Plan for Research, 07T/0055/2003).</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] “Altavista's Babel Fish Translation Service”</article-title>
          , http://babelfish.altavista.com/,
          <source>last accessed 12.08</source>
          .2004
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Brill</surname>
            <given-names>E.</given-names>
          </string-name>
          ,
          <article-title>"Some Advances in Transformation Based Part of Speech Tagging"</article-title>
          ,
          <source>proceedings of the Twelfth National Conference on Artificial Intelligence</source>
          ,
          <year>1994</year>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] “Eurowordnet:
          <article-title>Building a Multilingual Database with Wordnets for several European Languages</article-title>
          .”, http://www.let.uva.nl/ewn/, March, 1996
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.A.</given-names>
            <surname>Miller</surname>
          </string-name>
          . “
          <article-title>WordNet: A lexical database for English”</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>38</volume>
          (
          <issue>11</issue>
          ):
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          ,
          <year>1995</year>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>[5] “The Porter Stemming Algorithm” page maintained by Martin Porter</article-title>
          . http://www.tartarus.org/ ~martin/PorterStemmer/,
          <source>last accessed 12.08</source>
          .2004
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>[6] "The GNU Image-Finding Tool</article-title>
          .
          <source>GIFT 0.1.9"</source>
          , http://www.gnu.org/software/gift/,
          <source>last accessed 12.08</source>
          .2004
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7] “Translation Experts”, http://www.transexp.com,
          <source>last accessed 12.08</source>
          .2004
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Villena</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinez-Fernandez</surname>
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fombella</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>García-Serrano</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruíz-Cristina</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martínez</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goñi</surname>
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>GonzálezCristóbal J.C.</surname>
          </string-name>
          , “
          <article-title>Image Retrieval: the MIRACLE Approach”</article-title>
          ,
          <source>In CLEF 2003 Proceedings</source>
          , Springer-Verlag, to appear,
          <year>2003</year>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>