<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Miguel E. Ruiz</string-name>
          <email>meruiz@buffalo.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silvia B. Southwick</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>State University of New York at Buffalo School of Informatics Department of Library and Information Studies 534 Baldy Hall</institution>
          ,
          <addr-line>Buffalo NY</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2005</year>
      </pub-date>
      <abstract>
        <p>This work was part of SUNY at Buffalo's overall participation in cross-language retrieval of image collections (ImageCLEF). Our main goal was to explore the combination of Content-Based Image Retrieval (CBIR) and text retrieval of medical images that have clinical annotations in English, French and German. We used a system that combined the content-based image retrieval system GIFT and the well-known SMART system for text retrieval. Translations of English topics to French were performed by mapping the English text to UMLS concepts using the 2005 UMLS meta-thesaurus. Results show that combining both CBIR and Text retrieval yields significant improvements of retrieval performance.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The organizers of the conference provided a list of the top 1000 retrieved images returned by GIFT for
each of the 25 topics. We used this list as the output of the CBIR system. From the textual part of the
query we used only the English version provided in the query. This Text was processed with
MetaMap,. The current Version of UMLS includes vocabularies with translations in 13 languages
(Including French and German which are the ones of interest for this task). For these experiments we
decided to use the French terms present in UMLS as translations. We decided not to use German due to
the fact that the German documents come from the Pathopic collection and constitute a parallel
collection of English and German. Since our initial query is in English we consider redundant to try to
get German translations for the topics. However, German terms might be added if they rank high in the
automatic relevance feedback process. The corresponding English and French UMLS terms associated
with the concepts identified by MetaMap were added to the original English Queries. The text was
processed by SMART and the top 10 cases retrieved (together with the cases from top 10 image
retrieved by GIFT) are used for expand the query. Observe that this is a multilingual expansion since
terms in French, German or English could be added by the query expansion process. The images
associated with each case are given the same retrieval score from the text retrieval. Finally we
combined the image scores obtained from the CBIR system and from the text retrieval system. We
weighted this combination and the results indicate that a 3:1 weighting in favor of text retrieval achieve
the best performance in our system.</p>
      <p>Results
We submitted five runs that included different weights for the contribution of text and images, and
several variations on the number of terms used for expansion. The best combination of our official
results was obtained by weighting the text results 3 times higher than the visual results. The parameters
for the relevance feedback used the top 10 results (both images and text cases), and expand the queries
with the top 50 ranked terms using Rocchio’s formula (with α=8, β=64, γ=16). This combination
showed a significant improvement for retrieval performance (35% with respect to using text only and
150% using only image retrieval). Our runs were ranked 4th overall and our system was among the best
3 participating in the conference.</p>
    </sec>
    <sec id="sec-2">
      <title>Text only</title>
      <p>( UBimed_en-fr.T.Bl )</p>
    </sec>
    <sec id="sec-3">
      <title>Visual only</title>
    </sec>
    <sec id="sec-4">
      <title>Text and visual</title>
    </sec>
    <sec id="sec-5">
      <title>UBimed_en-fr.TI.1</title>
    </sec>
    <sec id="sec-6">
      <title>UBimed_en-fr.TI.2</title>
    </sec>
    <sec id="sec-7">
      <title>UBimed_en-fr.TI.3</title>
    </sec>
    <sec id="sec-8">
      <title>UBimed_en-fr.TI.4</title>
    </sec>
    <sec id="sec-9">
      <title>UBimed_en-fr.TI.5</title>
    </sec>
    <sec id="sec-10">
      <title>Parameters</title>
      <p>n=10, m=50
GIFT results</p>
      <p>MAP
0.1746
0.0864
Our preliminary analysis showed significant improvements of 35% over text-only and 173% over
queries that used only visual features. We still need to do the query by query analysis to try to identify
if there are specific types of queries for which our approach works better.</p>
      <p>Our future research plans include exploring in more detail the contribution of the translation using
MetaMap and UMLS, types of queries that benefit the most from either text or image retrieval, and
sensitivity of the system to different parameters.</p>
      <p>Acknowledgements
We want to thank the National Library of Medicine for their support of Miguel Ruiz as a visiting
faculty at the National Library of Medicine and for allowing us to use UMLS and MetaMap for this
research. We want to thank AT&amp;T for supporting this research by providing the necessary hardware
used in our experiments through a small grant for research at the UB School of Informatics.
Bibliography</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Effective Mapping of Biomedical Text to the UMLS Metathesaurus: The MetaMap Program</article-title>
          .
          <article-title>Paper presented at the AMIA</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Ruiz</surname>
            ,
            <given-names>M. E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Srikanth</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>UB at CLEF2004 Cross Language Medical Image Retrieval</article-title>
          .
          <article-title>Paper presented at the Fifth Workshop of the Cross-Language Evaluation Forum (CLEF</article-title>
          <year>2004</year>
          ), Bath, England.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Salton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>1971</year>
          ).
          <article-title>The SMART Retrieval System: Experiments in Automatic Document Processing</article-title>
          . Englewood Cliff, NJ: Prentice Hall.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>