<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UESTC at ImageCLEF 2011 Medical Retrieval Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hong Wu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chengbo Tian</string-name>
          <email>tianchengbo@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science and Engineering, University of Electronic Science and Technology of China</institution>
          ,
          <addr-line>611731 Chengdu</addr-line>
          ,
          <country country="CN">P. R. China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <abstract>
        <p>This paper describes methods and results archived by our research group at the ImageCLEF 2011 medical retrieval task. We performed two subtasks, ad-hoc retrieval and case-based retrieval, and only used text information for retrieval. In our work, a phrase-based retrieval model was adopted, and UMLS metathasaurus was used to expand query. The phrase-based model was implemented based on Indri search engine and their structured query language. For query expansion, the detected concepts and their direct children were used to append the structured query. Both phrases and medical concepts were identified with the help of the MetaMap program. The parameters of our approach were trained on the data of ImageCLEFmed 2010 ad-hoc retrieval subtask.</p>
      </abstract>
      <kwd-group>
        <kwd>Medical Retrieval</kwd>
        <kwd>Phrase-Based Model</kwd>
        <kwd>Indri</kwd>
        <kwd>Query Expansion</kwd>
        <kwd>MetaMap</kwd>
        <kwd>UMLS</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        This paper describes the second participation of the UESTC group at the ImageCLEF
medical retrieval task. In previous years, we tested a phrase-based approach for
medical retrieval. Phrases and subphrases were extracted with the help of MetaMap,
and the individual words in each phrase or subphrase were concatenated and used as
an indexing term in vector space model, together with single word terms. In this year,
we adopted a more principled and efficient phrase-based retrieval model, which was
implemented based on Indri search engine [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and their structured query language.
For query expansion, the concepts detected from the original text query and their
direct children were used to append the structured query. Our approach had gotten
promising results on the data of ImageCLEFmed 2010 ad-hoc retrieval subtask [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        ImageCLEFmed 2011 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] includes three types of tasks, ad-hoc retrieval,
casebased retrieval and modality classification. For the retrieval tasks, the dataset
contains 230,088 images from more than 55,000 articles published in online medical
journals. In the ad-hoc retrieval task, a set of 30 textual queries, each of which with
several sample images, are given, and the goal is to retrieve the images most relevant
to each topic. In the case-based task, a set of 10 case-based information requests are
given, and the goal is to retrieve the articles most relevant to the topic case.
      </p>
      <p>The remainder of this paper is organized as follows. Our approach is described in
section 2. And our submitted runs and results are presented in section 3, followed by
the conclusions in section 4.</p>
    </sec>
    <sec id="sec-2">
      <title>Phrase-Based Retrieval Model and Query Expansion</title>
      <p>
        To utilize phrase in information retrieval, there’re three steps: (1) Identify phrases in
query text, (2) Identify these phrases in document, (3) Combine phrase with
individual word in ranking function. In this work, the first step was performed with
the help of MetaMap, and the last two steps were implemented based on Indri search
engine [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and its structured query language. Indri is a scalable search engine that
inherits the inference net framework from InQuery and combines it with language
modeling approach to retrieval.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Phrase Identification</title>
        <p>
          In our approach, phrase identification was conducted with the help of MetaMap [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]
which is a tool to map biomedical text to concepts in the UMLS Metathesaurus [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
MetaMap first parses the text into phrases, and then performs intensive variant
generation on each phrase. After that, candidates are retrieved from the Metathesaurus
to match the variants. Finally, the candidates are evaluated by a mapping algorithm,
and the best candidates are returned as the mapped concepts.
        </p>
        <p>In this work, the concept mapping was restricted to three source vocabularies:
MeSH, SNOMED-CT, and FMA. And two phrase identification strategies were
explored. One is to use the phrases produced by the early step in MetaMap program,
and filter out the unwanted words in them, such as preposition, determiner etc.
Another is to consider the individual word or sequence words which mapped to
UMLS concept as a phrase.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Phrase Representation</title>
        <p>After identifying phrase in query text, the query phrases would be recognized
again within documents. Many phrase identification techniques only look at
contiguous sequences of words. But, the constituent words of a query phrase might
be several words apart, and even with different order when used within a document.
Fortunately, this problem can be easily solved by using operators in indri query
language. The query can be reformulated with the special operators to provide more
exact information about the relationship of terms in the original text query. Here, we
introduce some operators which are commonly used for representing phrase.
</p>
        <p>Ordered Window Operator : #N(T1...Tn) or #odN(T1...Tn)</p>
        <p>The terms within an ordered window operator must appear ordered with at most
N1 terms between adjacent terms in the document in order to contribute to the
document's belief score.
</p>
        <p>Unordered Window Operator: #uwN(T1 ... Tn)</p>
        <p>The terms contained in an unordered window operator must be found in any order
within a window of N words in order to contribute to the belief score of the document.
For example, the phrase “congestive heart failure” can be represented as
#1(congestive heart failure), which means that the phrase is recognized in document
only if the three constituent words are found in the right order and no other words
between them. It can also be represented as #uw6(congestive heart failure), which
means that the phrase is recognized in document only if the three constituent words
are found in any order within a window of 6 words.</p>
        <p>There are also some other operators related to our works. They are introduced as
follow.</p>
        <p>Combine Operator: #combine (T1 ... Tn)</p>
        <p>The terms or nodes contained in the combine operator are treated as having equal
influence on the final result. The belief scores provided by the arguments of the
combine operator are averaged to produce the belief score of the #combine node.</p>
        <p>The terms or nodes contained in the weight operator contribute unequally to the
final result according to the weight associated with each (Wi). The belief scores
provided by the arguments of the weight operator are weighted averaged to produce
the belief score of the #weight node. Taking #weight(1.0 dog 0.5 train) for example,
its belief score is 0.67 log( b(dog) ) + 0.33 log( b(train) ).
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Phrase-Based Retrieval Model</title>
        <p>When utilizing phrase in information retrieval, phrasal term should be combined
with word term in ranking function. Simply, we can use one ranking function for
word term, another for phrasal term, and the final ranking function is weight sum of
them.</p>
        <p>,
,
,
(1)
where Q and D stand for query and document respectively. f1(Q,D) is the ranking
function for word term, and f2(Q,D) is the ranking function for phrasal term. The
weights w1 and w2 can be tuned by experiment.</p>
        <p>With Indri search engine, this ranking function can be implemented easily with a
structured query and the inference network model. For example, the topic 8 in ad-hoc
track of imageCLEFmed 2011 is:
“x-ray images of a hip joint with prosthesis”</p>
        <p>With the second phrase identification strategy, the phrases are “x-ray images”, “hip
joint” and “prosthesis”. The query can be formulated as following,</p>
        <p>#weight(0.9 #combine(x ray images of a hip joint with prosthesis) 0.1
#combine(#uw8(x ray images) #uw8(hip joint) prosthesis) )</p>
        <p>The first #combine() in the structured query corresponds to f1(Q,D), and the second
one corresponds to f2(Q,D), while w1is 0.9 and w2 is 0.1.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Thesaurus-Assistant Query Expansion</title>
        <p>After the original query text is mapped to concepts in UMLS, the query can be
expanded with the mapped concept terms, their synonym, hierarchical or related term
information. Since the query terms of ordinary users tend to be general, we used the
preferred names of the mapped concepts and their direct children to expand query.
The adding terms are not necessarily important as the original ones, so weight can be
introduced. And the new ranking function is,
,
,
,
,
(2)
and f3(Q,D) is the ranking function for the concepts, and w3 is the corresponding
weight. The concepts can be represented in the same way as the phrases.</p>
        <p>To expand query with the mapped concept and their direct children, the preferred
names of these concepts are normalized and redundant names were erased. Taking
the topic 8 for example, the query expanded only with concepts can be as following.</p>
        <p>#weight(0.7 #combine(x-ray images of a hip joint with prosthesis) 0.1
#combine(#uw8(x-ray images) #uw8(hip joint) prosthesis) 0.2 #combine( radiography
#uw12(entire hip joint) prosthesis ) )
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments and Results</title>
      <p>The parameters wi of our approach were trained on the data of ImageCLEFmed
2010 ad-hoc retrieval subtask to maximize MAP with the constraint that sum of wi
equals to one. And the optimal parameters were used for both ad-hoc retrieval and
case-based retrieval.
3.1</p>
      <sec id="sec-3-1">
        <title>Ad-hoc Retrieval</title>
        <p>For ad-hoc retrieval, the caption of each image was used as document
representation. We tested tow phrase identification strategies. The first is represented
as “p1” which using the filtered phrase identified in the early step of MetaMap
program, and the second is represented as “p2” which using the word sequence as
phrase which corresponding to the mapped concept. We used #uwN() for representing
phrase and concept in the structured query, and chose N=k*n, where n is the number
of terms within the operator, and k is a free parameter. We tested two settings of k, 2
and 4, and denoted k=2 as sw, which standing for small window. Table 1 shows the
list of all 10 runs for ad-hoc retrieval. Table 2 shows the results of all runs. The
results indicate that the phrase-based retrieval model can improve the retrieval
performance, and query expansion can retrieve more relevant images but get lower
MAP. After all, the results were not competitive, the best of our runs only ranked
30th among all 64 automatic text runs.</p>
        <p>For case-based retrieval, we investigated two different document representations,
the first representation is named full which contains the full text, title and mesh terms
of an article, and the second one is called ac which contains the abstract, all image
captions, title and mesh terms of an article. We only experimented with p2 phrase
identification strategy and set k=4. Besides our phrase-based approach, we also tested
okapi retrieval model, and a pseudo relevance feedback method which add the
abstracts and mesh terms of the first two returned article to the original query. Table
3 shows the list of all 9 runs for case-based retrieval. Table 4 shows the results of all
runs for case-based retrieval. Our approach did not make success for case-based
retrieval, but the full document representation with indri search engine got the top
rank among all submissions in automatic text runs.
This paper describes our contribution to the ImageCLEF 2011 medical retrieval task.
We adopted a phrase-based retrieval model, and an UMLS-based query expansion.
For ad-hoc retrieval, we submitted 10 runs. The results were not competitive, but
indicate that phrase-based model and query expansion can improve the retrieval
performance. For ad-hoc retrieval, we submitted 9 runs. Our approach did not make
success for case-based retrieval, but the full document represent with indri search
engine got the top rank among all automatic text runs. We conjecture that the
unsatisfactory results of our approach at ImageCLEFmed 2011 are because the weight
parameters were trained on data of ImageCLEFmed 2010’s ad-hoc retrieval subtask,
and the two datasets are quite different.</p>
        <p>Acknowledgments. This research is partly supported by the National Science
Foundation of China under grants 60873185 and the Open Project Program of the
National Laboratory of Pattern Recognition (NLPR).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Strohman</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Metzler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turtle</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croft</surname>
          </string-name>
          , W. B.:
          <article-title>Indri: A language model-based search engine for complex queries</article-title>
          .
          <source>In: Proceedings of the International Conference on Intelligence Analysis</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tian</surname>
            <given-names>C.B.</given-names>
          </string-name>
          :
          <article-title>Thesaurus-Assistant Query Expansion for Context-based Medical Image Retrieval</article-title>
          , submitted to Pacific-Rim Conference on Multimedia (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kalpathy-Cramer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Müller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bedrick</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eggel</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , Garcia Seco de Herrera,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Tsikrika</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          :
          <article-title>The CLEF 2011 medical image retrieval and classification tasks</article-title>
          .
          <source>In: CLEF 2011 working notes</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          :
          <article-title>Effective Mapping of Biomedical text to the UMLS Metathesaurus: the MetaMap Program</article-title>
          .
          <source>In: Proceedings of the AMIA Symposium</source>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bodenreider</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The Unified Medical Language System (UMLS): integrating biomedical terminology</article-title>
          .
          <source>In: Nucleic Acids Research</source>
          <volume>32</volume>
          , pp.
          <fpage>267</fpage>
          --
          <lpage>270</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>