<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IPAL Knowledge-based Medical Image Retrieval in ImageCLEFmed 2006</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Caroline Lacoste</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jean-Pierre Chevallet</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joo-Hwee Lim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiong Wei</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Raccoceanu</string-name>
          <email>visdaniel@i2r.a-star.edu.sg</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Diem Le Thi Hoang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roxana Teodorescu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Measurement</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the contribution of IPAL group on the CLEF 2006 medical retrieval task (i.e. ImageCLEFmed). The main idea of our group is to incorporate medical knowledge in the retrieval system within a multimodal fusion framework. For text, this knowledge is in the Uni ed Medical Language System (UMLS) sources. For images, this knowledge is in semantic features that are learned from examples within structured learning framework. We propose to represent both image and text using UMLS concepts. The use of UMLS concepts allows the system to work at a higher semantic level and to standardize the semantic index of medical data, facilitating the communication between visual end textual indexing and retrieval. The results obtained with UMLS-based approaches show the potential of this conceptual indexing, especially when using a semantic dimension ltering, and the bene t of working within a fusion framework, leading to the best results of ImageCLEFmed 2006. We also test a visual retrieval system based on manual query design and visual task fusion. Even if it provides the best visual results, this purely visual retrieval provides poor results in comparison to the best textual approaches.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing|Indexing methods</kwd>
        <kwd>Thesauruses</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval|Retrieval Models</kwd>
        <kwd>Information ltering</kwd>
        <kwd>H</kwd>
        <kwd>2 [Database Management]</kwd>
        <kwd>H</kwd>
        <kwd>2</kwd>
        <kwd>4 System|Multimedia Database</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>retrieval methods could help researchers, lecturers, and student nd relevant images from large
repositories. Visual features not only allow the retrieval of cases with patients having similar
diagnoses but also cases with visual similarity but di erent diagnoses.</p>
      <p>
        Current CBIR systems [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] generally use primitive features such as color or texture [
        <xref ref-type="bibr" rid="ref17 ref18">17, 18</xref>
        ],
or logical features such as object and their relationships [
        <xref ref-type="bibr" rid="ref25 ref4">25, 4</xref>
        ] to represent images. Because they
do not use medical knowledge, such systems provide poor results in the medical domain. More
speci cally, the description of an image by low-level or medium-level features is not su cient to
capture the semantic content of a medical image. This loss of information is called the semantic
gap. In specialized systems, this semantic gap can be reduced leading to good retrieval results
[
        <xref ref-type="bibr" rid="ref11 ref20 ref6">11, 20, 6</xref>
        ]. Indeed, the more a retrieval application is specialized for a limited domain, the smaller
the gap can be narrowed by using domain knowledge.
      </p>
      <p>
        Among the limited research e orts of medical CBIR, classi cation or clustering driven feature
selection and weighting has received much attention as general visual cues often fail to be
discriminative enough to deal with more subtle, domain-speci c di erences and more objective ground
truth in the form of disease categories is usually available [
        <xref ref-type="bibr" rid="ref15 ref8">8, 15</xref>
        ]. In reality, pathology bearing
regions tend to be highly localized [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Hence, local features such as those extracted from segmented
dominant image regions approximated by best tting ellipses have been proposed [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. However,
it has been recognized that pathology bearing regions cannot be segmented out automatically for
many medical domains [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Hence it is desirable to have a medical CBIR system that represents
images in terms of semantic features, that can be learned from examples (rather than handcrafted
with a lot of expert input) and do not rely on robust region segmentation.
      </p>
      <p>
        The semantic gap can also be reduced by exploiting all sources of information. In
particular, mixing text and image information generally increases the retrieval performance [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
statistical methods are used for modeling the occurrence of document keywords and visual
characteristics. The proposed system is sensitive to the quality of the segmentation of the images. Other
initiatives to combine image and text analysis study the use of Latent Semantic Analysis (LSA)
techniques [
        <xref ref-type="bibr" rid="ref24 ref26">24, 26</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], the author applied the LSA method to features extracted from the two
media. The conclusion of this study is that combining the image and the text through the LSA
method is not always e cient. The usefulness of LSA is also not conclusive in [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. Conversely, a
simple late fusion of visual and textual indexes provides generally good results.
      </p>
      <p>
        In this paper, we present our work on medical image retrieval that is mainly based on the
incorporation of medical knowledge in the system within a fusion framework. For text, this
knowledge is in the Uni ed Medical Language System (UMLS) sources produced by NML1. For
images, this knowledge is in semantic features that are learned from examples and do not rely
on robust region segmentation. In order to manage large and complex sets of visual entities (i.e.,
high content diversity) in the medical domain, we developed a structured learning framework that
facilitates modular design and extract medical visual semantics. We developed two complementary
visual indexing approaches within this framework: a global indexing to access image modality,
and a local indexing to access semantic local features. This local indexing does not rely on region
segmentation but builds upon patch-based semantic detector [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>To bene t e ciently from both modalities, we propose to represent both image and text using
UMLS concepts in our principal retrieval system. The use of UMLS concepts allows our system to
work at a higher semantic level and to standardize the semantic index of medical data, facilitating
the communication between visual end textual indexing and retrieval. We propose several fusion
approaches and a visual modality ltering is designed to remove visually aberrant images according
to the query modality concept(s).</p>
      <p>Besides this UMLS-based system, we also investigate the potential of a closed visual retrieval
system where all queries are xed and manually designed (i.e. several examples are manually
selected to represent each query).</p>
      <p>Textual, visual, and mixed approaches derived from these two systems are evaluated on the
medical task of CLEF 2006 (i.e. imageCLEFmed).</p>
      <p>1National Library of Medicine -http://www.nlm.nih.gov/
UMLS is a good candidate as a knowledge base for medical image and text indexing. It is more
than a terminology base because terms are associated with concepts. There exists also di erent
type of links. The base is large (more than 50,000 concepts, 5.5 million of terms in 17 languages),
and is maintained by specialists with two updates a year. Unfortunately, UMLS is a merger of
di erent sources (thesaurus, terms lists), and is neither complete, nor consistent. In particular,
the links among concepts are not equally distributed. UMLS is a \meta thesaurus", i.e. a merger
of existing thesaurus. It is not an ontology, because there is no formal description of concepts, but
its large set of terms and variation restricted to medical domain only, enable us to experiment a
full scale conceptual indexing system. In UMLS, all concepts are assigned to at least one semantic
type from the Semantic Network. This provides consistent categorization of all concepts in the
meta-thesaurus at the relatively general level represented in the Semantic Network. This partially
solves the problem of merging existing thesaurus hierarchy during the merging process.</p>
      <p>
        Despite the large set of terms and terms variation available in UMLS, it still cannot cover all
possible (potentially in nite) term variation. So we need a concept identi cation tool that manages
terms variation. For English texts, we use MetaMap[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] provided by NLM. We have developed a
similar tool developed for French and German documents. This concept extraction tools do not
provide any disambiguation. We partially overcome this problem by manually ordering them by
thesaurus sources: we prefer source that strongly belong to medicine. For example, this enables
the identi cation of \x-ray" as radiography and not as the physical phenomenon (the wave) which
seldom appears in our documents. Concepts extraction is limited to noun phrase (i.e. verbs are
not treated).
      </p>
      <p>
        The extracted concepts are then organized in conceptual vectors, like a conventional vector
space IR model. We then use the same weighting scheme provided by our XIOTA indexing system
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>We tested six retrieval approaches based on this conceptual indexing and an approach -
corresponding to run IPAL Textual TDF - based on an indexing using MeSH2 terms.</p>
      <p>Each conceptual text retrieval approach uses a Vector Space Model (VSM) for representing
each document and a cosine similarity measure to compare the query index to the database
medical report. The tf idf measure is used to weight the concepts. The mapping text-concept is
separate for three languages, query vectors of concepts of three languages are fusioned, and used
for interrogation using the three indexing separately. Then, the three relevance status values are
fusioned together.</p>
      <p>One major criticism we have against VSM is the lack of structure of the query. VSM is known
2Medical Subject Headings, MeSH, is the controlled vocabulary thesaurus of the U.S. National Library of
Medicine. It is included in the meta-thesaurus UMLS.
to perform well using long textual queries but ignoring query structure. The ImageCLEFmed
2006 queries are rather short. Moreover, it seems obvious to us that it is the complete query that
should be solved and not only part of it. After query examination, we found out that queries
are implicitly structured according to some semantic types (e.g. anatomy, pathology, modality).
We call this the "semantic dimensions" of the query. Omitting correct answer to any of these
dimensions may lead to incorrect answers. Unfortunately VSM does not provide a way to ensure
answers to each dimension.</p>
      <p>To solve this problem, we decided to add a semantic dimension ltering step to the VSM, in
order to explicitly taking into account the query dimension structure. This extra ltering step
retains answers that incorporate at least one dimension. We use semantic structure on concepts
provided by UMLS. Semantic dimension of a concept is de ned by its UMLS semantic type,
grouped into semantic groups: Anatomy, Pathology and Modality. Only a conceptual indexing
and a structured meta-thesaurus like UMLS enable us to do such a semantic dimension ltering
(DF). This ltering discards noisy answers regarding to the dimension query semantic structure.
The corresponding run is \IPAL Textual CDF". We also test a similar dimension ltering based
MESH terms (run \IPAL Textual TDF"). In this case, the association between MeSH terms and
a dimension had to be done manually. According to Table 1, using UMLS concepts more than
terms improves the results of 2 Mean Average Precision (MAP) points (i.e. from 21% to 23%).</p>
      <p>Another solution to take into account query semantic structure is to re-weight answers
according to dimensions (DW). Here, Relevance Status Value output from VSM is multiplied by
the number of concepts matched with the query according to the dimensions. This simple
reweighting scheme strongly emphasizes the presence of maximum number of concepts related to
semantic dimensions. This re-weighting step implicitly do the previous dimension ltering (DF), as
the relevance value is multiplied by 0. According to our results in Table 1 this DW approach -
corresponding to the run \IPAL Textual CDW" - produces the best results of 2006 ImageCLEFmed
with 26% of MAP. This result outperforms any other classical textual indexing reported in
ImageCLEFmed 2006. Hence we have shown here the potential of conceptual indexing.</p>
      <p>In run \IPAL Textual CPRF", we tested Pseudo-Relevance Feedback (PRF). From the result
of late fusion of text-image retrieval results, three top relevant documents retrieved are taken
and all concepts of these documents are added into query for query expansion. Then, a dimension
ltering is applied. In fact, this run should have been classi ed in the mixed runs as we also use the
image information to have a better precision in the three rst images. This PRF approach improves
slightly the results obtained with a simple dimension ltering. This is principally due to the fact
errors can be present in the three rst documents, even with the best mixed retrieval result. Using
a manual Relevance Feedback (RF), we obtained a MAP of 25% that is 2 points higher than the
result obtained with a simple dimension ltering. In this last run - named \IPAL Textual CRF"
- a maximum of 4 top relevant documents were chosen by human judgment over 20 rst retrieved
image. All the concepts from these documents are added into the query for query expansion.</p>
      <p>We also tested document expansion using the UMLS semantic network. Based on UMLS
hierarchical relationships, each database concept is expanded by concepts positioned at a higher
level in the UMLS hierarchy and connected to this concept with respect to the semantic relation
\is a". The expanded concepts have a higher position than document concept in UMLS hierarchy.
For example a document indexed by the concept \molar teeth" would be also indexed by the
more general concept "teeth". This document would be thus retrieved if the user ask for a
teeth photography. This expansion does not seem relevant according to the Table 1 as the run
\IPAL Textual CDE" - that uses a document expansion technique and a dimension ltering - is 4
points below the simple dimension ltering.
3.1</p>
    </sec>
    <sec id="sec-2">
      <title>Visual Retrieval</title>
      <p>UMLS-based visual indexing and retrieval
In order to manage large and complex sets of visual entities in the medical domain, we developed a
structured learning framework to facilitate modular design and learning of medical semantics from
images. This framework allows to index images using VisMed terms, that are typical semantic
tokens characterized by a visual appearance in medical image regions. Each VisMed term is
expressed in the medical domain as a combination of UMLS concepts. In this way, we have a
common language to index both image and text, which facilitates the communication between
visual and textual indexing and retrieval. We developed two complementary indexing approaches
within this statistical learning framework:
a global indexing to access image modality (chest X-ray, gross photography of an organ,
microscopy, etc.);
a local indexing to access semantic local features that are related to modality, anatomy, and
pathology concepts.</p>
      <p>After a presentation of both approaches in Sections 3.1.1 and 3.1.2, retrieval procedures and
experimental results are given in Section 3.1.3.
3.1.1</p>
      <p>Global UMLS Indexing
The global UMLS indexing is based on a two level hierarchical classi er according to mainly
modality concepts. This modality classi er is learned from about 4000 images separated in 32
classes: 22 grey level modalities, and 10 color modalities. Each indexing term is characterized by
a UMLS modality concept (e.g. chest X-ray, gross photography of an organ) and, sometimes, a
spatial concept (e.g. axial, frontal, etc), or a color percept (color, grey). The training images come
from the CLEF database (about 2500 examples), from the IRMA3 database (about 300 examples),
and from the web (about 1200 examples). The training images from ImageCLEFmed database
was obtained from modality concept extraction using medical reports. A manual ltering step on
this extraction process to remove irrelevant examples had to be performed. We plan to automate
this ltering in the near future.</p>
      <p>
        The rst level of the classi er corresponds to a classi cation for grey level versus color images.
Indeed, some ambiguity can appear due to the presence of colored images, or the slightly blue or
green appearance of X-ray images. This rst classi er uses the rst three moments in the HSV
color space computed on the entire image. The second level corresponds to the classi cation of
modality UMLS concepts given that the image is in the grey or the color cluster. For the grey
level cluster, we use grey level histogram (32 bins), texture features (mean and variance of Gabor
coe cients for 5 scales and 6 orientations), and thumbnails (grey values of 16x16 resized image).
For the color cluster, we have adopted HSV histogram (125 bins), Gabor texture features, and
thumbnails. Zero-mean normalization [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] was applied to each feature . For each SVM classi er,
we adopted a RBF kernel:
where
= 212 and with a modi ed city-block distance:
exp(
jx
      </p>
      <p>yj2)
jx
yj =
1 XF jxf
F
f=1</p>
      <p>Nf
yf j
(1)
(2)
where x = fx1; :::; xF g and y = fy1; :::; yF g are feature vectors, xf ; yf are feature vectors of type
f, Nf is the feature vector dimension, and F is the number of feature types: F = 1 for the grey
versus color classi er, F = 3 for the conditional modality classi ers: color, texture, thumbnails.</p>
      <p>3http://phobos.imib.rwth-aachen.de/irma/index_en.php
where C and G denote the color and the grey level clusters respectively, and the conditional
probability P (MODijz; V ) is given by:
(3)
(4)
(5)
P (cjz; V ) =</p>
      <p>expDc(z)</p>
      <p>Pj2V expDj(z)
where Dc is the signed distance to the SVM hyperplane that separate class c from the other classes
of the cluster V .</p>
      <p>
        After learning - using SVM-Light software4 [
        <xref ref-type="bibr" rid="ref10 ref22">10, 22</xref>
        ] -, each database image z is indexed
according to modality given its low-level features zf . The indexes are the probability values given
by Equation (3).
3.1.2
      </p>
      <p>
        Local UMLS Indexing
To better capture the medical image content, we propose to extend the global modeling and
classi cation with local patch classi cation of local visual and semantic tokens (LVM terms).
Each LVM indexing term is expressed as a combination of Uni ed Medical Language System
(UMLS) concepts from Modality, Anatomy, and Pathology semantic types. A Semantic Patch
Classi er was designed to classify a patch according to the 64 LVM terms. In these experiments,
we have adopted color and texture features from patches (i. e. small image blocks) and a classi er
based on SVMs and the softmax function [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] given by Equation (4). The color features are the
three rst moments of the Hue, the Saturation, and the Value of the patch. The texture features
are the mean and variance of Gabor coe cients using 5 scales and 6 orientations. Zero-mean
normalization [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is applied to both the color and texture features. We adopted a RBF kernel with
modi ed city-block distance given by Equation (2). The training dataset is composed of 3631
patches extracted from 1033 images mostly coming from the web (921 images coming from the
web and 112 images from the ImageCLEFmed collection 0:2%).
      </p>
      <p>
        After learning, the LVM indexing terms are detected during image indexing from image patches
without region segmentation to form semantic local histograms. Essentially, an image is tessellated
into overlapping image blocks of size 40x40 pixels after size standardization. Each patch is then
classi ed into one of the 64 LVM terms using the Semantic Patch Classi er. An image containing
P overlapping patches is then characterized by the set of P LVM histograms and their respective
location in the image. An histogram aggregation per block gives the nal image index : M N
LVM histograms. Each bin of the histogram of a given block B corresponds to the probability of
a LVM term presence in this block. This probability is computed as follows:
We use = 1 in all our experiments. This just-in-time feature fusion within the kernel combines
the contribution of color, texture, and spatial features equally [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>The probability of a modality MODi for an image z is given by:</p>
      <p>P (MODijz) =</p>
      <p>P (MODijz; C)P (Cjz) if MODi 2 C</p>
      <p>P (MODijz; G)P (Gjz) if MODi 2 G
P (VMTijB) =</p>
      <p>Pz jz \ Bj P (VMTijz)</p>
      <p>Pz jz \ Bj
where B is a block of a given image, z denotes a path of the same image, jz \ Bj is the area of
the intersection between z and B, and P (VMTijz) is given by Equation (4). To facilitate spatial
aggregation and matching of image with di erent aspect ratios , we design 5 tiling templates,
namely M N = 3 1, 3 2, 3 3, 2 3, and 1 3 grids resulting in 3, 6, 9, 6, and 3 probability
vectors per image respectively.
3.1.3</p>
      <p>Visual retrieval using UMLS-based visual indexing
We propose three retrieval methods from query by example(s) based on the two UMLS-based
visual indexing. When several images are given in the query, the similarity between a database
image z with the query is given by the maximum value among the similarities between z and each
query image.</p>
      <p>The rst method - corresponding to run \IPAL Visual MC" - is based on the global indexing
scheme according the modality. An image is represented by a semantic histogram, each bin
corresponding to a modality probability. The distance between two images is given by the Manhattan
distance (i.e. city-block distance) between the two semantic histograms.</p>
      <p>
        The second method - corresponding to run \IPAL Visual SPC" - is based on the local UMLS
visual indexing. An image is then represented by M N semantic histograms. Given two images
represented as di erent grid patterns, we propose a exible tiling (FlexiTile) matching scheme to
cover all possible matches [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The distance between a query image and a database image is then
the mean of block by block distances on all the possible matches. The distance between two blocks
is given by the Manhattan distance between the two LocVisMed histograms.
      </p>
      <p>The last visual retrieval method - corresponding to run \IPAL Visual SPC+MC" - is the
fusion of the two rst approaches. This approach combines thus two complementary sources of
information, the rst concerning the general aspect of the image (global indexing according to
modality), the second concerning semantic local features with spatial information (local UMLS
indexing). The similarity to a query is given by the mean of the similarity to a query according
each index.</p>
      <p>The 2006 CLEF medical task was particularly di cult this year for purely visual approaches.
Indeed, the queries were at a hight semantic level for a general retrieval system. As a proof, the
best automatic visual result was less than 8% of MAP. Mixing the local and global indexing, gives
us the third place with 6% of MAP as showed in Table 25. We believe than we can improve these
results using also the textual query in the retrieval process. Indeed, besides the usual
similaritybased queries, our semantic indexing allow semantic-based query. Tests are in course on 2005 and
2006 medical tasks, providing promishing results.</p>
      <p>
        Manual Query Construction and Visual Task Fusion
To see how far we can go with a purely visual approach, we propose here a closed visual system
based on manual query construction and visual task fusion. This work is similar to what we
did in ImageCLEFmed 2005[
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. We fused retrieval results generated by systems using
multiplefeature representations and multiple retrieval systems. More speci cally, we used three types of
feature representations, i.e., \blob", \icon" and \blob+icon" and two retrieval systems, \SVM"
and \Dist". For each topic, we manually chose about 50 similar images which is used to form a
training set for SVM and to construct the query. All images are then represented by these features
respectively and passing either retrieval system.
      </p>
      <p>5The Mean Average Precision and the Recall-precision computed in imageCLEFmed for the run
IPAL Visual SPC+MC was 6:34% and 10:48% respectively because we only submitted - by error - the 25 rst
queries.</p>
      <p>In this year's attempts, we have submitted 10 runs based on the fusion of six sub-runs. The
generated sub-runs are denoted D1,D2,D3,D4, D5 and D6. D1,D2 and D3 use \Dist" retrieval
system but di erent features (D1 using \icon", D2 using \blob", D3 using \blob+icon"), D4 and
D5 use \SVM" with di erent features (D4 using \blob", D2 using \icon"), and D6 uses the
UMLSbased system presented in Section 3.1 (D6 corresponds to the run \IPAL Visual MC" that is based
on global UMLS indexing).</p>
      <p>Di erent from the work done in 2005, we also use the probability estimation of each image
about its modality that is given by Equation (3). Some of these sub-runs (D1-D6) are linearly
combined together to produce a score for each image. Each score may then be multiplied by
the probability and all the results are sorted to yield the nal retrieval ranking lists. The runs
\IPAL CMP D1D2D4D5D6", \IPAL CMP D1D2D3D4D5", \IPAL CMP D1D2D3D4D5D6"6, and
\IPAL CMP D1D2D4D5" used the probability estimations. We also applied a color lter to
remove those images whose number of color channels are less than that of the query images, except
for runs \IPAL D1D2D4D5D6" and \IPAL D1D2D4D5". The performance of these runs is given
in Table 3.</p>
      <p>Applying a probability estimation of each image about its modality and imaging anatomy
helps (compare \IPAL CMP D1D2D4D5D6" (MAP=0.1596) and \IPAL cfD1D2D4D5D6"
(MAP=0.155));
Color ltering generally improves performance, but not signi cantly;
Using D6 in the combination, the performance is improved. Compare \IPAL cfD1D2D4D5D6"
(MAP=0.155) versus \IPAL cfD1D2D4D5" (MAP = 0.1461) and \IPAL D1D2D4D5D6"
(MAP=0.1551) versus \IPAL D1D2D4D5" (MAP=0.1461).
4</p>
    </sec>
    <sec id="sec-3">
      <title>UMLS-based Mixed Retrieval</title>
      <p>We propose three types of fusion between text and images:
a simple late fusion (run \IPAL Cpt Im");
a fusion that uses a visual ltering according to modality concept(s) (runs \IPAL ModFDT Cpt Im",
\IPAL ModFST Cpt Im", \IPAL ModFDT TDF Im", and \IPAL ModFDT Cpt");
an early fusion of UMLS-based visual and textual indexes (run \IPAL MediSmart 1" and
\IPAL MediSmart 2").</p>
      <p>The rst fusion method is a late fusion of visual and textual similarity measures. The similarity
between a mixed query Q = (QI ; QT ) (QI : image(s), QT : text) and a couple composed of an image
and the associated medical report (I; R) is then given by:
(Q; I; R) =</p>
      <p>V (QI ; I)
max V (QI ; z)
z2DI
+ (1
)</p>
      <p>T (QT ; R)
max T (QT ; z)
z2DT
where V (QI ; I) denotes the visual similarity between the visual query QI and an image I,</p>
      <p>T (QT ; R) denotes the textual similarity between the textual query QT and the medical report
R, DI denotes the image database, and DT denotes the text database. The factor allows the
control of the weight of the textual similarity with respect to the image similarity. After some
experimentations on imageCLEFmed 2005, we choose = 0:7. In order to compare similarities
in the same range, each similarity is divided by the corresponding maximal similarity value on
the entire database. The result of the corresponding run, \\IPAL Cpt Im", given in Table 4 show
the good complementarity of the visual and textual indexing: from 26% for the textual retrieval
and 6% for the visual retrieval, the mixed retrieval provides 31% of MAP. The best results on
imageCLEFmed 2006 in terms of MAP and R-precision (i.e. precision after R retrieved images,
where R is the number of relevant images) were obtained with this simple late fusion.</p>
      <p>The second type of fusion exploits directly the UMLS index of images. Indeed, it is based on a
direct matching between concepts extracted from the textual query and conceptual image indexes.
This direct matching is done automatically with the use of the Uni ed Medical Language System.
More speci cally, a comparison between the query concepts related to modality and the image
modality index is done in order to remove all aberrant images. The decision rule is the following:
an image I is admissible for a query modality MODQ only if:</p>
      <p>P (MODQjI) &gt; (MODQ)
where (MODQ) is a threshold de ned for the modality MODQ. This decision rule de nes a set
of admissible images for a given modality MODQ: fI 2 DI : P (MODQjI) &gt; (MODQ)g. The nal
result is then the intersection of this set and the ordered set of images retrieved by any system.
This modality lter is particularly interesting for ltering textual retrieval results. Indeed, several
images of di erent modalities can be associated to the same medical report. The ambiguity is
thus removed when using a visual modality ltering. We test this approach with, rst, a xed
threshold for all modality (MODQ) = 0:15 (\ModFST") based on experimental tests on 2005
CLEF medical task, and, second, an adaptive threshold for each modality according a con dence
degree given to the classi er according this modality (\ModFDT"). The adaptive thresholding
performs slightly better than the constant thresholding (compare \IPAL ModFDT Cpt Im" and
\IPAL ModFST Cpt Im" in Table 4). In fact, we have over-estimated these thresholds for most
modalities. Indeed, when this modality ltering is applied to the late fusion results the results
6\IPAL CMP D1D2D3D4D5D6" corresponds to \IPAL CMP D1D2D3D4D5D" in ImageCLEFmed where the
last letter was missing
(6)
(7)
n
o
i
isc 0.4
e
r
P
0.7
0.6
0.5
0.3
0.2
0.1
IPAL_ModFDT_Cpt_Im</p>
      <p>IPAL_ModFDT_Cpt</p>
      <p>IPAL_Cpt_Im</p>
      <p>IPAL_Textual_CDW
IPAL_Visual_SPC+MC
decrease of 2 points (compare \IPAL ModFDT Cpt Im" and \IPAL Cpt Im"). That means that
this ltering not only removes aberrant images but also relevant images. This ltering
nevertheless increases the results of the purely textual retrieval approach from 26% to 27% (see
\IPAL ModFDT Cpt" in Table 4 and \IPAL Textual CDW" in Table 1). Moreover, this ltering
is relevant if the user - which is often the case - is more interested in the precision in the rst
retrieved images that in the mean average precision. Indeed, the Figure 1 shows that the adaptive
modality ltering on the late fusion results (\IPAL ModFDT Cpt Im") and even directly on the
textual results (\IPAL ModFDT Cpt") provides a better precision than the late fusion results
(\IPAL Cpt Im") for the rst retrieved documents (until 30 when applied on textual results, until
50 when applied on the mixed results).</p>
      <p>20
40
60
80 100 120
Number of documents
140
160
180
200</p>
      <p>
        We also submitted two runs concerning the early fusion of UMLS-based visual and textual
indexes. A Semantic level fuzzy cation algorithm takes into account the frequency, the localization,
the con dence and the source of the information [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Unfortunately, errors were found after
the submission. Using the corrected algorithm, we obtain 24% of mean average precision on
ImageCLEFmed 2006. We have to note that the dimension ltering and re-weighting used in the
other IPAL mixed runs are not applied here, which explains in part the di erence of precision. In
fact, this result is higher than the results obtained with runs that do not use this dimension ltering
(20% for mixed retrieval, 23% for textual retrieval). We currently develop clustering techniques
to improve the retrieval results. A fuzzy min-max boosted K-means clustering approach gives
promishing results on CASImage database.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>In this paper, we have proposed a medical image retrieval system that represents both texts and
images at a very high semantic level using concepts from the Uni ed Medical Language System.
Textual, visual, and mixed approaches derived from this system were evaluated on ImageCLEFmed
2006. A structured framework was proposed to bridge the semantic gap between low-level image
features and the semantic UMLS concepts. A closed visual system based on manual query
construction and visual task fusion was also tested to go as far as possible using a purely visual
approach. From the results on ImageCLEF 2006, we can conclude that the textual approaches
capture more easily the semantics of the medical queries, providing better results than purely
visual retrieval approaches. Indeed, the best visual approach in 2006 - that corresponds to a result
of our closed system - only provides 16% of MAP against 26% of MAP for the best textual results.
Moreover, the results show the potential of conceptual indexing, especially when using a semantic
dimension ltering: we obtained the best textual and mixed results in imageCLEF 2006 using
our UMLS-based system. The bene t of working in a fusion framework has been demonstrated.
Firstly, visual retrieval results are enhanced by the fusion of global and local similarities. Secondly,
mixing textual and visual information improves signi cantly the system performance. Besides
precision in the rst documents increases when using a visual modality ltering, allowing 68% of mean
precision on the 10 rst documents and 62% of mean precision for the 30 rst documents on the
30 queries of ImageCLEF 2006. We are currently investigating the potential of an early fusion
scheme using appropriate clustering methods. In the near future, we plan to use the LVM terms
from local indexing for semantics-based retrieval (i.e. cross-modal retrieval: processing textual
query on LVM-based image indexes). A visual ltering based on local information could also be
derived from the semantic local indexing.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Aronson</surname>
          </string-name>
          .
          <article-title>E ective mapping of biomedical text to the UMLS metathesaurus: The MetaMap program</article-title>
          .
          <source>In Proceedings of the Annual Symposium of the American Society for Medical Informatics</source>
          , pages
          <volume>17</volume>
          {
          <fpage>21</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Barnard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Duygulu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Forsyth</surname>
          </string-name>
          , N. de Freitas,
          <string-name>
            <given-names>D.M.</given-names>
            <surname>Blei</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.I</given-names>
            <surname>Jordan</surname>
          </string-name>
          .
          <article-title>Matching words and pictures</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>3</volume>
          :
          <fpage>1107</fpage>
          {
          <fpage>1135</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.M.</given-names>
            <surname>Bishop</surname>
          </string-name>
          .
          <article-title>Neural Networks for Pattern Recognition</article-title>
          . Clarendon Press, Oxford,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Carson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Belongie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Greenspan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Malik</surname>
          </string-name>
          . Blobworld:
          <article-title>Image segmentation using expectation-maximisation and its applications to image querying</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          ,
          <volume>24</volume>
          (
          <issue>8</issue>
          ):
          <volume>1026</volume>
          {
          <fpage>1038</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Jean-Pierre Chevallet</surname>
          </string-name>
          .
          <article-title>X-IOTA: An open XML framework for IR experimentation application on multiple weighting scheme tests in a bilingual corpus</article-title>
          .
          <source>Lecture Notes in Computer Science (LNCS)</source>
          ,
          <source>AIRS'04 Conference</source>
          , Beijing,
          <volume>3211</volume>
          :
          <fpage>263</fpage>
          {
          <fpage>280</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>W.W.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. C.</given-names>
            <surname>Alfonso</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.T.</given-names>
            <surname>Ricky</surname>
          </string-name>
          .
          <article-title>Knowledge-based image retrieval with spatial and temporal constructs</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <volume>10</volume>
          :
          <fpage>872</fpage>
          {
          <fpage>888</fpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Paul</given-names>
            <surname>Clough</surname>
          </string-name>
          , Henning Muller, Thomas Desealers, Michael Grubinger, Thomas Lehmann, Je ery Jensen, and
          <string-name>
            <given-names>William</given-names>
            <surname>Hersh</surname>
          </string-name>
          .
          <article-title>The CLEF 2005 automatic medical image annotation task</article-title>
          .
          <source>Springer Lecture Notes in Computer Science</source>
          . To appear.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.G.</given-names>
            <surname>Dy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.E.</given-names>
            <surname>Brodley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.C.</given-names>
            <surname>Kak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.S.</given-names>
            <surname>Broderick</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.M.</given-names>
            <surname>Aisen</surname>
          </string-name>
          .
          <article-title>Unsupervised feature selection applied to content-based retrieval of lung images</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          ,
          <volume>25</volume>
          (
          <issue>3</issue>
          ):
          <volume>373</volume>
          {
          <fpage>378</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Rui</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Mehrotra</surname>
          </string-name>
          .
          <article-title>Content-based image retrieval with relevance feedback in mars</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Image Processing</source>
          , pages
          <volume>815</volume>
          {
          <fpage>818</fpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Joachims</surname>
          </string-name>
          .
          <article-title>Learning to Classify Text using Support Vector Machines</article-title>
          . Kluwer,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Korn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sidiropoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Faloutsos</surname>
          </string-name>
          , E. Siegel, and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Protopapas</surname>
          </string-name>
          .
          <article-title>Fast and e ective retrieval of medical tumor shapes</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <volume>10</volume>
          :
          <fpage>889</fpage>
          {
          <fpage>904</fpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>T.M. Lehmann</surname>
          </string-name>
          et al.
          <article-title>Content-based image retrieval in medical applications</article-title>
          .
          <source>Methods Inf Med</source>
          ,
          <volume>43</volume>
          :
          <fpage>354</fpage>
          {
          <fpage>361</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.H.</given-names>
            <surname>Lim</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.-P</given-names>
            <surname>Chevallet. VisMed:</surname>
          </string-name>
          <article-title>a visual vocabulary approach for medical image indexing and retrieval</article-title>
          .
          <source>In Proceedings of the Asia Information Retrieval Symposium</source>
          , pages
          <volume>84</volume>
          {
          <fpage>96</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.H.</given-names>
            <surname>Lim</surname>
          </string-name>
          and
          <string-name>
            <surname>J.S. Jin.</surname>
          </string-name>
          <article-title>Discovering recurrent image semantics from class discrimination</article-title>
          .
          <source>EURASIP Journal of Applied Signal Processing</source>
          ,
          <volume>21</volume>
          :1{
          <fpage>11</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          et al.
          <article-title>Semantic based biomedical image indexing and retrieval</article-title>
          . In L. Shapiro,
          <string-name>
            <given-names>H.P.</given-names>
            <surname>Kriegel</surname>
          </string-name>
          , and R. Veltkamp, editors,
          <source>Trends and Advances in Content-Based Image and Video Retrieval</source>
          . Springer,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>H.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Michoux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bandon</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Geissbuhler</surname>
          </string-name>
          .
          <article-title>A review of content-based image retrieval systems in medical applications - clinical bene ts and future directions</article-title>
          .
          <source>International Journal of Medical Informatics</source>
          ,
          <volume>73</volume>
          :1{
          <fpage>23</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>W.</given-names>
            <surname>Niblack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Barber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Equitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Flickner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Glasman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Petkovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Yanker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Faloutsos</surname>
          </string-name>
          , and
          <string-name>
            <surname>G. Taubin.</surname>
          </string-name>
          <article-title>QBICproject: querying images by content, using color, texture, and shape</article-title>
          . In W. Niblack, editor,
          <source>Storage and Retrieval for Image and Video Databases</source>
          , volume
          <volume>1908</volume>
          , pages
          <fpage>173</fpage>
          {
          <fpage>187</fpage>
          .
          <string-name>
            <surname>SPIE</surname>
          </string-name>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Pentland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. W.</given-names>
            <surname>Picard</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Sclaro</surname>
          </string-name>
          . Photobook:
          <article-title>Tools for content-based manipulation of image databases</article-title>
          .
          <source>International Journal of Computer Vision</source>
          ,
          <volume>18</volume>
          :
          <fpage>233</fpage>
          {
          <fpage>254</fpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>D.</given-names>
            <surname>Racoceanu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lacoste</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Teodorescu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Vuillemenot</surname>
          </string-name>
          .
          <article-title>A semantic fusion approach between medical images and reports using UMLS</article-title>
          .
          <source>In Proceedings of the Asia Information Retrieval Symposium</source>
          (Special Session),
          <source>Singapore</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Chi-Ren</surname>
            <given-names>Shyu</given-names>
          </string-name>
          , Christina Pavlopoulou, Avinash C. Kak,
          <string-name>
            <given-names>Carla E.</given-names>
            <surname>Brodley</surname>
          </string-name>
          , and
          <string-name>
            <surname>Lynn</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Broderick</surname>
          </string-name>
          .
          <article-title>Using human perceptual categories for content-based retrieval from a medical image database</article-title>
          .
          <source>Computer Vision</source>
          and Image Understanding,
          <volume>88</volume>
          (
          <issue>3</issue>
          ):
          <volume>119</volume>
          {
          <fpage>151</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>A. W. M. Smeulders</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Worring</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Santini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Gupta</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Jain</surname>
          </string-name>
          .
          <article-title>Content-based image retrieval at the end of the early years</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          ,
          <volume>22</volume>
          :
          <fpage>1349</fpage>
          {
          <fpage>1380</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Vladimir</surname>
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Vapnik</surname>
          </string-name>
          .
          <source>The Nature of Statistical Learning Theory. Springer</source>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Xiong</surname>
            <given-names>Wei</given-names>
          </string-name>
          , Qiu Bo, Tian Qi, Xu Changsheng, Ong Sim-Heng, and
          <string-name>
            <given-names>Foong</given-names>
            <surname>Kelvin</surname>
          </string-name>
          .
          <article-title>Combining multilevel visual features for medical image retrieval in ImageCLEF 2005</article-title>
          . In Cross Language Evaluation Forum 2005 workshop, page 73, Vienna, Austria,
          <year>September 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>T.</given-names>
            <surname>Westerveld</surname>
          </string-name>
          .
          <article-title>Image retrieval : Content versus context</article-title>
          .
          <source>In Recherche d'Information Assistee par Ordinateur</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>J. K. Wu</surname>
            ,
            <given-names>A. Desai</given-names>
          </string-name>
          <string-name>
            <surname>Narasimhalu</surname>
            ,
            <given-names>B.M.</given-names>
          </string-name>
          <string-name>
            <surname>Mehtre</surname>
            ,
            <given-names>C.P.</given-names>
          </string-name>
          <string-name>
            <surname>Lam</surname>
            , and
            <given-names>Y.J.</given-names>
          </string-name>
          <string-name>
            <surname>Gao</surname>
          </string-name>
          .
          <article-title>CORE: a contentbased retrieval engine for multimedia information systems</article-title>
          .
          <source>Multimedia Systems</source>
          ,
          <volume>3</volume>
          :
          <fpage>25</fpage>
          {
          <fpage>41</fpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhao</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Grosky</surname>
          </string-name>
          .
          <article-title>Narrowing the semantic gap - improved text-based web document retrieval using visual features</article-title>
          .
          <source>IEEE Transactions on Multimedia</source>
          ,
          <volume>4</volume>
          (
          <issue>2</issue>
          ):
          <volume>189</volume>
          {
          <fpage>200</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>