<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NCTU-ISU's Evaluation for the User-Centered Search Task at ImageCLEF 2004</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pei-Cheng Cheng</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jen-Yuan Yeh</string-name>
          <email>jyyeh@cis.nctu.edu.tw</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hao-Ren Ke</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Been-Chian Chien</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wei-Pang Yang</string-name>
          <email>wpyang@cis.nctu.edu.tw</email>
          <email>wpyang@mail.ndhu.edu.tw</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer &amp; Information Science, National Chiao Tung University</institution>
          ,
          <addr-line>1001 Ta Hsueh Rd., Hsinchu, TAIWAN 30050, R.O.C</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science and Information Engineering, National University of Tainan 33</institution>
          ,
          <addr-line>Sec. 2, Su Line St., Tainan, TAIWAN 70005, R.O.C</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Information Engineering, I-Shou University 1</institution>
          ,
          <addr-line>Sec. 1, Hsueh Cheng Rd., Ta-Hsu Hsiang, Kaohsiung, TAIWAN 84001, R.O.C</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Department of Information Management, National Dong Hwa University 1</institution>
          ,
          <addr-line>Sec. 2, Da Hsueh Rd., Shou-Feng, Hualien, TAIWAN 97401, R.O.C</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University Library, National Chiao Tung University</institution>
          ,
          <addr-line>1001 Ta Hsueh Rd., Hsinchu, TAIWAN 30050, R.O.C</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>We participated in the user-centered search task at ImageCLEF 2004. In this paper, we proposed two interactive Cross-Language image retrieval systems - T_ICLEF and VCT_ICLEF. The first one is implemented with a practical relevance feedback approach based on textual information while the second one combines textual and image information to help users find a target image. The experimental results show that VCT_ICLEF has a better performance than T_ICLEF in almost all cases. Overall, VCT_ICLEF helps users find the image within a fewer iterations with a maximum of 2 iterations saved.</p>
      </abstract>
      <kwd-group>
        <kwd>Interactive search</kwd>
        <kwd>Cross-Language image retrieval</kwd>
        <kwd>Relevance feedback</kwd>
        <kwd>User behavior</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The ImageCLEF campaign under the CLEF1 (Cross-Language Evaluation Forum) is conducting a series of
evaluations on systems which are built to accept a query in a language and to find relevant images with captions
in different languages. In this year, three tasks2 are proposed based on different domains, scenarios, and
collections. They are 1) the bilingual ad hoc retrieval task, 2) the medical retrieval task, and 3) the user-centered
search task. The first is to perform bilingual retrieval against a photographic collection in which images are
accompanied with captions in English. The second is, given an example image, to find out similar images from a
medical image database which consists of images such as scans and x-rays. The last one aims to assess user
interaction for a known-item or target search.</p>
      <p>This paper concentrates on the user-centered search task. The task follows the scenario that a user is
searching with a specific image in mind, but without any key information about it. Previous work (e.g,
[Kushki04]) has shown that interactive search helps improve recall and precision in the retrieval task. However, in
this paper, the goal is to determine whether the retrieval system is being used in the manner intended by the
designers as well as to determine how the interface helps users reformulate and refined their search topics. We
proposed two systems: 1) T_ICLEF, and 2) VCT_ICLEF to address the task. T_ICLEF is a Cross-Language image
retrieval system, which is enhanced with the relevance feedback mechanism; VCT_ICLEF is practically
T_ICLEF but provides a color table which allows users to indicate color information about the target image.</p>
      <p>In the following sections, the design of our systems is first described. We then introduce the proposed
methods for the interactive search task, and present our results. Finally, we finish with a conclusion and a
discus1 The official website is available at http://clef.iei.pi.cnr.it:2002/.
2 Please refer to http://ir.shef.ac.uk/imageclef2004/ for further information.
sion of future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Overview of the Interactive Search Process</title>
      <p>The overview of the interactive search process is shown in Fig. 1. Given an initial query, Q = (QT , QI ) , in
which QT denotes a Chinese text query, and QI stands for a query image, the system performs the
Cross-Language image retrieval, and returns a set of “relevant” images to the user. The user then evaluates the
relevance of the returned images, and gives a relevance value to each of them. The process is called relevance
feedback. In our system, the user is only able to indicate as an image “non-relevant,” “neutral,” and “relevant.”
At the following stage, the system invokes the query reformulation process to derive a new query,
Q′ = (QT′ , QI′ ) . The new query is believed to be more corresponding to the user’s need. Finally, the system
performs again image retrieval according to Q′ . The process iterates until the user finds the target image.</p>
      <p>Q = (QT ,QI )</p>
      <p>Query</p>
      <p>Q′ = (QT′ , QI′ )</p>
      <sec id="sec-2-1">
        <title>Retrieved</title>
        <p>Images</p>
      </sec>
      <sec id="sec-2-2">
        <title>Relevant</title>
        <p>Images</p>
      </sec>
      <sec id="sec-2-3">
        <title>Reformulated Query</title>
      </sec>
      <sec id="sec-2-4">
        <title>Image</title>
        <p>Retrieval</p>
      </sec>
      <sec id="sec-2-5">
        <title>Relevance</title>
        <p>Feedback</p>
      </sec>
      <sec id="sec-2-6">
        <title>Query Reformulation Fig. 1. The overview of the user-centered search process.</title>
        <p>In this section, we describe how to create the representation for an image or a query, and how to compute the
similarity between an image and the query on the basis of their representations.
3.1</p>
        <p>Image/ Query Representations
We create both an image and a query representation by representing them as a vector in the vector space model
[Salton83]. First of all, we explain the symbols used in the following definitions of representations.
P = (PT , PI ) denotes an image where PT and PI stand for the captions of P and the image P respectively, and
Q = (QT , QI ) represents a query, which is defined as mentioned before. In our proposed approach, a textual
vector representation, such as PT and QT, is modeled in terms of three distinct features – term, category, and
temporal information while an image vector representation, for example, PI and QI, is represented with color
histogram.</p>
        <p>Textual Vector Representation
Let W (|W| = n) the set of significant keywords in the corpus, C (|C| = m) the set of categories which are defined
in the corpus, and Y (|Y| = k) the set of publication years of all images, for an image P, it textual vector
representation (i.e., PT) is defined as Eq. (1),</p>
        <p>PT =&lt; wt1 (PT ),..., wtn (PT ), wc1 (PT ),..., wcm (PT ), wy1 (PT ),..., wyk (PT ) &gt;
(1)
where the first n dimensions the weighting of a keyword ti in PT, which is measured by TF-IDF [Salton83], as
computed in Eq. (2); the following n+1 to n+m dimensions indicate whether P belongs to a category ci, which is
shown as Eq. (3); the final n+m+1 to n+m+k dimensions present whether P was published in yi, which is defined
as Eq. (4).</p>
        <p>wti (PT ) = tfti ,PT × log N
max tf nti
⎧1　if P belongs to ci ,
wci (PT ) = ⎨
⎩0　otherwise
⎧1　if P was published in yi ,
wyi (PT ) = ⎨
⎩0　otherwise
tfti ,PT
max tf
In Eq. (2),</p>
        <p>stands for the normalized frequency of ti in PT, maxtf is the maximum number of
occurrences of any keyword in PT, N indicates the number of images in the corpus, and nti denotes the number of
images in whose caption ti appears. Regarding Eq. (3) and Eq. (4), both of them compute the weighting of the
category and the temporal feature as a Boolean value. In other words, considering a category ci and a year yi,
wci (PT ) and wyi (PT ) are set to 1 if and only if P belongs to ci and P was published in yi respectively.
(2)
(3)
(4)
(5)
(6)
where wti (QT ) is the weighting of a keyword ti in QT, which is measured as Eq. (7), wci (QT ) indicates
whether there exists an e j ∈ AfterDisambiguity(QT ) and it also occurs in a category ci, which is shown as Eq.
(8), and wyi (QT ) presents whether there is an e j ∈ AfterDisambiguity(QT ) , ej is a temporal term, and ej
satisfies a condition caused by a predefined temporal operator.
3 An English/Chinese dictionary written by D. Gau, which is available at http://sourceforge.net/projects/pydict/.</p>
        <p>In the above, we introduce how to create a textual vector representation for PT. As for a query Q, one
problem is that since QT is given in Chinese, it is necessary to translate QT into English, which is the language used in
the image collection. We first perform the word segmentation process to obtain a set of Chinese words. For each
Chinese word, it is then translated into one or several corresponding English words by looking up it in a
dictionary. The dictionary that we use is pyDict3. Up to now, it is hard to determine the translation as correctly as
possible. We tend to keep all English translations in order not to lose the consideration of any correct word.</p>
        <p>Another problem is the so-called short query problem. A short query usually can not cover as many useful
search terms as possible because of the lack of sufficient words. We address this problem by performing the
query expansion process to add new terms to the original query. The additional search terms is taken from a
thesaurus – WordNet [Miller95]. For each English translation, we include its synonyms, hypernyms, and hyponyms
into the query.</p>
        <p>Now, it comes out a new problem. Assume AfterExpansion(QT ) = {e1,..., eh} the set of all English words
obtained after query translation and query expansion, it is obvious that AfterExpansion(QT ) may contain a lot
of words which are not correct translations or useful search terms. To resolve the translation ambiguity problem,
we exploit word co-occurrence relationships to determine final query terms. If the co-occurrence frequency of ei
and ej in the corpus is greater than a predefined threshold, both ei and ej are regarded as useful search terms for
monolingual image retrieval. So far, we have a set of search terms, AfterDisambiguity(QT ) , which is presented
as Eq. (5),</p>
        <p>AfterDisambiguity(QT ) = {ei , e j | ei , e j ∈ AfterExpansion(QT )</p>
        <p>&amp; ei , e j have a significant cooccurrence}</p>
        <p>After giving the definition of AfterDisambiguity(QT ) , for a query Q, its textual vector representation (i.e.,
QT) is defined in Eq. (6),</p>
        <p>QT =&lt; wt1 (QT ),..., wtn (QT ), wc1 (QT ),..., wcm (QT ), wy1 (QT ),..., wyk (QT ) &gt;
In Eq. (7), W is the set of significant keywords as defined before,
stands for the normalized
frequency of ti in AfterDisambiguity(QT ) , maxtf is the maximum number of occurrences of any keyword in
AfterDisambiguity(QT ) , N indicates the number of images in the corpus, and nti denotes the number of
images in whose caption ti appears. Similar to Eq. (3) and Eq. (4), both Eq. (8) and Eq. (9) compute the weighting
of the category and the temporal feature as a Boolean value.
tfti ,QT
max tf
wti (QT ) = ⎨⎪⎧ tfti ,QT × log N 　
⎪⎩ max tf nti
⎪⎩0　otherwise
⎧1　if ∃j, e j ∈ AfterDisambiguity(QT ) and e j occurs in ci ,
wci (QT ) = ⎨
⎩0　otherwise
⎧1　if QT contains" Y年以前," and yi is before Y,
⎪
⎪1　if QT contains " Y年之中," and yi is before Y,
wyi (QT ) = ⎨</p>
        <p>⎪1　if QT contains" Y年以後," and yi is before Y,
Color histogram [Swain91] is a basic method and has good performance for representing image content. The
color histogram method gathers statistics about the proportion of each color as the signature of an image. In our
work, the colors of an image are represented in the HSV (Hue/ Saturation/ Value) space, which is believed closer
to human perception than other models, such as RGB (Red/ Green/ Blue) or CMY (Cyan/ Magenta/ Yellow). We
quantize the HSV space into 18 hues, 2 saturations, and 4 values, with additional 4 levels of gray values; as a
result, there are a total of 148 (i.e., 18×2×4+4) bins. Let C (|C| = m) a set of colors (i.e., 148 bins), PI (QI) is
represented as Eq. (10), which models the color histogram H(PI) (H(QI)) as a vector, in which each bucket hci
counts the ratio of pixels of PI (QI) in color ci.</p>
        <p>Previous work models that each pixel is only assigned into a single color. Consider the following situation:
I1, I2 are two images, all pixels of I1 and I2 fall into ci and ci+1 respectively; I1 and I2 are indeed similar to each
other, but the similarity computed by the color histogram will regard them as different images. To address the
problem, we set an interval range δ to extend the color of each pixel and introduce the idea of a partial pixel as
shown in Eq. (11),</p>
        <p>To be mentioned, with regard to wyi (QT ) , three operators – BEFORE, IN, and AFTER – are defined to
take into account a query such as “1900 年以前拍攝的愛丁堡城堡的照片 (Pictures of Edinburgh Castle taken
before 1900),” which also concerns about the time dimension. Take, for example, the above query which targets
only images taken before 1900; a part of the textual vector of the above query about the temporal feature is given
in Table 1, it gives an idea that P1 will be retrieved since its publication year was in 1899 while P2 will not be
retrieved because of its publication year, 1901. Note that in this year, we only consider years for the temporal
feature. Hence, for a query like “1908 年四月拍攝的羅馬照片 (Photos of Rome taken in April 1908),” “四月
(April)” is treated as a general term, which only contributes its effect to the term feature.</p>
        <p>PI =&lt; hc1 (PI ),..., hcm (PI ) &gt; ,
QI =&lt; hc1 (QI ),..., hcm (QI ) &gt;
(7)
(8)
(9)
…
0
0
0
(10)
(12)
(13)
3.2</p>
        <p>Similarity Metric
While a query Q = (QT , QI ) and an image P = (PT , PI ) are represented in terms of a textual and an image
vector representation, we proposed two strategies to measure the similarity between the query and each image in the
collection. In the following, we briefly describe the proposed strategies: Strategy 1, which is exploited in the
system, T_ICLEF, only takes into account the textual similarity while Strategy 24, which combines the textual
and the image similarity, is employed in the system, VCT_ICLEF.
z
z</p>
        <p>Strategy 1 (T_ICLEF): Based on the textual similarity
PT =&lt; wt1 (PT ),..., wtn (PT ), wc1 (PT ),..., wcm (PT ), wy1 (PT ),..., wyk (PT ) &gt; ,
QT =&lt; wt1 (QT ),..., wtn (QT ), wc1 (QT ),..., wcm (QT ), wy1 (QT ),..., wyk (QT ) &gt; ,</p>
        <p>r r
Sim1(P, Q) = PrT ⋅ QrT</p>
        <p>| PT || QT |
Strategy 2 (VCT_ICLEF): Based on both the textual and the image similarity
Sim2 (P, Q) = α ⋅ Sim1(P, Q) + β ⋅ Sim3 (P, Q)
where</p>
        <p>PI =&lt; hc1 (PI ),..., hcm (PI ) &gt;,
QI =&lt; hc1 (QI ),..., hcm (QI ) &gt;,</p>
        <p>∑ min(hci (PI ), hci (QI ))
Sim3 (P, Q) = H (PI ) I H (QI ) = i
| H (QI ) |
∑ hci (QI )</p>
        <p>i
4 In our implementation, α is set to 0.7, and β is set to 0.3.
4.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Interactive Search Mechanism</title>
      <p>User Interface</p>
      <p>In the display area, a pull-down menu below each image assists users in feedback of the relevance of each
image. In fact, it is the color table shown in VCT_ICLEF which distinguishes the two systems. Users can
provide color information, which is considered as an image in the system, to help the system determine the best
query strategy. According to the experimental results, VCT_ICLEF has a better performance by exploiting color
information for searching.
As mentioned in Section 2, in the relevance feedback process, the user evaluates the relevance of the returned
images, and gives a relevance value (i.e., non-relevant, neutral, and relevant) to each of them. At the next stage,
the system performs query reformulation to modify the original query on the basis of the user’s relevance
judgments, and invokes again Cross-Language image retrieval based on the new query.</p>
      <p>Recall that we denote as the original query Q = (QT , QI ) and the new query Q′ = (QT′ , QI′ ) ; as for QT′ ,
we exploit a practical method, as shown in Eq. (14), for query reformulation. This mechanism, which has been
suggested by [Rocchio65], is achieved with a weighted query by adding useful information which is extracted
from relevant images as well as decreasing useless information which is derived from non-relevant images to the
original query. Regarding QI′ , it is computed as the centroid of the relevant images which is defined as the
average of the relevant images. We do not take into account the irrelevant images for QI′ since in our observations,
there is always a large difference among the non-relevant images in which we believe that adding the irrelevant
information to QI′ will make no contribution.</p>
      <p>In Eq. (14) and Eq. (15), α, β, γ ≥ 0 are parameters, REL and NREL stands for the sets of relevant ad irrelevant
images which are marked by the user.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation Results</title>
      <p>5.1</p>
      <p>The St. Andrews Collection
In this section, we present our evaluation results for the user-centered search task at ImageCLEF 2004.
At the ImageCLEF 2004, the St. Andrews Collection5 is used for the evaluation purpose in which the majority
of images (82%) are in black and white. It is indeed a subset of the St. Andrews University Library photographic
collection from which 10% (i.e., 28,133) images have been used. All images have an accompanying textual
description which is composed of 8 distinct fields6, including 1) Record ID, 2) Title, 3) Location, 4) Description, 5)
Date, 6) Photographer, 7) Categories, and 8) Notes. In total, the 28,133 captions consist of 44,085 terms and
1,348,474 word occurrences; the maximum caption length is 316 words, but on average 48 words in length. All
captions are written in British English, and around 81% of captions contain text in all fields. In most cases, the
caption is a grammatical sentence of about 15 words.
5.2</p>
      <p>The User-Centered Search Task
The goal of the user-centered search task is to study whether the retrieval system is being used in the manner
5 Please refer to http://ir.shef.ac.uk/imageclef2004/guide.pdf for a detail description.
6 In this paper, almost all fields were used for indexing, except for 1 and 8. The words were stemmed and
stop-words were removed as well.
intended by the system designers and how the interface helps users reformulate and refine their search requests,
given that a user searches with a specific image in mind but without knowing key information thereby requiring
them to describe the image instead. At the ImageCLEF 2004, the interactive task is using an experimental
procedure similar to iCLEF 20037.</p>
      <p>In brief, given two interactive Cross-Language image retrieval systems – T_ICLEF and VCT_ICLEF, and
the 16 topics shown in Fig. 5, 8 users are asked to test each system with 8 topics. For a system/topic combination,
a total of 4 searchers will test the system. Users are given a maximum of 5 mins only to find each image. Topics
and systems will be presented to the user in combinations following a latin-square design to ensure user/topic
and system/topic interactions are minimized. Moreover, in order to know the search strategies used by searchers,
we also conducted an interview after the task.
There are 8 people involved in the task, including 5 male and 3 female searchers. The average age of them is
23.5, with the youngest of 22 and the oldest of 26. Three of them major in computer science, two major in social
science, and the others are librarians. In particular, three searchers have experiences in participating in projects
about image retrieval. All of them have an average of 3.75 years (with a minimum of 2 years and a maximum of
5 years) accessing online search services, specifically, in Web search. In average, they search about 4 times a
week, with a minimum of once and a maximum of 7 times. However, only a half of them have experiences in
using image search services, such as Google images search. In addition, all of them are native speaker of Chinese,
but all learned English before. Most of them report that his/her English ability is acceptable or good.
5.4</p>
      <p>Results
We are interested in which system helps searchers find a topic image efficiently. We summarize the average
number of iterations8 and the average time spent by a searcher for each topic in Fig. 6. In the figure, it does not
give information in the case that all searchers did not find the target image. (For instance, regarding topic 2, all
searchers failed to complete the task by using T_ICLEF within the definite time.) The figure shows that overall
VCT_ICLEF helps users find the image within a fewer iterations with a maximum of 2 iterations saved. For
top7 Please refer to http://terral.lsi.uned.es/iCLEF/2003/guidelines.htm for further information.
8 Please note that our system does not have a efficient performance; since for each iteration it spent about 1
minute to retrieve relevant images, approximately 5 iterations is performed within the time limit.
ics 2, 5, 7, 11, 15 and 16, no searcher can find the image by making use of T_ICLEF. Furthermore, with regard to
topics 10 and 12, VCT_ICLEF has a worse performance. In our observations, it is because that most images
(82%) in the corpus are in black and white, once the user gives imprecise color information, VCT_ICLEF needs
to cost more iterations to find the image consequently.</p>
      <p>Table 2 presents the number of searchers who failed to find the image for each topic. It is clear that
VCT_ICLEF outperforms T_ICLEF in almost all cases. Considering topic 3, we believe that it is caused by the
same reason we mentioned above for topics 10 and 12. Finally, we give a summary of our proposed systems in
Table 3. The table illustrates that while considering those topics that at least one search completed the task,
T_ICLEF cost additional 0.4 iterations and 76.47 secs. Besides, by using VCT_ICLEF, on average, 89% of
searchers successfully found the image while by using T_ICLEF, there is around 56.25% of searchers who did
not fail the task.
300
250
200
sd
con150
e
S
100
50
0
T_ICLEF
VCT_ICLEF</p>
      <p>T_ICLEF
VCT_ICLEF
4.5
3.5
sno2.5
it
tIrea 2
1.5
0.5
4
3
1
0
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16</p>
      <p>Topics
(a) Average iterations
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16</p>
      <p>Topics
(b) Average time</p>
      <p>To show the effect of color information used in VCT_ICLEF, we take Fig. 3 and Fig. 4 for example.
Regarding topic 6, the query used was “燈塔 (Lighthouse).” For T_ICLEF, it returned a set of images
corresponding to the query; however, the target image could not be found in the top 80 images. Since topic 6 is a color
image, while we searched the image with color information by using VCT_ICLEF, the image was found in the first
iteration. We conclude that color information can help a user indicate to the system what he is searching for. For
an interactive image retrieval system, it is necessary to provide users not only an interface to issue a textual
query but also an interface to indicate the system the visual information of the target.</p>
      <p>Finally, we give an example as Fig. 7 to show that the proposed interactive mechanism works effectively.
The query is “長鬍鬚的男人 (A man with a beard)”. The evaluated system is T_ICLEF; after 2 iterations of
relevance feedback, it is obviously that we can improve the result by our feedback method.</p>
      <p>(a) Results after the initial search</p>
      <p>(b) Results after 2 feedback iterations
In our survey of search strategies exploited by searchers, we found that 5 searchers thought that additional color
information about the target image was helpful to indicate the system what they really wanted. Four searchers
preferred to search the image with a text query first, even though by using VCT_ICLEF. They then considered
color information for the next iteration in the situation that the target image was in color but the system returned
images all in black and white. When searching for a color image, 3 searchers preferred to use color information
first. Moreover, 2 searchers hoped that in the future, users can provide a textual query to indicate color
information, such as “黃色 (Yellow).” Finally, to be mentioned, in our systems, the user is allowed to provide a query
consisting of temporal conditions. However, since it is hard to decide in which year the image was published, no
one used a query in which temporal conditions were contained.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We participated in the user-centered search task at ImageCLEF 2004. In this paper, we proposed two interactive
Cross-Language image retrieval systems – T_ICLEF and VCT_ICLEF. The first one is implemented with a
practical relevance feedback approach based on textual information while the second one combines textual and
image information to help users find a target image. The experimental results show that VCT_ICLEF has a better
performance than T_ICLEF in almost all cases. Overall, VCT_ICLEF helps users find the image within a fewer
iterations with a maximum of 2 iterations saved.</p>
      <p>In the future, we plan to investigate user behaviors to understand in which cases users prefer a textual query
as well as in which situations users prefer to provide visual information for searching. Besides, we also intend to
implement a SOM (Self-Organizing Map) [Kohonen98] on image clustering, which we believe that it can
provide an effective browsing interface to help searchers find a target image.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Kohonen98]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kohonen</surname>
          </string-name>
          , “The
          <string-name>
            <surname>Self-Organizing</surname>
            <given-names>Map</given-names>
          </string-name>
          ,” Neurocomputing, Vol.
          <volume>21</volume>
          , No.
          <fpage>1</fpage>
          -
          <issue>3</issue>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Kushki04]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kushki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Androutsos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. N.</given-names>
            <surname>Plataniotis</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Venetsanopoulos</surname>
          </string-name>
          , “
          <article-title>Query Feedback for Interactive Image Retrieval,” IEEE Transactions on Circuits and Systems for Video Technology</article-title>
          , Vol.
          <volume>14</volume>
          , No.
          <volume>5</volume>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Miller95]
          <string-name>
            <given-names>G.</given-names>
            <surname>Miller</surname>
          </string-name>
          , “
          <article-title>WordNet: A Lexical Database for English,” Communications of the ACM</article-title>
          , pp.
          <fpage>39</fpage>
          -
          <lpage>45</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Rocchio65]
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Rocchio</surname>
          </string-name>
          , and G. Salton, “Information Search Optimization and Iterative Retrieval Techniques,
          <source>” Proc. of AFIPS 1965 FJCC</source>
          , Vol.
          <volume>27</volume>
          ,
          <string-name>
            <surname>Pt</surname>
          </string-name>
          . 1,
          <string-name>
            <surname>Spartan</surname>
            <given-names>Books</given-names>
          </string-name>
          , New York, pp.
          <fpage>293</fpage>
          -
          <lpage>305</lpage>
          ,
          <year>1965</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Salton83]
          <string-name>
            <given-names>G.</given-names>
            <surname>Salton</surname>
          </string-name>
          , and
          <string-name>
            <surname>M. J. McGill</surname>
          </string-name>
          , “Introduction to Modern Information Retrieval,”
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          ,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>[Swain91] M. J. Swain</surname>
            , and
            <given-names>D. H.</given-names>
          </string-name>
          <string-name>
            <surname>Ballard</surname>
          </string-name>
          , “Color Indexing”,
          <source>International Journal of Computer Vision</source>
          , Vol.
          <volume>7</volume>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>32</lpage>
          ,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>