<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the CLEF 2005 Interactive Track</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julio Gonzalo</string-name>
          <email>julio@lsi.uned.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul Clough</string-name>
          <email>p.d.clough@she</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Vallin</string-name>
          <email>vallin@itc.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Interactive Information Retrieval, Cross-Language Information Retrieval</institution>
          ,
          <addr-line>Question Answering, Image retrieval</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The CLEF Interactive Track (iCLEF) is devoted to the comparative study of userinclusive cross-language search strategies. In 2005, we have studied two cross-language search tasks: retrieval of answers and retrieval of annotated images. In both tasks, no further translation or post-processing is needed after performing the tasks to fulfill the information need. In the interactive Question Answering task, users are asked to find the answer to a number of questions in a foreign-language document collection, and write the answers in their own native language. In the interactive image retrieval task, a picture is shown to the user, and then the user is asked to find the picture in the collection. This paper summarizes the task design, experimental methodology, and the results obtained by the research groups participating in the track.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>4 [Information Systems Applications]</kwd>
        <kwd>H</kwd>
        <kwd>4</kwd>
        <kwd>m Miscellaneous</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>In CLEF 2005, user studies have consolidated the two research issues studied in CLEF 2004 as
pilot tasks: cross-language question answering and known-item image search.</p>
      <p>In the interactive Question Answering task, users are asked to find the answer to a number
of questions in a foreign-language document collection, and write the answers in their own native
language. Subjects must use two interactive search assistants (which are to be compared), pairing
questions and systems according to a latin-square design to filter out question and user effects.
For this task, we have used a subset of the ad-hoc QA testbed, including questions, collections
and evaluation methodology.</p>
      <p>In the interactive image retrieval task, a picture is shown to the user, and then the user
is asked to find the picture in the collection. This was chosen as a realistic task (finding stuff
I’ve seen before) in which visual features could also play an important role (users are given a
picture instead of a written description of what they have to look for). The target data is the St.
Andrews’ collection (as used in the ad-hoc image CLEF task), in which images are annotated in
English with a number of rich metadata descriptions. Again, each participant group was expected
to compare two different search assistants, combining users, queries and systems according to a
latin-square desing to filter out query and user effects.</p>
      <p>The remainder of this paper describes the experimental design and the results obtained by the
research groups for each of these tasks.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Image Retrieval task</title>
      <p>The ImageCLEF interactive search task provides user–centered evaluation of cross–language image
retrieval systems. In cross–language image search, the object to be retrieved is an image. This
is appealing as a CLIR task because often (depending on the user and query) the object to be
retrieved (i.e. the image) can be assumed to be language-independent, i.e. there is no need for
further translation when presenting results to the user. This makes a good introductory task
to CLIR, requiring only query translation to bridge the language gap between the user’s query
(source) language, and the language used to annotate the images (target language).</p>
      <p>Image retrieval can be purely visual in the case of query–by–example (QBE) which is entirely
language–independent, but this assumes the user wants to perform a visual search (e.g. find
me images which appear visually similar to the one provided). However, users may also want to
search for images starting with text-based queries (e.g. Web image search) requiring that texts are
associated with the target image collection. For CLIR, the language of the texts used to annotate
the images should not affect retrieval, i.e. a user should be able to query the images in their native
language making the target language transparent. Effective cross–language image retrieval will
involve both text–based and content–based IR (CBIR) methods in conjunction with translation.</p>
      <p>The main areas of study for a cross–language image retrieval assistant include:
 How well a system supports user query formulation for images with associated texts (e.g.
captions or metadata) written in a language different from the native language of the users.
This is also an opportunity to study how the images themselves could also be used as part
of the query formulation process.
 How well a system supports query re–formulation, e.g. the support of positive and
negative feedback to improve the user’s search experience, and how this affects retrieval. This
aims to address issues such as how visual and textual features can be combined for query
reformulation/expansion.
 How well a system allows users to browse the image collection. This might include support
for summarising results (e.g. grouping images by some pre-assigned categorization scheme or
by visual feature such as shape, colour or texture). Browsing becomes particularly important
in a CLIR system when query translation fails and returns irrelevant or no results.
 How well a system presents the retrieved results to the user to enable the selection of relevant
images. This might include how the system presents the caption to the user (particularly
if they are not familiar with the language of the text associated with the images, or some
of the specific and colloquial language used in the captions) and investigate the relationship
between the image and caption for retrieval purposes.</p>
      <p>The interactive image retrieval task in 2004 concentrated on query re–formulation and this
has been the focus of experiments in 2005 also, together with the presentation of search results.
Groups were not set a specific retrieval goal to enable some degree of flexibility.
2.1</p>
      <sec id="sec-2-1">
        <title>Experimental Procedure</title>
        <p>Participants were required to compare two interactive cross–language image retrieval systems (one
intended as a baseline) that differ in the facilities provided for interactive retrieval. For example,
comparing the use of visual versus textual features in query formulation and refinement. As a
crosslanguage image retrieval task, the initial query was required to be in a language different from the
collection (i.e. not English) and translated into English for retrieval. Any text displayed to the user
was also required to be translated into the user’s source language. This might include captions,
summaries, pre-defined image categories etc. ImageCLEF used a within–subject experimental
design: users were required to test both interactive systems.</p>
        <p>
          The same search task as 2004 was used: given an image (not including the caption) from the
St Andrews collection of historic photographs, the goal for the searcher is to find the same image
again using a cross–language image retrieval system. This models the situation in which a user
searches with a specific image in mind (perhaps they have seen it before) but without knowing
key information thereby requiring them to describe the image instead, e.g. searches for a familiar
painting whose title and painter are unknown (i.e. a high precision task or target search [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]).
        </p>
        <p>The interactive ImageCLEF task is run similar to iCLEF 2003 using a similar experimental
procedure. However, because of the type of evaluation (i.e. whether known items are found or
not), the experimental procedure for iCLEF 2004 (Q&amp;A) is also very relevant and we make use of
both iCLEF procedures. The user–centered search task required groups to recruit a minimum of 8
users (native speakers in the source language) to complete 16 search tasks (8 per system). Images
which users were required to find are shown in Fig. 1. Users are given a maximum of 5 mins only
to find each image. Topics and systems were presented to the user in combinations following a
latin–square design to ensure minimisation of user/topic and system/topic interactions.
(a)
(b)
(c)
(d)
(e)
(f)
(g)
(h)
(i)
(j)
(k)
(l)
(m)
(n)
(o)
(p)</p>
        <p>Participants were encouraged to make use of questionnaires to obtain feedback from the user
about their level of satisfaction with the system and how useful the interfaces were for retrieval. To
measure the effectiveness and efficiency with which a cross–language image retrieval search could
be performed, participants were asked to submit the following information: whether the user could
find the intended image or not (mandatory), the time taken to find the image (mandatory), the
number of steps/iterations required to reach the solution (e.g. the number of clicks or the number
of queries - optional), and the number of images displayed to the user (optional).
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Participating Groups</title>
        <p>
          Although 11 groups signed up for the interactive task, only 2 groups submitted results: Miracle
and the University of Sheffield. Miracle compared the same interface but using Spanish (European)
versus English versions [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. The focus of the experiment was whether it is better to use an AND
operator to group terms of multi-word queries (in the English system) or combine terms using an
OR operator (in the Spanish system). Their aim was to compare whether it is better to use English
queries with terms conjuncted (which have to be precise and use the exact vocabulary - maybe
difficult for a specialised domain like historical Scottish photographs) or to use the disjunction of
terms in Spanish and have the option of relevance feedback (a more “fuzzy” and noisy search but
which doesn’t require precise vocabulary and exact translations). Their objective was to test the
similarity of retrieval performance using both approaches.
        </p>
        <p>
          Sheffield compared 2 interfaces with the same source language (Italian): one displaying search
results as a list, the other organizing retrieved images into a hierarchy of text concepts displayed
on the interface as an interactive menu [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. The aim of the experiment was to determine the
usefulness of grouping results using concept hierarchies and investigate translation issues in cross–
language image search. Queries were translated using Babelfish and the entire user interface also
translated to provide a working system in Italian.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Results and Discussion</title>
        <p>Given only two submissions, conclusions that can be deduced from the interactive task are limited.
However, the findings of individual groups were interesting and we summarise their main results
to highlight the effectiveness of selected approaches. Miracle found results to be similar for both
systems evaluated: English (69% of images found; 102 secs. average search time), Spanish (66%
of images found; 113 secs. average search time). Based on investigation of the results and
observation of users, a number of interesting points are made: that domain-specific terminology causes
problems for cross–language searches (and therefore impacts far greater on queries with a
conjunction of terms). In addition, translated Spanish query terms did not match caption terms also
causing vocabulary mismatch. From questionnaires, users preferred the English version because
the conjunction of terms often gave results users expected (i.e. a set of documents containing all
query terms). Miracle also observed users extracting words from captions to further refine their
search and user’s commented on differences between the expected results of a search for a given
keyword and those actually obtained. Users were also allowed to continue searching after the
allotted time and in most cases found the relevant image in a short time (less than 1 minute).</p>
        <p>The experiments undertaken by Sheffield also highlighted some interesting search strategies by
users and problems with the concept hierarchies and interface for cross–language image retrieval.
Quantitative results were similar using both a list of images and a menu generated from the
concept hierarchies: list (53% of images found; 113 secs. average search time) and menu (47%
images found; 139 secs. average search time). Overall users of the Sheffield systems found 82/128
relevant images and users of the Miracle system 86/128 images. The experiments undertaken
by Sheffield observed negative effects on search, generation of the concept hierarchy and results
display due to translation errors such as mis-translations and un-translated terms. Although based
on effectiveness the menu appears to offer no difference compared to presenting results as a list,
users preferred the menu (75% vs. 25% for the list) indicating this approach to be an engaging and
interesting feature. In particular users liked the compact representation of search results offered
by the menu compared to the ranked list.
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Question Answering task</title>
      <sec id="sec-3-1">
        <title>Experiment Design</title>
        <p>Participating teams performed an experiment by constructing two conditions (identified as
“reference” and “contrastive”), formulating a hypothesis that they wished to test, and using a common
evaluation design to test that hypothesis. Human subjects were in groups of eight (i.e.,
experiments could be run with 8, 16, 24, or 32 subjects). Each subject conducted 16 search sessions.
A search session is uniquely identified by three parameters: the human subject performing the
search, the search condition tested by that subject (reference or contrastive), and the question
to be answered. Each team used different subjects, but the questions, the assignment of
questions to searcher-condition pairs, and the presentation order were common to all experiments. A
latin-square matrix design was adopted to establish a set of presentation orders for each subject
that would minimize the effect of user-specific, question-specific and order-related factors on the
quantitative task effectiveness measures that were used. The remainder of this section explains
the details of this experiment design.
3.1.1</p>
        <sec id="sec-3-1-1">
          <title>Question set</title>
          <p>Questions were selected from the CLEF 2005 QA question set in order to facilitate insightful
comparisons between automatic and interactive experiments that were evaluated under similar
conditions. The criteria to select questions was similar to those used in iCLEF 2004:
 Answers should not be known in advance by the human subjects; this restriction
resulted in elimination of a large fraction of the initial question set.
 Given that the question set had to be necessarily small, we wanted to avoid NIL questions
(i.e., questions with no answer. Ideally, it should be possible to find an answer to every
question in any collection that a participating team might elect to search.
 We focused on four question types to avoid excessive sparseness in the question set: two
question types that called for named entities as answers (person and organization) and
two question types that called for temporal or quantitative measures (time and measure).
The additional restriction of having answers in the largest number of languages forced us to
include also some other questions.</p>
          <p>The final set of sixteen questions, plus four additional questions for user training, are shown
in Table 1.
3.1.2</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Latin-Square Design</title>
          <p>
            One factor that makes reliable evaluation of interactive systems challenging is that once a user
has searched for the answer to a question in one condition, the same question cannot be used
with the other condition (formally, the learning effect would likely mask the system effect). We
adopt a within-subjects study design, in which the condition seen for each user-topic pair is
varies systematically in a balanced manner using a latin square, to accommodate this. This same
approach has been used in the Text Retrieval Conference (TREC) interactive tracks [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] and in
past iCLEF evaluations [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]. Table 2 shows the presentation order used for each experiment..
3.1.3
          </p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Evaluation Measures</title>
          <p>In order to establish some degree of comparability, we chose to follow the design of the automatic
CL-QA task in CLEF-2005 as closely as possible. Thus, we used the same assessment rules, the
same assessors and the same evaluation measures as the CLEF QA task:
 Human subjects were asked to designate a supporting document for each answer (we
eliminated the exceptions allowed last year, as for instance building an answer from the
information in two documents, because in practice no user exploited these alternative possibilities).
 Users were allowed to record their answers in whatever language was appropriate to the
study design in which they were participating. For example, users with no knowledge of the
document language would generally be expected to record answers in the question language.
Participating teams were asked to hand-translate answers into the document language after
completion of the experiment in such cases in order to facilitate assessment.
 Answers were assessed by the same assessors that assessed the automatic CL-QA results
for CLEF 2005. The same answer categories were used in iCLEF as in the automatic
CL-QA track: correct (valid, supported answer), unsupported (valid but not supported by
the designated document(s)), non-exact or incorrect . The CLEF CL-QA track guidelines
at http://clef-qa.itc.it/2005/guidelines.html provide additional details on the definition of
these categories.
 We reported the same official effectiveness measures as the CLEF-2005 CL-QA track. Strict
accuracy (the fraction of correct answers) and lenient accuracy (the fraction of correct plus
unsupported answers) were reported for each condition. Complete results were reported to
each participating team by user, question and condition to allow more detailed analyses to
be conducted locally.
3.1.4</p>
        </sec>
        <sec id="sec-3-1-4">
          <title>Suggested User Session</title>
          <p>We set a maximum search time of five minutes per question, but allowed our human subjects to
move on to the next question after recording an answer and designating supporting document(s)
even if the full five minutes had not expired. We established the following typical schedule for
each 3-hour session:</p>
          <p>Orientation
Initial questionnaire
Training on both systems
Break
Searching in the first condition (8 topics)
System questionnaire
Break
Searching in the second condition (8 topics)
System questionnaire
Final questionnaire
10 minutes
5 minutes
30 minutes
10 minutes
40-60 minutes
5 minutes
10 minutes
40-60 minutes
5 minutes
10 minutes</p>
          <p>Half of the users saw condition A (the reference condition) first, the other half saw condition B
first. Participating teams were permitted to alter this schedule as appropriate to their goals. For
example, teams that chose to run each subject separately to permit close qualitative assessment
by a trained observer might choose to substitute a semi-structured exit interview for the final
questionnaire. Questionnaire design was not prescribed, but sample questionnaires were made
available to participating teams on the iCLEF Web site (http://nlp.uned.es/iCLEF/).
3.2</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Experiments</title>
        <p>
          Three groups submitted results:
University of Alicante. This group investigated how much context is needed to recognize
answers accurately with a low-medium knowledge of the document language [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Their baseline
system shows whole passages (maximum context) to users, while the experimental system
shows only a clause (minimum context). Both systems highlight query terms, synonyms of
query terms and candidate answers to facilitate the task.
        </p>
        <p>
          University of Salamanca. Their focus has been exploring the use of free on-line machine
translation programs for query formulation and presentation of results [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Both systems compared
permit entering the query either in the user language or in the target language; in the first
case, machine translation is applied to the query before searching the collection. In the
reference system, results are displayed without translation; the contrastive system permits
translating passages. Users were classified as having “poor” or “good” foreign language skills
in four experiments, Spanish to English and Spanish to French.
        </p>
        <p>
          UNED. This team has compared searching full documents with searching single sentences [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>Both systems highlight fragments of the appropriate answer type to help locating the answer.
In addition, the contrastive system filters out sentences which do not contain expressions of
the appropriate answer type.</p>
        <p>Group</p>
        <p>Users</p>
        <p>Docs</p>
        <p>Experiment Condition</p>
        <p>Accuracy
Strict Lenient
Alicante
Alicante
Salamanca
Salamanca
Salamanca
Salamanca
Salamanca
Salamanca
Salamanca
Salamanca
UNED
UNED</p>
        <p>ES
ES
ES
ES
ES
ES
ES
ES
ES
ES
ES
ES</p>
        <p>EN
EN
EN
EN
EN
EN
FR
FR
FR
FR
EN
EN
full passages
clauses
good lang. skills / no translation
good lang. skills / translation
poor lang. skills / no translation
poor lang. skills / translation
good lang. skills / no translation
good lang. skills / translation
poor lang. skills / no translation
poor lang. skills / translation
documents
sentences with answer type filter
Although iCLEF experiments continue producing interesting research results, which may have a
substantial impact on the way effective cross-language search assistants are built, participation
in this track has remain low across the five years of existence of the track. Interactive studies,
however, remain as a recognized necessity in most CLEF tracks.</p>
        <p>In order to find an explanation for this apparent contradiction, a questionnaire was created to
establish reasons for low participation in the interactive ImageCLEF task and sent to all
ImageCLEF participants. Seven participants returned their questionnaires and, out of these, 6 stated
(the 7th participated in interactive ImageCLEF) their reason for not participating was lack of
time, 5 lack of local resources and 4 that interactive experiments involved too much set-up time.
Interactive experiments consume resources which many groups do not have.</p>
        <p>We can think of a number of measures to solve this problem:
 Lowering the cost of participation. One approach is to provide a common task in which all
groups participate, or use a shared multilingual document collection which can be accessed
via an API, e.g. Flickr, Yahoo! or Google. This is only a partial solution, because the highest
cost comes from recruting, training and monitorizing users for the searching sessions. An
alternative is devising an experiment design in which search interfaces are deployed in real
working environments, and then study the search logs of real users with real needs. This
is a less controlled environment which could, nevertheless, provide a wealth of information
about why and how users search in a cross-language manner.
 Adding value to the experimental setting. For instance, if we could work with online
multilingual collections which have large user communities, setting up cross-language search
interfaces for them has the additional appeal of being able to provide demonstrations which
turn into useful web services for a significant set of web users.</p>
        <p>We are currently contemplating the possibility of using a large-scale, web-based image database,
such as Flickr (www.flickr.com), for iCLEF experiments. The Flickr database contains over five
million images freely accesible via web, daily updated by a large number of users and available for
all web users. These images are annotated by the authors with freely chosen keywords in a naturally
multilingual manner: most authors use keywords in their native language, some combine more than
one language. In addition, photographs have titles, descriptions, colaborative annotations, and
comments in many languages. Participating groups would have the opportunity of building search
interfaces not only for testing/demo purposes, but also to offer a useful web service with many
potential users.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>We are indebted to Richard Sutcliffe and Christelle Ayache for taking care of the Q&amp;A assessments;
Victor Peinado and Fernando Lo´pez for their assistance with submission processing, Javier Artiles
for maintaining the iCLEF web page, and Jianqiang Wang for creating the Systran translations
that were made available to the iCLEF teams. Thanks also to Daniela Petrelli for the fruitful
discussions about the iCLEF experimental design, and to Carol Peters and Doug Oard for their
support. This work has been partially supported by the Spanish Government, project
R2D2Syembra (TIC2003-07158-C04-02).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>. P.</given-names>
            <surname>Clough</surname>
          </string-name>
          , H. Mu¨eller, T. Desealers, M. Grubinger, t. Lehmann,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jensen</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Hersh</surname>
          </string-name>
          .
          <article-title>The CLEF 2005 Cross-Language Image Retrieval Track</article-title>
          , in this volume.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>. J.</given-names>
            <surname>Cox</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Omohundro</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Yianilos</surname>
          </string-name>
          . Pichunter:
          <article-title>Bayesian relevance feedback for image retrieval</article-title>
          .
          <source>Proceedings of the 13th International Conference on Pattern Recognition</source>
          ,
          <volume>3</volume>
          :
          <fpage>361</fpage>
          -
          <lpage>369</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>William</given-names>
            <surname>Hersh</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Over</surname>
          </string-name>
          . TREC-
          <article-title>9 interactive track report</article-title>
          .
          <source>In The Ninth Text Retrieval Conference (TREC-9)</source>
          ,
          <year>November 2000</year>
          . http://trec.nist.gov.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V.</given-names>
            <surname>Peinado</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <article-title>Lo´pez-</article-title>
          <string-name>
            <surname>Ostenero</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalo</surname>
            and
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Verdejo</surname>
          </string-name>
          . UNED at iCLEF 2005:
          <article-title>automatic highlighting of potential answers In this volume</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Navarro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Moreno-Monteagudo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Noguera</surname>
          </string-name>
          , S. Va´zquez,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <article-title>LLopis and A</article-title>
          . Montoyo. “
          <article-title>How much context do you need?” An experiment about the context size in Interactive Cross-Language Question Answering</article-title>
          . In this volume.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Douglas</surname>
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Oard</surname>
            and
            <given-names>Julio</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalo</surname>
          </string-name>
          .
          <article-title>The CLEF 2003 interactive track</article-title>
          . In Carol Peters, editor,
          <source>Proceedings of the Fourth Cross-Language Evaluation Forum</source>
          .
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Petrelli</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Clough</surname>
          </string-name>
          , P.D.
          <article-title>Concept Hierarchy across Languages in Text-Based Image Retrieval: A User Evaluation</article-title>
          , In this volume.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Villena-Rom</surname>
          </string-name>
          ´an, R., Crespo-Garc´ıa, R.M., and Gonza´lez-Crist´obal, J.C.
          <article-title>Boolean Operators in Interactive Search</article-title>
          , in this volume.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zazo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Figuerola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Alonso</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Ferna</surname>
          </string-name>
          <article-title>´ndez</article-title>
          . iCLEF 2005 at REINA-USAL:
          <article-title>Use of Free On-line Machine Translation Programs for Interactive Cross-Language Question Answering</article-title>
          . In this volume.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>