<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>iCLEF 2004 Track Overview: Interactive Cross-Language Question Answering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julio Gonzalo</string-name>
          <email>julio@lsi.uned.es</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Douglas W. Oardy</string-name>
          <email>oard@glue.umd.edu</email>
        </contrib>
      </contrib-group>
      <abstract>
        <p>For the 2004 Cross-Language Evaluation Forum (CLEF) interactive track (iCLEF), ve participating teams used a common evaluation design to assess the ability of interactive systems of their own design to support the task of nding speci c answers to narrowly focused questions in a collection of documents written in a language di erent from the language in which the questions were expressed. This task is an interactive counterpart to the fully automatic cross-language question answering task at CLEF) 2003 and 2004. This paper describes the iCLEF 2004 evaluation design, outlines the experiments conducted by the participating teams, and presents some initial results from analysis of o cial evaluation measures that were reported to each participating team.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The design of systems to support information access depends on three fundamental factors: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
the user's task, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) the way in which the system will be used to achieve that task, and (3) the
nature of the information being searched. In the Cross-Language Evaluation Forum, it is assumed
that the information being searched is expressed in a di erent natural language (e.g., Spanish)
than that chosen by the user to express their information needs to the system (e.g. English). In
the CLEF interactive track (iCLEF), it is further assumed that the user will engage in an iterative
search process using a system that is designed to support human-system interaction. In 2001,
2003, and 2003, iCLEF modeled the user's task as nding documents that were topically relevant
to a written statement of the information need. In 2004 iCLEF adopted a new task; to nd speci c
answers to narrowly focused questions.
      </p>
      <p>
        The iCLEF evaluations have two fundamental goals: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) to explore evaluation design, and
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) to permit contrastive evaluation of alternative system designs. These goals are somewhat in
tension; the rst inspires us to try new tasks, while the second would bene t from stability and
continuity in the task design. Over the rst three years of iCLEF, our focus was on progressive
re nement of the evaluation design for a consistent task ( nding topically relevant documents),
and substantial progress resulted. Individual teams can continue to use the evaluation design that
were developed at iCLEF over those three years, and evaluation resources that were produced
over that period (e.g., o cial and interactive topical relevance judgments) can be of continuing
value to both CLEF participants and to teams that subsequently begin to work on cross-language
information retrieval.
      </p>
      <p>
        When selecting a new task for iCLEF this year, we considered two options: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) cross-language
question answering, and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) cross-language image retrieval. Ultimately, we selected cross-language
question answering because there was a broader base of prior work on the evaluation of fully
automated question answering systems to which we could compare our results. The Image CLEF
track did, however, also explore the design of an interactive image retrieval task this year. We
therefore achieved the best of both worlds, with the opportunity to learn about evaluation design
for both tasks. Readers interested in interactive image retrieval should consult the Image CLEF
overview paper in this volume. In this paper, we focus on interactive Cross-Language Question
Answering (CL-QA). The next section describes the iCLEF 2004 CL-QA experiment design. That
is followed by sections describing the experiments and providing an overview of the results obtained
by the participating teams. The paper concludes with some thoughts about future directions for
iCLEF.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Experiment Design</title>
      <p>Participating teams performed an experiment by constructing two conditions (identi ed as
\reference" and \contrastive"), formulating a hypothesis that they wished to test, and using a common
evaluation design to test that hypothesis. Human subjects were in groups of eight (i.e.,
experiments could be run with 8, 16, 24, or 32 subjects). Each subject conducted 16 search sessions.
A search session is uniquely identi ed by three parameters: the human subject performing the
search, the search condition tested by that subject (reference or contrastive), and the question
to be answered. Each team used di erent subjects, but the questions, the assignment of
questions to searcher-condition pairs, and the presentation order were common to all experiments. A
latin-square matrix design was adopted to establish a set of presentation orders for each subject
that would minimize the e ect of user-speci c, question-speci c and order-related factors on the
quantitative task e ectiveness measures that were used. The remainder of this section explains
the details of this experiment design.
2.1</p>
      <sec id="sec-2-1">
        <title>Question set</title>
        <p>Question selection proved to be challenging. We adopted the following guidelines to guide our
choice of questions:</p>
        <p>We selected only questions from the CLEF 2004 QA question set in order to facilitate
insightful comparisons between automatic and interactive experiments that were evaluated
under similar conditions.</p>
        <p>The largest number of questions that could be accommodated in three hours were needed
in order to maximize the reliability of the quantitative measures of task e ectiveness. Our
experience in previous years suggests that three hours is about the longest we can expect
subjects to participate in a single day, and extending an experiment across multiple days
would adversely a ect the practicality of recruiting an adequate number of subjects. We
chose to allow up to ve minutes for each search. Once training time was accounted for, this
left time for 16 questions during the experiment itself.</p>
        <p>Answers should not be known in advance by the human subjects. This restriction
proved to be particularly challenging in view regardless of the breadth of cultural
backgrounds that we expected among the participating teams in this international evaluation,
resulting in elimination of a large fraction of the CLEF 2004 QA set (e.g., \What is the
frequency unit?," \Who is Simon Peres?" and \What are Japanese suicide pilots called?,").
Two types of questions were found to be more often compatible with this restriction:
temporal questions (e.g., \When was the Convention on the Rights of the Child adopted?") and
measure questions (e.g., \How many illiterates are there in the world?" or \How much does
the world population increase each year?").</p>
        <p>Given that the question set had to be necessarily small, we wanted to avoid NIL questions
(i.e., questions with no answer. Ideally, it should be possible to nd an answer to every
question in any collection that a participating team might elect to search. Ultimately, we
found that we had to limit this restriction to presence in both the Spanish and English
collections in order to get a su ciently large number of questions from which to choose.
Together, these cover four of the ve experiments that were run (the fth used the French
collection).</p>
        <p>A small set of question cannot have a representative number of questions for each question
type. To avoid averaging over tiny sets of di erent types of questions, we decided to focus on
four question types. The CLEF QA set includes eight question types: location (e.g., \In
what city is St Peter's Cathedral?"), manner (e.g., \How did Jimi Hendrix die?"), measure
(e.g., \How much does the world population increase each year?"), object (e.g., \What is
the Antarctic continent covered with?"), organization (e.g., \What is the Mossad?"),
person (e.g., \Who is Michael Jackson married to?"), time (e.g., \When was the Cyrillic
alphabet introduced?"), and other (e.g., \What is a basic ingredient of Japanese cuisine?").
We selected two question types that called for named entities as answers (person and
organization) and two question types that called for temporal or quantitative measures
(time and measure) and sought to balance those four types of questions in the nal set.
Some iCLEF 2004 question types call for de nitions rather than succinct facts (e.g., \What
is the INCB?"). We decided to omit de nition questions because we felt that evaluation
might be di cult in an interactive setting (e.g., a user might combine information found
in documents with their own background knowledge and then create answers in their own
writing style that could not be judged using the same criteria as automatic QA systems).</p>
        <p>The nal set of sixteen questions, plus four additional questions for user training, are shown
in Table 1.
One factor that makes reliable evaluation of interactive systems challenging is that once a user
has searched for the answer to a question in one condition, the same question cannot be used
with the other condition (formally, the learning e ect would likely mask the system e ect). We
adopt a within-subjects study design, in which the condition seen for each user-topic pair is
varies systematically in a balanced manner using a latin square, to accommodate this. This same
approach has been used in the Text Retrieval Conference (TREC) interactive tracks [1] and in
past iCLEF evaluations [2]. Table 2 shows the presentation order used for each experiment..
In order to establish some degree of comparability, we chose to follow the design of the automatic
CL-QA task in CLEF-2004 as closely as possible. Thus, we used the same assessment rules, the
same assessors and the same evaluation measures as the CLEF QA task:</p>
        <p>Human subjects were asked to designate a supporting document for each answer. Automatic
CL-QA systems were required to designate exactly one such document, but for iCLEF we
also allowed the designation of zero or two supporting documents:
{ We anticipated the possibility that people might construct an answer from information
found in more than one document. Users were therefore allowed to mark either one
or two supporting documents for an answer. When two documents were designated,
assessors were instructed to determine whether both documents together supported the
answer.
{ Upon expiration of the search time, users might wish to record an answer even though
time would no longer be available to identify a supporting document. In such cases,
we allowed users to write an answer with no supporting document. Assessors were
instructed to judge such an answer to be correct if and only if that answer had been
found by some automatic CLEF CL-QA system.</p>
        <p>Users were not encouraged to use either option, and in practice there were very few cases in
which they were used.</p>
        <p>Users were allowed to record their answers in whatever language was appropriate to the
study design in which they were participating. For example, users with no knowledge of the
document language would generally be expected to record answers in the question language.
Participating teams were asked to hand-translate answers into the document language after
completion of the experiment in such cases in order to facilitate assessment.</p>
        <p>Answers were assessed by the same assessors that assessed the automatic CL-QA results
for CLEF 2004. The same answer categories were used in iCLEF as in the automatic
CL-QA track: correct (valid, supported answer), unsupported (valid but not supported by
the designated document(s)), non-exact or incorrect. The CLEF CL-QA track guidelines at
http://clef-qa.itc.it/2004/guidelines.html provide additional details on the de nition of these
categories. Assessment in CLEF is distributed geographically on the basis of the document
language, so some variation in the degree of strictness of the assessment across languages is
natural. For iCLEF 2004, assessors reported that they sometimes held machines to a higher
standard than they applied in the case of fully automated systems. For example, \July
25" was accepted as an answer to \When did the attack at the Saint-Michel underground
station in Paris occur?" for fully automatic systems (because the year was not stated in the
supporting document), but it was scored as inexact for iCLEF because the assessor believed
that the user should have been able to infer the correct year from the date of the article.
We reported the same o cial e ectiveness measures as the CLEF-2004 CL-QA track. Strict
accuracy (the fraction of correct answers) and lenient accuracy (the fraction of correct plus
unsupported answers) were reported for each condition. Complete results were reported to
each participating team by user, question and condition to allow more detailed analyses to
be conducted locally.
2.4</p>
      </sec>
      <sec id="sec-2-2">
        <title>Suggested User Session</title>
        <p>We set a maximum search time of ve minutes per question, but allowed our human subjects to
move on to the next question after recording an answer and designating supporting document(s)
even if the full ve minutes had not expired. We established the following typical schedule for
each 3-hour session:</p>
        <sec id="sec-2-2-1">
          <title>Orientation Initial questionnaire Training on both systems Break</title>
          <p>Searching in the rst condition (8 topics)
System questionnaire
Break
Searching in the second condition (8 topics)
System questionnaire
Final questionnaire
10 minutes
5 minutes
30 minutes
10 minutes
40-60 minutes
5 minutes
10 minutes
40-60 minutes
5 minutes
10 minutes</p>
          <p>Half of the users saw condition A (the reference condition) rst, the other half saw condition B
rst. Participating teams were permitted to alter this schedule as appropriate to their goals. For
example, teams that chose to run each subject separately to permit close qualitative assessment
by a trained observer might choose to substitute a semi-structured exit interview for the nal
questionnaire. Questionnaire design was not prescribed, but sample questionnaires were made
available to participating teams on the iCLEF Web site (http://nlp.uned.es/iCLEF/).
3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <p>Five groups submitted results: The Swedish Institute of Computer Science (SICS) from Sweden,
the University of Alicante, the University of Salamanca and UNED from (Spain), and the
University of Maryland from the USA. Four of the ve groups had previously participated in iCLEF (the
University of Salamanca joined the track this year). Somewhat surprisingly, all of the participants
used interactive CLIR systems of fairly conventional designs; none adapted existing QA systems
to support this task. In the remainder of this section, we brie y describe the experiment run at
each site.</p>
      <p>Alicante. The experiment compared two passage retrieval systems. In both systems, the query
was formulated in Spanish, automatically translated into English before passage retrieval,
and then passages were shown to the users in English (untranslated). The reference system
also showed ontological concepts for the query and the passage, ranking passages with the
same concepts as the query higher. The contrastive system showed syntactic-semantic
patterns (SSP) for the query and for each verb in the passage. The hypothesis being tested was
that for users with low English skills, it would be more useful to nd the answer through
SSPs than through the whole passage.</p>
      <p>Maryland. Two types of summaries were compared. The rst was an indicative summary
consisting of three sentence snippets sampled from the beginning, the middle, and the end of
a document that each contain at least one query term. That type of summary aims to
provide users with a concise overview of the document in order to permit rapid judgments
of relevance. The second was an informative summary with one longer passage
automatically selected by the system. Both systems used variants of the UMD MIRACLE interactive
CLIR system, and the hypothesis being tested was that informative summaries would be
more useful that indicative summaries for this task. Maryland was also interested in
studying search behavior (query formulation, query re nement, user-assisted query translation,
relevance judgment, and stopping criteria) for interactive CL-QA . The experiment involved
eight native English speakers searching Spanish documents to answer questions written in
English.</p>
      <p>UNED. The UNED hypothesis was that a passage retrieval system that ltered out paragraphs
that did not contain expressions of and appropriate type (named entities, dates or quantities,
depending on the question) could outperform a baseline consisting of a standard information
retrieval system (Inquery) that indexed and displayed Systran translations of the documents
(i.e. performing monolingual searches over the translated collection). A second research goal
was to establish a strong baseline for interactive CL-QA to be compared with automatic
CLQA in the context of CLEF.</p>
      <p>Salamanca The Salamanca team experimented with a passage retrieval system in which machine
translation was used to translate the query. They tested whether the possibility of
ondemand access to a full documents would be more useful for CL-QA than display of a
passage alone. Both systems included suggestion of query expansion terms; another goal of
the experiment was to determine whether users would take advantage of that possibility in
a question answering task.</p>
      <p>SICS. SICS explored the e ect of interactive query expansion using paired users (working on
different questions) that could communicate within the pair (e.g., to discuss system operation
or vocabulary selection). Additional research goals were to explore the nature of
communication within pairs and the the e ect of a \bookmark" capability on user con dence in the
reported result. The SICS experiment was monolingual, with French questions and French
documents; the human subjects were all native speakers of Swedish with moderate skills in
French.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results and Discussion</title>
      <p>In this section, we present the o cial results, draw comparisons with comparable results from the
CLEF-2004 CL-QA track, and describe some issues that arose with the assessment of submitted
answers.
4.1</p>
      <p>O</p>
      <p>cial results</p>
      <p>Three of the ve experiments yielded di erences in strict accuracy of approximately 1
answer out of 16 (0.0625% absolute), suggesting that the magnitude of detectable di erences
with this experiment design is likely appropriate for the types of hypotheses being tested.
Conformation of this result must, however, await the results of statistical signi cance tests
(e.g., analysis of variance) at each site.</p>
      <p>Five of the ten tested conditions yielded strict accuracy above 0.50, indicating that the
interactive CL-QA task is certainly feasible. There may still be room for improvement,
however; even in the best condition, more than 30% of the answers were either incorrect,
inexact, or unsupported. Inter-assessor agreement studies would be needed, however, before
we can quantify the magnitude of the further improvement that could be reliably measured
with this experiment design.</p>
      <p>Remarkably, the system used in the condition that yielded the highest strict accuracy (0.69)
was one of the simplest baselines: a standard document retrieval system performing
mono</p>
      <sec id="sec-4-1">
        <title>Maryland Maryland</title>
      </sec>
      <sec id="sec-4-2">
        <title>UNED UNED</title>
      </sec>
      <sec id="sec-4-3">
        <title>SICS SICS</title>
      </sec>
      <sec id="sec-4-4">
        <title>Alicante Alicante</title>
      </sec>
      <sec id="sec-4-5">
        <title>Salamanca Salamanca EN EN</title>
        <p>ES
ES
FR
FR
ES
ES
ES
ES</p>
        <p>ES
ES
EN
EN
FR
FR
EN
EN
EN
EN
indicative summaries
informative summaries
doc. retrieval + Systran
passage ret. + entity lter
baseline
contrastive
ontological concepts
syntactic/semantic patterns
only passages
passages + full documents
lingual searches over machine translation results. This suggests that when user interaction
is possible, relatively simple systems designs may su ce for CL-QA tasks.</p>
        <p>No evidence is yet available regarding the utility of more sophisticated question answering
techniques (e.g., question reformulation or nding candidate answers in side collections)
for interactive CL-QA because all iCLEF 2004 experiments employed fairly standard
crosslanguage information retrieval techniques.</p>
        <p>Readers are referred to the papers submitted by the participating teams for analyses of results
from speci c experiments.
4.2</p>
        <sec id="sec-4-5-1">
          <title>Comparison with CLEF QA results</title>
          <p>English was the only document language for which multiple iCLEF experiment results were
submitted, so we have chosen to focus our comparison with the CLEF 2004 CL-QA track on cases in
which English documents were used. Results from 13 automatic systems were submitted to the
CLEF-2004 CL-QA track for English documents. We compared the results of the six iCLEF 2004
conditions in which English documents were used (two conditions from each of three experiments)
with the results from those 13 automatic runs.</p>
          <p>Participating teams in the CLEF-2004 CL-QA track automatically found answers to 200
questions, of which 14 were common to iCLEF. Table 4 compares the results of the automatic systems
on these 14 questions with the results of the interactive conditions on all 16 topics.1</p>
          <p>Most of the interactive conditions yielded strict accuracy results that were markedly better
than the fully automatic systems on these questions. These large di erences cannot be explained
by the omission of two questions in the case of the automatic systems; correct answers to those
two questions would increase the strict accuracy of the best automatic system from 0.36 to 0.44,
which is nowhere near the strict accuracy of 0.69 achieved by the best interactive condition. Nor
could language di erences alone be used to explain the large observed di erences between the best
interactive and automatic systems since the same trend is present over the ve question languages
that were tried with the automatic systems.</p>
          <p>1Removal of two topics from the interactive results would unbalance some conditions, so the interactive results
include the e ect of two questions that were not assessed for the automatic systems.
question
docs</p>
          <p>Run
Automatic Systems (14 questions)
irst042iten
irst041iten
dfki041deen
bgas041bgen
lire042fren
dltg041fren
edin041deen
edin042fren
lire041fren
dltg042fren
edin042deen
edin041fren
hels041 en
IRST
IRST
DFKI
BGAS
LIRE
DLTG
EDIN
EDIN
LIRE
DLTG
EDIN
EDIN
HELS</p>
        </sec>
      </sec>
      <sec id="sec-4-6">
        <title>Average</title>
      </sec>
      <sec id="sec-4-7">
        <title>UNED</title>
        <p>UNED
Salamanca
Salamanca
Alicante
Alicante</p>
      </sec>
      <sec id="sec-4-8">
        <title>Average IT IT DE</title>
        <p>BG
FR
FR
DE
FR
FR
FR
DE
FR
FI
ES
ES
ES
ES
ES
ES</p>
        <p>EN
EN
EN
EN
EN
EN
EN
EN
EN
EN
EN
EN
EN
EN
EN
EN
EN
EN
EN</p>
        <p>Accuracy
strict lenient</p>
        <p>This observed di erence is particularly striking in view of our expectation that the question
types that we chose for the interactive evaluation would be particularly well suited to the
application automated techniques because the answers could be found literally in most cases. It seems
reasonable to expect that the gap would be proportionally larger for more questions types that
required a greater degree of inference.</p>
        <p>It is also notable that human subjects received a larger relative bene t from lenient rather than
strict scoring. The automatic results in Table 4 cannot accurately reveal di erences smaller than
1 answer out of 14 (0.07). But half of the six interactive experiments exhibited di erences at least
that large, while only one of the eight (non-zero) automatic systems showed such a di erence. We
interpret this as an indication that lenient accuracy re ects characteristics of an answer than may
be more prevalent in human question answering than in automatic question answering.
4.3</p>
        <sec id="sec-4-8-1">
          <title>The Assessment Process</title>
          <p>Richard Sutcli e and Alessandro Vallin, who coordinated the iCLEF assessment process for
English, o ered the following observations about the process:</p>
          <p>Users made more elaborate inferences than machines. For example:
Q: When did Latvia gain independence?
answer:
was judged correct even though the document said \(..)breakup of the Soviet Union (..) in
1991 ". In this case, the user inferred that Latvia was part of the Soviet Union. Another
example of this e ect is:
Q: When did Lenin die? answer: January 20 1924
The document states that \Friday is the 70th anniversary of Lenin's death." As it is dated
on Saturday, 22 January 1994, the user could could have inferred the date. Of course, the
user might also make mistakes that a machine would not; in this case, the date calculated
by the user was o by one day, leading the answer to be scored as wrong.</p>
          <p>Sometimes inexact answers were provided when a more complete answer could be inferred.
For example,
Q: When did the attack at the Saint-Michel underground station in Paris occur?
answer: July 25
In this case, the user gave an incomplete answer \July 25," but the date of the document
could have been used to accurately infer the year in which the event occurred. This could
re ect a system limitation (the date of the document may not have been displayed to the
user), or it may simply re ect a misunderstanding of the desired degree of completeness in
the answer.</p>
          <p>The option to designate more than one supporting document was used only 9 times out of
the 384 answers provided in the three experiments for which EN was the target language.
In none of those 9 cases was it used correctly (i.e., no inference using combined information
from both documents was appropriate). This suggests that this option may add an unhelpful
degree of complexity to the evaluation process.</p>
          <p>People were more creative than machines regarding what constitutes a valid answer. For
example, they might select \hundreds" as an answer, while automatic systems may fail to
recognize such an imprecise expression as a possible answer.</p>
          <p>Manual translation of the answers into the document language after completion of the
experiment introduced errors in a few cases. For example, a Spanish user correctly answered
\15 mil millones de dolares," but it was translated with a typo \$15 billions" and therefore
judged as inexact. When detected, these mistakes were corrected prior to generation of
the o cial results (since it was not our objective to assess the manual answer translation
process).
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Plans</title>
      <p>
        The iCLEF 2004 evaluation contributed a new evaluation design and results from ve experiments
in three language pairs with a total of 640 search sessions. The only similar evaluation of interactive
question answering that we are aware of was the TREC-9 interactive track [1]. The iCLEF 2004
evaluation di ers from the TREC-9 interactive track in two key ways: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) iCLEF 2004 is focused
on a cross-language task, while the TREC-9 interactive track focused on a monolingual task; and
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) iCLEF 2004 used questions and measures that facilitate comparison with an evaluation of
automatic QA systems while the TREC-9 interactive track used more complex question types and
document-oriented evaluation measures.
      </p>
      <p>The iCLEF 2004 evaluations have already made a number of speci c contributions, including:
Developing a methodology to study user-inclusive aspects of CL-QA,
Demonstrating that the accuracy of automatic QA systems is presently far below the
accuracy that a typical user can obtain using a cross-language information retrieval system of
fairly conventional design, and
Establishing an initial baseline for the interactive CL-QA task, with a median across 8 tested
conditions of about 50% strict accuracy for ve-minute searches.</p>
      <p>Much remain to be done, of course. Further analysis will be required before we are able to
apportion the judged errors between the search and translation technologies embedded in the
present systems. Moreover, we are now operating in a region where inter-assessor agreement
studies will soon be needed if we are to avoid pursuing putative improvements that extend beyond
our ability to measure their e ect. Finally, there is a large design space that remains to be
explored; no participating team has yet tried advanced techniques of the type normally used in
fully automatic CL-QA systems in interactive systems.</p>
      <p>Perhaps the most important legacy of iCLEF 2004 will be the discussions that it sparks about
new directions for information retrieval research. How can we craft an evaluation venue that will
attract participants with interests in both interactive and automatic CL-QA? What can we learn
from the CLEF-2004 CL-QA evaluation that would help us design interactive CL-QA evaluations
that re ect real application scenarios with greater delity? Given the accuracy achieved by
interactive systems with a limited investment of the user's time, what applications do we see for fully
automated CL-QA systems? With iCLEF 2004, we have gained a glimpse of these questions about
our future, and we're looking forward to discussing them when we meet in Bath this September!
6</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The authors would like to thank Alessandro Vallin and Richard Sutcli e for serving as our liaison
to the CLEF 2004 CL-QA track, for their help with assessments, and for sharing with us their
insights into the assessment process. We are also grateful to Christelle Ayache for help with the
French assessments, to Victor Peinado for le processing, to Fernando Lopez and Javier Artiles for
creating the iCLEF 2004 Web pages, and to Jianqiang Wang for creating the Systran translations
that were made available to the iCLEF teams.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>William</given-names>
            <surname>Hersh</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Over</surname>
          </string-name>
          . TREC-
          <article-title>9 interactive track report</article-title>
          .
          <source>In The Ninth Text Retrieval Conference (TREC-9)</source>
          ,
          <year>November 2000</year>
          . http://trec.nist.gov.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Douglas</surname>
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Oard</surname>
            and
            <given-names>Julio</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalo</surname>
          </string-name>
          .
          <article-title>The CLEF 2003 interactive track</article-title>
          . In Carol Peters, editor,
          <source>Proceedings of the Fourth Cross-Language Evaluation Forum</source>
          .
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>