<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bernardo  Magnini</string-name>
          <email>magnini@itc.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Danilo  Giampiccolo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pamela  Forner</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christelle  Ayache</string-name>
          <email>ayache@elda.fr</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valentin  Jijkoun</string-name>
          <email>jijkoun@science.uva.nl</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Petya  Osenova</string-name>
          <email>petya@bultreebank.org</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anselmo Peñas</string-name>
          <email>anselmo@lsi.uned.es</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paulo Rocha</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bogdan Sacaleanu</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>and Richard Sutcliffe</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2006</year>
      </pub-date>
      <abstract>
        <p>Having being proposed for the fourth time, the QA at CLEF track has confirmed a still raising interest from the  research community, recording a constant increase both in the number of participants and submissions.  In  2006,  two  pilot  tasks,  WiQA  and  AVE,  were  proposed  beside  the  main  tasks,  representing  two  promising  experiments for the future of QA.  Also  in  the  main  task  some  significant  innovations  were  introduced,  namely  list  questions  and  requiring  text  snippet(s) to support the exact answers. Although this had an impact on the work load of the organizers both to  prepare  the  question  sets  and  especially  to  evaluate  the  submitted  runs,  it  had  no  significant  influence  on  the  performance  of  the  systems,  which  registered  a  higher  Best  accuracy  than  in  the  previous  campaign,  both  in  monolingual and bilingual tasks.  In  this  paper  the  preparation  of  the  test  set  and  the  evaluation  process  are  described,  together  with  a  detailed  presentation  of  the  results  for  each  of  the  languages.  The  pilot  tasks  WiQA  and  AVE  will  be  presented  in  dedicated articles.  1  Intr oduction </p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Inspired by previous TREC evaluation campaigns, QA tracks have been proposed at CLEF since 2003. During 
these  years,  the  effort  of  the  organisers  has  been  focused  on  two  main  issues.  One  issue  was  to  offer  an 
evaluation  exercise  characterised  by  cross­linguality,  covering  as  many  languages  as  possible.  From  this 
perspective, major attention has been given to European languages, adding at least one new language each year, 
but  keeping  the  offer  open  to  languages  from  all­over  the  world,  as  the  use  of  Indonesian  shows.  The  other 
important issue was to maintain a balance between the established procedure inherited by the TREC campaigns 
and innovation. This allowed newcomers to join the competition and, at the same time, offered “veterans” more 
challenges.  Following  these  principles,  in  QA@CLEF  2006  two  pilot  tasks,  namely  WiQA  and  Answer 
Validation Exercise (AVE), were proposed together with a main task. As far as the latter is concerned, the most 
significant  innovation  was  the  introduction  of  lIST  questions,  which  had  also  been  considered  for  previous 
competitions, but had previously been avoided due to the problems that their selection and assessment implied. 
Other important innovations consisted in the possibility to return more than one answer per question, and by the 
request to  provide text  snippets together  with  the  docid  to  support  the  exact  answer.  All  these  changes implied 
the necessity  of introducing new evaluation measures, which would account also for List and multiple answers. 
Nevertheless,  the  evaluation  process  proved  to  be  more  complicated  than  expected,  partly  because  of  the 
excessive  workload that multiple answers represented for groups already in charge for a larger number of runs.
As  a  consequence,  some  groups,  like  the  Spanish  and  the  English  ones,  could  only  correct  one  answer  per 
question, which decreased the possibility of comparisons between runs. 
As a general remark, it can be said that the positive trend in participation registered in the previous  campaigns 
was confirmed, and 10 new participants joined the competition from Europe, Asia and America. 
As reflected in the results, systems' performance improved considerably, with the Best Accuracy increasing from 
64% to 68% in the monolingual tasks, and, more significantly, from 39% to 49% in the bilingual ones. 
This paper describes the preparation process and presents the results of the QA track at CLEF 2006. In section 2, 
the  task  is  described  in  detail. The  different  phases  of  the  Gold  Standard  preparation  are exposed  in  section  3. 
After a quick presentation of the participants in section 4, the evaluation procedure and the results are reported 
respectively in  section  5  and  6.  In  section  7,  some  final  considerations  are  given  about  this  campaign  and  the 
future of QA@CLEF. 
2 </p>
    </sec>
    <sec id="sec-2">
      <title>Tasks </title>
      <p>In 2006  campaign,  the  procedure  consolidated in  previous  competitions  was  used.  Accordingly,  there  was  a 
main  task  (which  was  comprehensive  of  a  monolingual  task  and  several  cross­language  sub­tasks),  and two 
pilot tasks described below: 
1. </p>
      <p>
        WiQ A:  developed  by  Maarten  de  Rijke.  The  purpose  of  the  WiQA  pilot  is  to  see  how  IR  and  NLP 
techniques can be effectively used to help readers and authors of Wikipages access information spread 
throughout Wikipedia rather than stored locally on the pages.[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] 
2.  Answer  Validation Exer cise (AVE): A voluntary exercise to promote the development and evaluation 
of subsystems aimed at validating the correctness of the answers given by a QA system. The basic idea 
is that once a pair [answer + snippet] is returned by a  QA system, a hypothesis is built by turning the 
pair  [question  +  answer]  into  the  affirmative  form.  If  the  related  text  (a  snippet  or  a  document) 
semantically entails this hypothesis, then the answer is expected to be correct. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] 
Two specific papers in the present Working Notes are dedicated to these pilot tasks. More detailed information, 
together with the results, can be found there. 
In  addition  to  the  tasks  proposed  during  the  actual  competition,  a  "time­constrained"  QA  exercise  will  be 
proposed by the University of Alicante during the CLEF 2006 Workshop. In order to evaluate the ability of QA 
systems  to  retrieve  answers  in  real  time,  the  participants  will  be  given  a  time  limit  (e.g.  one  or  two  hours)  in 
which to answer a set of questions. These question sets are different and smaller than those provided in the main 
task (e.g. 15­25 questions). The initiative is aimed towards providing a more realistic scenario for a QA exercise. 
The main task was basically the same as in previous campaigns. Some new ideas were implemented in order to 
make the competition more challenging. The participating systems were fed a set of 200 questions, which could 
be about:
·  facts or events (F­actoid questions);
·  definitions of people, things or organisations (D­efinition questions);
·  lists of people, objects or data (L­ist questions). 
The systems were then asked to return from one to ten exact answers. “Exact” meant that neither more nor less 
than the  information required  is  given. The  answer needed  to  be supported  by  the  docid  of  the  document(s) in 
which the exact answer was found, and by one to ten text snippets which gave the actual context of it. 
The  text  snippets  were  to  be  put  one  next to  the  other,  separated  by  a  tab.  The  snippets  were  substrings  of  the 
specified documents. They should provide enough context to  justify the exact answer suggested. Snippets for a 
given  response  had  to  be  a  set  of  sentences  of  not  more  than  500  bytes  in  total  (although  for  example  the 
Portuguese  group  accepted  –  and  actually  preferred  –  length  to  be  specified  in  sentences).  There  were  no 
particular restrictions  on the  length  of  an answer­string,  but  unnecessary  pieces  of  information  were  penalized, 
since the answer was marked as ineXact. Since Definition questions may have long strings as answers, they were 
(subjectively) assessed mainly on their informativity and usefulness, and not on exactness. The tasks were both:
·  monolingual,  where  the  language  of  the  question  (Source  language)  and  the  language  of  the  news 
collection (Target language) were the same;
·  cross­lingual,  where  the  questions  were  formulated  in  a  language  different  from  that  of  the  news 
collection.
      </p>
      <p>TARGET  LANGUAGES  (corpus and answers) </p>
      <p>BG  DE  EN  ES  FR  IT  NL  PT 
S
O
U
R
C
E
 
L
A
N
G
U
A
G
E
S
(q 
u
e
s
t
i
o
n
s
)
 </p>
      <p>BG 
DE 
EN 
ES 
FR 
IN
IT 
NL 
PT 
PL </p>
      <p>RO 
Eleven  source  languages  were  considered,  namely,  Bulgarian,  Dutch  ,  English,  French,  German,  Indonesian, 
Italian, Polish , Portuguese,  Romanian and Spanish. Note the loss of Finnish, and the introduction of Polish and 
Romanian  with  respect  to  last  year.  All  these  languages  were  also  considered  as  target  languages,  except  for 
Indonesian,  Polish and  Romanian.  These  three  languages  had no news  collection  available  for  the  queries.  As 
was done for Indonesian in the previous two campaigns, the English question set was translated into Indonesian 
(IN),  Polish  (PL)  and  Romanian  (RO),  and  the  German  question  set  into  Romanian  (RO).  Only  the  bilingual 
tasks  IN­EN,  PL­EN,  RO­EN  and  RO­DE  were  activated.  In  the  case  of  IN­EN,  PL­EN,  and  RO­EN,  the 
questions  were  posed  in  the respective  language  (i.e.  IN,  PL,  RO),  while  the answers  were  retrieved  from  the 
English  collection.  In  the  RO­DE  case,  the  question  was  made  in  Romanian,  whilst  the  answer  was  retrieved 
from the German collection. 
As shown in Table 1, 24 tasks were proposed and divided in:
·  7 Monolingual ­i.e. Bulgarian (BG), German (DE), Spanish (ES), French (FR), Italian (IT), Dutch (NL), 
and Portuguese (PT);
·  17 Cross­lingual. 
and let organizing groups decide on how to assess the answers to these different kinds of questions. </p>
      <sec id="sec-2-1">
        <title>Other innovations were:</title>
        <p>·  the input format, where the type of question (F,D,L) was no longer indicated;
·  and the result format, where up to a maximum of ten answers per question was allowed, with one to ten 
text snippets supporting the exact answer. 
Following the procedure established in previous campaigns, initially each organising group (one for each Target 
language) was assigned a number of topics taken from the CLEF IR track on which candidates’ questions were
PERIOD 
2002 
2002 
1994 
1994 
1995 
Initially,  100  questions  were  selected  in  each  of  the  source  languages,  distributed  between  Factoid,  Definition 
and List questions. 
Factoid questions are fact­based questions, asking for the name of a person, a location, the extent of something, 
the day on which something happened, etc. The following 6 answer types for factoids were considered:
-  PERSON (e.g. "Who was Lisa Marie Presley's father ?")
-  TIME (e.g. "What year did the Second World War finish? ")
-  LOCATION (e.g. "What is the capital of Japan? ")
-  ORGANIZATION (e.g. "What party did Hitler belong to? ")
-  MEASURE (e.g. "How many monotheistic religions are there in the world? ")
-  OTHER, i.e. everything else that does not fit into the other five categories (e.g. "What is the most­read </p>
        <p>Italian daily newspaper? ") 
Definition questions, i.e. questions like "What/Who is X?",  were divided into the following categories:
·  PERSON  ­i.e.  questions  asking  for  the  role,  job,  and/or  important  information  about  someone  (e.g. 
"Who is Lisa Marie Presley? ");
·  ORGANIZATION ­i.e. questions asking for the mission, full name, and/or important information about 
an organization (e.g. "What is Amnesty International? " or "What is the FDA?");
·  OBJECT  ­i.e.  questions  asking  for  the  description  or  function  of  objects  (e.g.  “What  is  a  Swiss  army 
knife? ”, “What is a router? ”);
·  OTHER ­i.e.  question  asking  for the  description  of  natural phenomena,  technologies,  legal  procedures 
etc. (e.g. “What is a tsunami? ”, “What is DSL? ”, “What is impeachment? ”). 
The  last  two  categories  were  especially  added  to  reduce  the  numbers  of  definition  questions  which  may  be 
answered  very  easily  (such  as  acronyms  concerning  organizations,  which  are  usually  answered  rendering  the 
abbreviation  in  full,  and  people’s  job­description,  which  are  usually  found  as  appositions  of  proper  names  in 
news text). 
As  mentioned  above,  questions  that require  a  list  of  items  as  answers,  were introduced  for  the  first  time.  (e.g. 
Name works by Tolstoy). 
Among these three categories, a number of NIL questions, i.e. questions that do not have any known answer in 
the target document collection, were distributed. They are important because a good QA system should identify 
them, instead of returning wrong answers. 
Three  different  types  of  temporal restriction  –  a temporal  specification  that  provides  important  information  for 
the retrieval of the correct answer, were associated to a certain number of F, D, L, more specifically:
-  restriction by DATE (e.g. "Who was the US president in 1962?"; “ Who was Berlusconi in 1994?” )
-  restriction by PERIOD (e.g. "How many cars were sold in Spain between 1980 and 1995?")
-  restriction  by  EVENT  (e.g.  "Where  did  Michael  Milken  study  before  enrolling  in  the  university  of </p>
        <p>Pennsylvania? ")
The distribution of the questions among these categories is described in Table 4. 
Each  of  the  question  sets  was  then  translated  into  English,  so  that  each  group  could  choose  additional  100 
questions from those proposed by the others and translate them in their own languages. At the end, each source 
language had 200 questions, which were collected in an XML document. Unlike in the previous campaigns, the 
questions  were  not  translated  in  all  the  languages  due  to  time  constraints,  and  the  Gold  Standard  contained 
questions in multiple languages only for activated tasks. Since Indonesian, Polish and Romanian  did not have a 
data  collection  of  their  own,  the  English  question  set  was  translated,  so  that  the  cross­lingual  subtasks  IN­EN, 
PL­EN and RO­EN were made available. As not all questions had been previously translated, a translation of the 
target  language  question  sets  into  the  source  languages  was  needed  for  cross­language  sub­tasks  which  had  at 
least one registered participant. 
4 </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Pa r ticipants </title>
      <p>The number of participants has constantly grown over the years [see Table 5]. In fact, about ten new groups have 
joined the competition each year, and in 2006 a total of 30 participants was reached. 
The introduction of list questions, the possibility to return multiple answers, and the requirement of supporting 
the  answers  with  snippets  of  texts  from  the  relevant  documents  made  the  evaluation  process  more  difficult. 
Moreover, in some languages the large amount of data requiring assessment made it impossible  for the judging 
panels  to  correct  more  than  one  answer  per  question.  Therefore,  only  the  first  answers  were  evaluated  in  runs 
that had English and Spanish as a target. In all other cases at least the first three answers were evaluated.
Considering these issues, it was decided to follow the procedure utilised during the previous campaign. The files 
submitted by the participants in all tasks were manually judged by native speakers. Each language coordination 
group guaranteed the evaluation of at least one answer per question. 
If  a  group  decided  to  assess  more  than  one  answer  per  question,  the answers  were  assessed  in  the  order  they 
occurred in the submission file and the same number was applied to all questions, and all the runs assessed by 
the group. The exact answer (i.e. the shortest string of words which is supposed to provide the exact amount of 
information to answer the question) was assessed as:
· 
· 
· 
· </p>
      <p>
        R (Right) if correct;
W (Wrong) if incorrect;
X (ineXact) if contained less or more information than that required by the query;
U (Unsupported) if either the docid was missing or wrong, or the supporting snippet did not contain the 
exact answer. 
Most  assessor­groups  managed  to  guarantee  a  second  judgement  of  all  the  runs,  with  a  good  average  inter­ 
assessor  agreement.  As  far  as  the  evaluation  measures  are  concerned,  the  list  questions  had  to  be  scored 
separately,  and  different  groups  returned  a  different  number  of  answers  for  originally  meant  Factoid  and 
Definition questions. As a consequence, we decided to provide the following measures:
·  accuracy, as the main evaluation score, defined as the average of SCORE(q) over all 200 questions q;
·  the  mean  reciprocal  rank  (MRR)  over  N  assessed  answers  per  question.  That  is,  the  mean  of  the 
reciprocal of the rank of the first correct label over all questions;
·  the K1 measure used in earlier QA@CLEF campaigns [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
·  the  Confident  Weighted  Score  (CWS)  designed  for  systems  that  give  only  one  answer  per  question. 
Answers are in a decreasing order of confidence and CWS rewards systems that give correct answers at 
the top of the ranking [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] 
Although some  other kinds of measures have  been  proposed and used in CLEF 2005, such as a more detailed 
analysis/breakdown of bad answers by the Portuguese group .[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], they were not considered this year. Also, issues 
like providing more accurate description of what X means: too much or too little were only distinguished by the 
Portuguese assessors, argued for i.a. in Rocha and Santos [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. 
6  Result s
80 
70 
60 
50 
40 
30 
20 
10 
      </p>
      <p>0 
Best 
5
9
,
8
4 6
6
Here below a more detailed analyses of the results in each language follows, giving more specific information on 
the  performances  of  systems  in  the  single  sub­tasks  and  on  the  different  types  of  questions,  providing  the 
relevant statistics and comments.</p>
    </sec>
    <sec id="sec-4">
      <title>6.1  Bulgar ian as Ta r get </title>
      <p>At CLEF 2006 Bulgarian was addressed as a target language for the second time. This year there was no change 
in the number of the participants ­­ again two groups took part in the monolingual evaluation task with Bulgarian 
as a target language: BTB at Linguistic Modelling Laboratory, Sofia and The Joint Research Centre (JRC), Ispra. 
Three  runs  altogether  were  submitted  –  one  by  the  first  group  and  two  by  the  second  group  with  insignificant 
difference between them. The 2006 results are presented in Table 8 below. First, the correct answers in numbers 
and  percentage  are  given  (Right)  per  run.  Then  the  wrong  (W),  inexact  (X)  and  unsupported  answers  (U)  are 
shown  in numbers.  Further,  the number  of  the  factoids  (F),  temporally  restricted  questions  (T),  definitions  (D) 
and list questions (L) are given. Also, the percentage of the correct answers per each type is registered in Table 
8.  NIL  questions  are  presented  as  the  number  of  correctly  and  wrongly  returned  answers  by  the  systems  with 
NIL  marking.  It  is  obvious  that  the  systems  returned  NIL  answer  also  when  they  could  not  detect  a  possibly 
existing  answer  in  the  corpus  themselves.  In  our  opinion,  the  present  NIL  marking  might  be  divided  into  two 
labels: NIL = no answer in the corpus is existing and CANNOT = the system itself cannot find an answer. In this 
way  the  evaluation  would  be  more  realistic.  Main  reciprocal  rank  score  is  provided  in  the  last  column  of  the 
table. 
As it can be seen, this year the first system performs better. However, its overall accuracy is slightly worse with 
respect to the 2005 best accuracy result, achieved then by IRST, Trento. Now it is 26.60 %, while in 2005 it was 
27.50 %. However, the BTB 2005 year result was significantly improved. Both systems ‘crashed’ at temporally 
restricted  questions  with  no  single  match  (see  the  empty  slots  in  the  table).  It  is  a  step  back  from  2005,  when 
both  systems  had  some  hits,  best  of  which  scored  17.65  %.  List  questions  are  also  very  poorly  answered  (1 
correct answer per run). 
The  only  outperforming  results  in  comparison  with  the  last  year  are  the  following:  the  improvement  of  the 
definition type answers (from 42 % to 55.81 %) and the raise  of the main reciprocal rank score (from 0.160 to 
0.2660). 
The introduction of the snippet support proved out to be a good idea. There was only 1 unsupported answer in all 
three runs. 
The interannotator agreement was very high due to two reasons: first, the number of the answered questions was 
not  very  high,  and  second,  there  were  strict  guidelines  for  the  interpretation  of  the  answers,  based  on  our  last 
year experience. 
In spite of the somewhat controversial results from the participating systems this year, there is a lot of potential 
in  the  task  of  Bulgarian  as  a  target  language  in  several  aspects:  investing  in  the  development  of  the  present 
systems and creating new systems. We hope that Bulgarian will become even more attractive as an EU language. </p>
    </sec>
    <sec id="sec-5">
      <title>6.2  Dutch as Tar get </title>
      <p>This year three teams that took part in the CLEF QA track used Dutch as the target language: the University of 
Amsterdam, the University of Groningen and the University of Roma – 3, with six runs submitted in total: three 
Dutch monolingual and three crosslingual (English to Dutch). All runs were assessed by two assessors, with the 
overall  inter­assessor  agreement  0.96.  For  creating  the  gold  standard  for  Dutch,  the  assessments  were 
automatically  reconciled  in  favour  of  more  lenient  assessments:  for  example,  in  case  the  same  answer  was 
assessed  as  W  (incorrect)  by  one  assessor  and  as  X  (inexact)  by  another, the  X  judgement  was  included  in  the 
gold standard. The results of the evaluation of the six runs are provided in Tables 9 and 10. The columns labelled 
Right, W, X and U give the results for factoid, definition and temporally restricted questions.
An interesting thing to notice about this year’s task is that the overall scores of the systems are lower, compared 
to the last year’s numbers (44% and 50% of correct answers to factoid questions last year). This year’s questions 
were created by annotators who were explicitly instructed to think of “harder” questions, that is, involving 
paraphrases and some limited general knowledge reasoning. It would be interesting to compare the performance 
of this year’s systems on last year’s questions to the previous results of the campaign. </p>
    </sec>
    <sec id="sec-6">
      <title>6.3  English as Tar get </title>
      <p>Cr eation of Q uestions. The question for creation of the questions was very similar to last year and is now a well 
understood  procedure. This  year  it  was required  to  store  supporting  snippets  for  the reference  answers  but  this 
was not difficult and is well worth the trouble. As previously,  we  were requested to set Temporarily  Restricted 
questions  and  to  distribute  these  in  a  prescribed  way  over  the  various  Factoid  question  types  (PERSON, 
LOCATION etc). We achieved our quotas but this was extremely difficult to accomplish and we do not feel the 
time spent is worthwhile as the addition of temporal restrictions more than doubles the time taken to generate the 
questions.  On  the  other  hand,  as  the  restrictions  are  frequently  synthetic  in  nature,  our  knowledge  of  how  to 
solve these important questions does not necessarily advance from year to year. 
Searching for Definition questions (or indeed any questions beyond Factoids) is always very interesting work but 
the method of evaluation was not clarified this year. So, while the topics  we  selected do  follow the guidelines, 
we were not required to (or indeed able to) state at generation time exactly what a complete and correct answer 
should look like. In consequence we can not conclude much from an analysis of the answers returned by systems 
to such questions. 
Summar y  Statistics  for   all  the  Runs.  Overall,  thirteen  cross­lingual  runs  with  English  as  a  target  were 
submitted. The results are  shown  in.  Ten  groups  participated  in  seven  languages,  French,  German,  Indonesian, 
Italian, Romanian, Polish and Spanish. There were three groups for French, two  for Spanish and one for all the 
rest. 
Results  Analysis.  There  were  three  main  types  of  question  this  year,  Factoids,  Definitions  and  Lists  and  we 
consider the results over these types as well as considering the best scores overall. The most indicative measure 
overall is a simple count of correct answers and this is what we have used. For the 150 Factoids the best system 
was utjp061plen (Polish­English) with 132 correct. This is by far the best and is vastly higher than last year. By 
comparison,  the  top  five  are  utjp061plen  (132),  lire062fren  (39),  lire061fren  (33),  dltg061fren  (32)  and 
aliv061esen (29). The other results are not greatly different from last year. The top result of 132/150 amounts to 
88%. The next best result of 39/150 is 26%. 
For  the  40  definitions,  the  picture  is  similar.  The  top  five  results  are  utjp061plen  (32),  aliv062esen  (11), 
lire061fren (10), aliv061esen (9), lire062fren (9) and dfki061deen (8). Again, the top result is far higher than the 
rest amounting to 32/40 i.e. 80% with the next being 11/40 i.e. 28%.
For each of the ten list questions, a system could return up to ten candidate answers. Considering both a simple 
count of correct answers and the P@N score achieved, the top  five results by count are utjp061plen (18, 0.65), 
uaic061roen (10, 0.11), lire061fren (9, 0.09), irst061iten (8, 0.16), lire062fren (8, 0.08) and dfki061deen (6, 0.2). 
By either score, utjp061plen is the best while the ordering of the rest differs for the P@N score: utjp061plen (18, 
0.65), dfki061deen (6, 0.2) irst061iten (8, 0.16), uaic061roen (10, 0.11), lire061fren (9, 0.09) and lire062fren (8, 
0.08). 
There  were  considerable  practical  problems  with  the  assessment  of  runs  this  year.  Firstly,  several  runs  used 
invalid run tags. Secondly two of the runs were answering the questions in a completely different order! Thirdly, 
one  question  in  these  two  runs  was  different  from  the  question  being  answered  by  the  other  systems  in  that 
position. Fourthly, one run had the fields in the wrong order. Fifthly  one run used NULL instead of NIL  while 
another  run  used  nil.  Luckily  we  spotted  problems  2  and  3  and  were  able  to  correct  them  and  indeed  all  the 
others but this was extremely time consuming and difficult. 
As  in  all  previous  years  the  runs  were  anonymised  by  a  third  party  so  none  of  the  assessors  knew  either  the 
origin of a run or the original source language. 
This  year  it  had  been  decided  to  allow  multiple  answers  to  Factoid  and  Definition  questions  (up  to  ten  per 
question).  The  rationale  for  this  was  never  quite  clear  since  the  whole  objective  of  Question  Answering  (as 
against Information Retrieval) is to return only the right answer. Even in cases where there are genuinely several 
right  answers  (a  rare  situation  in  our  carefully  designed  question  sets)  a  system  should  still  return  a  correct 
answer  in  the  first  place.  For  this  reason  and  due  to  our  limited  time  and  resources,  we  only  judged  the  first 
answer returned to Factoid and Definition questions. For List questions, all candidate answers were judged, as is 
normal at TREC. 
For  the  questions  double  judged,  we  measured  the  agreement  level.  There  were  149  differences  over  thirteen 
runs  of  100  questions.  This  amounts  to  149/1300 i.e.  11% disagreement  or 89% agreement. The  overall  figure 
for last year was 93%. </p>
    </sec>
    <sec id="sec-7">
      <title>1  This result is still under verification.</title>
      <p>Concerning the judgement process itself, Factoids and Lists did not present a problem as we were very familiar 
with them. On the other hand Definitions were in the same state as last year in that they had been included in the 
task without a suitable evaluation prodedure having been defined. In consequence we used the same approach as 
last  year:  If  an  answer  contained  information  relevant  to  the  question  and  also  contained  no  irrelevant 
information,  it  was  judged  R  if  supported,  and  U  otherwise.  If  both  relevant  and  irrelevant  information  was 
present it was judged X. Finally, if no relevant information was present, the answer was judged W. 
Comment  and  Conclusions.  The  number  of  runs  judged  (13)  was  similar  to  last  year  (12).  However,  three 
source languages were introduced: Indonesian, Polish and Romanian. The results themselves  were also broadly 
similar with the exception of the Polish run which was vastly higher on all question types. 
Definition  questions  remained  in  the  same  unspecified  state  as  previously.  This  means  that  we  have  not  been 
successful  in  stretching  the  boundaries  of  question  answering  beyond  Factoids  which  are  now  very  well 
understood. This is a great pity as the extraction of useful 'definition type' information on a topic is a very useful 
task for groups to study but it is one which needs to be carefully quantified. 
The  introduction  of  snippets  was  very  helpful  at  question  generation  time  and  also  invaluable  for  judging  the 
answers. Snippets are a great step forward for CLEF and are the most significant development for the QA Track 
this year. </p>
    </sec>
    <sec id="sec-8">
      <title>6.4  Fr ench as Tar get </title>
      <p>This year (as last year) seven groups took part in evaluation tasks using French as target language: four French 
groups:  Laboratoire  d’Informatique  d’Avignon  (LIA),  CEA­List,  Université  de  Nantes  (LINA)  and  Synapse 
Développement; one Spanish group: Universitat Politécnica de Valencia; one Japanese group; and one American 
group: LCC. 
In total, 15 runs have been returned by the participants: eight monolingual runs (FR­to­FR) and seven bilingual 
runs (6 EN­to­FR, 1 PT­to­FR). 
It appears that the number of participants for the French task is the same that last year but it’s the first time there 
are non­European participants. This shows there is a new major interest for the French as target language. 
Two groups submitted four runs, two other groups submitted two runs and three groups submitted only one run. 
This  year  and  for  the  first  time,  the  participants  could  return  up  to  10  answers  per  question.  A  major  part  of 
participants  returned  only  one  answer  per  question,  only  three  groups  returned  more  than  one  answer  per 
question. For these three groups, ELDA (Evaluation and Language resources Distribution Agency) assessed the 
three first answers for Factual, Definition and Temporally restricted questions. 
For the monolingual task, the best system returned 67.89 % of correct answers (overall accuracy in 1st rank). We 
can  observe  this  system  obtained  better  results  for  definition  questions  (83.33  %)  than  for  Factoid  questions 
(63.51 %). 
The  LIA’  system,  which reached  the  second  position in this  task, returned  46.32  %  of  correct answers  (overall 
accuracy in 1st rank). We can also observe the difference between the results for the Factual questions and the 
results  for  the  Definition  questions: 37.84  %  of  correct answers  for  the  Factual  and  76.19  %  for  the  Definition 
questions. 
For  the  bilingual  task,  the  best  system  obtained  45.26  %  of  correct  answers  as  opposed  to  34.74  %  of  correct 
answers for the LIA’ system. 
We  can  remark  that  the  best  system  for  the  bilingual  task  (EN­to­FR)  obtained  worse  results  than  the  second 
system for the monolingual task. 
This year, before the assessment, the French assessors determined some rules to face up to problems encountered 
the last year. 
Concerning Temporally restricted questions for example, to assess an answer as “Correct”, the date, the period 
or the event had to be present in the document returned by the systems. 
They  decided  also  to  check  separately,  at  the  end  of  the  assessment,  some  questions  which  seemed  difficult  to 
them, to make sure that each answer had received the same “treatment” during the evaluation. 
The  main  problem  encountered  this  year,  was  related  to  the  assessment  of  the  List  questions.  This  was  a  new 
kind of questions this year and participants followed different ways to answer to these questions. Some systems 
returned  a  list  of  answers  in  a  same  line;  others  returned  an  answer  per  line.  ELDA  evaluated  these  answers 
according  to  each run  (if  a  line  contained  one  of  correct  answers  or  all  the  correct  answers,  these  answers  had 
been assessed as “Correct”.</p>
      <p>The best system obtained 5 correct answers out of 10 List questions in total. 
We can observe that the results for the List questions were not very relevant because of not much questions and 
not much rules. 
In  conclusion,  this  year,  a  system  obtained  “excellent”  results.  Synapse  Développement  obtained  129  correct 
answers out of 200 (as opposed to 128 last year). 
This system is the best system for the French language. This year, it’s again the dominant system. 
In  addition,  we  can  observe  the  same  great interest in  Question  Answering  from the  European (and now  non­ 
European) research community for the tasks using French as target language. 
6.5  Ger man as Ta r get 
Three research groups submitted runs for evaluation in the track having German as target language: The German 
Research Center for Artificial Intelligence (DFKI), FernUniversität Hagen (FUHA) and The Institute for Natural 
Language Processing in Stuttgart (IMS). All of them provided system runs for the monolingual scenario and just 
one group (DFKI) submitted runs for the cross­language English­German scenario. Two assessors with different 
profiles  conducted  the  evaluation:  a  native  German  speaker  with  little  knowledge  of  QA  systems  and  a 
researcher with advanced knowledge of QA systems and a good command of German. Compared to the previous 
editions  of  the  evaluation  forum, this  year  an increase  in the  performance  of  an  aggregated  virtual  system  for 
both monolingual and cross­language tasks was registered, as well as for the cross­language best system’s result 
(Figure 4). Given the increased complexity of the task (no question type provided, supporting snippets required) 
and of questions (definition and list), the stability of the best monolingual results can be considered also a gain in 
terms of performance. 
2006 
2005 </p>
      <p>2004 </p>
      <sec id="sec-8-1">
        <title>Figur e 4: Results evolution </title>
        <p>Except for FUHA, the other two groups provided more than one possible answer per question, of which only the 
first three were manually evaluated. In order to come up with a measure of performance for systems providing 
several  answers  per  question,  Mean  Reciprocal  Rank  (MRR)  over  right  answers  has  been  considered  for  this 
purpose. 
Two things can be concluded from the answer distribution of Table 14: first, there are a fair number of inexact 
and unsupported answers that show performance could be improved with a better answer extraction; second, the 
fair  number  of  right  answers  among  the  second  and  third  ranked  positions  indicate  that  there  is  still  place  for 
improvements with a more focused answer selection. 
Run ID 
# 
The details of systems’ results can be seen in Table 15, in which the performance measures has been computed 
only for the first ranked answers to each question, except for the list questions. Interesting to observe is that none 
of the systems managed to correctly respond any temporal question. 
Table  16  describes  the  inter­rater  disagreement  on  the  assessment  of  answers  in  terms  of  question  and  answer 
disagreement.    Question  disagreement  reflects  the  number  of  questions  on  which  the  assessors  delivered 
different judgments and answer disagreement is a figure of the total number of answers disagreed on. Along the 
total  figures  for  both  types  of  disagreement,  a  breakdown  at  the  question  type  level  (Factoid,  Definition,  List) 
and at the  assessment  value  level  (ineXact,  Unsupported, Wrong/Right) is  listed. The  answer  disagreements  of 
type  Wrong/Right  are trivial  errors  during  the assessment  process  when a right  answers  was  considered  wrong 
by mistake and the other way around, while those of type X or U reflect different judgments whereby an assessor 
considered an answer inexact or unsupported while the other marked it as right or wrong.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>6.6  Italian as Tar get </title>
      <p>Two  groups  participated  in  the  Italian monolingual  task,  ITC­irst  and  the  Universidad  Politécnica  de  Valencia 
(UPV); while one group, the Università La Sapienza di Roma, participated in the cross­language EN­IT task. In 
total, five runs were submitted.</p>
      <p>437 
476 
198 
198 
432 
436 
405 
402 
B
i
l
i
n
g
u
a
l</p>
      <p>Figur e 4: Best and Aver age per for mance in the Monolingual and Bilingual tasks 
For the first time a cross­language task with Italian as target was chosen to test a participating system. 
The best performance in the monolingual task was obtained by the UPV, which achieved an accuracy of 28.19%. 
Almost the same result was recorded last year (see Figure 4). The average accuracy in the monolingual task was 
26.41%, which is an improvement of more than 2% with respect to last year’s results. 
The accuracy in the bilingual task was 17.02%, achieved by both submitted runs. 
During the years the overall accuracy has steadily decreased starting from a 25.17% in the 2004, we reached a 
24.08% in the 2005 and 22.06% this year. This could be partly due to newcomers – who usually get lower scores 
– and first experiments with bilingual tasks.
7 
6 
4 
5 
0 
0 
5 
6 
0 
3 
0 
0 
2 
2 
1 
0 
35 
28 
12 
13 
15 
17 
26 
27 
28 
19 
8 
8 
13 
15 
20 
21 
M
o
n
o
28,19
26,41
44 
40 
11 
12 
30 
28 
33 
35 
17,02</p>
      <p>B
i
l
i
n
g
u
a
l
From the results shown in Table 17, it can be seen that the Universidad Politécnica de Valencia (UPV) submitted 
two runs in the monolingual task and achieved the best overall performance. The accuracy over Definition and 
Factoid  questions  ranged  from  26.83%  to  29.27%.  ITC­Irst  submitted  one  run,  and  achieved  much  better 
accuracy  over  Factoid  questions  (25.00%) than  over  Definition questions  (17.07%).  As  previously  mentioned, 
the  Università  La  Sapienza  di  Roma  submitted  two  runs  in  the  cross­language  EN­IT  tasks,  performing much 
better in the Definition questions (24.39%) than in the Factoid questions (15.28%). 
As far as List questions are concerned, all participating systems performed rather poorly, with a P@N ranging 
from  0.08  to  0.17.  This  implies  that  a  more  in­depth  research  on  these  questions  and  the  measures  for  their 
evaluation is still needed. 
Temporally restricted questions represented a challenge for the systems, which generally achieved a lower than 
average accuracy in this sub­category. The Universidad Politécnica de Valencia achieved the best performance 
of 23.68% (see Table 18). 
The evaluation process did not presented particular problems, although it was more demanding than usual 
because of the necessity to check the supporting text snippet. All runs were anyway assessed by two judges. The 
inter­assessor agreement was averagely 90,14 %, most disagreement being between U and X. A couple of cases 
of disagreement between R and W were due just to trivial mistakes. </p>
    </sec>
    <sec id="sec-10">
      <title>6.7  Por tuguese as Ta r get </title>
      <p>This year five research groups took part in tasks with Portuguese as target language, submitting ten runs: seven 
in  the  monolingual  task,  two  with  English  as  source,  and  one  with  Spanish.  Two  new  groups  joined  for 
Portuguese:  University  of  Porto,  and  Brazilian  NILC,  while  LCC  participated  with  an  English­Portuguese  run 
only. Universidade de Évora did not participate this year. </p>
      <p>Table  19  presents  the  overall  results  concerning  the  188  non­list  questions.  We  present  values  both 
taking  into  account  only  the  first  answer  to  each  question,  and  –  for  the  only  system  where  this  makes  any 
We  also  provide  in  Table  20  the  overall  accuracy  considering  (and  evaluating)  independently  all  different 
answers provided by the systems. </p>
      <p>X+ 
(#) 
7 
6 
1 
0 
6 
0 
2 
3 
2 
3 
2 </p>
      <p>R
(#) 
50 
46 
0 
3 
134 
36 
42 
A virtual run, called combination, was included in Table 21 and computed as follows: if any of the participating 
systems  found  a  right  answer,  it  is  considered  right  in  the  combination  run.  Ideally,  this  combination  run 
measures  the  potential  achievement  of  cooperation  among  all  participants.  However,  for  Portuguese  this 
combination  does  not  significantly  outperform  the  best  performance:  Priberam  alone  corresponds  to  92.4%  of 
the combination run. </p>
      <p>We  have  also  analysed  the  size  in  words  of  both  answers and  justification  snippets,  as  displayed  in  Table 
22.  (Computations  were  made  excluding  NIL  answers.)  Interestingly,  Priberam  provided  the  shortest 
justifications. 
In Table 23, we compare the accuracy of the systems for the 22 temporally restricted questions in the Portuguese 
question set with their scores for non­temporally restricted ones and their overall performance.
Finally, a total of twelve questions were defined by the organization as requiring a list as proper answer. The fact 
that  the  systems  had  to  find  out  whether multiple  or  single answers  were  expected  was  a  new  feature  this  year 
and  was  not  conveniently  handled  by  most  systems.  In  fact,  two  systems  (Priberam  and  NILC)  completely 
ignored  this  and  provided  a  single  answer  to  every  question,  while  two  other  systems,  although  attempting  to 
deal with list questions, seemed to fail in appropriately identifying them: RAPOSA (UPorto) provided multiple 
answers  only  to  non­list  questions,  and  Esfinge  produced  12  answers  for  ten  questions.  In  fact,  only  LCC 
presented  multiple  answers  systematically,  yielding  an  average  of  7.32  answers  per  question,  while  no  other 
group exceeded 1.1. </p>
      <p>We  believe  further  study  should  be  devoted  to  the  list  questions  for  the  next  years,  since  a  distinction 
between closed lists and open lists, although acknowledged, was not properly taken into consideration. We have 
thus chosen to handle all these questions alike, assigning them the following accuracy score: number of correct 
answers (where X counted as ½) divided by the sum of the number of existing answers in the collections and the 
number of wrong distinct answers provided by the system. The results are displayed in Table 24. </p>
      <p>For  the  case  of  closed  lists  (where  "one"  answer  might  bring  all  answers,  such  as  "Lituânia,  Estónia  e 
Letónia"), we still counted the number of answers individually (3). </p>
      <sec id="sec-10-1">
        <title>Table 24: Results for  List questions </title>
        <p>Known 
Question  answers </p>
        <p>esfg 
061ptpt </p>
        <p>esfg 
062ptpt </p>
        <p>nilc 
061ptpt </p>
        <p>nilc 
062ptpt </p>
        <p>prib 
061ptpt 
uporto 
061ptpt 
uporto 
062ptpt </p>
        <p>esfg  lcc  esfg 
061enpt  061enpt  061espt 
205 
399 
400 
759 
770 
784 
785 
786 
795 
score 
3 
3 
3 
3 
3 
5 
3 
3 
5 
0/1 
0/1 
The  participation  at  the  Spanish  as  Target  subtask  is  still  growing.  Nine  groups,  two  more  than  the  last  year, 
submitted 17 runs: 12 monolingual, 3 from English, 1from French and 1 from Portuguese. Table 25 and Table 26 
show  the  summary  of  systems  results  for  monolingual  and  cross­lingual  respectively.  The  number  of  Right,
105 
102 
80 
72 
70 
57 
56 
56 
41 
39 
37 
27 </p>
        <p>W 
%  </p>
        <p>
          X 
52,50  86  4 
51,00  86  3 
40,00  112  3 
36,00  105  15 
35,00  119  5 
28,50  123  6 
28,00  123  8 
28,00  132  6 
20,50  148  4 
19,50  146  6 
%  
F 
# 
5 
9 
5 
8 
6 
14 
13 
6 
7 
9 
%  T 
[108] 
%  D 
[40] 
55,56  30,00 
47,22  35,00 
32,41  25,00 
38,89  22,50 
37,04  25,00 
27,78  25,00 
29,63  22,50 
26,85  25,00 
21,30  17,50 
16,67  17,50 
Wrong  (W),  Inexact  (X)  and  Unsupported  (U)  answers.  Tables  show  also  the  accuracy  (in  percentage)  of 
factoids (F), factoids with temporal restriction (T), definitions (D) and list questions (L). Best values are marked 
in  bold  face.  Best  performing  systems  have  improved  their  performance  (as  seen  in  Figure  5),  mainly  with 
respect to factoids. However, performance when the question has a temporal restriction didn’t vary significantly. 
Last year, the answering of definitions with respect to persons and organizations was almost solved. In spite of 
the  fact  that  this  year  the  set  of  definition  questions  was  more  realistic  systems  have  improved  slightly  their 
performance. 
List questions have been introduced this year so they deserve some attention regarding their evaluation. We have 
differentiated  two  types  of  list  questions:  conjunctive  and  disjunctive  (as  presented  in  [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]).  Conjunctive  list 
questions are asking for a set of items and they are Right  if all the items are present in the answer. For example, 
“ Nombre los tres Beatles que siguen vivos”  (Name the three Beatles alive). Disjunctive list questions are asking 
for  an  undetermined  number  of  items.  For  example,  “ Nombre  luchadores  de  Sumo”   (Name  Sumo  fighters). 
Only the first answer of each system has been evaluated in both cases. 
        </p>
        <p>Right 
%  </p>
        <p>W 
# </p>
        <p>X 
# </p>
        <p>U  %  F 
#  [108] 
%  T 
[40] 
Regarding the NIL questions, Table 25 and 26 show the harmonic mean (F) of precision (P) and recall (R). The 
best performing systems have increased again their performance (see Table 27) in NIL questions. The correlation 
efficient  r  between  the  self­score  and  the  correctness  of  the  answers  has  been  increased  in  the  majority  of 
systems, although results are not good enough yet. 
This year a supporting text snippet was requested. For this reason, we have evaluated the systems  capability to 
extract  the  answer  when  the  snippet  contains  it. The  last  column  of Tables  25  and  26  shows  the  percentage  of 
cases where the correct answer was correctly extracted. This information is very useful to diagnose if the lack of 
performance is due to the passage retrieval or to the answer extraction.</p>
        <p>100,00 
90,00 
80,00 
70,00 
60,00 
50,00 
40,00 
30,00 
20,00 
10,00 
0,00 </p>
        <p>42,00 
32,50 </p>
        <p>52,5 
24,50 
24,50 </p>
        <p>55,56 
31,11 29,66 
80,00 </p>
        <p>83,33 
70,00 
2003 
2004 
2005 
2006 
Best Overall Acc. % </p>
        <p>Best in Factoids % </p>
        <p>Best in Definitions % </p>
        <p>Figur e 5: Evolution of best per for ming systems 2003­2006 
Regarding Cross­Lingual runs,  it is  worth to  mention  that  Priberam has  achieved  in  the  Portuguese  to  Spanish 
task a result comparable to the monolingual runs. </p>
        <p>Table 27: Evolution of best r esults in NIL questions </p>
      </sec>
      <sec id="sec-10-2">
        <title>Year   F­measur e </title>
        <p>2003  0,25 
2004  0,30 
2005  0,38 
2006  0,46 
All the  answers have  been  assessed  anonymously  considering  all  systems’  answers  simultaneously  question  by 
question.  The  inter­annotator  agreement  was  evaluated  over  985  answers  assessed  by  the  two  judges.  Only  a 
2.5% of the judgements were different and the resulting kappa value was 0.93. 
7 </p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>Conclusions </title>
      <p>The  QA  track  at  CLEF  2006  has  once  again  demonstrated  the  interest  for  Question  Answering  in  languages 
other than English. In fact, both the number of participants and runs submitted has grown, following the positive 
trends  of  the  previous  campaign. Equally  positive  was  the  fact  that, despite the  loss  of  Finnish,  two additional 
languages from Eastern Europe have been added, strengthening the cross­linguality of QA@CLEF. 
The  balance  between  tradition  and  innovations  –i.e  the  introduction  of  list  questions  and  supporting  text 
snippets­  has  proved  to  be  a  good  solution,  which  allows  both  new­comers  and  veterans  to  test  their  systems 
against  adequately  challenging  tasks  and,  at  the  same  time,  to  make  a  comparison  with  previous  exercises. 
Generally speaking, the results recorded an improvement in performance, with best accuracy significantly higher 
than in previous campaigns both in monolingual and bilingual tasks. 
As far as the organisation of the campaign is concerned, the introduction of new elements such as list questions 
and  supporting  snippets  has  implied  a  significant  increase  of  work  both  in  the  question  collection  and  in  the 
evaluation  phase,  which  was  particularly  demanding  for  language  groups  which  had  a  great  number  of 
participants.  A  better  distribution  of  the  workload  and  solutions  to  speed  up  the  evaluation  process,  also  with 
automatic assessment of part of the submissions will be essential in next campaigns. 
A future perspective of QA is certainly outlined by the two pilot tasks offered in 2006­i.e. AVE and WiQa­, the 
latter in particular representing a significant step toward a more realistic scenario, where queries are carried out 
on the Web. For these reasons, a quick integration of these experiments into the main task is hoped for.
The authors would like to thank Donna Harman for her valuable feedback and advice, and Diana Santos for her 
precious contribution in the organization of the campaign and the revision of this paper. 
Paulo  Rocha  is  thankful  to  the  many  useful  comments  and  overall  discussion  with  Diana  Santos  for  the 
Portuguese part. 
Paulo  Rocha  was  supported  by  the  Portuguese  Fundação  para  a  Ciência  e  Tecnologia  within  the  Linguateca 
project, through grant POSI/PLP/43931/2001, co­financed by POSI. 
Bogdan Sacaleanu was supported by the German Federal Ministry of Education and Research (BMBF) through 
the projects HyLaP and COLLATE II. </p>
    </sec>
    <sec id="sec-12">
      <title>Refer ences </title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.  QA@CLEF 2006 Organizing Committee. Guidelines 
          <year>2006</year>
          . http://clef­qa.itc.it/guidelines.html 
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>2.  WiQA Website: http://ilps.science.uva.nl/WiQA/ </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>3.  AVE Website: http://nlp.uned.es/QA/AVE/ </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.  Herrera, J., Peñas A., Verdejo, F.: Question answering pilot task at CLEF 
          <year>2004</year>
          . In: Peters, C., Clough, P.,  Gonzalo,  J.,  Jones,  Gareth 
          <string-name>
            <surname>J.F.</surname>
          </string-name>
          ,  Kluck,  M.,  Magnini,  B.  (eds.):  Multilingual  Information  Access  for  Text,  Speech  and  Images.  Lecture  Notes  in  Computer  Science,  Vol. 
          <volume>3491</volume>
          .  Springer­Verlag,  Berlin  Heidelberg  New York  (
          <year>2005</year>
          ) 
          <fpage>581</fpage>
          -
          <lpage>590</lpage>
           
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.  Magnini, B.,
          <string-name>
            <surname>Vallin</surname>
          </string-name>
          , A., Ayache, C., Erbach, G., Peñas, A., de Rijke, M., Rocha, P., Simov, K., Sutcliffe, R.:  Overview of the CLEF 
          <year>2004</year>
           Multilingual Question Answering Track. In: Peters, C., Clough, P., Gonzalo, J.,  Jones,  Gareth 
          <string-name>
            <surname>J.F.</surname>
          </string-name>
          ,  Kluck,  M.,  Magnini,  B.  (eds.):  Multilingual  Information  Access  for  Text,  Speech  and  Images.  Lecture  Notes  in  Computer  Science,  Vol. 
          <volume>3491</volume>
          .  Springer­Verlag,  Berlin  Heidelberg  New  York  (
          <year>2005</year>
          ) 
          <fpage>371</fpage>
          ­
          <lpage>391</lpage>
           
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.  Rocha,  Paulo  &amp; 
          <string-name>
            <surname>Diana</surname>
          </string-name>
            Santos:  CLEF: 
          <article-title>Abrindo  a  porta  à  participa  internacional  em  avaliação  de  RI  do  português</article-title>
          .  In:  Diana  Santos  (ed.):
          <article-title>  Avaliação  conjunta:  um  novo  paradigma  no  processamento  computacional da língua portuguesa</article-title>
          . IST Press, Lisbon, 
          <year>2006</year>
           (in press). 
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.  Santos,  D.,  Rocha,  P.:  The  Key  to  the  First  CLEF  with  Portuguese:  Topics,  Questions  and  Answers  in  CHAVE.  In:  Peters,  C.,  Clough,  P.,  Gonzalo,  J.,  Jones,  Gareth 
          <string-name>
            <surname>J.F.</surname>
          </string-name>
          ,  Kluck,  M.,  Magnini,  B.  (eds.):  Multilingual  Information  Access  for  Text,  Speech  and  Images.  Lecture  Notes  in  Computer  Science,  Vol. 
          <volume>3491</volume>
          . Springer­Verlag, Berlin Heidelberg New York  (
          <year>2005</year>
          ) 
          <fpage>821</fpage>
          ­
          <lpage>832</lpage>
          . 
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.  Spark Jones, K.: Is 
          <article-title>question answering a rational task?</article-title>
           In: Bernardi, R., Moortgat, M. (eds): Questions and  Answers:  Theoretical  and  Applied  Perspectives.  Second  CoLogNETElsNET  Symposium.  Amsterdam  (
          <year>2003</year>
          ) 
          <fpage>24</fpage>
          -
          <lpage>35</lpage>
           
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.  Vallin, Alessandro,  Bernardo Magnini, Danilo Giampiccolo, Lili Aunimo, Christelle Ayache, Petya Osenova, Anselmo  Peñas, Maarten de Rijke, Bogdan Sacaleanu, Diana Santos, Richard Sutcliffe: Overview of the CLEF 2005 Multilingual  Question  Answering  Track.  In:  Cross  Language  Evaluation  Forum:  Working  Notes  for  the  CLEF  2005  Workshop  (CLEF 
          <year>2005</year>
          ) (Vienna, Áustria, 
          <fpage>21</fpage>
          ­
          <lpage>23</lpage>
           September 
          <year>2005</year>
          . 
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.  Voorhees,  E.  M.
          <article-title>:  Overview  of  the  TREC </article-title>
          <year>2002</year>
            Question  Answering  Track.  In:  Voorhees,  E.  M.  and  Buckland,  L.  P.  (eds),  Proceedings  of  the  Eleventh  Text  Retrieval  Conference  (TREC 
          <year>2002</year>
            NIST  Special  Publication 
          <fpage>500</fpage>
          ­
          <lpage>251</lpage>
          , Washington DC (
          <year>2002</year>
          ) 
          <volume>115</volume>
           
          <fpage>123</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>