<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>OVERVIEW OF THE CLEF 2007 MULTILINGUAL QUESTION ANSWERING TRACK</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Danilo Giampiccolo</string-name>
          <email>giampiccolo@celct.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anselmo Peñas</string-name>
          <email>anselmo@lsi.uned.es</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christelle Ayache</string-name>
          <email>ayache@elda.fr</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dan Cristea</string-name>
          <email>dcristea@info.uaic.ro</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pamela Forner</string-name>
          <email>forner@celct.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valentin Jijkoun</string-name>
          <email>jijkoun@science.uva.nl</email>
          <xref ref-type="aff" rid="aff7">7</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Petya Osenova</string-name>
          <email>petya@bultreebank.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paulo Rocha</string-name>
          <email>Paulo.Rocha@alfa.di.uminho.pt</email>
          <xref ref-type="aff" rid="aff8">8</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bogdan Sacaleanu</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Richard Sutcliffe</string-name>
          <email>richard.sutcliffe@ul.ie</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>BTB</institution>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>CELCT</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy (</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>DFKI</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>DLTG, University of Limerick</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Departamento de Lenguajes y Sistemas Informáticos, UNED</institution>
          ,
          <addr-line>Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>ELDA/ELRA</institution>
          ,
          <addr-line>Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>Faculty of Computer Science, University “Al. I. Cuza” of Iaşi, Romania Institute for Computer Science, Romanian Academy</institution>
          ,
          <addr-line>Iaşi</addr-line>
          ,
          <country country="RO">Romania</country>
        </aff>
        <aff id="aff7">
          <label>7</label>
          <institution>Informatics Institute, University of Amsterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff8">
          <label>8</label>
          <institution>Linguateca</institution>
          ,
          <addr-line>SINTEF ICT</addr-line>
          ,
          <country>Norway and Portugal</country>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>23</lpage>
      <abstract>
        <p>The fifth QA campaign at CLEF, the first having been held in 2006. was characterized by continuity with the past and at the same time by innovation. In fact, topics were introduced, under which a number of Question-Answer pairs could be grouped in clusters, containing also co-references between them. Moreover, the systems were given the possibility to search for answers in Wikipedia. In addition to the main task, two other tasks were offered, namely the Answer Validation Exercise (AVE), which continued last year's successful pilot, and QUAST, aimed at evaluating the task of Question Answering in Speech Transcription. As general remark, it must be said that the task proved to be more difficult than expected, as in comparison with last year's results the Best Overall Accuracy dropped from 49,47% to 41,75% in the multilingual subtasks, and, more significantly, from 68,95% to 54% in the monolingual subtasks. The fifth QA campaign at CLEF [1], the first having been held in 2003, was characterized by continuity with the past, maintaining the focus on cross-linguality and covering as many European languages as possible (with the addition of Indonesian); and by innovation 1) by introducing a number of Question-Answer pairs, grouped in clusters, which referred to a same topic and which contained co-references between them, and 2) by giving the possibility to search for answers in Wikipedia. In this way, the newcomers had the possibility to test themselves with the classic task, and those who had participated in the previous campaigns had a new challenging factor to test their systems. In addition to the main task, two other tasks were offered, namely the Answer Validation Exercise (AVE), which continued last year's successful pilot, and the Question Answering for Speech Transcripts (QAST), aimed at evaluating the task of Question Answering in Speech Transcription. In the following sections, the main task and its preparation will be described. A presentation of the participants and the runs submitted will be also given, together with a description of the evaluation method and the results achieved.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Following the procedure consolidated in previous years, in the 2007 campaign several different tasks were
proposed:
1. a main task, divided into several monolingual and bi-lingual sub-tasks;
2. the Answer Validation Exercise (AVE), which continued the successful experiment proposed in 2006.</p>
      <p>
        Systems were required to emulate human assessment of QA responses and decide whether an Answer to
a Question is correct or not according to a given Text. Participating systems were given a set of triplets
(Question, Answer, Supporting Text) and they had to return a boolean value for each triplet. Results
were evaluated against the QA human assessments [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ];
3. the QA Answering on Speech Transcripts (QAST), a pilot task which aimed at providing a framework in
which factual. Relevant points of this pilot were:
a. Comparing the performances of the systems dealing with both types of transcriptions.
b. Measuring the loss of each system due to the state of the art ASR technology.
c. In general, motivating and driving the design of novel and robust factual QA architectures for
automatic speech transcriptions [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The AVE and QAST tasks are described in details in dedicated papers in this Working Notes.
As far as the main task is concerned, the consolidated procedure was followed, although some relevant
innovations were introduced.</p>
      <p>The systems were given a set of 200 questions -which could concern facts or events (F-actoid questions),
definitions of people, things or organisations (D-efinition questions), or lists of people, objects or data (L-ist
questions)- and were asked to return one exact answer, where exact meant that neither more nor less than the
information required was given. Following the example of TREC, this year the exercise consisted of
topicrelated questions, i.e. clusters of questions which were related to the same topic and possibly contained
coreferences between one question and the others. Neither the question types (F, D, L) or the topics were given to
the participants.</p>
      <p>The answer needed to be supported by the docid of the document in which the exact answer was found, and by
portion(s) of text, which provided enough context to support the correctness of the exact answer. Supporting
texts could be taken from different sections of the relevant documents, and had to sum up to a maximum of 700
bytes. There were no particular restrictions on the length of an answer-string, but unnecessary pieces of
information were penalized, since the answer was marked as ineXact. As in previous years, the exact answer
could be exactly copied and pasted from the document, even if it was grammatically incorrect (e.g.: inflectional
case did not match the one required by the question). Anyway, this year systems were also allowed to use NL
generation in order to correct morpho-syntactical inconsistencies (e.g., in German, changing "dem Presidenten"
into "der President" if the question implies that the answer is in Nominative case), and to introduce grammatical
and lexical changes (e.g., QUESTION: What nationality is X? TEXT: X is from the Netherlands =&gt; EXACT
ANSWER: Dutch).
As customary in recent campaigns, a monolingual English (EN) task was not available as it seems to have been
already thoroughly investigated in TREC campaigns. English was still both source and target language in the
cross-language tasks.</p>
      <p>
        As the format is concerned, this year both input and output files were formatted as an XML file (for more details
see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]).
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Test Set Preparation</title>
      <p>The procedure followed to prepare the test set was much different from that used in the previous campaigns.
First at all, each organizing group, responsible for a target language, freely chose a number of topics. For each
topic, one to four questions were generated. Topics could be not only named entities or events, but also other
categories such as objects, natural phenomena, etc. (e.g. George W. Bush; Olympic Games; notebooks;
hurricanes; etc.). The set of ordered questions were related to the topic as follows:</p>
      <p>The topic was named either in the first question or in the first answer
The following questions can contain co-references to the topic expressed in the first question/answer
pair.</p>
      <p>Topics were not given in the test set, but could be inferred from the first question/answer pair. For example, if
the topic was George W. Bush, the cluster of questions related to it could have been:</p>
      <sec id="sec-2-1">
        <title>Q1: Who is George W. Bush?</title>
      </sec>
      <sec id="sec-2-2">
        <title>Q2: When was he born?</title>
      </sec>
      <sec id="sec-2-3">
        <title>Q3: Who is his wife?</title>
        <sec id="sec-2-3-1">
          <title>The Table 3: Document collections used in CLEF 2007.</title>
          <p>The questions in the set were numbered from 1 to 200, with no indication about whether they were part of a
cluster belonging to the same topic.</p>
          <p>Another major innovation of this year’s campaign concerned the corpora at which the questions were aimed at.
In fact, beside the data collections composed of news articles provided by ELRA/ELDA, also Wikipedia was
considered, capitalizing on the experience of the WiQA pilot task proposed in 2006. The Wikipedia pages in the
target languages, as found in the version of the Wikipedia of November, 2006 could be used. XML and the
HTML versions were available for download, even though any other versions of the Wikipedia files could be
used as long as they dated back to the end of November / beginning of December 2006. All the answers to the
questions had to be taken from "actual entries" or articles of Wikipedia pages - the ones whose filenames
normally correspond to the topic of the article. Other types of data (“image”, “discussion”, “category”,
“template”, “revision histories”, any files with user information, and any “meta-information” pages), had to be
excluded.</p>
          <p>As far as the question types are concerned, as in previous years of QA@CLEF, the three following categories
were still considered:
a) Factoid questions, fact-based questions, asking for the name of a person, a location, the extent of something,
the day on which something happened, etc.</p>
          <p>We consider the following 8 answer types for factoids:</p>
          <p>PERSON, e.g.</p>
          <p>TIME, e.g.</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>Q: Who was called the “Iron-Chancellor”?</title>
      </sec>
      <sec id="sec-2-5">
        <title>A: Otto von Bismarck.</title>
      </sec>
      <sec id="sec-2-6">
        <title>Q: What year was Martin Luther King murdered? A: 1968.</title>
      </sec>
      <sec id="sec-2-7">
        <title>LOCATION, e.g. Q: Which town was Wolfgang Amadeus Mozart born in?</title>
      </sec>
      <sec id="sec-2-8">
        <title>A: Salzburg.</title>
      </sec>
      <sec id="sec-2-9">
        <title>ORGANIZATION, e.g. Q: What party does Tony Blair belong to?</title>
      </sec>
      <sec id="sec-2-10">
        <title>A: Labour Party.</title>
      </sec>
      <sec id="sec-2-11">
        <title>MEASURE, e.g. Q: How high is Kanchenjunga? A: 8598m.</title>
      </sec>
      <sec id="sec-2-12">
        <title>COUNT, e.g. Q: How many people died during the Terror of Pol Pot?</title>
      </sec>
      <sec id="sec-2-13">
        <title>A: 1 million.</title>
      </sec>
      <sec id="sec-2-14">
        <title>OBJECT, e.g. Q: What does magma consist of?</title>
      </sec>
      <sec id="sec-2-15">
        <title>A: Molten rock.</title>
        <p>OTHER, i.e. everything that does not fit into the other categories above.</p>
      </sec>
      <sec id="sec-2-16">
        <title>Q: Which treaty was signed in 1979?</title>
      </sec>
      <sec id="sec-2-17">
        <title>A: Israel-Egyptian peace treaty.</title>
        <p>b) Definition questions, questions such as "What/Who is X?", and are divided into the following subtypes:
PERSON, i.e. questions asking for the role/job/important information about someone,</p>
      </sec>
      <sec id="sec-2-18">
        <title>Q: Who is Robert Altmann?</title>
      </sec>
      <sec id="sec-2-19">
        <title>A: Film maker.</title>
        <p>ORGANIZATION, i.e. questions asking for the mission/full name/important information about an
organization, e.g.</p>
      </sec>
      <sec id="sec-2-20">
        <title>Q: What is the Knesset?</title>
      </sec>
      <sec id="sec-2-21">
        <title>A: Parliament of Israel.</title>
        <p>OBJECT, i.e. questions asking for the description/function of objects, e.g.</p>
      </sec>
      <sec id="sec-2-22">
        <title>Q: What is Atlantis?</title>
      </sec>
      <sec id="sec-2-23">
        <title>A: Space Shuttle.</title>
        <p>OTHER, i.e. question asking for the description of natural phenomena, technologies, legal procedures
etc., e.g.</p>
      </sec>
      <sec id="sec-2-24">
        <title>Q: What is Eurovision?</title>
      </sec>
      <sec id="sec-2-25">
        <title>A: Song contest.</title>
      </sec>
      <sec id="sec-2-26">
        <title>Q: Name all the airports in London, England.</title>
      </sec>
      <sec id="sec-2-27">
        <title>A: Gatwick, Stansted, Heathrow, Luton and City.</title>
        <p>c) closed list questions: i.e. questions that require one answer containing a determined number of items, e.g:
As only one answer was allowed, all the items had to be presented in sequence, one next to the other, in one
document of the target collections.
Besides, all types of questions could contain a temporal restriction, i.e. a temporal specification that provided
important information for the retrieval of the correct answer, for example:</p>
        <p>Some questions could have no answer in the document collection, and in that case the exact answer was "NIL"
and the answer and support docid fields were left empty. A question is assumed to have no right answer when
neither human assessors nor participating systems could find one.</p>
        <p>The distribution of the questions among these categories is described in Table 4.</p>
        <p>Each of the question sets was finally then translated into English, so that each group could translate another set
into their own language, when preparing the cross-lingual data sets which had been activated.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Participants</title>
      <p>After years of constant growth, the number of participants has decreased in 2007 [see Table 5]..
The geographical distribution has anyway remained almost the same, recording a new entry of a group from
Australia. No participants took part to any Bulgarian tasks.
Also the number of submitted runs has decreased sensibly, from a total of 77 registered last year to 22 (see table
6). As in previous campaigns, a larger number of people chose to participate in the monolingual tasks, which
once again demonstrated to be more approachable.</p>
      <p>•
•
•
•
•
•</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>No changes were made as far the evaluation process is concerned- Human judges assessed the exact answer (i.e.
the shortest string of words which is supposed to provide the exact amount of information to answer the
question) as:</p>
      <p>R (Right) if correct;
W (Wrong) if incorrect;
X (ineXact) if contained less or more information than that required by the query;
U (Unsupported) if either the docid was missing or wrong, or the supporting snippet did not contain the
exact answer.</p>
      <p>
        Most assessor-groups managed to guarantee a second judgment of all the runs. As regards the evaluation
measures the following measures:
accuracy, as the main evaluation score, defined as the average of SCORE(q) over all 200 questions q;
the K1 measure[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]:
      </p>
      <p>K1(sys) = r∈answers (sys )</p>
      <p># questions</p>
      <p>score(r ) • eval (r )</p>
      <p>
        K1(sys )∈ IR ∧ K1(sys )∈ [
        <xref ref-type="bibr" rid="ref1">− 1,1</xref>
        ]
where:
score (r) is the confidence score assigned by the system to the answer r and eval(r) depends on the
judgment given by the human assessor.
      </p>
      <p>eval (r ) = { 1 if (r ) is judged
− 1 in other
cases
as correct
•</p>
      <p>
        K1(sys) = 0 is established as a baseline.
the Confident Weighted Score (CWS), designed for systems that give only one answer per question.
Answers are in a decreasing order of confidence and CWS rewards systems that give correct answers at
the top of the ranking [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] .
6
      </p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>As far as accuracy is concerned, scores were generally far lower this year than usual, as Figure 1 shows. In
detail, Best accuracy in the monolingual task decreased by almost 15 points, passing from last year’s 68.95% to
54%, while Best accuracy in cross-language tasks passed from 49.47% to 41.75% recording.
As far as average performances are concerned, this year a neat decrease has been recorded in the biligual tasks,
which went from 22.8% to 10.9%. This was due also due to the presence of systems which participated for the
first time, achieving very low score in tasks which are quite difficult also for veterans.</p>
      <p>As a general remark, it can be said that the new factors introduced this year appear to have had an impact on the
performances of the systems. As more than one participant has noticed, there has been not enough time to adjust
the systems to the new requirements.</p>
      <p>Here below a more detailed analyses of the results in each language follows, giving more specific information on
the performances of systems in the single sub-tasks and on the different types of questions, providing the
relevant statistics and comments.
41,5</p>
      <p>29</p>
    </sec>
    <sec id="sec-6">
      <title>6.1 Dutch as Target</title>
      <p>For the Dutch subtask of the CLEF 2007 QA task, three annotators generated 200 questions organized in 78
groups so that there were 16 groups with one question, 21 groups with two, 22 with three and 19 groups with
four questions. Among the 200 questions 156 were factoids, 28 definitions and 16 list questions. In total, 41
questions had temporal restrictions. Table XXX below shows the distributions of topic types for groups and
expected answer types for questions.
Annotators were asked to create questions with answers either in Dutch Wikipedia or in the Dutch newspaper
corpus, as well as questions without known answers. Of 200 questions, 186 had answers in Wikipedia, and 14 in
the newspaper corpus. Annotators did not create NIL questions.
This year, two teams took part in the QA track with Dutch as the target language: the University of Amsterdam
and the University of Groningen. The latter submitted both monolingual and crosslingual (English to Dutch)
runs. The 5 submitted runs were assessed independently by 3 Dutch native speakers in such a way that each
question group was assessed by at least two assessors. In case of conflicting assessments, assessors were asked
to discuss the judgements and come to an agreement. Most of the occured conflicts were due to difficulties in
distinguishing between inexact and correct answers. Table 7 below shows the evaluation results for the five
submitted runs (three monolingual and two cross-lingual). The table shows the number of Right, Wrong, ineXact
and Unsupported answers, as well as the percentage of correctly answered Factoids, Temporally restricted
questions, Definition and List questions.</p>
      <p>The best monolingual run (gron072NLNL) achieved accuracy of 25.5%, which is slightly less that the best
results in the 2006 edition of the QA task. The same tendency holds for the performance on factoid and
definition questions. We interpret this an indication of the increased difficulty of the task due the newly
introduced Wikipedia collection.</p>
      <p>One of the runs contained as many as 23 unsupported answers—this might indicate a bug in the system.</p>
    </sec>
    <sec id="sec-7">
      <title>6.2 English as Target</title>
      <p>Creation of questions. This year the questions set were radically different from last year. Instead of 200
independent questions, we were required to devise questions in groups. Each group had a declared topic (e.g.
"Polygraph") but unlike in TREC, this topic was not communicated to the participants. As at CLEF last year, the
type of question (e.g. definition, factoid or list) was not declared to participants either.</p>
      <p>Run
cind071fren
cind072fren
csui071inen
dfki071deen
dfki071esen
mqaf071nlen
mqaf072nlen
wolv071roen</p>
      <p>R
#
26
26
20
14
5
0
0
28</p>
      <p>W
#
171
170
175
178
189
200
200
166</p>
      <p>X
#
1
2
4
6
4
0
0
2</p>
      <p>U
#
2
2
1
2
2
0
0
4
160 Factoids (in groups) were requested, together with 30 definitions and ten lists. The numbers of temporally
restricted factoids and questions with NIL answers was at our discretion. In the end we submitted 161 factoids,
30 definitions and nine lists. In previous years we have been obliged to devise a considerable number of
temporally restricted questions and this has proved very difficult to do with the majority of them being very
contrived and artificial. For this reason it was intended to set no such questions this year. However, one
reasonable one was spotted during the data entry process and so was flagged as such. Two others were also
flagged accidentally during data entry. Unfortunately, therefore, the statistics can not tell us anything about
temporally restricted questions.</p>
      <p>Concerning NIL questions, we have long argued that they tell us very little about the performance of a system
unless it can report the reason why there is no answer. For example, this is a useful system:
Q: Who is the Queen of France?
A: France is a Republic!
By contrast, answering NIL would not tell us whether there was an answer which was simply not found, or
whether no answer in fact exists. Another important point following from this is that NIL questions artificially
boost the performance of a system which returns many NIL answers. For these reasons we decided not to include
any questions with NIL anwers. However, we would like to see ‘Queen of France’ answers being returned in
future workshops.</p>
      <p>The grouped nature of the questions had a considerable effect on their difficulty; instead of a series of ‘trivia’
type questions, each with a simple, clear answer, a single topic was effectively investigated in much more detail.
To achieve the goals set by the organisers it was necessary to find topics about which several questions could be
asked and then to devise as many questions as possible from that topic. Each task was surprisingly hard, and an
inevitable consequence was that the questions are much harder this year than in previous years. We had no wish
to set especially difficult or convoluted questions, but unfortunately this arose as a side-effect of the new
procedures.</p>
      <p>The requirement for related questions on a topic necessarily implies that the questions will refer to common
concepts and entities within the domain in question. In a series of questions this is accomplished by co-reference
– a well known phenomenon within Natural Language Processing which nevertheless has not been a major
factor in the success of QA systems at previous CLEF workshops. The most common form is anaphoric
reference to the topic declared implicitly in the first question, e.g.:
Q: What is a Polygraph?
Q: When was it invented?
However, other forms of co-reference occurred in the questions. Here is an example:
Q: Who wrote the song "Dancing Queen"?
Q: How many people were in the group?
Here the group refers to the category of entity into which the answer to the first question is known by the
questioner to belong. However, the QA system does not know this and has to infer it, a task which can be very
complex and indirect, especially where the topic is concealed from the participants.</p>
      <p>In addition to the issue of question grouping, it was decided at a very late stage to use not only the two
collections from last year (the LA Times and Glasgow Herald) but also the English Wikipedia. The latter is
extremely large and greatly increases the task complexity for the participants in terms of both indexing and IR
searching. In addition, some questions had to be heavily qualified in order to reduce the ambiguity introduced by
alternative readings in the Wikipedia. Here is an example:
Q: What is the "KORG" on which Niky Orellana is a soccer commentator?
Thirdly, we should bear in mind that the Wikipedia varies considerably in size depending on the language, with
the English one being by far the largest. We have not controlled for this fact in CLEF and the consequence could
be that the addition of Wikipedia had a greater effect on difficulty for English than it did for other languages.
Summary Statistics. Eight cross-lingual runs with English as target were submitted this year, as compared with
thirteeen for last year. Five groups participated in six source languages, Dutch, French, German, Indonesian,
Romanian and Spanish. DFKI submitted runs for two source languages, German and Spanish, while all other
groups worked in only one. Cindi Group and Macquarie University both submitted two runs for a language pair
(French-English and Dutch-English respectively) but unfortunately there was no language for which more than
one group submitted a run. This means that no direct comparisons can be made between QA systems this year,
because the task being solved by each was different.</p>
      <p>Assessment Procedure. An XML format was used for the submission of runs this year, by constrast with
previous years when fairly similar plain text formats were adopted. This meant that our evaluation tools were no
longer usable. However, last year we also participated in the evaluation of the WiQA task organised by
University of Amsterdam. For this they developed an excellent web-based tool which was subsequently adapted
for this year’s Dutch CLEF evaluations. We are extremely grateful to Martin de Rijke and Valentin Jijkoun for
allowing us to use it and for setting it up in Amsterdam especially for us. It allows multiple assessors to work
independently, shows runs anonymised, allows all answers to a particular question to be judged at the same time
(like the TREC software), and includes the supporting snippets for each submitted answer as well as the ‘correct’
(reference) answer. It also shows inter-assessor disagreement, and, once this has been eliminated, can produce
the assessed runs in the correct XML format. Overall, this software worked perfectly for us and saved us a
considerable amount of time.</p>
      <p>All answers were double-judged. The first assessor was Richard Sutcliffe and the second was Udo Kruschwitz
from University of Essex to whom we are indebted for his invaluable help. Where assessors differed, the case
was discussed between us and a decision taken. We measured the agreement level by two methods. For
Agreement 1 we take agreement on each group of 8 answers to a question as a whole as either exactly the same
for both assessors or not exactly the same. This is a very strict measure. There were disagreements for 30
questions out of the 200, i.e. 15%, which equates to an agreement level of 85%.</p>
      <p>For Agreement Level 2 we taking each decision made on one of the eight answers to a question and count how
many decisions were the same for both assessors and how many were not the same. There were 39 differences of
decision and a total of 1600 decisions (200 questions by eight runs). This is 2.4%, which equates to an agreement
level of 97.6%. This is the measure we used in previous years. Last year the agreement level was 89% and the
previous year it was 93%. We conclude from these figures that the assessment of our CLEF runs is quite
accurate and that double judging is sufficient.</p>
      <p>Results Analysis. As in previous years there were three types of question within the question groups, Factoids,
Definitions and Lists. Considering all question types together, the best performance is University of
Wolverhampton with 28 R and 2 X, (14% strict or 15% lenient) closely followed by the CINDI Group at
Concordia University with 26 R and 1 X (13% strict or 13.50% lenient). Note that these systems are working on
different tasks (RO-EN and FR-EN respectively) as noted above, so the results are not directly comparable. The
best performance last year for English targets was 25.26%. Nevertheless, considering the extreme difficulty of
the questions, this represents a remarkable achievement for these systems.</p>
      <p>For Factoids alone, the best system was CINDI (FR-EN) at 11.18% followed by University of Indonesia
(INEN) with 10.56%. For Definitions the best result was University of Wolverhampton (RO-EN) with 43.33%
correct, followed equally by CINDI (FR-EN) and DFKI (DE-EN) both with 23.33%. It is interesting that this
year the best Definition score is almost four times the best Factoid score, whereas last year they were nearly
equal. One reason for this may be that the definitions either occurred first in a group of questions or on their own
in a ‘singleton’ group. This was not specifically intended but seems to be a consequence of the relationship
between Factoids and Definitions, namely that the latter are somehow epistemologically prior to the former1. In
consequence, Definitions may be more simply phrased than Factoids and in particular may avoid co-reference in
the vast majority of cases.</p>
      <p>Nine lists questions were set but only CINDI was able to answer any of them correctly (11.11% accuracy).
(University of Indonesia was ineXact on one list question.) Perhaps the problem here was recognising the list
question in the first place – unlike at TREC they are not explicitly flagged. We believe this is not necessarily
reasonable since in a real dialogue a questioner would surely make it quite clear whether they expected a list of
answers or just one. They would not come up with a list question out of the blue.
1 Perhaps it is just a consequence of setting too many undergraduate examination papers!</p>
    </sec>
    <sec id="sec-8">
      <title>6.3 French as Target</title>
      <p>This year two groups took part in evaluation tasks using French as target language: one French group: Synapse
Développement ; and one American group: Language Computer Corporation (LCC).</p>
      <p>In total, only two runs have been returned by the participants: one monolingual run (FR-to-FR) from Synapse
Développement and one bilingual run (EN-to-FR) from LCC.</p>
      <p>It appears that the number of participants for the French task has clearly decreased this year, certainly due to the
many changes that appeared in the 2007 Guidelines for the participants: adding to a large new answer source
(Wikipedia 2006) and adding to a large number of topic-related questions, i.e. clusters of questions which are
related to the same topic and possibly contain anaphoric references between one question and the other
questions. These changes explain certainly the cause of the strong decrease of participation this year.
Three types of questions were proposed: factual, definition and closed list questions. The participating teams
could return one exact answer per question and up to two runs. Some questions (10%) had no answer in the
document collection, and in this case the exact answer is "NIL".
80
70
60
50
40
30
20
10
0
24,5
64
67,89
54
17</p>
      <p>49,47
39,5
41,75
The French test set was composed of 200 questions: 163 Factual (F), 27 Definition (D) and 10 closed List
questions (L). Among these 200 questions, 41 were Temporally restricted questions (T).</p>
      <p>The accuracy has been calculated over all the answers of F, D, T and L questions and also the Confidence
Weighted Score (CWS) and the K1 measure.</p>
      <p>For the monolingual task, the Synapse Développement’ system returned 108 correct answers i.e. 54 % of correct
answers (as opposed to 67,89 % last year).</p>
      <p>For the bilingual task, the LCC’s system returned 81 correct answers i.e. 41,75 % of correct answers (as opposed
to 49,47 % for the best bilingual system last year).</p>
      <p>We can observe that the two systems obtained different results according to the answer types. The monolingual
system obtained better results for Definition questions (74,07 %) than for Factoid (52,76 %) and Temporally
questions (46,34 %) whereas the bilingual system obtained better results for Temporally (46,34 %) and Factoid
questions (44,17 %) than for Definition questions (22,22 %).</p>
      <p>We can note that the bilingual system has not returned NIL answer, whereas the monolingual one returned 40
NIL answers (out of 9 expected NIL answers in the French test set). As there were only 9 NIL answers in the
French test set and as the monolingual system returned 40 NIL answers, his final score is not very high (even if
this system returned the 9 expected correct NIL answers).</p>
      <p>The main difficulties encountered this year by the systems were the new type of questions: topic-related
questions and the adding of a new large answer source (Wikipedia 2006). The participants had to adapt their
system in a few weeks to be able to deal with this new type of questions.</p>
      <p>Moreover, larger is the corpus, more difficult is the expected exact answer to be extracted from the corpus source
for a system (even if very often, there are several possible answers in the corpus).</p>
      <p>In conclusion, despite the important changes in the Guidelines for the participants, the monolingual system
obtained the best results of all the participants at CLEF@QA track this year (108 correct answers out of 200).
We can note that the American group (LCC) participated only for the second time in the Question Answering
track using French in target and has already obtained good results that can let us imagine it will improve again in
the future. In addition, we can still observe this year the increasing interest in Question Answering for the tasks
using French as target language from the non-European research community due to the second participation of
an American team.</p>
    </sec>
    <sec id="sec-9">
      <title>6.4 German as Target</title>
      <p>Two research groups submitted runs for evaluation in the track having German as target language: The German
Research Center for Artificial Intelligence (DFKI) and the Fern Universität Hagen (FUHA). Both provided
system runs for the monolingual scenario and just DFKI submitted runs for the cross-language English-German
and Portuguese-German scenario. The assessment was conducted by two native German speakers with fair
knowledge of information access systems. Compared to the previous editions of the evaluation forum, this year a
decrease in the accuracy of the best performing system and of an aggregated virtual system for both monolingual
and cross-language tasks was registered.</p>
      <sec id="sec-9-1">
        <title>Best Mono</title>
        <p>30
42.33
43.5
34.01
Aggregated Cross</p>
        <p>Run
dfki071dedeM
fuha071dedeM
fuha072dedeM
dfki071endeC
dfki071ptdeC</p>
        <p>W
#
121
146
164
144
180
The details of systems’ results can be seen in Table 11. There were no NIL questions tested in this year’s
evaluation. The results submitted by DFKI did not provide a normalized value for the confidence score of an
answer and therefore both CWS and K1 values could not be computed.</p>
        <p>The number of topics covered by the questions was of 116 distributed as it follows: 69 topics consisting of 1
question, 19 topics of 2 related questions, and each 19 topics of 3 and 4 related questions. The most frequents
topic types were PERSON (40), OBJECT (33) and ORGANIZATION (23). As regards the source of the
answers, 101 questions from 68 topics asked for information out of the CLEF document collection and the rest
of 99 from 48 topics for information from Wikipedia. The distribution of the topics over the document
collections (CLEF vs. Wikipedia) is as follows: 53 vs. 16 topics of 1 question, 4 vs. 15 topics of each 2 and 3
questions and 7 vs. 2 topics of 4 questions.</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>6.5 Italian as Target</title>
      <p>Only one group took part in this year to the monolingual Italian task, i.e. FBK-irst, submitting only one run. The
results are shown in table 13.
As Figure 4 shows, the results were much lower than both best and average performances in monolingual Italian
tasks in the achieved in the previous campaigns.
The Italian question consisted of 147 factoid questions, 41 definition questions and 12 list questions. 38
questions contained a temporal restriction, and 11 had no answer in the Gold Standard. In the Gold Standard,
108 answers were retrieved from Wikipedia, the remains from the news collections.</p>
      <p>The submitted run was assessed by two judges; the inter-annotator agreement was 92,5%.</p>
      <p>The system achieved low accuracy in all types of questions, performing anyway better in factoids questions.
Definition questions, with 2,63% of accuracy and list questions, for which no correct answer was retrieved, were
particularly challenging. A relevant number of questions (about 6%) was judged unsupported, meaning that the
correct answer was retrieved by the system, which did not provided enough context to support it.</p>
    </sec>
    <sec id="sec-11">
      <title>6.6 Portuguese as Target</title>
      <p>Six research groups took part in tasks with Portuguese as target language, submitting eigth runs: seven in the
monolingual task, and one with English as source; unlike last year, no group presented Spanish as source. One
new group (INESC) participated this. The group of University of Évora (UE) returned this year, while the group
from NILC, the sole Brazilian group to take part to date, was absent.</p>
      <p>Again, Priberam presented the best result for the third year in a row; the group of the University of Évora wasn’t
however far behind. As last year, we added the classification X-, meaning incomplete, while keeping the
classification X+ for answers with extra text or other kinds of inexactness. In Table 3 we present the overall
results.
A direct comparison with last year’s results is not fully possible, due to the existance of multiple questions to
each topic. Therefore, 14 presents results regarding the first question of each topic, which we believe is more
readily comparable to the results of previous years.
As it can be seen, the removal of subsequent questions to each topic doesn’t cause a big change on the overal
results, apart from a clear improvement by Priberam. On the whole, compared to last year (Vallin et al., 2007),
Priberam saw a slight drop on its results, Raposa (FEUP) a clear improvement from an admitedly low level,
Esfinge (SINTEF) a clear drop, and LCC kept last year’s levels. Senso (UE) shows a marked improvement since
its last participation in 2005. We leave it to the participants to comment on whether it might have been caused by
harder questions or changes (or lack thereof) in the systems.</p>
      <p>Question 94 was reclassified as NIL due to a spelling error, and question 135 because of the use of a word with a
rare meaning. On the other hand, one system saw through that rare meaning, providing a correct answer; we
decided to keep the question as NIL, considering correct both the system’s answer and any NIL answer from
other systems. The same system also found a correct answer to a NIL question, not discovered during the
question creating process; that question was therefore reclassified as non-NIL. In the end, there were 13 NIL
questions.</p>
      <p>Table 16 shows the results for each answer type of definition questions, while Table 17 shows the results for
each answer type of factoid questions (including list questions). As it can be seen, four out of six systems
perform clearly better when it comes to definitions than to factoids. This may well have been helped by the use
of Wikipedia texts, where a large proportion of articles begin with a definition.
We included in both Table 16 and in Table 17 a virtual run, called combination, in which one question is
considered correct if at least one participating system found a valid answer. The objective of this combination
run is to show the potential achievement when combining the capacities of all the participants. The combination
run can be considered, somehow, state-of-the-art in monolingual Portuguese question answering. The system
with best results, Priberam, answered correctly 72.7% the questions with at least one correct answer, not as
dominating as last year. Despite being a bilingual run, LCC answered correctly 14 questions not answered by
any of the monolingual systems.</p>
      <p>In Table 18, we present some values concerning answer and snippet size (in number of words).</p>
      <p>Run</p>
      <p>Average
answer
size
2.8
2.4
Temporally restricted questions: Table 19 presents the results of the 20 temporally restricted questions. As in
previous years, the effectiveness of the systems to answer those questions is visibly lower than for non-TRQ
questions (and indeed several systems only answered correctly question 160, which is a NIL TRQ).
List questions: a total of twelve questions were defined as list questions; unlike last year, all these questions
were closed list factoids, with two to twelve answers each2. The results were, in general, weak, with UE and
LCC getting two correct answers, Priberam five, and all other system zero. There was a single case of
incomplete answer (i.e., answering some elements of the list only), but it was judged W since, besides
incomplete, it was also unsupported.</p>
    </sec>
    <sec id="sec-12">
      <title>6.7 Romanian as Target</title>
      <p>At CLEF 2007 Romanian was addressed as a target language for the first time, based on the collection of
Wikipedia Romanian pages frozen in November 2006, and as a source language for the second time, using the
English news collection (Los Angeles Times, 1994 and Glasgow Herald, 1995) and the Wikipedia English
pages.</p>
      <p>Creation of Questions. The creation of the questions was realized at the Faculty of Computer Science, Al.I.
Cuza University of Iasi. The group3 was very well instructed with respect to this task, using the Guidelines for
Question Generation and based on a good feedback received from the organizers at IRST4. The final 200 created
questions are distributed according to table 20.
2 There were some open list questions as well, but they were classified and evaluated as ordinary factoids.
3 Three Computational Linguistics Master students: Anca Onofraşc, Ana-Maria Rusu, Cristina Despa, supervised
and working in collaboration with the two organizers
4 Without the help received from Danilo Giampicolo and Pamela Forner, we wouldn’t have solved all our
problems.</p>
      <sec id="sec-12-1">
        <title>FACTOID</title>
      </sec>
      <sec id="sec-12-2">
        <title>DEFINITION</title>
      </sec>
      <sec id="sec-12-3">
        <title>LIST</title>
        <p>NIL</p>
      </sec>
      <sec id="sec-12-4">
        <title>QUESTIONS</title>
        <p>PERSON
Most difficulties in this task were raised by deciding on the supporting snippets, especially for questions
belonging to the same topic. We found unnatural to include answers through “copy-paste” from the text: if the
question requires an answer in the Nominative case, but the text includes the answer in the Genitive case, then
we had to include the Genitive in the answer, even if it is more natural to have the answer in Nominative.
Participants. This year two Romanian groups took part in the monolingual task with Romanian as a target
language: the Faculty of Computer Science from the Al. I. Cuza University of Iasi, and the Research Institute for
Artificial Intelligence from the Romanian Academy, Bucharest. Three runs were submitted – one by the first
group and two by the second group, with the differences between them due to the way they treated the
questionprocessing and the answer-extraction. The 2007 results are presented in Tables 21 below. One system with
Romanian as a source language and English as target was submitted by the Computational Linguistics Group
from the University of Wolverhampton, United Kingdom.</p>
        <p>NIL</p>
      </sec>
      <sec id="sec-12-5">
        <title>RETURNED NIL correct</title>
        <p>12
30
30
100
54
54
5
7
7</p>
        <p>CWS
All three systems “crashed” on the LIST questions. The NIL questions are hard to classify, starting from the
question-classifier (the classifier should “know” that the QA system has no possibility, no knowledge to find the
answer). It would be better to have a clear separation between the NIL answers due to impossibility to find
answer and the NIL answers classified as such by the system. None of the three systems could handle the
questions related under one same topic: the systems returned at most the answer to the first question in a topic.
Assessment Procedure. Due to time restrictions, all three runs where judged by only one assessor at the Faculty
of Computer Science in Iasi, so an inter-annotator agreement was not possible. Based on the Guidelines, all three
systems were judged in parallel. The same evaluation criteria, especially with respect to the UNSUPPORTED
and INEXACT answers, were used.</p>
      </sec>
    </sec>
    <sec id="sec-13">
      <title>6.8 Spanish as Target</title>
      <p>
        The participation at the Spanish as Target subtask has decreased from 9 groups in 2006 to 5 groups this year. All
the runs were monolingual. We think that the changes in the task (linked questions and wikipedia) led to a lower
participation and worse overall results because systems could not be tuned on time. Table 22 shows the summary
of systems results with the number of Right (R), Wrong (W), Inexact (X) and Unsupported (U) answers. The
table shows also the accuracy (in percentage) of factoids (F), factoids with temporal restriction (T), definitions
(D) and list questions (L). Best values are marked in bold face. All the runs were assessed by two assessors.
Only a 1.5% of the judgements were different and the resulting kappa value was 0,966, which corresponding to
“almost perfect” assessment [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Best performing systems have obtained worse results than last year due mainly to the low performance in
answering linked questions (15% of the questions) and due to the questions with answer only in Wikipedia.
Table 23 shows that considering only self-contained questions (the first one of each topic group) the results are
closer to the ones obtained last year. In fact the accuracy for the linked questions is less than 20%.
Regarding NIL questions, Table 25 shows the harmonic mean (F) of precision and recall for self-contained,
linked and all questions. The best performing system has decreased their overall performance with respect to the
last edition (see Table 26) in NIL questions. However, the performance considering only self-contained
questions is closer to the one obtained last year.
2006
2007
      </p>
      <p>
        The correlation coefficient r between the self-score and the correctness of the answers (shown in Table 27)
has been similar to the obtained last year, being not good enough yet, and explaining the low results in CWS and
K1 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] measures.
      </p>
      <p>Since a supporting snippet is requested in order to assess the correctness of the answer, we have evaluated
the systems capability to extract the answer when the snippet contains it. The first column of table 27 shows the
percentage of cases where the correct answer was present in the snippet and correctly extracted. This information
is very useful to diagnose if the lack of performance is due to the passage retrieval or to the answer extraction
process. As shown in the table, the best systems are also better in the task of answer extraction, whereas the rest
of systems still have a lot of room for improvement.
This year the task was changed considerably and this affected the general level of results and also the level of
participation in the task. The grouped questions could be regarded as more realistic and more searching but in
consequence they were much more difficult. The policy of not declaring the question type means that if this is
deduced incorrectly then the answer is bound to be wrong. Moreover, the policy of not even declaring the topic
of a question group, but leaving it implicit (usually within the first question) means that if a system infers the
topic wrongly, then all questions in the group will be answered wrongly. This should be probably re-considered,
as it is not ‘realistic’. In a real dialogue, if a question is answered inappropriately we do not dismiss all
subsequent answers from that person, we simply re-phrase the question instead. The level of ambiguity
concerning question type in a real dialogue is not fixed at some arbitrary value but varies according to many
factors which the questioner estimates. In CLEF we are not modelling this process at all accurately and this
affects the validity of our results. Finally, co-reference has now entered CLEF. This is interesting and useful but
it might be preferable if we could separate the effect of co-reference resolution from other factors in analysing
results. This could be done by marking up the co-references in the question corpus and allowing participants to
use this information under certain circumstances.</p>
    </sec>
    <sec id="sec-14">
      <title>Acknowledgments</title>
      <p>A special thank to Bernardo Magnini (FBK-irst, Trento, Italy), who has given his precious advise and valueble
support at many levels for the preparation and realization of the QA track at CLEF 2007.
Anselmo Peñas has been partially supported by the Spanish Ministry of Science and Technology within the
Text-Mess-INES project (TIN2006-15265-C06-02).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. QA@CLEF website: http://clef-qa.itc.it/</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>2. AVE Website: http://nlp.uned.es/QA/ave/.</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>3. QAST Website: http://www.lsi.upc.edu/~qast/</mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>4. QART Website: http://gplsi.dlsi.ua.es/qart/</mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. QA@
          <article-title>CLEF 2007 Organizing Committee</article-title>
          .
          <source>Guidelines</source>
          <year>2007</year>
          . http://clef-qa.itc.it/2007/download/QA@
          <article-title>CLEF07_Guidelines-for-Participants</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Herrera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peñas</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verdejo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Question answering pilot task at CLEF 2004</article-title>
          . In: Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Gareth J.F.</given-names>
            ,
            <surname>Kluck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Magnini</surname>
          </string-name>
          ,
          <string-name>
            <surname>B</surname>
          </string-name>
          . (eds.):
          <article-title>Multilingual Information Access for Text</article-title>
          ,
          <source>Speech and Images. Lecture Notes in Computer Science</source>
          , Vol.
          <volume>3491</volume>
          . Springer-Verlag, Berlin Hidelberg New York (
          <year>2005</year>
          )
          <fpage>581</fpage>
          -
          <lpage>590</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Landis</surname>
          </string-name>
          and
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Koch</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>The measurements of observer agreement for categorical data</article-title>
          .
          <source>Biometrics</source>
          ,
          <volume>33</volume>
          :
          <fpage>159</fpage>
          -
          <lpage>174</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>