<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CLEF 2006: Ad Hoc Track Overview</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giorgio M. Di Nunzio</string-name>
          <email>dinunzio@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <email>ferro@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Mandl</string-name>
          <email>mandl@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carol Peters</string-name>
          <email>carol.peters@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>General Terms Experimentation</institution>
          ,
          <addr-line>Performance, Measurement, Algorithms</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ISTI-CNR</institution>
          ,
          <addr-line>Area di Ricerca - 56124 Pisa -</addr-line>
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Information Science, University of Hildesheim -</institution>
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We describe the objectives and organization of the CLEF 2006 ad hoc track and discuss the main characteristics of the tasks offered to test monolingual, bilingual, and multilingual textual document retrieval systems. The track was divided into two streams. The main stream offered mono- and bilingual tasks using the same collections as CLEF 2005: Bulgarian, English, French, Hungarian and Portuguese. The second stream, designed for more experienced participants, offered the so-called ”robust task” which used test collections from previous years in six languages (Dutch, English, French, German, Italian and Spanish) with the objective of privileging experiments which achieve good stable performance over all queries rather than high average performance. The document collections used were taken from the CLEF multilingual comparable corpus of news documents. The performance achieved for each task is presented and a statistical analysis of results is given.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The ad hoc retrieval track is generally considered to be the core track in the
Cross-Language Evaluation Forum (CLEF). The aim of this track is to promote
the development of monolingual and cross-language textual document retrieval
systems. The CLEF 2006 ad hoc track was structured in two streams. The main
stream offered monolingual tasks (querying and finding documents in one
language) and bilingual tasks (querying in one language and finding documents in
another language) using the same collections as CLEF 2005. The second stream,
designed for more experienced participants, was the ”robust task”, aimed at
finding documents for very difficult queries. It used test collections developed in
previous years.</p>
      <p>The Monolingual and Bilingual tasks were principally offered for
Bulgarian, French, Hungarian and Portuguese target collections. Additionally, in the
bilingual task only, newcomers (i.e. groups that had not previously participated
in a CLEF cross-language task) or groups using a “new-to-CLEF” query
language could choose to search the English document collection. The aim in all
cases was to retrieve relevant documents from the chosen target collection and
submit the results in a ranked list.</p>
      <p>The Robust task offered monolingual, bilingual and multilingual tasks using
the test collections built over three years: CLEF 2001 - 2003, for six languages:
Dutch, English, French, German, Italian and Spanish. Using topics from three
years meant that more extensive experiments and a better analysis of the results
were possible. The aim of this task was to study and achieve good performance
on queries that had proved difficult in the past rather than obtain a high average
performance when calculated over all queries.</p>
      <p>In this paper we describe the track setup, the evaluation methodology and the
participation in the different tasks (Section 2), present the main characteristics
of the experiments and show the results (Sections 3 - 5). Statistical testing
is discussed in Section 6 and the final section provides a brief summing up.
For information on the various approaches and resources used by the groups
participating in this track and the issues they focused on, we refer the reader to
the other papers in the Ad Hoc section of the Working Notes.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Track Setup</title>
      <p>The ad hoc track in CLEF adopts a corpus-based, automatic scoring method
for the assessment of system performance, based on ideas first introduced in the
Cranfield experiments in the late 1960s. The test collection used consists of a
set of “topics” describing information needs and a collection of documents to be
searched to find those documents that satisfy these information needs.
Evaluation of system performance is then done by judging the documents retrieved
in response to a topic with respect to their relevance, and computing the recall
and precision measures. The distinguishing feature of CLEF is that it applies
this evaluation paradigm in a multilingual setting. This means that the criteria
normally adopted to create a test collection, consisting of suitable documents,
sample queries and relevance assessments, have been adapted to satisfy the
particular requirements of the multilingual context. All language dependent tasks
such as topic creation and relevance judgment are performed in a distributed
setting by native speakers. Rules are established and a tight central
coordination is maintained in order to ensure consistency and coherency of topic and
relevance judgment sets over the different collections, languages and tracks.
2.1</p>
      <sec id="sec-2-1">
        <title>Test Collections</title>
        <p>Different test collections were used in the ad hoc task this year. The main (i.e.
non-robust) monolingual and bilingual tasks used the same document collections
as in Ad Hoc last year but new topics were created and new relevance assessments
made. As has already been stated, the test collection used for the robust task
was derived from the test collections previously developed at CLEF. No new
relevance assessments were performed for this task.</p>
        <p>Documents. The document collections used for the CLEF 2006 ad hoc tasks are
part of the CLEF multilingual corpus of newspaper and news agency documents
described in the Introduction to these Proceedings.</p>
        <p>In the main stream monolingual and bilingual tasks, the English, French and
Portuguese collections consisted of national newspapers and news agencies for
the period 1994 and 1995. Different variants were used for each language. Thus,
for English we had both US and British newspapers, for French we had a national
newspaper of France plus Swiss French news agencies, and for Portuguese we
had national newspapers from both Portugal and Brazil. This means that, for
each language, there were significant differences in orthography and lexicon over
the sub-collections. This is a real world situation and system components, i.e.
stemmers, translation resources, etc., should be sufficiently flexible to handle
such variants. The Bulgarian and Hungarian collections used in these tasks were
new in CLEF 2005 and consist of national newspapers for the year 20021. This
has meant using collections of different time periods for the ad-hoc mono- and
bilingual tasks. This had important consequences on topic creation. Table 1
summarizes the collections used for each language.</p>
        <p>The robust task used test collections containing data in six languages (Dutch,
English, German, French, Italian and Spanish) used at CLEF 2001, CLEF 2002
and CLEF 2003. There are approximately 1.35 million documents and 3.6
gigabytes of text in the CLEF 2006 ”robust” collection. Table 2 summarizes the
collections used for each language.
1 It proved impossible to find national newspapers in electronic form for 1994 and/or
1995 in these languages.
Topics Topics in the CLEF ad hoc track are structured statements representing
information needs; the systems use the topics to derive their queries. Each topic
consists of three parts: a brief “title” statement; a one-sentence “description”; a
more complex “narrative” specifying the relevance assessment criteria.</p>
        <p>Sets of 50 topics were created for the CLEF 2006 ad hoc mono- and bilingual
tasks. One of the decisions taken early on in the organization of the CLEF ad
hoc tracks was that the same set of topics would be used to query all
collections, whatever the task. There were a number of reasons for this: it makes it
easier to compare results over different collections, it means that there is a single
master set that is rendered in all query languages, and a single set of relevance
assessments for each language is sufficient for all tasks. However, in CLEF 2005
the assessors found that the fact that the collections used in the CLEF 2006 ad
hoc mono- and bilingual tasks were from two different time periods (1994-1995
and 2002) made topic creation particularly difficult. It was not possible to create
time-dependent topics that referred to particular date-specific events as all
topics had to refer to events that could have been reported in any of the collections,
regardless of the dates. This meant that the CLEF 2005 topic set is somewhat
different from the sets of previous years as the topics all tend to be of broad
coverage. In fact, it was difficult to construct topics that would find a limited
number of relevant documents in each collection, and consequently a - probably
excessive - number of topics used for the 2005 mono- and bilingual tasks have a
very large number of relevant documents.</p>
        <p>For this reason, we decided to create separate topic sets for the two different
time-periods for the CLEF 2006 ad hoc mono- and bilingual tasks. We thus
created two overlapping topic sets, with a common set of time independent
topics and sets of time-specific topics. 25 topics were common to both sets while
25 topics were collection-specific, as follows:
- Topics C301 - C325 were used for all target collections
- Topics C326 - C350 were created specifically for the English, French and
Portuguese collections (1994/1995)</p>
        <p>- Topics C351 - C375 were created specifically for the Bulgarian and
Hungarian collections (2002).</p>
        <p>This meant that a total of 75 topics were prepared in many different languages
(European and non-European): Bulgarian, English, French, German,
Hungarian, Italian, Portuguese, and Spanish plus Amharic, Chinese, Hindi, Indonesian,
Oromo and Telugu. Participants had to select the necessary topic set according
to the target collection to be used.</p>
        <p>Below we give an example of the English version of a typical CLEF topic:
&lt;top&gt; &lt;num&gt; C302 &lt;/num&gt;
&lt;EN-title&gt; Consumer Boycotts &lt;/EN-title&gt; &lt;
EN-desc&gt; Find documents that describe or discuss the impact of consumer
boycotts. &lt;/EN-desc&gt;
&lt;EN-narr&gt; Relevant documents will report discussions or points of view on
the efficacy of consumer boycotts. The moral issues involved in such
boycotts are also of relevance. Only consumer boycotts are relevant,
political boycotts must be ignored. &lt;/EN-narr&gt; &lt;/top&gt;</p>
        <p>For the robust task, the topic sets used in CLEF 2001, CLEF 2002 and
CLEF 2003 were used for evaluation. A total of 160 topics were collected and
split into two sets: 60 topics used to train the system, and 100 topics used for
the evaluation. Topics were available in the languages of the target collections:
English, German, French, Spanish, Italian, Dutch.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Participation Guidelines</title>
        <p>To carry out the retrieval tasks of the CLEF campaign, systems have to build
supporting data structures. Allowable data structures include any new structures
built automatically (such as inverted files, thesauri, conceptual networks, etc.)
or manually (such as thesauri, synonym lists, knowledge bases, rules, etc.) from
the documents. They may not, however, be modified in response to the topics,
e.g. by adding topic words that are not already in the dictionaries used by their
systems in order to extend coverage.</p>
        <p>Some CLEF data collections contain manually assigned, controlled or
uncontrolled index terms. The use of such terms has been limited to specific
experiments that have to be declared as “manual” runs.</p>
        <p>Topics can be converted into queries that a system can execute in many
different ways. CLEF strongly encourages groups to determine what constitutes
a base run for their experiments and to include these runs (officially or
unofficially) to allow useful interpretations of the results. Unofficial runs are those
not submitted to CLEF but evaluated using the trec eval package. This year
we have used the new package written by Chris Buckley for the Text REtrieval
Conference (TREC) (trec eval 7.3) and available from the TREC website.</p>
        <p>As a consequence of limited evaluation resources, a maximum of 12 runs each
for the mono- and bilingual tasks was allowed (no more than 4 runs for any one
language combination - we try to encourage diversity). We accepted a maximum
of 4 runs per group and topic language for the multilingual robust task. For
biand mono-lingual robust tasks, 4 runs were allowed per language or language
pair.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Relevance Assessment</title>
        <p>
          The number of documents in large test collections such as CLEF makes it
impractical to judge every document for relevance. Instead approximate recall values
are calculated using pooling techniques. The results submitted by the groups
participating in the ad hoc tasks are used to form a pool of documents for each
topic and language by collecting the highly ranked documents from all
submissions. This pool is then used for subsequent relevance judgments. The stability
of pools constructed in this way and their reliability for post-campaign
experiments is discussed in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] with respect to the CLEF 2003 pools. After calculating
the effectiveness measures, the results are analyzed and run statistics produced
and distributed. New pools were formed in CLEF 2006 for the runs submitted
for the main stream mono- and bilingual tasks and the relevance assessments
were performed by native speakers. Instead, the robust tasks used the original
pools and relevance assessments from CLEF 2003.
        </p>
        <p>
          The individual results for all official ad hoc experiments in CLEF 2006 are
given in the Appendix at the end of the on-line Working Notes prepared for the
Workshop [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
2.4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Result Calculation</title>
        <p>
          Evaluation campaigns such as TREC and CLEF are based on the belief that
the effectiveness of Information Retrieval Systems (IRSs) can be objectively
evaluated by an analysis of a representative set of sample search results. For
this, effectiveness measures are calculated based on the results submitted by the
participant and the relevance assessments. Popular measures usually adopted for
exercises of this type are Recall and Precision. Details on how they are calculated
for CLEF are given in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. For the robust task, we used different measures, see
below Section 5.
2.5
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Participants and Experiments</title>
        <p>As shown in Table 3, a total of 25 groups from 15 different countries submitted
results for one or more of the ad hoc tasks - a slight increase on the 23 participants
of last year. Table 4 provides a breakdown of the number of participants by
country.</p>
        <p>A total of 296 experiments were submitted with an increase of 16% on the
254 experiments of 2005. On the other hand, the average number of submitted
runs per participant is nearly the same: from 11 runs/participant of 2005 to 11.7
runs/participant of this year.</p>
        <p>Participants were required to submit at least one title+description (“TD”)
run per task in order to increase comparability between experiments. The large
majority of runs (172 out of 296, 58.11%) used this combination of topic fields,
78 (26.35%) used all fields, 41 (13.85%) used the title field, and only 5 (1.69%)
used the description field. The majority of experiments were conducted using
automatic query construction (287 out of 296, 96.96%) and only in a small fraction
of the experiments (9 out 296, 3.04%) have queries been manually constructed
from topics. A breakdown into the separate tasks is shown in Table 5(a).</p>
        <p>Fourteen different topic languages were used in the ad hoc experiments. As
always, the most popular language for queries was English, with French second.
The number of runs per topic language is shown in Table 5(b).
3</p>
        <p>
          Main Stream Monolingual Experiments
Monolingual retrieval was offered for Bulgarian, French, Hungarian, and
Portuguese. As can be seen from Table 5(a), the number of participants and runs
for each language was quite similar, with the exception of Bulgarian, which had
a slightly smaller participation. This year just 6 groups out of 16 (37.5%)
submitted monolingual runs only (down from ten groups last year), and 5 of these
groups were first time participants in CLEF. This year, most of the groups
submitting monolingual runs were doing this as part of their bilingual or multilingual
system testing activity. Details on the different approaches used can be found
in the papers in this section of the working notes. There was a lot of detailed
work with Portuguese language processing; not surprising as we had four new
groups from Brazil in Ad Hoc this year. As usual, there was a lot of work on the
development of stemmers and morphological analysers ([
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], for instance, applies
a very deep morphological analysis for Hungarian) and comparisons of the pros
and cons of so-called ”light” and ”heavy” stemming approaches (e.g. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]). In
contrast to previous years, we note that a number of groups experimented with
NLP techniques (see, for example, papers by [6], and [7]).
3.1
        </p>
      </sec>
      <sec id="sec-2-6">
        <title>Results</title>
        <p>Main Stream Bilingual Experiments
The bilingual task was structured in four subtasks (X → BG, FR, HU or PT
target collection) plus, as usual, an additional subtask with English as target
language restricted to newcomers in a CLEF cross-language task. This year, in
this subtask, we focussed in particular on non-European topic languages and in
particular languages for which there are still few processing tools or resources
were in existence. We thus offered two Ethiopian languages: Amharic and Oromo;
two Indian languages: Hindi and Telugu; and Indonesian. Although, as was to</p>
        <p>Ad−Hoc Monolingual Bulgarian track Top 5 Participants − Interpolated Recall vs Average Precision
0%
0%
10%
20%
30%
40% 50% 60%</p>
        <sec id="sec-2-6-1">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%
Fig. 2. Monolingual French
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-2">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-3">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%
Fig. 4. Monolingual Portuguese
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-4">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%</p>
        </sec>
        <sec id="sec-2-6-5">
          <title>Ad−Hoc Monolingual Portuguese track Top 5 Participants − Interpolated Recall vs Average Precision</title>
          <p>unine [UniNEpt1; MAP 45.52%; Pooled]
hummingbird [humPT06tde; MAP 45.07%; Not Pooled]
alicante [30okapiexp; MAP 43.08%; Not Pooled]
rsi−jhu [95aplmopttd5; MAP 42.42%; Not Pooled]
u.buffalo [UBptTDrf1; MAP 40.53%; Pooled]
be expected, the results are not particularly good, we feel that experiments of
this type with lesser-studied languages are very important (see papers by [8], [9],
[10])</p>
          <p>While these results are very good for the well-established-in-CLEF languages,
and can be read as state-of-the-art for this kind of retrieval system, at a first
glance they appear very disappointing for Bulgarian and Hungarian. However,
Track
we have to point out that, unfortunately, this year only one group submitted
cross-language runs for Bulgarian and Hungarian and thus it does not make
much sense to draw any conclusions from these, apparently poor, results for
these languages. It is interesting to note that when Cross Language Information
Retrieval (CLIR) system evaluation began in 1997 at TREC-6 the best CLIR
systems had the following results:
– EN → FR: 49% of best monolingual French IR system;
– EN → DE: 64% of best monolingual German IR system.
The robust task was organized for the first time at CLEF 2006. The evaluation
of robustness emphasizes stable performance over all topics instead of high
average performance [11]. The perspective of each individual user of an information
retrieval system is different from the perspective taken by an evaluation
initiative. The user will be disappointed by systems which deliver poor results for
some topics whereas an evaluation initiative rewards systems which deliver good
average results. A system delivering poor results for hard topics is likely to be
considered of low quality by a user although it may reach high average results.
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-6">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%</p>
          <p>Ad−Hoc Bilingual track, French target collection(s) Top 5 Participants − Interpolated Recall vs Average Precision
100%
Ad−Hoc Bilingual track, Bulgarian target collection(s) Top 5 Participants − Interpolated Recall vs Average Precision
100%</p>
          <p>daedalus [Experiment bgFSbgWen2S; MAP 17.39%; Pooled]
90%
80%
70%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-7">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%
Fig. 6. Bilingual French
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-8">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%
Ad−Hoc Bilingual track, Portuguese target collection(s) Top 5 Participants − Interpolated Recall vs Average Precision
100%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-9">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%</p>
          <p>Fig. 8. Bilingual Portuguese
90%
80%
70%
ion 60%
s
i
c
e
reP 50%
g
a
r
veA 40%
30%
20%
10%
rsi−jhu [aplbiinen5; MAP 32.57%; Pooled]
depok [UI_td_mt; MAP 26.71%; Not Pooled]
ltrc [OMTD; MAP 25.04%; Pooled]
celi [CELItitleNOEXPANSION; MAP 23.97%; Not Pooled]
dsv [DsvAmhEngFullNofuzz; MAP 22.78%; Not Pooled]</p>
          <p>Ad−Hoc Bilingual track, English target collection(s) Top 5 Participants − Interpolated Recall vs Average Precision
100%
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-10">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%
The robust task has been inspired by the robust track at TREC where it ran at
TREC 2003, 2004 and 2005. A robust evaluation stresses performance for weak
topics. This can be done by using the Geometric Average Precision (GMAP) as
a main indicator for performance instead of the Mean Average Precision (MAP)
of all topics. Geometric average has proven to be a stable measure for robustness
at TREC [11]. The robust task at CLEF 2006 is concerned with the multilingual
aspects of robustness. It is essentially an ad-hoc task which offers mono-lingual
and cross-lingual sub tasks.</p>
          <p>During CLEF 2001, CLEF 2002 and CLEF 2003 a set of 160 topics (Topics
#41 - #200) was developed for these collections and relevance assessments were
made. No additional relevance judgements were made this year for the robust
task. However, the data collection was not completely constant over all three
CLEF campaigns which led to an inconsistency between relevance judgements
and documents. The SDA 95 collection has no relevance judgements for most
topics (#41 - #140). This inconsistency was accepted in order to increase the
size of the collection. One participant reported that exploiting the knowledge
would have resulted in an increase of approximately 10% in MAP [12]. However,
participants were not allowed to use this knowledge. The results of the original
submissions for the data sets were analyzed in order to identify the most
difficult topics. This turned out to be an impossible task. The difficulty of a topic
varies greatly among languages, target collections and tasks. This confirms the
finding of the TREC 2005 robust task where the topic difficulty differed greatly
even for two different English collections. It was found that topics are not
inherently difficult but only in combination with a specific collection [13]. Topic
difficulty is usually defined by low MAP values for a topic. We also considered
a low number of relevant documents and high variation between systems as
indicators for difficulty. Consequently, the topic set for the robust task at CLEF
2006 was arbitrarily split into two sets. Participants were allowed to use the
available relevance assessments for the set of 60 training topics. The remaining
100 topics formed the test set for which results are reported. The participants
were encouraged to submit results for training topics as well. These runs will be
used to further analyze topic difficulty. The robust task received a total of 133
runs from eight groups listed in Table 5(a).</p>
          <p>Most popular among the participants were the mono-lingual French and
English tasks. For the multi-lingual task, four groups submitted ten runs. The
bilingual tasks received fewer runs. A run using title and description was
mandatory for each group. Participants were encouraged to run their systems with the
same setup for all robust tasks in which they participated (except for language
specific resources). This way, the robustness of a system across languages could
be explored.</p>
          <p>Effectiveness scores for the submissions were calculated with the GMAP
which is calculated as the n-th root of a product of n values. GMAP was
computed using the version 8.0 of trec eval2 program. In order to avoid undefined
results, all precision score lower than 0.00001 are set to 0.00001.
2 http://trec.nist.gov/trec_eval/trec_eval.8.0.tar.gz</p>
          <p>Track Participant Rank</p>
          <p>1st 2nd 3rd 4th 5th
Dutch hummingbird daedalus colesir</p>
          <p>MAP 51.06% 42.39% 41.60%
GMAP 25.76% 17.57% 16.40%</p>
          <p>Run humNL06Rtde nlFSnlR2S CoLesIRnlTst
English hummingbird reina dcu daedalus colesir</p>
          <p>MAP 47.63% 43.66% 43.48% 39.69% 37.64%
GMAP 11.69% 10.53% 10.11% 8.93% 8.41%</p>
          <p>Run humEN06Rtde reinaENtdtest dcudesceng12075 enFSenR2S CoLesIRenTst
French unine hummingbird reina dcu colesir</p>
          <p>MAP 47.57% 45.43% 44.58% 41.08% 39.51%
GMAP 15.02% 14.90% 14.32% 12.00% 11.91%</p>
          <p>Run UniNEfrr1 humFR06Rtde reinaFRtdtest dcudescfr12075 CoLesIRfrTst
German hummingbird colesir daedalus</p>
          <p>MAP 48.30% 37.21% 34.06%
GMAP 22.53% 14.80% 10.61%</p>
          <p>Run humDE06Rtde CoLesIRdeTst deFSdeR2S
Italian hummingbird reina dcu daedalus colesir</p>
          <p>MAP 41.94% 38.45% 37.73% 35.11% 32.23%
GMAP 11.47% 10.55% 9.19% 10.50% 8.23%</p>
          <p>Run humIT06Rtde reinaITtdtest dcudescit1005 itFSitR2S CoLesIRitTst
Spanish hummingbird reina dcu daedalus colesir</p>
          <p>MAP 45.66% 44.01% 42.14% 40.40% 40.17%
GMAP 23.61% 22.65% 21.32% 19.64% 18.84%
Run humES06Rtde reinaEStdtest dcudescsp12075 esFSesR2S CoLesIResTst</p>
          <p>Ad−Hoc Robust Monolingual English track Top 5 Participants − Interpolated Recall vs Average Precision
100%
90%
80%
70%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-11">
          <title>Interpolated Recall</title>
          <p>0%
0%
10%
20%
30%
40% 50% 60%</p>
          <p>Interpolated Recall
70%
80%
90%
100%
hummingbird [humEN06Rtde; MAP 47.63%; Not Pooled]
reina [reinaENtdtest; MAP 43.66%; Not Pooled]
dcu [dcudesceng12075; MAP 43.48%; Not Pooled]
daedalus [enFSenR2S; MAP 39.69%; Not Pooled]
colesir [CoLesIRenTst; MAP 37.64%; Not Pooled]
100%
90%
80%
70%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-12">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%</p>
          <p>Fig. 12. Robust Monolingual French.</p>
          <p>Ad−Hoc Robust Monolingual German track Top 5 Participants − Interpolated Recall vs Average Precision
hummingbird [humDE06Rtde; MAP 48.30%; Not Pooled]
colesir [CoLesIRdeTst; MAP 37.21%; Not Pooled]
daedalus [deFSdeR2S; MAP 34.06%; Not Pooled]
Ad−Hoc Robust Monolingual French track Top 5 Participants − Interpolated Recall vs Average Precision
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-13">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-14">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%</p>
          <p>Fig. 14. Robust Monolingual Italian.</p>
          <p>Ad−Hoc Robust Monolingual Spanish track Top 5 Participants − Interpolated Recall vs Average Precision
100%
90%
80%
70%
ion 60%
s
i
c
e
r
P 50%
e
g
a
r
e
vA 40%
30%
20%
10%</p>
          <p>Ad−Hoc Robust Monolingual Italian track Top 5 Participants − Interpolated Recall vs Average Precision
hummingbird [humIT06Rtde; MAP 41.94%; Not Pooled]
reina [reinaITtdtest; MAP 38.45%; Not Pooled]
dcu [dcudescit1005; MAP 37.73%; Not Pooled]
daedalus [itFSitR2S; MAP 35.11%; Not Pooled]
colesir [CoLesIRitTst; MAP 32.23%; Not Pooled]
hummingbird [humES06Rtde; MAP 45.66%; Not Pooled]
reina [reinaEStdtest; MAP 44.01%; Not Pooled]
dcu [dcudescsp12075; MAP 42.14%; Not Pooled]
daedalus [esFSesR2S; MAP 40.40%; Not Pooled]
colesir [CoLesIResTst; MAP 40.17%; Not Pooled]
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-15">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%
– X → ES: 80.88% of best monolingual Spanish IR system;
– X → NL: 69.27% of best monolingual Dutch IR system.
Some participants relied on the high correlation between the measure and
optimized their systems as in previous campaigns. However, several groups worked
specifically at optimizing for robustness. The SINAI system took an approach
which has proved successful at the TREC robust task, expansion with terms
gathered from a web search engine [14]. The REINA system from the
University of Salamanca used a heuristic to determine hard topics during training.
Ad−Hoc Robust Bilingual track, Dutch target collection(s) Top 5 Participants − Interpolated Recall vs Average Precision
100%</p>
          <p>daedalus [nlFSnlRLfr2S; MAP 35.37%; Not Pooled]
90%
80%
70%
ion 60%
s
i
c
e
r
P 50%
e
g
a
r
e
vA 40%
30%
20%
10%
daedalus [deFSdeRSen2S; MAP 29.16%; Not Pooled]
colesir [CoLesIRendeTst; MAP 25.24%; Not Pooled]
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-16">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%
Fig. 17. Robust Bilingual German
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-17">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%</p>
          <p>Fig. 16. Robust Bilingual Dutch
Ad−Hoc Robust Bilingual track, German target collection(s) Top 5 Participants − Interpolated Recall vs Average Precision
100%
Ad−Hoc Robust Bilingual track, Spanish target collection(s) Top 5 Participants − Interpolated Recall vs Average Precision
100%
reina [reinaIT2EStdtest; MAP 36.93%; Not Pooled]
dcu [dcuitqydescsp12075; MAP 33.22%; Not Pooled]
daedalus [esFSesRLit2S; MAP 26.89%; Not Pooled]
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-18">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%</p>
          <p>Fig. 18. Robust Bilingual Spanish
Ad−Hoc Robust Multilingual track Top 5 Participants − Interpolated Recall vs Average Precision
jaen [ujamlrsv2; MAP 27.85%; Not Pooled]
daedalus [mlRSFSen2S; MAP 22.67%; Not Pooled]
colesir [CoLesIRmultTst; MAP 22.63%; Not Pooled]
reina [reinaES2mtdtest; MAP 19.96%; Not Pooled]
90%
80%
70%
ion 60%
s
i
c
e
r
P 50%
e
g
a
r
e
vA 40%
30%
20%
10%
0%
0%
10%
20%
30%
40% 50% 60%</p>
        </sec>
        <sec id="sec-2-6-19">
          <title>Interpolated Recall</title>
          <p>70%
80%
90%
100%</p>
          <p>Subsequently, different expansion techniques were applied [15]. Hummingbird
experimented with other evaluation measures than those used in the track [16].
The MIRACLE system tried to find a fusion scheme which had a positive effect
on the robust measure [17].
6
When the goal is to validate how well results can be expected to hold beyond a
particular set of queries, statistical testing can help to determine what differences
between runs appear to be real as opposed to differences that are due to sampling
issues. We aim to identify runs with results that are significantly different from
the results of other runs. “Significantly different” in this context means that
the difference between the performance scores for the runs in question appears
greater than what might be expected by pure chance. As with all statistical
testing, conclusions will be qualified by an error probability, which was chosen
to be 0.05 in the following. We have designed our analysis to follow closely the
methodology used by similar analyses carried out for TREC [18].</p>
          <p>We used the MATLAB Statistics Toolbox, which provides the necessary
functionality plus some additional functions and utilities. We use the ANalysis Of
VAriance (ANOVA) test. ANOVA makes some assumptions concerning the data
be checked. Hull [18] provides details of these; in particular, the scores in
question should be approximately normally distributed and their variance has to be
approximately the same for all runs. Two tests for goodness of fit to a normal
distribution were chosen using the MATLAB statistical toolbox: the Lilliefors
test [19] and the Jarque-Bera test [20]. In the case of the CLEF tasks under
analysis, both tests indicate that the assumption of normality is violated for
most of the data samples (in this case the runs for each participant).</p>
          <p>In such cases, a transformation of data should be performed. The
transformation for measures that range from 0 to 1 is the arcsin-root transformation:
arcsin √x
which Tague-Sutcliffe [21] recommends for use with precision/recall measures.</p>
          <p>
            Table 11 shows the results of both the Lilliefors and Jarque-Bera tests before
and after applying the Tague-Sutcliffe transformation. After the transformation
the analysis of the normality of samples distribution improves significantly, with
some exceptions. The difficulty to transform the data into normally distributed
samples derives from the original distribution of run performances which tend
towards zero within the interval [
            <xref ref-type="bibr" rid="ref1">0,1</xref>
            ].
          </p>
          <p>
            In the following sections, two different graphs are presented to summarize
the results of this test. All experiments, regardless of topic language or topic
fields, are included. Results are therefore only valid for comparison of individual
pairs of runs, and not in terms of absolute performance. Both for the ad-hoc
and robust tasks, only runs where significant differences exist are shown; the
remainder of the graphs can be found in the Appendices [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ].
          </p>
          <p>The first graph shows participants’ runs (y axis) and performance obtained
(x axis). The circle indicates the average performance (in terms of Precision)
while the segment shows the interval in which the difference in performance is
not statistically significant.</p>
          <p>The second graph shows the overall results where all the runs that are
included in the same group do not have a significantly different performance. All
runs scoring below a certain group perform significantly worse than at least
the top entry of the group. Likewise all the runs scoring above a certain group
perform significantly better than at least the bottom entry in that group. To
determine all runs that perform significantly worse than a certain run, determine
the rightmost group that includes the run, all runs scoring below the bottom
entry of that group are significantly worse. Conversely, to determine all runs
that perform significantly better than a given run, determine the leftmost group
that includes the run. All runs that score better than the top entry of that group
perform significantly better.
7</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>We have reported the results of the ad hoc cross-language textual document
retrieval track at CLEF 2006. This track is considered to be central to CLEF as</p>
      <p>Ad−Hoc Monolingual French track − Tukey T test with "top group" highlighted
Fig. 20. Ad-Hoc Monolingual French. Experiments grouped according to the
Tukey T Test.</p>
      <p>UniNEhu2</p>
      <p>UniNEhu1
02aplmohutd4
02aplmohutdn4</p>
      <p>UniNEhu3
30dfrexp
plain2
plain
8dfrexp
0.4
0.45
0.5</p>
      <p>0.55 0.6
arcsin(sqrt(Mean average precision))
Fig. 21. Ad-Hoc Monolingual Hungarian. Experiments grouped according to the
Tukey T Test.
CELItitleCwnCascadeExpansion</p>
      <p>CELItitleNOEXPANSION
CELIdescNOEXPANSION
CELItitleLisaExpansion
DsvAmhEngFul Nofuzz</p>
      <p>CELItitleCwnExpansion
s
tn CELItitleNOEXPANSIONboost
e
m
ir
xep CELIdescLisaExpansion
E</p>
      <p>CELIdescLisaExpansionboost
CELIdescCwnCascadeExpansion
CELIdescNOEXPANSIONboost
CELItitleLisaExpansionboost</p>
      <p>CELIdescCwnExpansion
aplbi nen5
aplbi nen4
UI_td_mt
OMTD
OMTDN
UI_title_mt</p>
      <p>OMT
UI_td_dic</p>
      <p>UI_title_dic
DsvAmhEngFul Weighted</p>
      <p>DsvAmhEngFul</p>
      <p>UI_td_dicExp
DsvAmhEngTitle
UI_title_dicExp</p>
      <p>HNTD
HNT
TETD</p>
      <p>TET
UI_td_prl
UI_title_prl
−0.1
0
Fig. 22. Ad-Hoc Bilingual English. Experiments grouped according to the Tukey
T Test.</p>
      <p>Ad−Hoc Bilingual track, Portuguese target collection(s) − Tukey T test with "top group" highlighted
UniNEBipt2
UniNEBipt1
UniNEBipt3
ptFSptSfr3S
aplbiesptn
aplbiesptd
aplbienptn
aplbienptd</p>
      <p>QMUL06e2p10b
s
tne UBen2ptTDNrf3
m
ir
pe UBen2ptTDNrf2
x
E</p>
      <p>UBen2ptTDNrf1
ptFSptSen3S
UBen2ptTDrf3
UBen2ptTDrf2
ptFSptLes3S
UBen2ptTDrf1
ptFSptSen2S
XLDBBiRel16qe20k
XLDBBiRel32qe10k
XLDBBiRel32qe20k
XLDBBiRel16qe10k
Fig. 23. Ad-Hoc Bilingual Portuguese. Experiments grouped according to the
Tukey T Test.</p>
      <p>Ad−Hoc Robust Monolingual German track − Tukey T test with "top group" highlighted
humDE06Rtde
humDE06Rtd
dex011deRFW3FS3
s
t
n
e
irmdex021deRFW3FS3
e
p
x
E</p>
      <p>deFSdeR3S
CoLesIRdeTst
deFSdeR2S
0.55
Fig. 24. Robust Monolingual German. Experiments grouped according to the
Tukey T Test.</p>
      <p>Ad−Hoc Robust Monolingual Dutch track − Tukey T test with "top group" highlighted
s
t
n
e
m
i
r
e
p
x
E
humNL06Rtde
humNL06Rtd</p>
      <p>nlFSnlR4S
nlynlRFS3456
nlx011nlRFW4FS4</p>
      <p>CoLesIRnlTst
nlFSnlR2S
Fig. 25. Robust Monolingual Dutch. Experiments grouped according to the Tukey
T Test.</p>
      <p>Ad−Hoc Robust Bilingual track, Spanish target collection(s) − Tukey T test with "top group" highlighted
reinaIT2EStdtest
dcuitqydescsp12075
esx011esRLitFW3FS3
s
t
n
e
m
i
r
e
p
x
E
esx021esRLitFW3FS3
esFSesRLit3S
reinaIT2ESttest
esFSesRLit2S
0.4
Fig. 26. Robust Bilingual Spanish. Experiments grouped according to the Tukey
T Test.</p>
      <p>Ad−Hoc Robust Multilingual track: ONLY experiments with TEST topics − Tukey T test with "top group" highlighted
ujamlrsv2
ujamllr
ujamlblr
ml5XRSFSen4S
s
t
n
e
irmml6XRSFSen4S
e
p
x
E
ml4XRSFSen4S</p>
      <p>mlRSFSen2S
CoLesIRmultTst
reinaES2mtdtest
reinaES2mttest
0.35
0.4</p>
      <p>0.45 0.5 0.55
arcsin(sqrt(Mean average precision))
Fig. 27. Robust Multilingual. Experiments grouped according to the Tukey T Test.
for many groups it is the first track in which they participate and provides them
with an opportunity to test their systems and compare performance between
monolingual and cross-language runs, before perhaps moving on to more complex
system development and subsequent evaluation. However, the track is certainly
not just aimed at beginners. It also gives groups the possibility to measure
advances in system performance over time. In addition, each year, we also include
a task aimed at examining particular aspects of cross-language text retrieval.
This year, the focus was examining the impact of ”hard” topics on performance
in the ”robust” task.</p>
      <p>Thus, although the ad hoc track in CLEF 2006 offered the same target
languages for the main mono- and bilingual tasks as in 2005, it also had two new
focuses. Groups were encouraged to use non-European languages as topic
languages in the bilingual task. We were particularly interested in languages for
which few processing tools were readily available, such as Amharic, Oromo and
Telugu. In addition, we set up the ”robust task” with the objective of providing
the more expert groups with the chance to do in-depth failure analysis.</p>
      <p>Finally, it should be remembered that, although over the years we vary the
topic and target languages offered in the track, all participating groups also
have the possibility of accessing and using the test collections that have been
created in previous years for all of the twelve languages included in the CLEF
multilingual test collection. The test collections for CLEF 2000 - CLEF 2003 are
about to be made publicly available on the Evaluations and Language resources
Distribution Agency (ELDA) catalog3.
3 http://www.elda.org/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Braschler</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>CLEF 2003 - Overview of results</article-title>
          . In Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kluck</surname>
          </string-name>
          , M., eds.:
          <source>Comparative Evaluation of Multilingual Information Access Systems: Fourth Workshop of the Cross-Language Evaluation Forum (CLEF 2003) Revised Selected Papers, Lecture Notes in Computer Science (LNCS) 3237</source>
          , Springer, Heidelberg, Germany (
          <year>2004</year>
          )
          <fpage>44</fpage>
          -
          <lpage>63</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Di</given-names>
            <surname>Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.M.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          , N.:
          <article-title>Appendix A. Results of the Core Tracks</article-title>
          . In Nardi,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Vicedo</surname>
          </string-name>
          , J.L., eds.
          <source>: Working Notes for the CLEF 2006 Workshop</source>
          , Published Online (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Braschler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>CLEF 2003 Methodology and Metrics</article-title>
          . In Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kluck</surname>
          </string-name>
          , M., eds.:
          <source>Comparative Evaluation of Multilingual Information Access Systems: Fourth Workshop of the Cross-Language Evaluation Forum (CLEF 2003) Revised Selected Papers, Lecture Notes in Computer Science (LNCS) 3237</source>
          , Springer, Heidelberg, Germany (
          <year>2004</year>
          )
          <fpage>7</fpage>
          -
          <lpage>20</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Hal´acsy, P.:
          <article-title>Benefits of Deep NLP-based Lemmatization for Information Retrieval</article-title>
          . In Nardi,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Vicedo</surname>
          </string-name>
          , J.L., eds.
          <source>: Working Notes for the CLEF 2006 Workshop</source>
          , Published Online (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Moreira</given-names>
            <surname>Orengo</surname>
          </string-name>
          , V.:
          <article-title>A Study on the use of Stemming for Monolingual Ad-Hoc Portuguese</article-title>
          . Information
          <string-name>
            <surname>Retrieval</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>