<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <article-id pub-id-type="doi">10.2455/TUKEY</article-id>
      <title-group>
        <article-title>CLEF 2007: Ad Hoc Track Overview</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giorgio M. Di Nunzio</string-name>
          <email>dinunzio@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <email>ferro@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Mandl</string-name>
          <email>mandl@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carol Peters</string-name>
          <email>carol.peters@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>General Terms Experimentation</institution>
          ,
          <addr-line>Performance, Measurement, Algorithms</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ISTI-CNR, Area di Ricerca</institution>
          ,
          <addr-line>Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Information Science, University of Hildesheim</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2007</year>
      </pub-date>
      <abstract>
        <p>We describe the objectives and organization of the CLEF 2007 ad hoc track and discuss the main characteristics of the tasks offered to test monolingual and cross-language textual document retrieval systems. The track was divided into two streams. The main stream offered mono- and bilingual tasks on target collections for central European languages (Bulgarian, Czech and Hungarian). Similarly to last year, a bilingual task encouraging system testing with non-European languages against English documents was also offered; this year, particular attention was given to Indian languages. The second stream, designed for more experienced participants, offered mono- and bilingual ”robust” tasks with the objective of privileging experiments which achieve good stable performance over all queries rather than high average performance. These experiments re-used CLEF test collections from previous years in three languages (English, French, and Portuguese). The performance achieved for each task is presented and a statistical analysis of results is given.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The ad hoc retrieval track is generally considered to be the core track in the
Cross-Language Evaluation Forum (CLEF). The aim of this track is to promote
the development of monolingual and cross-language textual document retrieval
systems. Similarly to last year, the CLEF 2007 ad hoc track was structured in
two streams. The main stream offered mono- and bilingual retrieval tasks on
target collections for central European languages plus a bilingual task
encouraging system testing with non-European languages against English documents.
The second stream, designed for more experienced participants, was the
”robust task”, aimed at finding documents for very difficult queries. It used test
collections developed in previous years.</p>
      <p>The Monolingual and Bilingual tasks were principally offered for
Bulgarian, Czech and Hungarian target collections. Additionally, a bilingual task was
offered to test querying with non-European language queries against an English
target collection. As a result of requests from a number of Indian research
institutes, a special sub-task for Indian languages was offered with topics in Bengali,
Hindi, Marathi, Tamil and Telugu. The aim in all cases was to retrieve relevant
documents from the chosen target collection and submit the results in a ranked
list.</p>
      <p>The Robust task proposed mono- and bilingual experiments using the test
collections built over the last six CLEF campaigns. Collections and topics in
English, Portuguese and French were used. The goal of the robust analysis is
to improve the user experience with a retrieval system. Poor performing topics
are more serious for the user than performance losses in the middle and upper
interval. The robust task gives preference to systems which achieve a minimal
level for all topics. The measure used to assure this, is the geometric mean over
all topics. The robust task intends to evaluate stable performance over all topics
instead of high average performance.</p>
      <p>This was the first year since CLEF began that we have not offered a
Multilingual ad hoc task (ie searching a target collection in multiple languages).</p>
      <p>In this paper we describe the track setup, the evaluation methodology and the
participation in the different tasks (Section 2), present the main characteristics
of the experiments and show the results (Sections 3 - 5). Statistical testing
is discussed in Section 6 and the final section provides a brief summing up.
For information on the various approaches and resources used by the groups
participating in this track and the issues they focused on, we refer the reader to
the other papers in the Ad Hoc section of these Working Notes.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Track Setup</title>
      <p>The ad hoc track in CLEF adopts a corpus-based, automatic scoring method
for the assessment of system performance, based on ideas first introduced in the
Cranfield experiments in the late 1960s. The test collection used consists of a
set of “topics” describing information needs and a collection of documents to be
searched to find those documents that satisfy these information needs.
Evaluation of system performance is then done by judging the documents retrieved
in response to a topic with respect to their relevance, and computing the recall
and precision measures. The distinguishing feature of CLEF is that it applies
this evaluation paradigm in a multilingual setting. This means that the criteria
normally adopted to create a test collection, consisting of suitable documents,
sample queries and relevance assessments, have been adapted to satisfy the
particular requirements of the multilingual context. All language dependent tasks
such as topic creation and relevance judgment are performed in a distributed
setting by native speakers. Rules are established and a tight central
coordination is maintained in order to ensure consistency and coherency of topic and
relevance judgment sets over the different collections, languages and tracks.
2.1</p>
      <sec id="sec-2-1">
        <title>Test Collections</title>
        <p>Different test collections were used in the ad hoc task this year. The main stream
used national newspaper documents from 2002 as the target collections, creating
sets of new topics and making new relevance assessments. The robust task reused
existing CLEF test collections and did not create any new topics or make any
fresh relevance assessments.</p>
        <p>Documents. The document collections used for the CLEF 2007 ad hoc tasks are
part of the CLEF multilingual corpus of newspaper and news agency documents
described in the Introduction to these Proceedings.</p>
        <p>In the main stream monolingual and bilingual tasks, Bulgarian, Czech,
Hungarian and English national newspapers for 2002 were used. Much of this data
represented new additions to the CLEF multilingual comparable text corpora:
Czech is a totally new language in the ad hoc track although it was introduced
into the speech retrieval track last year; the Bulgarian collection was expanded
with the addition of another national newspaper, and in order to have
comparable data for English, we acquired a new American-English collection: Los
Angeles Times 2002. Table 1 summarizes the collections used for each language.</p>
        <p>The robust task used test collections containing news documents for the
period 1994-1995 in three languages (English, French, and Portuguese) used in
CLEF 2000 through CLEF 2006. Table 2 summarizes the collections used for
each language.</p>
        <p>Topics Topics in the CLEF ad hoc track are structured statements representing
information needs; the systems use the topics to derive their queries. Each topic
consists of three parts: a brief “title” statement; a one-sentence “description”; a
more complex “narrative” specifying the relevance assessment criteria.</p>
        <p>Sets of 50 topics were created for the CLEF 2007 ad hoc mono- and bilingual
tasks. All topic sets were created by native speakers. One of the decisions taken
early on in the organization of the CLEF ad hoc tracks was that the same set
of topics would be used to query all collections, whatever the task. There were
a number of reasons for this: it makes it easier to compare results over different
collections, it means that there is a single master set that is rendered in all
query languages, and a single set of relevance assessments for each language is
sufficient for all tasks. In CLEF 2006 we deviated from this rule as we were
using document collections from two distinct periods (1994/5 and 2002) and
created partially separate (but overlapping) sets with a common set of
timeindependent topics and separate sets of time-specific topics. As we had expected
this really complicated our lives as we had to build more topics and had to
specify very carefully which topic sets were to be used against which document
collections1. We determined not to repeat this experience this year and thus only
used collections from the same time period.</p>
        <p>We created topics in both European and non-European languages. European
language topics were offered for Bulgarian, Czech, English, French, Hungarian,
Italian and Spanish. The non-European languages were prepared according to
demand from participants. This year we had Amharic, Chinese, Indonesian, Oromo
plus the group of Indian languages: Bengali, Hindi, Marathi, Tamil and Telugu.</p>
        <p>The provision of topics in unfamiliar scripts did lead to some problems. These
were not caused by encoding issues (all CLEF data is encoded using UTF-8) but
rather by errors in the topic sets which were very difficult for us to spot. Although
most such problems were quickly noted and corrected, and the participants were
informed so that they all used the right set, one did escape our notice: the title
of Topic 430 in the Czech set was corrupted and systems using Czech thus did
not do well with this topic. It should be remembered, however, that an error is
one topic does not really impact significantly on the comparative results of the
systems. The topic will, however, be corrected for future use.</p>
        <p>
          This year topics have been identified by means of a Digital Object Identifier
(DOI)2 of the experiment [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] which allows us to reference and cite them. Below
we give an example of the English version of a typical CLEF 2007 topic:
&lt;top lang="en"&gt;
&lt;num&gt;10.2452/401-AH&lt;/num&gt;
&lt;title&gt;Euro Inflation&lt;/title&gt;
&lt;desc&gt;Find documents about rises in prices after the introduction of the
Euro.&lt;/desc&gt;
1 This is something that anyone reusing the CLEF 2006 ad hoc test collection needs
to be very careful about.
2 http://www.doi.org/
&lt;narr&gt;Any document is relevant that provides information on the rise of
prices in any country that introduced the common European
currency.&lt;/narr&gt;
&lt;/top&gt;
        </p>
        <p>For the robust task, the topic sets from CLEF 2001 to 2006 in English,
French and Portuguese were used. For English and French, which have been
part of CLEF for more time, training topics were offered and a set of 100 topics
were used for testing. For Portuguese, no training topics were possible and a set
of 150 test topics was used.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Participation Guidelines</title>
        <p>To carry out the retrieval tasks of the CLEF campaign, systems have to build
supporting data structures. Allowable data structures include any new structures
built automatically (such as inverted files, thesauri, conceptual networks, etc.)
or manually (such as thesauri, synonym lists, knowledge bases, rules, etc.) from
the documents. They may not, however, be modified in response to the topics,
e.g. by adding topic words that are not already in the dictionaries used by their
systems in order to extend coverage.</p>
        <p>Some CLEF data collections contain manually assigned, controlled or
uncontrolled index terms. The use of such terms is limited to specific experiments that
have to be declared as “manual” runs.</p>
        <p>Topics can be converted into queries that a system can execute in many
different ways. CLEF strongly encourages groups to determine what constitutes
a base run for their experiments and to include these runs (officially or
unofficially) to allow useful interpretations of the results. Unofficial runs are those
not submitted to CLEF but evaluated using the trec eval package. This year
we have used the new package written by Chris Buckley for the Text REtrieval
Conference (TREC) (trec eval 8.0) and available from the TREC website3.</p>
        <p>As a consequence of limited evaluation resources, a maximum of 12 runs each
for the mono- and bilingual tasks was allowed (no more than 4 runs for any one
language combination - we try to encourage diversity). For bi- and monolingual
robust tasks, 4 runs were allowed per language or language pair.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Relevance Assessment</title>
        <p>The number of documents in large test collections such as CLEF makes it
impractical to judge every document for relevance. Instead approximate recall values
are calculated using pooling techniques. The results submitted by the groups
participating in the ad hoc tasks are used to form a pool of documents for each
topic and language by collecting the highly ranked documents from selected runs
according to a set of predefined criteria. Traditionally, the top 100 ranked
documents from each of the runs selected are included in the pool; in such a case we</p>
        <sec id="sec-2-3-1">
          <title>3 http://trec.nist.gov/trec_eval/</title>
          <p>say that the pool is of depth 100. This pool is then used for subsequent relevance
judgments. After calculating the effectiveness measures, the results are analyzed
and run statistics produced and distributed.</p>
          <p>
            The stability of pools constructed in this way and their reliability for
postcampaign experiments is discussed in [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ] with respect to the CLEF 2003 pools.
New pools were formed in CLEF 2007 for the runs submitted for the main stream
mono- and bilingual tasks. Instead, the robust tasks used the original pools and
relevance assessments from previous CLEF campaigns.
          </p>
          <p>The main criteria used when constructing these pools were:
– favour diversity among approaches adopted by participants, according to the
descriptions of the experiments provided by the participants;
– choose at least one experiment for each participant in each task, chosen
among the experiments with highest priority as indicated by the participant;
– add mandatory title+description experiments, even though they do not have
high priority;
– add manual experiments, when provided;
– for bilingual tasks, ensure that each source topic language is represented.</p>
          <p>One important limitation when forming the pools is the number of documents
to be assessed. We estimate that assessors can judge from 60 to 100 documents
per hour, providing binary judgments: relevant / not relevant. This is actually an
optimistic estimate and shows what a time-consuming and resource expensive
task human relevance assessment is. This limitation impacts strongly on the
application of the criteria above - and implies that we are obliged to be flexible
in the number of documents judged per selected run for individual pools.</p>
          <p>
            This meant that this year, in order to create pools of more-or-less equivalent
size (approx. 20,000 documents), the depth of the Bulgarian, Czech and
Hungarian pools varied: 60 for Czech and 80 for Bulgarian and Hungarian, rather
than the depth of 100 originally used to judge TREC ad hoc experiments4. In
his paper in these working notes, Tomlinson [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] makes some interesting
observations in this respect. He claims that on average, the percentage of relevant items
assessed was less than 60% for Czech, 70% for Bulgarian and 85% for
Hungarian. However, as Tomlinson also points out, it has already been shown that test
collections created in this way do normally provide reliable results, even if not
all relevant documents are included in the pool.
          </p>
          <p>When building the pool for English, in order to respect the above criteria and
also to obtain a pool depth of 60, we had to include more than 25,000 documents.
Even so, as can be seen from Table 3, it was impossible to include very many
runs - just one monolingual and one bilingual run for each set of experiments. We
will certainly be performing some post-workshop stability tests on these pools.</p>
          <p>The box plot of Figure 1 compares the distributions of the relevant
documents across the topics of each pool for the different ad hoc pools; the boxes
4 Tests made on NTCIR pools in previous years have suggested that a depth of 60
in normally adequate to create stable pools, presuming that a sufficient number of
runs from different systems have been included
l
o
o
P
are ordered by decreasing mean number of relevant documents per topic. As can
be noted, Bulgarian, Czech, and Hungarian distributions appear similar, even
though the Czech and Hungarian ones are slightly more asymmetric towards
topics with a greater number of relevant documents. On the other hand, the
English distribution presents a greater number of relevant documents per topic,
with respect to the other distributions, and is quite asymmetric towards topics
with a greater number of relevant documents. All the distributions show some
upper outliers, i.e. topics with a great number of relevant document with
respect to the behaviour of the other topics in the distribution. These outliers are
probably due to the fact that CLEF topics have to be able to retrieve relevant
documents in all the collections; therefore, they may be considerably broader in
one collection compared with others depending on the contents of the separate
datasets. Thus, typically, each pool will have a different set of outliers.</p>
          <p>Table 3 reports summary information on the 2007 ad hoc pools used to
calculate the results for the main monolingual and bilingual experiments. In
particular, for each pool, we show the number of topics, the number of runs
submitted, the number of runs included in the pool, the number of documents
in the pool (relevant and non-relevant), and the number of assessors.
2.4</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>Result Calculation</title>
        <p>Evaluation campaigns such as TREC and CLEF are based on the belief that
the effectiveness of Information Retrieval Systems (IRSs) can be objectively
evaluated by an analysis of a representative set of sample search results. For
this, effectiveness measures are calculated based on the results submitted by the
participants and the relevance assessments. Popular measures usually adopted
for exercises of this type are Recall and Precision. Details on how they are
Pool size
Pool size
Pool size
Pool size
19,441 pooled documents
– 18,429 not relevant documents
– 1,012 relevant documents
50 topics
13 out of 18 submitted experiments
Pooled Experiments – monolingual: 11 out of 16 submitted experiments
– bilingual: 2 out of 2 submitted experiments
Assessors</p>
        <p>4 assessors
Czech Pool (DOI 10.2454/AH-CZECH-CLEF2007)
20,607 pooled documents
– 19,485 not relevant documents
– 762 relevant documents
50 topics
19 out of 29 submitted experiments
Pooled Experiments – monolingual: 17 out of 27 submitted experiments
– bilingual: 2 out of 2 submitted experiments
Assessors</p>
        <p>4 assessors
English Pool (DOI 10.2454/AH-ENGLISH-CLEF2007)
24,855 pooled documents
– 22,608 not relevant documents
– 2,247 relevant documents
50 topics
20 out of 104 submitted experiments
18,704 pooled documents
– 17,793 not relevant documents
– 911 relevant documents
50 topics
14 out of 21 submitted experiments
Pooled Experiments – monolingual: 10 out of 31 submitted experiments
– bilingual: 10 out of 73 submitted experiments
Assessors</p>
        <p>5 assessors</p>
        <p>Hungarian Pool (DOI 10.2454/AH-HUNGARIAN-CLEF2007)
Pooled Experiments – monolingual: 12 out of 19 submitted experiments
– bilingual: 2 out of 2 submitted experiments
Assessors</p>
        <p>
          6 assessors
calculated for CLEF are given in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. For the robust task, we used different
measures, see below Section 5.
        </p>
        <p>
          The individual results for all official ad hoc experiments in CLEF 2007 are
given in the Appendix at the end of these Working Notes [
          <xref ref-type="bibr" rid="ref5 ref6">5,6</xref>
          ].
As shown in Table 4, a total of 22 groups from 12 different countries submitted
results for one or more of the ad hoc tasks - a slight decrease on the 25
participants of last year. Table 5 provides a breakdown of the number of participants
by country.
        </p>
        <p>A total of 235 runs were submitted with a decrease of about 20% on the 296
runs of 2006. The average number of submitted runs per participant also slightly
decreased: from 11.7 runs/participant of 2006 to 10.6 runs/participant of this
year.</p>
        <p>Participants were required to submit at least one title+description (“TD”)
run per task in order to increase comparability between experiments. The large
majority of runs (138 out of 235, 58.72%) used this combination of topic fields,
50 (21.28%) used all fields, 46 (19.57%) used the title field, and only 1 (0.43%)
used the description field. The majority of experiments were conducted using
automatic query construction (230 out of 235, 97.87%) and only in a small fraction
of the experiments (5 out 237, 2.13%) were queries been manually constructed
from topics. A breakdown into the separate tasks is shown in Table 6(a).</p>
        <p>Fourteen different topic languages were used in the ad hoc experiments. As
always, the most popular language for queries was English, with Hungarian
second. The number of runs per topic language is shown in Table 6(b).
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Main Stream Monolingual Experiments</title>
      <p>Monolingual retrieval focused on central-European languages this year, with
tasks offered for Bulgarian, Czech and Hungarian. Eight groups presented results
for 1 or more of these languages. We also requested participants in the
bilingualto-English task to submit one English monolingual run, but only in order to
provide a baseline for their bilingual experiments and in order to strengthen the
English pool for relevance assessment5.</p>
      <p>Five of the participating groups submitted runs for all three languages. One
group was unable to complete its Bulgarian experiments, submitting results for
just the other two languages. The two groups from the Czech Republic only
submitted runs for Czech. From the graphs and from 7, it can be seen that the
best performing groups were more-or-less the same for each language and that
the results did not greatly differ. It should be noted that these are all veteran
participants with much experience at CLEF.
5 Ten groups submitted runs for monolingual English. We have included a graph
showing the top 5 results but it must be remembered that the systems submitting these
were actually focusing on the bilingual part of the task.
Participant Institution Country
alicante U.Alicante - Languages&amp;CS Spain
bohemia U.W.Bohemia Czech Republic
bombay-ltrc Indian Inst. Tech. India
budapest-acad Informatics Lab Hungary
colesun COLESIR &amp; U.Sunderland Spain
daedalus Daedalus &amp; Spanish Univ. Consortium Spain
depok U.Indonesia Indonesia
hildesheim U.Hildesheim Germany
hyderabad International Institute of Information Technology (IIIT) India
isi Indian Statistical Institute India
jadavpur Jadavpur University India
jaen U.Jaen-Intell.Systems Spain
jhu-apl Johns Hopkins University Applied Physics Lab United States
kharagpur IIT-Kharagpur-CS India
msindia Microsoft India India
nottingham U.Nottingham United Kingdom
opentext Open Text Corporation Canada
prague Charles U., Prague Czech Republic
reina U.Salamanca Spain
stockholm U. Stockholm Sweden
unine U.Neuchatel-Informatics Switzerland
xldb U.Lisbon Portugal</p>
      <p>
        As usual in the CLEF monolingual task, the main emphasis in the
experiments was on stemming and morphological analysis. The group from University
of Neuchatel, which had the best overall performances for all languages, focused
very much on stemming strategies, testing both light and aggressive stemmers
for the Slavic languages (Bulgarian and Czech). For Hungarian they worked on
decompounding. This group also compared performances obtained using
wordbased and 4-gram indexing strategies [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Another of the best performers,
JHUAPL, normally uses an n-gram approach. Unfortunately, we have not received a
paper yet from this group so cannot comment on their performance. The other
group with very good performance for all languages was Opentext. This group
also compared 4-gram results against results using stemming for all three
languages. They found that while there could be large impacts on individual topics,
there was little overall difference in average performance. Their experiments also
confirmed past findings that indicate that blind relevance feedback can be
detrimental to results, depending on the evaluation measures used [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The results
of the statistical tests given towards the end of this paper show that the best
results of these three groups did not differ significantly.
      </p>
      <p>
        The group from Alicante also achieved good results testing query expansion
techniques [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], while the group from Kolkata compared a statistical stemmer
against a rule-based stemmer for both Czech and Hungarian [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Czech is a
morphologically complex language and the two Czech only groups both used
approaches involving morphological analysis and lemmatization [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>Track</p>
    </sec>
    <sec id="sec-4">
      <title>Main Stream Bilingual Experiments</title>
      <p>The bilingual task was structured in three tasks (X → BG, CS, or HU target
collection) plus a task for non-European topic languages against an English
target collection. A special sub-task testing Indian languages against the English
collection was also organised in response to requests from a number of research
groups working in India. For the bilingual to English task, participating groups
also had to submit an English monolingual run, to be used both as baseline
and also to reinforce the English pool. All groups participating in the Indian
Ad−Hoc Monolingual Bulgarian Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
unine [Experiment UNINEBG4; MAP 44.22%; Pooled]
jhu−apl [Experiment APLMOBGTD4; MAP 36.57%; Not Pooled]
opentext [Experiment OTBG07TDE; MAP 35.02%; Not Pooled]
alicante [Experiment IRNBUEXP2N; MAP 29.81%; Pooled]
daedalus [Experiment BGFSBG2S; MAP 27.19%; Pooled]
50%</p>
      <p>Recall
10%
20%
30%
40%
60%
70%
80%
90%
100%</p>
      <p>Ad−Hoc Monolingual Czech Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
unine [Experiment UNINECZ4; MAP 42.42%; Pooled]
jhu−apl [Experiment APLMOCSTD4; MAP 35.86%; Not Pooled]
opentext [Experiment OTCS07TDE; MAP 34.84%; Not Pooled]
prague [Experiment PRAGUE01; MAP 34.19%; Pooled]
daedalus [Experiment CSFSCS2S; MAP 32.03%; Pooled]
0%
0%
10%
20%
30%
40%
50%
Recall
60%
70%
80%
90%
100%</p>
      <p>Fig. 3. Monolingual Czech
Ad−Hoc Monolingual Hungarian Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
unine [Experiment UNINEHU4; MAP 47.73%; Pooled]
opentext [Experiment OTHU07TDE; MAP 43.34%; Not Pooled]
alicante [Experiment IRNHUEXP2N; MAP 40.09%; Not Pooled]
jhu−apl [Experiment APLMOHUTD5; MAP 39.91%; Pooled]
daedalus [Experiment HUFSHU2S; MAP 34.99%; Pooled]
100%
90%
80%
70%
60%
10%
n
o
i
ics 50%
e
r
P
90%
80%
70%
60%
40%
30%
20%
10%
0%</p>
      <p>0%
0%
0%
10%
20%
30%
40%
languages sub-task also had to submit at least one run in Hindi (mandatory)
plus runs in other Indian languages (optional).</p>
      <p>We were disappointed to only receive runs from one participant for the X →
BG, CS, or HU tasks. Furthermore, the results were quite poor; as this group
normally achieves very good performance, we suspect that these runs were
probably corrupted in some way. For this reason, we have decided to disregard them
as being of little significance. Therefore, in the rest of this section, we only
comment briefly on the X → EN results.</p>
      <p>We received runs using the following topic languages: Amharic, Chinese,
Indonesian and Oromo plus, for the Indian sub-task, Bengali, Hindi, Marathi
and Telugu6.</p>
      <p>For many of these languages few processing tools or resources are available.
It is thus very interesting to see what measures the participants adopted to
overcome this problem. Unfortunately, there has not yet been time to read the
submitted reports from each group and, here below, we give just a first cursory
glance at some of the approaches and techniques adopted. We will provide a
more in-depth analysis at the workshop.</p>
      <p>The top performance in the bilingual task was obtained by an Indonesian
group; they compared different translation techniques: machine translation using
Internet resources, transitive translation using bilingual dictionaries and French
and German as pivot languages, and lexicons derived from parallel corpus
created by translating all the CLEF English documents into Indonesian using a
commercial MT system. They found that they obtained best results using the
MT system together with query expansion [12].</p>
      <p>The second placed group used Chinese for their queries and a dictionary based
translation technique. The experiments of this group concentrated on developing
new strategies to address two well-known CLIR problems: translation ambiguity,
and coverage of the lexicon [13]. The work by [14] which used Amharic as the
topic language also paid attention to the problems of sense disambiguation and
out-of-vocabulary terms.</p>
      <p>The third performing group also used Indonesian as the topic language;
unfortunately we have not received a paper from them so far so cannot comment
on their approach. An interesting paper, although slightly out of the task as the
topic language used was Hungarian was [15]. This group used a machine
readable dictionary approach but also applied Wikipedia data to eliminate unlikely
translations according to the conceptual context. The group testing Oromo used
linguistic and lexical resources developed at their institute; they adopted a
bilingual dictionary approach and also tested the impact of a light stemmer for Afaan
Oromo on their performance with positive results [16].</p>
      <p>The groups using Indian topic languages tested different approaches. The
group from Kolkata submitted runs for Bengali, Hindi and Telugu to English
using a bilingual dictionary lookup approach [17]. They had the best
performance using Telugu probably because they carried out some manual tasks
during indexing. A group from Bangalore tested a statistical MT system trained on
6 Although topics had also been requested in Tamil, in the end they were not used.
parallel aligned sentences and a language modelling based retrieval algorithm
for a Hindi to English system [18]. The group from Bombay had the best overall
performances; they used bilingual dictionaries for both Hindi and Marathi to
English and applied term-to-term cooccurrence statistics for sense
disambiguation [19]. The Hyderabad group also used bilingual lexicons for query translation
from Hindi and Telugu to English together with a variant of the TFIDF
algorithm and a hybrid boolean formulation for the queries to improve ranking [20].
Interesting work was done by the group from Kharagpur which submitted runs
for Hindi and Bengali. They attempted to overcome the lack of resources for
Bengali by using phoneme-based transliterations to generate equivalent English
queries from Hindi and Bengali topics [21].
7 Since for the other bilingual tasks only one participant submitted experiments, only
the graphs for bilingual English are reported
100%
90%
80%
70%
60%
last year for two well-established CLEF languages: French and Portuguese, when
the equivalent figures were 93.82% and 90.91%, respectively.</p>
      <sec id="sec-4-1">
        <title>4.2 Indian To English Subtask Results</title>
        <p>Table 9 shows the best results for the Indian sub-task. The performance
difference between the best and the last (up to 6) placed group is given (in terms
of average precision). The first set of rows regard experiments for the
mandatory topic language: Hindi; the second set of rows report experiments where the
source language is one of other Indian languages.</p>
        <p>It is interesting to note that in both sets of experiments, the best performing
participant is the same. In the second set, we can note that for three (Hindi,
Marathi, and Telegu) out of the four Indian languages used the performances of
the top groups are quite similar.</p>
        <p>The best performance for the Indian sub-task is 76.12% of the best
bilingual English system (achieved by veteran CLEF participants) and 67.06% of the
monolingual baseline, which is quite encouraging for a new task with languages
where encoding issues and linguistic resources make the task difficult. This is in
fact comparable with the performances of some newly introduced European
languages. For example, we can compare them to those for Bulgarian and Hungarian
in CLEF 2006:
– X → BG: 52.49% of best monolingual Bulgarian IR system;
– X → HU: 53.13% of best monolingual Hungarian IR system.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Robust Experiments</title>
      <p>The robust task ran for the second time at CLEF 2007. It is an ad-hoc retrieval
task based on data of previous CLEF campaigns. The evaluation approach is
modified and a different perspective is taken. The robust task emphasizes the
difficult topics by a non-linear integration of the results of individual topics
into one result for a system [22,23]. By doing this, the evaluation results are
interpreted in a more user oriented manner. Failures and very low results for
some topics hurt the user experience with a retrieval system. Consequently, any
system should try to avoid these failures. This has turned out to be a hard task
[24]. Robustness is a key issue for the transfer of research into applications. The
robust task rewards systems which achieve a minimal performance level for all
topics.</p>
      <p>In order to do this, the robust task uses the geometric mean of the average
precision for all topics (GMAP) instead of the mean average of all topics (MAP).
This measure has also been used at a roust track at the Text Retrieval Conference
(TREC) where robustness was explored for monolingual English retrieval [23].
At CLEF, robustness is evaluated for monolingual and bilingual retrieval for
several European languages.</p>
      <p>The robust task at CLEF exploits data created for previous CLEF editions.
Therefore, a larger data set can be used for the evaluation. A larger number of</p>
      <p>Language Target Collection
English LA Times 1994</p>
      <p>Training Topic DOIs Test Topic DOIs
10.2452/41-AH–10.2452/200-AH 10.2452/251-AH–10.2452/350-AH
French Le Monde 1994</p>
      <p>SDA 1994
Portuguese Pu´blico 1995
topics allows a more reliable evaluation [25]. A secondary goal of the robust task
is the definition of larger data sets for retrieval evaluation.</p>
      <p>As described above, the CLEF2007 robust task offered three languages often
used in previous CLEF campaigns: English, French and Portuguese. The data
used has been developed during CLEF 2001 through 2006. Generally, the topics
from CLEF 2001 until CLEF 2003 were training topics whereas the topics
developed between 2004 and 2006 were the test topics on which the main evaluation
measures are given.</p>
      <p>Thus, the data used in the robust task in 2007 is different from the set defined
for the roust task at CLEF 2006. The documents which need to be searched are
articles from major newspapers and news providers in the three languages. Not
all collections had been offered consistently for all CLEF campaigns, therefore,
not all collections were integrated into the robust task. Most data from 1995 was
omitted in order to provide a homogeneous collection. However, for Portuguese,
for which no training data was available, only data from 1995 was used. Table
10 shows the data for the robust task.</p>
      <p>The robust task attracted 63 runs submitted by 7 groups (CLEF 2006: 133
runs from 8 groups). Effectiveness scores were calculated with the version 8.0
of the program which provides the Mean Average Precision (MAP), while the
Geometric Average Precision (GMAP) was calculated using DIRECT version
2.0.</p>
      <sec id="sec-5-1">
        <title>Robust Monolingual Results</title>
        <p>Table 11 shows the best results for this task. The performance difference
between the best and the last (up to 5) placed group is given (in terms of average
precision).</p>
        <p>The results cannot be compared to the results of the CLEF 2005 and CLEF
2006 campaign in which the same topics were used because a smaller collection
had to be searched.</p>
        <p>Figures from 7 to 9 compare the performances of the top participants of the
Robust Monolingual.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Robust Bilingual Results</title>
        <p>Table 12 shows the best results for this task. The performance difference
between the best and the last (up to 5) placed group is given (in terms of average
precision). All the experiments where from English to French.</p>
        <p>Track Rank Participant
– X → FR: 85.05% of best monolingual French IR system;</p>
        <p>Figure 10 compares the performances of the top participants of the Robust
Bilingual task.</p>
      </sec>
      <sec id="sec-5-3">
        <title>Approaches Applied to Robust Retrieval</title>
        <p>The REINA system applied different measures of robustness during the training
phase in order to optimize the performance. A local query expansion technique
added terms. The CoLesIR system experimented with n-gram based translation
for bi-lingual retrieval which requires no languages specific components. SINAI
tried to increase the robustness of the results by expanding the query with an
external knowledge source. This is a typical approach in order to obtain additional
Ad−Hoc Robust Monolingual French Test Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
unine [Experiment UNINEFR1; MAP 42.13%; Not Pooled]
reina [Experiment REINAFRTDET; MAP 38.04%; Not Pooled]
jaen [Experiment UJARTFR1; MAP 34.76%; Not Pooled]
daedalus [Experiment FRFSFR22S; MAP 29.91%; Not Pooled]
hildesheim [Experiment HIMOFRBRF2; MAP 27.31%; Not Pooled]
100%
90%
80%
70%
60%
n
o
i
ics 50%
e
r
P
90%
80%
70%
60%
40%
30%
20%
10%
0%
0%
50%</p>
        <p>Recall
10%
20%
30%
40%
60%
70%
80%
90%
100%</p>
        <p>Fig. 8. Robust Monolingual English.
0%
0%
10%
20%
30%
40%
50%
Recall
60%
70%
80%
90%
100%</p>
        <p>Ad−Hoc Robust Monolingual English Test Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
reina [Experiment REINAENTDNT; MAP 38.97%; Not Pooled]
daedalus [Experiment ENFSEN22S; MAP 37.78%; Not Pooled]
hildesheim [Experiment HIMOENBRFNE; MAP 5.88%; Not Pooled]
100%
90%
80%
70%
60%
n
o
i
ics 50%
e
r
P
90%
80%
70%
60%
40%
30%
20%
10%
0%</p>
        <p>0%
Ad−Hoc Robust Monolingual Portuguese Test Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
reina [Experiment REINAPTTDNT; MAP 41.40%; Not Pooled]
jaen [Experiment UJARTPT1; MAP 24.74%; Not Pooled]
daedalus [Experiment PTFSPT2S; MAP 23.75%; Not Pooled]
xldb [Experiment XLDBROB16_10; MAP 1.21%; Not Pooled]
0%
0%
10%
20%
30%
40%
50%
Recall
60%
70%
80%
90%
100%</p>
        <p>Ad−Hoc Robust Bilingual Test Task, French target collection(s) Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
reina [Experiment REINAE2FTDNT; MAP 35.83%; Not Pooled]
unine [Experiment UNINEBILFR1; MAP 33.50%; Not Pooled]
colesun [Experiment EN2FRTST4GRINTLOGLU001; MAP 22.87%; Not Pooled]
50%</p>
        <p>Recall
10%
20%
30%
40%
60%
70%
80%
90%
100%</p>
        <p>Fig. 10. Robust Bilingual French
query terms and avoid zero hits in case of out of vocabulary problems. Contrary
to standard query expansion techniques, the new terms form a second query and
results of both initial and second query are integrated under a logistic fusion
strategy. The Daedalus group submitted experiments wit the Miracle system.
BM25 weighting without blind relevance feedback was applied. For descriptions
of all the robust experiments, see the Robust section in these Working Notes.
6
When the goal is to validate how well results can be expected to hold beyond a
particular set of queries, statistical testing can help to determine what differences
between runs appear to be real as opposed to differences that are due to sampling
issues. We aim to identify runs with results that are significantly different from
the results of other runs. “Significantly different” in this context means that
the difference between the performance scores for the runs in question appears
greater than what might be expected by pure chance. As with all statistical
testing, conclusions will be qualified by an error probability, which was chosen
to be 0.05 in the following. We have designed our analysis to follow closely the
methodology used by similar analyses carried out for TREC [26].</p>
        <p>We used the MATLAB Statistics Toolbox, which provides the necessary
functionality plus some additional functions and utilities. We use the ANalysis Of
VAriance (ANOVA) test. ANOVA makes some assumptions concerning the data
be checked. Hull [26] provides details of these; in particular, the scores in
question should be approximately normally distributed and their variance has to be
approximately the same for all runs. Two tests for goodness of fit to a normal
distribution were chosen using the MATLAB statistical toolbox: the Lilliefors
test [27] and the Jarque-Bera test [28]. In the case of the CLEF tasks under
analysis, both tests indicate that the assumption of normality is violated for
most of the data samples (in this case the runs for each participant).</p>
        <p>In such cases, a transformation of data should be performed. The
transformation for measures that range from 0 to 1 is the arcsin-root transformation:
arcsin √x
which Tague-Sutcliffe [29] recommends for use with precision/recall measures.</p>
        <p>
          Table 13 shows the results of both the Lilliefors and Jarque-Bera tests before
and after applying the Tague-Sutcliffe transformation. After the transformation
the analysis of the normality of samples distribution improves significantly, with
some exceptions. The difficulty to transform the data into normally distributed
samples derives from the original distribution of run performances which tend
towards zero within the interval [
          <xref ref-type="bibr" rid="ref1">0,1</xref>
          ].
        </p>
        <p>In the following sections, two different graphs are presented to summarize
the results of this test. All experiments, regardless of topic language or topic
fields, are included. Results are therefore only valid for comparison of individual
pairs of runs, and not in terms of absolute performance. Both for the ad-hoc</p>
        <p>
          Track LF
Monolingual Bulgarian 8
Monolingual Czech 4
Monolingual English 22
Monolingual Hungarian 6
Bilingual Bulgarian 0
Bilingual Czech 0
Bilingual English 6
Bilingual Hungarian 0
Robust Monolingual English 3
Robust Monolingual French 2
Robust Monolingual Portuguese 0
Robust Bilingual French 0
and robust tasks, only runs where significant differences exist are shown; the
remainder of the graphs can be found in the Appendices [
          <xref ref-type="bibr" rid="ref5 ref6">5,6</xref>
          ].
        </p>
        <p>The first graph shows participants’ runs (y axis) and performance obtained
(x axis). The circle indicates the average performance (in terms of Precision)
while the segment shows the interval in which the difference in performance is
not statistically significant.</p>
        <p>The second graph shows the overall results where all the runs that are
included in the same group do not have a significantly different performance. All
runs scoring below a certain group perform significantly worse than at least
the top entry of the group. Likewise all the runs scoring above a certain group
perform significantly better than at least the bottom entry in that group. To
determine all runs that perform significantly worse than a certain run, determine
the rightmost group that includes the run, all runs scoring below the bottom
entry of that group are significantly worse. Conversely, to determine all runs
that perform significantly better than a given run, determine the leftmost group
that includes the run. All runs that score better than the top entry of that group
perform significantly better.</p>
        <p>UNINEBG4
UNINEBG1
UNINEBG2</p>
        <p>UNINEBG3
APLMOBGTD4</p>
        <p>OTBG07TDE
APLMOBGTDN4
s
t
n
e
irm APLMOBGTD5
e
p
x
E</p>
        <p>OTBG07TD</p>
        <p>OTBG07T
IRNBUEXP2N</p>
        <p>BGFSBG2S
OTBG07TDNZ
IRNBUEXP3
IRNBUEXP2
IRNBUNEXP
0.35
0.4
Tukey T Test (DOI 10.2455/TUKEY T TEST.9633F4C12FC97855A347B6AF7DD9B7A5).
UNINECZ2</p>
        <p>UNINECZ1
APLMOCSTD4
OTCS07TDE</p>
        <p>PRAGUE02
APLMOCSTD5</p>
        <p>PRAGUE01
OTCS07TD</p>
        <p>CSFSCS2S
APLMOCSTDN4</p>
        <p>PRAGUE03
OTCS07T</p>
        <p>PRAGUE04
OTCS07TDNZ</p>
        <p>IRNCZEXP2N
UWB_TD_L_BM25
UWB_TDN_L_BM25
UWB_TDN_W_BM25</p>
        <p>IRNCZEXP3
IRNCZEXP2
IRNCZNEXP</p>
        <p>CZTD</p>
        <p>ISICL</p>
        <p>ISICZNS
UWB_TD_W_RAW_FB
0.1
0.2
0.3</p>
        <p>Ad−Hoc Monolingual English Task (ONLY for English Pool creation and AH−BILI−X2EN participants) − Tukey T test with "top group" highlighted
IITB_MONO_TITLE_DESC</p>
        <p>APLMOENTDN4
APLMOENTD5</p>
        <p>MONOT
ENTD_OMENG07</p>
        <p>UIQTDMONO</p>
        <p>APLMOENTD4
LM_ALL_MONO_TD_1000_POSSCORES
2007_LM_ALL_MONO_1000_POSSCORES
s
t
n
e
m
ir
e
xp
E</p>
        <p>RBLM_ALL_MONO_TD_1000_POSSCORES
2007_RBLM_ALL_MONO_1000_POSSCORES</p>
        <p>IITB_MONO_TITLE
ENTDN_OMENG07</p>
        <p>UIQTMONO
OMENG07E
OMENG07
MONOTD</p>
        <p>ENGMONO</p>
        <p>ENGLISHTITLEDESC
ENGLISHTITLEDESCNARR</p>
        <p>UIQDMONO
MONOEN6
MONOEN2
MONOEN4
MONOEN3
MONOEN5</p>
        <p>MONOEN1</p>
        <p>ENGLISHTITLE
DSV_ENG_MONO_LONG
DSV_ENG_MONO_RUN0</p>
        <p>AHMONOENR1
0.2
0.3
0.4
0.5 0.6 0.7
arcsin(sqrt(Average Precsion))
0.8
0.9</p>
        <p>1</p>
        <p>s
t
n
e
m
ir
e
p
x
E</p>
        <p>UNINEHU4
UNINEHU2
UNINEHU1
UNINEHU3</p>
        <p>OTHU07TDE
APLMOHUTDN4
APLMOHUTD5
IRNHUEXP2N</p>
        <p>OTHU07TD</p>
        <p>IRNHUEXP3
APLMOHUTD4</p>
        <p>IRNHUEXP2
HUFSHU2S
IRNHUNEXP</p>
        <p>OTHU07T
OTHU07TDNZ
YASSTDHUN</p>
        <p>YASSHUN
ISIDWLDHSTEMGZ
0.2
0.3
0.4
0.5 0.6 0.7
arcsin(sqrt(Average Precsion))
0.8
0.9</p>
        <p>1
Fig. 14. Ad-Hoc Monolingual Hungarian. Experiments grouped according to the
UIQTDTOGGLEFB10D10T
UIQTTOGGLEFB10D10T</p>
        <p>UIQTDTOGGLE
UIQTDDESYNFB10D10T</p>
        <p>GRAWOTD
UIQTDINTERSECTIONUNIONSYNFB5D10T</p>
        <p>UIQTDDESYNFB5D10T</p>
        <p>APLBI DENTDS
UIQTDINTERSECTIONUNIONSYNFB10D10T</p>
        <p>UIQTTOGGLE
UIQTDDESYNFB10D5T</p>
        <p>GRAUWOTD</p>
        <p>GRAUWOT
APLBI DENTD5
APLBI DENTD4</p>
        <p>GRAWOT
OMTD07</p>
        <p>OMTDN07</p>
        <p>APLBI DENTDW</p>
        <p>I TB_HINDI_TITLEDESC_DICE
UIQTDINTERSECTIONUNIONSYN</p>
        <p>GRAWTD</p>
        <p>COOTD
I TB_HINDI_TITLEDESC_PMI</p>
        <p>OMT07
ALLTRANSTD
GRAUWTD
GRAUWT
GRAWT</p>
        <p>COOT
I TB_HINDI_TITLE_DICE</p>
        <p>ALLTRANST
tsen TETD
rm 2007_LM_ALL_CROSS_1000_POSSCORES
i
xpe I TB_MAR_TITLE_DICE
E2007_RBLM_ALL_CROSS_1000_POSSCORES
RBLM_ALL_CROSS_TD_1000_POSSCORES</p>
        <p>NOST_OMTDN07</p>
        <p>I TB_HINDI_TITLE_PMI
LM_ALL_CROSS_TD_1000_POSSCORES</p>
        <p>I TB_MAR_TITLE_PMI</p>
        <p>BILING_WIKI6
BILING_WIKI2
BILING_WIKI4
BILING_WIKI1
BILING_WIKI3
BILING_WIKI5</p>
        <p>HITD
BILING_6
BILING_4
BILING_2
BILING_5
BILING_1</p>
        <p>BILING_3
AHBILITE2ENR1</p>
        <p>AHBILIHI2ENR1
DSV_AMH_BLNG_RUN8</p>
        <p>AHBILIBN2ENR1
DSV_AMH_BLNG_RUN6
DSV_AMH_BLNG_RUN7
DSV_AMH_BLNG_RUN4
DSV_AMH_BLNG_RUN2
DSV_AMH_BLNG_RUN5</p>
        <p>BENGALITITLEDESC
BENGALITITLEDESCNARR
DSV_AMH_BLNG_RUN1
DSV_AMH_BLNG_RUN3
HINDITITLEDESCNARR</p>
        <p>HINDITITLE</p>
        <p>BENGALITITLE
UIQTDENGFB5D5T10D19T</p>
        <p>HINDITITLEDESC
UIQTDENGPSEUDOTRANS20D14T
0
Fig. 15. Ad-Hoc Bilingual English. The figure shows the Tukey T Test (DOI
10.2455/TUKEY T TEST.4E800CB597F00B0A2CEF6D50479867EF).</p>
        <p>Ad−Hoc Robust Monolingual English Test Task − Tukey T test with "top group" highlighted
REINAENTDT
REINAENTDNT
REINAENTDET</p>
        <p>ENFSEN22S</p>
        <p>REINAENTT
s
t
n
e
rm REINAENTET
i
e
p
x
E
HIMOENBRFNE
HIMOENBRF2
HIMOENBRF1</p>
        <p>HIMOENNE
HIMOENBASE
Tukey T Test (DOI 10.2455/TUKEY T TEST.90F31E2BBDA52E383201421CD2207C37).</p>
        <p>Ad−Hoc Robust Monolingual French Test Task − Tukey T test with "top group" highlighted</p>
        <p>UNINEFR1
REINAFRTDET</p>
        <p>UNINEFR2
REINAFRTDNT</p>
        <p>REINAFRTDT
s
t
n
e
irm UJARTFR1
e
p
x
E</p>
        <p>REINAFRTET
REINAFRTT
FRFSFR22S
HIMOFRBRF2
HIMOFRBRF
HIMOFRBASE</p>
        <p>Ad−Hoc Robust Monolingual Portuguese Test Task − Tukey T test with "top group" highlighted</p>
        <p>Ad−Hoc Robust Bilingual Test Task, French target collection(s) − Tukey T test with "top group" highlighted
REINAE2FTDNT
REINAE2FTDET
REINAE2FTDT
REINAE2FTET
UNINEBILFR1</p>
        <p>REINAE2FTT
EN2FRTST4GRINTLOGLU001
EN2FRTST4GRINTDICEU001</p>
        <p>EN2FRTST4GRINTMIU010</p>
        <p>Experiment DOI
10.2415/AH-ROBUST-BILI-X2FR-TEST-CLEF2007.REINA.REINAE2FTDNT
10.2415/AH-ROBUST-BILI-X2FR-TEST-CLEF2007.REINA.REINAE2FTDET
10.2415/AH-ROBUST-BILI-X2FR-TEST-CLEF2007.REINA.REINAE2FTDT
10.2415/AH-ROBUST-BILI-X2FR-TEST-CLEF2007.REINA.REINAE2FTET
10.2415/AH-ROBUST-BILI-X2FR-TEST-CLEF2007.UNINE.UNINEBILFR1
10.2415/AH-ROBUST-BILI-X2FR-TEST-CLEF2007.REINA.REINAE2FTT
10.2415/AH-ROBUST-BILI-X2FR-TEST-CLEF2007.COLESUN.EN2FRTST4GRINTLOGLU001
10.2415/AH-ROBUST-BILI-X2FR-TEST-CLEF2007.COLESUN.EN2FRTST4GRINTDICEU001
Groups
X
X X
X X
X X
X X</p>
        <p>X</p>
        <p>X
X
Test (DOI 10.2455/TUKEY T TEST.D284127041AE7A69919B3C09EBA3769F).</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>We have reported the results of the ad hoc cross-language textual document
retrieval track at CLEF 2007. This track is considered to be central to CLEF as
for many groups it is the first track in which they participate and provides them
with an opportunity to test their systems and compare performance between
monolingual and cross-language runs, before perhaps moving on to more
complex system development and subsequent evaluation. This year, the monolingual
task focused on central European languages while the bilingual task included an
activity for groups that wanted to use non-European topic languages and
languages with few processing tools and resources. Each year, we also include a task
aimed at examining particular aspects of cross-language text retrieval. Again this
year, the focus was examining the impact of ”hard” topics on performance in
the ”robust” task.</p>
      <p>The paper also describes in some detail the creation of the pools used for
relevance assessment this year. We still have to do stability tests on these pools;
the results will be published in the CLEF 2007 post-workshop Proceedings.</p>
      <p>Although there was quite a good participation in the monolingual Bulgarian,
Czech and Hungarian tasks and the experiments report some interesting work
on stemming and morphological analysis, we were very disappointed by the lack
of participation in bilingual tasks for these languages. On the other hand, the
interest in the task for non-European topic languages was encouraging and the
results reported can be considered positively. We are currently undecided about
the future of the main mono- and cross-language tasks in the ad hoc track; this
will be a topic for discussion at the breakout session during the workshop.</p>
      <p>The robust task has analyzed the performance of systems for older CLEF
data under a new perspective. A larger data set which allows a more reliable
comparative analysis of systems was assembled. Systems needed to avoid low
performing topics. Their success was measured with the geometric mean (GMAP)
which introduces a bias on poor performing topics. Results for the robust task
for mono-lingual retrieval or English, French and Portuguese as well as for
bilingual retrieval from English to French are reported. Robustness can also be
interpreted as the fitness of a system under a variety of conditions. The
definition on what robust retrieval means has to continue. All participants in CLEF
2007 are invited to engage in the discussion of the future of the robust task.</p>
      <p>The test collections for CLEF 2000 - CLEF 2003 are now publicly available on
the Evaluations and Language resources Distribution Agency (ELDA) catalog8.
8</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>We should like to acknowledge the enormous contribution of the groups
responsible for topic creation and relevance assessment. In particulat, we thank the group
responsible for the work on Bulgarian led by Kiril Simov and Petya Osenova,</p>
      <sec id="sec-7-1">
        <title>8 http://www.elda.org/</title>
        <p>the group responsible for Czech led by Pavel Pecina9, and the group responsible
for Hungarian led by Tam´as V´aradi and Gergely Botty´an. These groups worked
very hard under great pressure in order to complete the heavy load of relevance
assessments in time.
9 Ministry of Education of the Czech Republic, project MSM 0021620838
12. Hayurani, H., Sari, S., Adriani, M.: Evaluating Language Resources for CLEF
2007. In Nardi, A., Peters, C., eds.: Working Notes for the CLEF 2007 Workshop,
http://www.clef-campaign.org/ [last visited 2007, September 5] (2007)
13. Zhou, D., Truran, M., Brailsford, T.: Ambiguity and Unknown Term Translation in
CLIR. In Nardi, A., Peters, C., eds.: Working Notes for the CLEF 2007 Workshop,
http://www.clef-campaign.org/ [last visited 2007, September 5] (2007)
14. Argaw, A.A.: Amharic-English Information Retrieval with Pseudo Relevance
Feedback. In Nardi, A., Peters, C., eds.: Working Notes for the CLEF 2007 Workshop,
http://www.clef-campaign.org/ [last visited 2007, September 5] (2007)
15. Sch¨onhofen, P., Benczu´r, A., B´ır´o, I., Csalog´any, K.: Performing Cross-Language
Retrieval with Wikipedia. In Nardi, A., Peters, C., eds.: Working Notes for
the CLEF 2007 Workshop, http://www.clef-campaign.org/ [last visited 2007,
September 5] (2007)
16. Tune, K.K., Varma, V.: Oromo-English Information Retrieval Experiments at
CLEF 2007. In Nardi, A., Peters, C., eds.: Working Notes for the CLEF 2007
Workshop, http://www.clef-campaign.org/ [last visited 2007, September 5] (2007)
17. Bandyopadhyay, S., Mondal, T., Naskar, S.K., Ekbal, A., Haque, R., Godavarthy,
S.R.: Bengali, Hindi and Telugu to English Ad-hoc Bilingual task at CLEF 2007.
In Nardi, A., Peters, C., eds.: Working Notes for the CLEF 2007 Workshop, http:
//www.clef-campaign.org/ [last visited 2007, September 5] (2007)
18. Jagarlamudi, J., Kumaran, A.: Cross-Lingual Information Retrieval System for
Indian Languages. In Nardi, A., Peters, C., eds.: Working Notes for the CLEF
2007 Workshop, http://www.clef-campaign.org/ [last visited 2007, September
5] (2007)
19. Chinnakotla, M.K., Ranadive, S., Bhattacharyya, P., Damani, O.P.: Hindi and
Marathi to English Cross Language Information Retrieval at CLEF 2007. In Nardi,
A., Peters, C., eds.: Working Notes for the CLEF 2007 Workshop, http://www.
clef-campaign.org/ [last visited 2007, September 5] (2007)
20. Pingali, P., Varma, V.: IIIT Hyderabad at CLEF 2007 – Adhoc Indian Language
CLIR task. In Nardi, A., Peters, C., eds.: Working Notes for the CLEF 2007
Workshop, http://www.clef-campaign.org/ [last visited 2007, September 5] (2007)
21. Mandal, D., Dandapat, S., Gupta, M., Banerjee, P., Sarkar, S.: Bengali and Hindi
to English Cross-language Text Retrieval under Limited Resources. In Nardi,
A., Peters, C., eds.: Working Notes for the CLEF 2007 Workshop, http://www.
clef-campaign.org/ [last visited 2007, September 5] (2007)
22. Robertson, S.: On GMAP: and Other Transformations. In Yu, P.S., Tsotras, V.,
Fox, E.A., Liu, C.B., eds.: Proc. 15th International Conference on Information and
Knowledge Management (CIKM 2006), ACM Press, New York, USA (2006) 78–83
23. Voorhees, E.M.: The TREC Robust Retrieval Track. SIGIR Forum 39 (2005)
11–20
24. Savoy, J.: Why do Successful Search Systems Fail for Some Topics. In Cho, Y.,
Wan Koo, Y., Wainwright, R.L., Haddad, H.M., Shin, S.Y., eds.: Proc. 2007 ACM
Symposium on Applied Computing (SAC 2007). ACM Press, New York, USA
(2007) 872–877
25. Sanderson, M., Zobel, J.: Information Retrieval System Evaluation: Effort,
Sensitivity, and Reliability. In Baeza-Yates, R., Ziviani, N., Marchionini, G., Moffat,
A., Tait, J., eds.: Proc. 28th Annual International ACM SIGIR Conference on
Research and Development in Information Retrieval (SIGIR 2005), ACM Press, New
York, USA (2005) 162–169</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Paskin</surname>
          </string-name>
          , N., ed.:
          <source>The DOI Handbook - Edition 4.4</source>
          .1.
          <string-name>
            <surname>International</surname>
            <given-names>DOI</given-names>
          </string-name>
          Foundation (IDF). http://dx.doi.org/10.1000/186 [last visited
          <year>2007</year>
          ,
          <source>August</source>
          <volume>30</volume>
          ] (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Braschler</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>CLEF 2003 - Overview of results</article-title>
          . In Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kluck</surname>
          </string-name>
          , M., eds.:
          <source>Comparative Evaluation of Multilingual Information Access Systems: Fourth Workshop of the Cross-Language Evaluation Forum (CLEF 2003) Revised Selected Papers, Lecture Notes in Computer Science (LNCS) 3237</source>
          , Springer, Heidelberg, Germany (
          <year>2004</year>
          )
          <fpage>44</fpage>
          -
          <lpage>63</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Tomlinson</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Sampling Precision to Depth 10000:
          <article-title>Evaluation Experiments at CLEF 2007</article-title>
          . In Nardi,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Peters</surname>
          </string-name>
          , C., eds.
          <source>: Working Notes for the CLEF 2007 Workshop</source>
          , http://www.clef-campaign.
          <source>org/ [last visited</source>
          <year>2007</year>
          , September 5] (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Braschler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>CLEF 2003 Methodology and Metrics</article-title>
          . In Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kluck</surname>
          </string-name>
          , M., eds.:
          <source>Comparative Evaluation of Multilingual Information Access Systems: Fourth Workshop of the Cross-Language Evaluation Forum (CLEF 2003) Revised Selected Papers, Lecture Notes in Computer Science (LNCS) 3237</source>
          , Springer, Heidelberg, Germany (
          <year>2004</year>
          )
          <fpage>7</fpage>
          -
          <lpage>20</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Di</given-names>
            <surname>Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.M.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          , N.:
          <string-name>
            <surname>Appendix</surname>
            <given-names>A</given-names>
          </string-name>
          :
          <article-title>Results of the Core Tracks - Ad-hoc Bilingual and Monolingual Tasks</article-title>
          . In
          <string-name>
            <surname>Nardi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
          </string-name>
          , C., eds.
          <source>: Working Notes for the CLEF 2007 Workshop</source>
          , http://www.clef-campaign.
          <source>org/ [last visited</source>
          <year>2007</year>
          , September 5] (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Di</given-names>
            <surname>Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.M.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          , N.:
          <string-name>
            <surname>Appendix</surname>
            <given-names>B</given-names>
          </string-name>
          :
          <article-title>Results of the Core Tracks - Ad-hoc Robust Bilingual and Monolingual Tasks</article-title>
          . In
          <string-name>
            <surname>Nardi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
          </string-name>
          , C., eds.
          <source>: Working Notes for the CLEF 2007 Workshop</source>
          , http://www.clef-campaign.
          <source>org/ [last visited</source>
          <year>2007</year>
          , September 5] (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dolamic</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savoy</surname>
          </string-name>
          , J.:
          <article-title>Stemming Approaches for East European Languages</article-title>
          . In
          <string-name>
            <surname>Nardi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
          </string-name>
          , C., eds.
          <source>: Working Notes for the CLEF 2007 Workshop</source>
          , http: //www.clef-campaign.
          <source>org/ [last visited</source>
          <year>2007</year>
          , September 5] (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Noguera</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Llopis</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Applying Query Expansion techniques to Ad Hoc Monolingual tasks with the IR-n system</article-title>
          . In
          <string-name>
            <surname>Nardi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
          </string-name>
          , C., eds.
          <source>: Working Notes for the CLEF 2007 Workshop</source>
          , http://www.clef-campaign.
          <source>org/ [last visited</source>
          <year>2007</year>
          , September 5] (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Majumder</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitra</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pal</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Hungarian and Czech Stemming using YASS</article-title>
          . In Nardi,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Peters</surname>
          </string-name>
          , C., eds.
          <source>: Working Notes for the CLEF 2007 Workshop</source>
          , http: //www.clef-campaign.
          <source>org/ [last visited</source>
          <year>2007</year>
          , September 5] (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Cˇeˇska</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Charles University at CLEF 2007
          <string-name>
            <surname>Ad-Hoc Track</surname>
            . In Nardi,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
          </string-name>
          , C., eds.
          <source>: Working Notes for the CLEF 2007 Workshop</source>
          , http://www. clef-campaign.
          <source>org/ [last visited</source>
          <year>2007</year>
          , September 5] (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Ircing</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Mu¨uller, L.:
          <article-title>Czech Monolingual Information Retrieval Using Off-TheShelf Components - the University of West Bohemia at CLEF 2007 Ad-Hoc track</article-title>
          . In Nardi,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Peters</surname>
          </string-name>
          , C., eds.
          <source>: Working Notes for the CLEF 2007 Workshop</source>
          , http: //www.clef-campaign.
          <source>org/ [last visited</source>
          <year>2007</year>
          , September 5] (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>