<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CLEF 2009 Ad Hoc Track Overview: TEL &amp; Persian Tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <email>ferro@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carol Peters</string-name>
          <email>carol.peters@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>General Terms Experimentation</institution>
          ,
          <addr-line>Performance, Measurement, Algorithms</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ISTI-CNR, Area di Ricerca</institution>
          ,
          <addr-line>Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2009</year>
      </pub-date>
      <abstract>
        <p>The 2009 Ad Hoc track was to a large extent a repetition of last year's track, with the same three tasks: Tel@CLEF, Persian@CLEF, and Robust-WSD. In this first of the two track overviews, we describe the objectives and results of the TEL and Persian tasks and provide some statistical analyses. Categories and Subject Descriptors H.3 [Information Storage and Retrieval]: H.3.1 Content Analysis and Indexing; H.3.3 Information Search and Retrieval; H.3.4 [Systems and Software]: Performance evaluation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>From 2000 - 2007, the ad hoc track at CLEF used exclusively collections of
European newspaper and news agency documents1. In 2008 it was decided to
change the focus and to introduce document collections in a different genre
(bibliographic records from The European Library - TEL2), a non-European
language (Persian), and an IR task that would appeal to the NLP community
(robust retrieval on word-sense disambiguated data). The 2009 Ad Hoc track
has been to a large extent a repetition of last year’s track, with the same three
tasks: Tel@CLEF, Persian@CLEF, and Robust-WSD. An important objective
has been to ensure that for each task a good reusable test collections is created.
1 Over the years, this track has built up test collections for monolingual and cross
language system evaluation in 14 European languages.
2 See http://www.theeuropeanlibrary.org/
In this first of the two track overviews we describe the activities of the TEL and
Persian tasks3.</p>
      <p>TEL@CLEF: This task offered monolingual and cross-language search on
library catalog. It was organized in collaboration with The European Library
and used three collections derived from the catalogs of the British Library, the
Biblioth´eque Nationale de France and the Austrian National Library. The
underlying aim was to identify the most effective retrieval technologies for searching
this type of very sparse multilingual data. In fact, the collections contained
records in many languages in addition to English, French or German. The task
presumed a user with a working knowledge of these three languages who wants
to find documents that can be useful for them in one of the three target catalogs.</p>
      <p>Persian@CLEF: This activity was coordinated again this year in
collaboration with the Database Research Group (DBRG) of Tehran University. We chose
Persian as the first non-European language target collection for several reasons:
its challenging script (a modified version of the Arabic alphabet with elision of
short vowels) written from right to left; its complex morphology (extensive use
of suffixes and compounding); its political and cultural importance. The task
used the Hamshahri corpus of 1996-2002 newspapers as the target collection
and was organised as a traditional ad hoc document retrieval task. Monolingual
and cross-language (English to Persian) tasks were offered.</p>
      <p>In the rest of this paper we present the task setup, the evaluation
methodology and the participation in the two tasks (Section 2). We then describe the
main features of each task and show the results (Sections 3 and 4). The final
section provides a brief summing up. For information on the various approaches
and resources used by the groups participating in the two tasks and the issues
they focused on, we refer the reader to the papers in the relevant Ad Hoc sections
of these Working Notes.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Track Setup</title>
      <p>As is customary in the CLEF ad hoc track, again this year we adopted a
corpusbased, automatic scoring method for the assessment of the performance of the
participating systems, based on ideas first introduced in the Cranfield
experiments in the late 1960s [5]. The tasks offered are studied in order to effectively
measure textual document retrieval under specific conditions. The test
collections are made up of documents, topics and relevance assessments. The topics
consist of a set of statements simulating information needs from which the
systems derive the queries to search the document collections. Evaluation of system
performance is then done by judging the documents retrieved in response to
a topic with respect to their relevance, and computing the recall and precision
measures. The pooling methodology is used in order to limit the number of
manual relevance assessments that have to be made. As always, the distinguishing
3 As the task design was the same as last year, much of the task set-up section is a
repetition of a similar section in our CLEF 2008 working notes paper.
feature of CLEF is that it applies this evaluation paradigm in a multilingual
setting. This means that the criteria normally adopted to create a test collection,
consisting of suitable documents, sample queries and relevance assessments, have
been adapted to satisfy the particular requirements of the multilingual context.
All language dependent tasks such as topic creation and relevance judgment are
performed in a distributed setting by native speakers. Rules are established and
a tight central coordination is maintained in order to ensure consistency and
coherency of topic and relevance judgment sets over the different collections,
languages and tracks.
2.1</p>
      <sec id="sec-2-1">
        <title>The Documents</title>
        <p>As mentioned in the Introduction, the two tasks used different sets of documents.</p>
        <p>The TEL task used three collections:
– British Library (BL); 1,000,100 documents, 1.2 GB;
– Biblioth´eque Nationale de France (BNF); 1,000,100 documents, 1.3 GB;
– Austrian National Library (ONB); 869,353 documents, 1.3 GB.</p>
        <p>We refer to the three collections (BL, BNF, ONB) as English, French and
German because in each case this is the main and expected language of the
collection. However, each of these collections is to some extent multilingual and
contains documents (catalog records) in many additional languages.</p>
        <p>The TEL data is very different from the newspaper articles and news agency
dispatches previously used in the CLEF ad hoc track. The data tends to be very
sparse. Many records contain only title, author and subject heading information;
other records provide more detail. The title and (if existing) an abstract or
description may be in a different language to that understood as the language of
the collection. The subject heading information is normally in the main language
of the collection. About 66% of the documents in the English and German
collection have textual subject headings, in the French collection only 37%. Dewey
Classification (DDC) is not available in the French collection; negligible (¡0.3%)
in the German collection; but occurs in about half of the English documents
(456,408 docs to be exact).</p>
        <p>Whereas in the traditional ad hoc task, the user searches directly for a
document containing information of interest, here the user tries to identify which
publications are of potential interest according to the information provided by
the catalog card. When we designed the task, the question the user was presumed
to be asking was “Is the publication described by the bibliographic record
relevant to my information need?”</p>
        <p>The Persian task used the Hamshahri corpus of 1996-2002 newspapers as
the target collection. This corpus was made available to CLEF by the Data
Base Research Group (DBRG) of the University of Tehran. Hamshahri is one
of the most popular daily newspapers in Iran. The Hamshahri corpus consists
of 345 MB of news texts for the years 1996 to 2002 (corpus size with tags is
564 MB). This corpus contains more than 160,000 news articles about a variety
of subjects and includes nearly 417000 different words. Hamshahri articles vary
between 1KB and 140KB in size4.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Topics</title>
        <p>Topics in the CLEF ad hoc track are structured statements representing
information needs. Each topic typically consists of three parts: a brief “title” statement; a
one-sentence “description”; a more complex “narrative” specifying the relevance
assessment criteria. Topics are prepared in xml format and uniquely identified
by means of a Digital Object Identifier (DOI)5.</p>
        <p>For the TEL task, a common set of 50 topics was prepared in each of the 3
main collection languages (English, French and German) plus this year also in
Chinese, Italian and Greek in response to specific requests. Only the Title and
Description fields were released to the participants. The narrative was employed
to provide information for the assessors on how the topics should be judged. The
topic sets were prepared on the basis of the contents of the collections.</p>
        <p>In ad hoc, when a task uses data collections in more than one language,
we consider it important to be able to use versions of the same core topic set
to query all collections. This makes it easier to compare results over different
collections and also facilitates the preparation of extra topic sets in additional
languages. However, it is never easy to find topics that are effective for several
different collections and the topic preparation stage requires considerable
discussion between the coordinators for each collection in order to identify suitable
common candidates. The sparseness of the data makes this particularly difficult
for the TEL task and leads to the formulation of topics that were quite broad in
scope so that at least some relevant documents could be found in each collection.
A result of this strategy is that there tends to be a considerable lack of evenness
of distribution in relevant documents. For each topic, the results expected from
the separate collections can vary considerably. An example of a TEL topic is
given in Figure 1.</p>
        <p>For the Persian task, 50 topics were created in Persian by the Data Base
Research group of the University of Tehran, and then translated into English.
The rule in CLEF when creating topics in additional languages is not to produce
literal translations but to attempt to render them as naturally as possible. This
was a particularly difficult task when going from Persian to English as cultural
differences had to be catered for. An example of a CLEF 2009 Persian topic is
given in Figure 2.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Relevance Assessment</title>
        <p>The number of documents in large test collections such as CLEF makes it
impractical to judge every document for relevance. Instead approximate recall values
are calculated using pooling techniques. The results submitted by the groups
4 For more information, see http://ece.ut.ac.ir/dbrg/hamshahri/
5 http://www.doi.org/
&lt;?xml version="1.0" encoding="UTF-8" standalone="no"?&gt;
&lt;topic&gt;
&lt;identifier&gt;10.2452/711-AH&lt;/identifier&gt;
&lt;title lang="zh"&gt; &lt;/title&gt;
&lt;title lang="en"&gt;Deep Sea Creatures&lt;/title&gt;
&lt;title lang="fr"&gt;Créatures des fonds océaniques&lt;/title&gt;
&lt;title lang="de"&gt;Kreaturen der Tiefsee&lt;/title&gt;
&lt;title lang="el"&gt; &lt;/title&gt;
&lt;title lang="it"&gt;Creature delle profondità oceaniche&lt;/title&gt;
&lt;description lang="zh"&gt;
&lt;/description&gt;
&lt;description lang="en"&gt;</p>
        <p>Find publications about any kind of life in the depths
of any of the world's oceans.
&lt;/description&gt;
&lt;description lang="fr"&gt;</p>
        <p>Trouver des ouvrages sur toute forme de vie dans les
profondeurs des mers et des océans.
&lt;/description&gt;
&lt;description lang="de"&gt;</p>
        <p>Finden Sie Veröffentlichungen über Leben und</p>
        <p>Lebensformen in den Tiefen der Ozeane der Welt.
&lt;/description&gt;
&lt;description lang="el"&gt;
&lt;/description&gt;
&lt;description lang="it"&gt;</p>
        <p>Trova pubblicazioni su qualsiasi forma di vita nelle
profondità degli oceani del mondo.</p>
        <p>
          &lt;/description&gt;
&lt;/topic&gt;
participating in the ad hoc tasks are used to form a pool of documents for each
topic and language by collecting the highly ranked documents from selected runs
according to a set of predefined criteria. One important limitation when forming
the pools is the number of documents to be assessed. Traditionally, the top 100
ranked documents from each of the runs selected are included in the pool; in
such a case we say that the pool is of depth 100. This pool is then used for
subsequent relevance judgments. After calculating the effectiveness measures,
the results are analyzed and run statistics produced and distributed. The
stability of pools constructed in this way and their reliability for post-campaign
experiments is discussed in [
          <xref ref-type="bibr" rid="ref3">3] with respect to the CLEF 2003</xref>
          pools.
        </p>
        <p>The main criteria used when constructing the pools in CLEF are:
&lt;?xml version="1.0" encoding="UTF-8" standalone="no"?&gt;
&lt;topic&gt;
&lt;identifier&gt;10.2452/641-AH&lt;/identifier&gt;
&lt;title lang="en"&gt;Pollution in the Persian Gulf&lt;/title&gt;
&lt;title lang="fa"&gt; &lt;/title&gt;
&lt;description lang="en"&gt;</p>
        <p>Find information about pollution in the Persian Gulf and the causes.
&lt;/description&gt;
&lt;description lang="fa"&gt;
&lt;/description&gt;
&lt;narrative lang="en"&gt;</p>
        <p>Find information about conditions of the Persian Gulf with respect to
pollution; also of interest is information on the causes of pollution
and comparisons of the level of pollution in this sea against that of
other seas.
&lt;/narrative&gt;
&lt;narrative lang="fa"&gt;
&lt;/narrative&gt;
&lt;/topic&gt;
– favour diversity among approaches adopted by participants, according to the
descriptions of the experiments provided by the participants;
– choose at least one experiment for each participant in each task, chosen
among the experiments with highest priority as indicated by the participant;
– add mandatory title+description experiments, even though they do not have
high priority;
– add manual experiments, when provided;
– for bilingual tasks, ensure that each source topic language is represented.</p>
        <p>From our experience in CLEF, using the tools provided by the DIRECT
system [1], we find that for newspaper documents, assessors can normally judge
from 60 to 100 documents per hour, providing binary judgments: relevant / not
relevant. Our estimate for the TEL catalog records is higher as these records
are much shorter than the average newspaper article (100 to 120 documents
per hour). In both cases, it can be seen what a time-consuming and resource
expensive task human relevance assessment is. This limitation impacts strongly
on the application of the criteria above - and implies that we are obliged to be
flexible in the number of documents judged per selected run for individual pools.</p>
        <p>This year, in order to create pools of more-or-less equivalent size, the depth
of the TEL English, French, and German pools was 606. For each collection, we
included in the pool two monolingual and one bilingual experiments from each
participant plus any documents assessed as relevant during topic creation.
6 Tests made on NTCIR pools in previous years have suggested that a depth of 60
in normally adequate to create stable pools, presuming that a sufficient number of
runs from different systems have been included.</p>
        <p>As we only had a relatively small number of runs submitted for Persian, we
were able to include documents from all experiments, and the pool was created
with a depth of 80.</p>
        <p>These pool depths were the same as those used last year. Given the resources
available, it was not possible to manually assess more documents. For the CLEF
2008 ad hoc test collections, Stephen Tomlinson reported some sampling
experiments aimed at estimating the judging coverage [10]. He found that this tended
to be lower than the estimates he produced for the CLEF 2007 ad hoc
collections. With respect to the TEL collections, he estimated that at best 50% to
70% of the relevant documents were included in the pools - and that most of
the unjudged relevant documents were for the 10 or more queries that had the
most known answers. Tomlinson has repeated these experiments for the 2009
TEL and Persian data [9]. Although for two of the four languages concerned
(German and Persian), his findings were similar to last year’s estimates, for the
other two languages (English and French) this year’s estimates are substantially
lower. These findings need further investigation. They suggest that if we are to
continue to use the pooling technique, we would perhaps be wise to do some
more exhaustive manual searches in order to boost the pools with respect to
relevant documents. We also need to consider more carefully other techniques for
relevance assessment in the future such as, for example, the method suggested
by Sanderson and Joho [8] or Mechanical Turk [2].</p>
        <p>Table 1 reports summary information on the 2009 ad hoc pools used to
calculate the results for the main monolingual and bilingual experiments. In
particular, for each pool, we show the number of topics, the number of runs
submitted, the number of runs included in the pool, the number of documents
in the pool (relevant and non-relevant), and the number of assessors.</p>
        <p>The box plot of Figure 3 compares the distributions of the relevant documents
across the topics of each pool for the different ad hoc pools; the boxes are ordered
by decreasing mean number of relevant documents per topic.</p>
        <p>As can be noted, TEL French and German distributions appear similar and
are slightly asymmetric towards topics with a greater number of relevant
documents while the TEL English distribution is slightly asymmetric towards topics
with a lower number of relevant documents. All the distributions show some
upper outliers, i.e. topics with a greater number of relevant document with
respect to the behaviour of the other topics in the distribution. These outliers are
probably due to the fact that CLEF topics have to be able to retrieve relevant
documents in all the collections; therefore, they may be considerably broader in
one collection compared with others depending on the contents of the separate
datasets.</p>
        <p>For the TEL documents, we judged for relevance only those documents that
are written totally or partially in English, French and German, e.g. a catalog
record written entirely in Hungarian was counted as not relevant as it was of no
use to our hypothetical user; however, a catalog record with perhaps the title and
a brief description in Hungarian, but with subject descriptors in French, German
or English was judged for relevance as it could be potentially useful. Our assessors</p>
        <sec id="sec-2-3-1">
          <title>Pool size</title>
        </sec>
        <sec id="sec-2-3-2">
          <title>Pool size</title>
        </sec>
        <sec id="sec-2-3-3">
          <title>Pool size</title>
        </sec>
        <sec id="sec-2-3-4">
          <title>Pool size</title>
          <p>26,190 pooled documents
– 23,663 not relevant documents
– 2,527 relevant documents
50 topics
31 out of 89 submitted experiments
Pooled Experiments – monolingual: 22 out of 43 submitted experiments
– bilingual: 9 out of 46 submitted experiments</p>
        </sec>
        <sec id="sec-2-3-5">
          <title>Assessors</title>
          <p>4 assessors
TEL French Pool (DOI 10.2454/AH-TEL-FRENCH-CLEF2009)
21,971 pooled documents
– 20,118 not relevant documents
– 1,853 relevant documents
50 topics
21 out of 61 submitted experiments
Pooled Experiments – monolingual: 16 out of 35 submitted experiments
– bilingual: 5 out of 26 submitted experiments</p>
        </sec>
        <sec id="sec-2-3-6">
          <title>Assessors</title>
          <p>1 assessor
TEL German Pool (DOI 10.2454/AH-TEL-GERMAN-CLEF2009)
25,541 pooled documents
– 23,882 not relevant documents
– 1,559 relevant documents
50 topics
21 out of 61 submitted experiments
23,536 pooled documents
– 19,072 not relevant documents
– 4,464 relevant documents
50 topics
20 out of 20 submitted experiments
Pooled Experiments – monolingual: 16 out of 35 submitted experiments
– bilingual: 5 out of 26 submitted experiments</p>
        </sec>
        <sec id="sec-2-3-7">
          <title>Assessors</title>
          <p>2 assessors</p>
          <p>Persian Pool (DOI 10.2454/AH-PERSIAN-CLEF2009)
Pooled Experiments – monolingual: 17 out of 17 submitted experiments
– bilingual: 3 out of 3 submitted experiments
Assessors
23 assessors
0 D
4 t
1 n
a
v
e
l
e
120 rbe</p>
          <p>R
f
o
m
u
9
0
0
2
F
E
L
C
−
A
M
R
E
G
−
L
E
T
−
A
/
4
5
4
2
.
0
1
0
6
2
0
4
2
0
2
2
0
0
2
0
8
1
0
6
1
0
0
1
0
8
0
6
0
4
0
2
0
9
0
0
2
F
E
L
C
−
A
I
S
R
P
−
L
E
T
−
A
/
4
5
4
2
.
0
1
9
0
0
2
F
E
L
C
−
S
I
L
G
E
−
L
E
T
−
A
/
4
5
4
2
.
0
1
9
0
0
2
F
E
L
C
−
H
C
N
E
R
F
−
L
E
T
−
H
A
/
4
5
4
2
.
0
1
N</p>
          <p>H</p>
          <p>N
E</p>
          <p>N
H</p>
          <p>H</p>
          <p>H
had no additional knowledge of the documents referred to by the catalog records
(or surrogates) contained in the collection. They judged for relevance on the
information contained in the records made available to the systems. This was
a non trivial task due to the lack of information present in the documents.
During the relevance assessment activity there was much consultation between
the assessors for the three TEL collections in order to ensure that the same
assessment criteria were adopted by everyone.</p>
          <p>As shown in the box plot of Figure 3, the Persian distribution presents a
greater number of relevant documents per topic with respect to the other
distributions and is slightly asymmetric towards topics with a number of relevant
documents. In addition, as can be seen from Table 1, it has been possible to
sample all the experiments submitted for the Persian tasks. This means that there
were fewer unique documents per run and this fact, together with the greater
number of relevant documents per topic suggests either that all the systems were
using similar approaches and retrieval algorithms or that the systems found the
Persian topics quite easy.</p>
          <p>The relevance assessment for the Persian results was done by the DBRG
group in Tehran. Again, assessment was performed on a binary basis and the
standard CLEF assessment rules were applied.
Evaluation campaigns such as TREC and CLEF are based on the belief that
the effectiveness of Information Retrieval Systems (IRSs) can be objectively
evaluated by an analysis of a representative set of sample search results. For
this, effectiveness measures are calculated based on the results submitted by the
participants and the relevance assessments. Popular measures usually adopted
for exercises of this type are Recall and Precision. Details on how they are
calculated for CLEF are given in [4].</p>
          <p>
            The individual results for all official Ad-hoc TEL and Persian experiments in
CLEF 2009 are given in the Appendices of the CLEF 2009 Working Notes [6,7].
You can also access them online at:
– Ad-hoc TEL:
• monolingual English: http://direct.dei.unipd.it/DOIResolver.do?
type=task&amp;id=AH-TEL-MONO-EN-CLEF2009
• bilingual English: http://direct.dei.unipd.it/DOIResolver.do?type=
task&amp;id=AH-TEL-BILI-X
            <xref ref-type="bibr" rid="ref2">2EN-CLEF2009</xref>
            • monolingual French: http://direct.dei.unipd.it/DOIResolver.do?
type=task&amp;id=AH-TEL-MONO-FR-CLEF2009
• bilingual French: http://direct.dei.unipd.it/DOIResolver.do?type=
task&amp;id=AH-TEL-BILI-X
            <xref ref-type="bibr" rid="ref2">2FR-CLEF2009</xref>
            • monolingual German: http://direct.dei.unipd.it/DOIResolver.do?
type=task&amp;id=AH-TEL-MONO-DE-CLEF2009
• bilingual German: http://direct.dei.unipd.it/DOIResolver.do?type=
task&amp;id=AH-TEL-BILI-X2DE-CLEF2009
          </p>
          <p>Participant
aeb
celi
chemnitz
cheshire
cuza
hit
inesc
karlsruhe
opentext
qazviniau
trinity
trinity-dcu
weimar</p>
        </sec>
        <sec id="sec-2-3-8">
          <title>Participant</title>
          <p>jhu-apl
opentext
qazviniau
unine</p>
        </sec>
        <sec id="sec-2-3-9">
          <title>Ad hoc TEL participants Institution Country</title>
          <p>Athens Univ. Economics &amp; Business Greece
CELI Research srl Italy
Chemnitz University of Technology Germany
U.C.Berkeley United States
Alexandru Ioan Cuza University Romania
HIT2Lab, Heilongjiang Inst. Tech. China
Tech. Univ. Lisbon Portugal
Univ. Karlsruhe Germany
OpenText Corp. Canada
Islamic Azaz Univ. Qazvin Iran
Trinity Coll. Dublin Ireland
Trinity Coll. &amp; DCU Ireland
Bauhaus Univ. Weimar Germany</p>
          <p>Ad hoc Persian participants</p>
          <p>Institution Country
Johns Hopkins Univ. USA
OpenText Corp. Canada
Islamic Azaz Univ. Qazvin Iran</p>
          <p>
            U.Neuchatel-Informatics Switzerland
– Ad-hoc Persian:
• monolingual Farsi: http://direct.dei.unipd.it/DOIResolver.do?type=
task&amp;id=AH-PERSIAN-MONO-FA-CLEF2009
• bilingual German: http://direct.dei.unipd.it/DOIResolver.do?type=
task&amp;id=AH-PERSIAN-BILI-X
            <xref ref-type="bibr" rid="ref2">2FA-CLEF2009</xref>
            2.5
          </p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>Participants and Experiments</title>
        <p>As shown in Table 2, a total of 13 groups from 10 countries submitted official
results for the TEL task, while just four groups participated in the Persian task.</p>
        <p>A total of 231 runs were submitted with an average number of submitted
runs per participant of 13.5 runs/participant.</p>
        <p>Participants were required to submit at least one title+description (“TD”)
run per task in order to increase comparability between experiments. The large
majority of runs (216 out of 231, 93.50%) used this combination of topic fields, 2
(0.80%) used all fields7, 13 (5.6%) used the title field. All the experiments were
conducted using automatic query construction. A breakdown into the separate
tasks and topic languages is shown in Table 3.</p>
        <p>Seven different topic languages were used in the ad hoc experiments. As
always, the most popular language for queries was English, with German second.
However, it must be noted that English topics were provided for both the TEL
7 The narrative field was only offered for the Persian task.
and the Persian tasks. It is thus hardly surprising that English is the most used
language in which to formulate queries. On the other hand, if we look only at
the bilingual tasks, the most used source languages were German and French.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>TEL@CLEF</title>
      <p>The objective of this activity was to search and retrieve relevant items from
collections of library catalog cards. The underlying aim was to identify the most
effective retrieval technologies for searching this type of very sparse data.
3.1</p>
      <sec id="sec-3-1">
        <title>Tasks</title>
        <p>Two subtasks were offered: Monolingual and Bilingual. In both tasks, the aim
was to retrieve documents relevant to the query. By monolingual we mean that
the query is in the same language as the expected language of the collection.
By bilingual we mean that the query is in a different language to the expected
language of the collection. For example, in an EN → FR run, relevant documents
(bibliographic records) could be any document in the BNF collection (referred
to as the French collection) in whatever language they are written. The same
is true for a monolingual FR → FR run - relevant documents from the BNF
collection could actually also be in English or German, not just French.</p>
        <p>Ten of the thirteen participating groups attempted a cross-language task; the
most popular being with the British Library as the target collection. Six groups
submitted experiments for all six possible official cross-language combinations.
In addition, we had runs submitted to the English target with queries in Greek,
Chinese and Italian.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Results.</title>
      </sec>
      <sec id="sec-3-3">
        <title>Monolingual Results</title>
        <p>Table 4 shows the top five groups for each target collection, ordered by mean
average precision. The table reports: the short name of the participating group;
the mean average precision achieved by the experiment; the DOI of the
experiment; and the performance difference between the first and the last participant.
Figures 4, 6, and 8 compare the performances of the top participants of the TEL
Monolingual tasks.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Bilingual Results</title>
        <p>– X → EN: 99.07% of best monolingual English IR system;
– X → FR: 94.00% of best monolingual French IR system;
– X → DE: 90.06% of best monolingual German IR system.</p>
        <p>These figures are very encouraging, especially when compared with the results
for last year for the same TEL tasks:
– X → EN: 90.99% of best monolingual English IR system;
– X → FR: 56.63% of best monolingual French IR system;</p>
        <p>Ad−Hoc TEL Monolingual English Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
inesc [Experiment RUN11; MAP 40.84%; Not Pooled]
chemnitz [Experiment CUT_11_MONO_MERGED_EN_9_10; MAP 40.71%; Not Pooled]
trinity [Experiment TCDENRUN2; MAP 40.35%; Pooled]
hit [Experiment MTDD10T40; MAP 39.36%; Pooled]
trinity−dcu [Experiment TCDDCUEN3; MAP 36.96%; Not Pooled]
0%
0%
10%
20%
30%
40%
50%
Recall
60%
70%
80%
90%
100%
Ad−Hoc TEL Bilingual English Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
chemnitz [Experiment CUT_13_BILI_MERGED_DE2EN_9_10; MAP 40.46%; Pooled]
hit [Experiment XTDD10T40; MAP 35.27%; Not Pooled]
trinity [Experiment TCDDEENRUN3; MAP 35.05%; Not Pooled]
trinity−dcu [Experiment TCDDCUDEEN1; MAP 33.33%; Not Pooled]
karlsruhe [Experiment DE_INDEXBL; MAP 32.70%; Not Pooled]
Ad−Hoc TEL Monolingual French Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
karlsruhe [Experiment INDEXBL; MAP 27.20%; Not Pooled]
chemnitz [Experiment CUT_19_MONO_MERGED_FR_17_18; MAP 25.83%; Not Pooled]
inesc [Experiment RUN12; MAP 25.11%; Not Pooled]
opentext [Experiment OTFR09TDE; MAP 24.12%; Not Pooled]
celi [Experiment CACAO_FRBNF_ML; MAP 23.61%; Not Pooled]
Ad−Hoc TEL Bilingual French Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
chemnitz [Experiment CUT_24_BILI_EN2FR_MERGED_LANG_SPEC_REF_CUT_17; MAP 25.57%; Not Pooled]
karlsruhe [Experiment EN_INDEXBL; MAP 24.62%; Not Pooled]
cheshire 9[E0x%periment BIENFRT2FB; MAP 16.77%; Not Pooled]
trinity [Experiment TCDDEFRRUN2; MAP 16.33%; Not Pooled]
weimar [Experiment CLESA169283ENINFR; MAP 14.51%; Pooled]
80%
Ad−Hoc TEL Monolingual German Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
opentext [Experiment OTDE09TDE; MAP 28.68%; Not Pooled]
chemnitz [Experiment CUT_3_MONO_MERGED_DE_1_2; MAP 27.89%; Not Pooled]
inesc [Experiment RUN12; MAP 27.85%; Not Pooled]
trinity−dcu [Experiment TCDDCUDE3; MAP 26.86%; Not Pooled]
trinity [Experiment TCDDERUN1; MAP 25.77%; Not Pooled]
0%
0%
10%
20%
30%
40%
50%
Recall
60%
70%
80%
90%
100%
90%
80%
70%
60%
Ad−Hoc TEL Bilingual German Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
chemnitz [Experiment CUT_5_BILI_MERGED_EN2DE_1_2; MAP 25.83%; Pooled]
trinity [Experiment TCDENDERUN3; MAP 19.35%; Not Pooled]
karlsruhe [Experiment EN_INDEXBL; MAP 16.46%; Not Pooled]
weimar [Experiment COMBINEDFRINDE; MAP 15.75%; Not Pooled]
cheshire [Experiment BIENDET2FBX; MAP 11.50%; Not Pooled]
– X → DE: 53.15% of best monolingual German IR system.</p>
        <p>In particular, it can be seen that there is a considerable improvement in
performance for French and German This will be commented in the following
section.</p>
        <p>The monolingual performance figures for all three tasks are quite similar to
those of last year but as these are not absolute values, no real conclusion can be
drawn from this.
As stated in the introduction, the TEL task this year is a repetition of the task set
last year. A main reason for this was to create a good reusable test collection with
a sufficient number of topics; another reason was to see whether the experience
gained and reported in the literature last year, and the opportunity to use last
year’s test collection as training data, would lead to differences in approaches
and/or improvements in performance this year. Although we have exactly the
same number of participants this year as last year, only five of the thirteen 2009
participants also participated in 2008. These are the groups tagged as Chemnitz,
Cheshire, Karlsruhe, INESC-ID and Opentext. The last two of these groups only
tackled monolingual tasks. These groups all tend to appear in the top five for
the various tasks. In the following we attempt to examine briefly the approaches
adopted this year, focusing mainly on the cross-language experiments.</p>
        <p>In the TEL task in CLEF 2008, we noted that all the traditional approaches
to monolingual and cross language retrieval were attempted by the different
groups. Retrieval methods included language models, vector-space and
probabilistic approaches, and translation resources ranged from bilingual dictionaries,
parallel and comparable corpora to on-line MT systems and Wikipedia. Groups
often used a combination of more than one resource. What is immediately
noticeable in 2009 is that, although similarly to last year a number of different
retrieval models were tested, there is a far more uniform approach to the
translation problem.</p>
        <p>Five of the ten groups that attempted cross-language tasks used the Google
Translate functionality, while a sixth used the LEC Power Translator [13].
Another group also used an MT system combining it with concept-based techniques
but did not disclose the name of the MT system used [16]. The remaining three
groups used a bilingual term list [17], a combination of resources including on-line
and in house developed dictionaries [19], and Wikipedia translation links [18]. It
is important to note that four out of the five groups in the bilingual to English
and bilingual to French tasks and three out of five for the bilingual to German
task used Google Translate, either on its own or in combination with another
technique. One group noted that topic translation using a statistical MT
system resulted in about 70% of the mean average precision (MAP) achieved when
using Google Translate [20]. Another group [11] found that the results obtained
by simply translating the query into all the target languages via Google gave
results that were comparable to a far more complex strategy known as
CrossLanguage Explicit Semantic Analysis, CL-ESA, where the library catalog records
and the queries are represented in a multilingual concept space that is spanned
by aligned Wikipedia articles. As this year’s results were significantly better
than last year’s, can we take this as meaning that Google is going to solve the
cross-language translation resource quandary?</p>
        <p>Taking a closer look at three groups that did consistently well in the
crosslanguage tasks we find the following. The group that had the top result for
each of the three tasks was Chemnitz [15]. They also had consistently good
monolingual results. Not surprisingly, they appear to have a very strong IR
engine, which uses various retrieval models and combines the results. They used
Snowball stemmers for English and French and an n-gram stemmer for German.
They were one of the few groups that tried to address the multilinguality of the
target collections. They used the Google service to translate the topic from the
source language to the four most common languages in the target collections,
queried the four indexes and combined the results in a multilingual result set.
They found that their approach combining multiple indexed collections worked
quite well for French and German but was disappointing for English.</p>
        <p>Another group with good performance, Karlsruhe [16], also attempted to
tackle the multilinguality of the collections. Their approach was again based on
multiple indexes for different languages with rank aggregation to combine the
different partial results. They ran language detectors on the collections to
identify the different languages contained and translated the topics to the languages
recognized. They used Snowball stemmers to stem terms in ten main languages,
fields in other languages were not preprocessed. Disappointingly, a baseline
consisting of a single index without language classification and a topic translated
only to the index language achieved similar or even better results. For the
translation step, they combined MT with a concept-based retrieval strategy based on
Explicit Semantic Analysis and using the Wikipedia database in English, French
and German as concept space.</p>
        <p>A third group that had quite good cross-language results for all three
collections was Trinity [12]. However, their monolingual results were not so strong.
They used a language modelling retrieval paradigm together with a document
re-ranking method which they tried experimentally in the cross-language
context. Significantly, they also used Google Translate. Judging from the fact that
they did not do so well in the monolingual tasks, this seems to be the probable
secret of their success for cross-language.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Persian@CLEF</title>
      <p>This activity was again coordinated in collaboration with the Data Base Research
Group (DBRG) of Tehran University.</p>
      <p>The activity was organised as a typical ad hoc text retrieval task on
newspaper collections. Two tasks were offered: monolingual retrieval; cross-language
retrieval (English queries to Persian target) and 50 topics were prepared (see
section 2.2). For each topic, participants had to find relevant documents in the
collection and submit the results in a ranked list.</p>
      <p>Table 3 provides a breakdown of the number of participants and submitted
runs by task and topic language.
– X → FA: 5.50% of best monolingual Farsi IR system.</p>
      <p>This appears to be a very clear indication that something went wrong with
the bilingual system that has been developed. These results should probably be
discounted.</p>
      <p>Ad−Hoc TEL Monolingual Persian Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%
jhu−apl [Experiment JHUFASK41R400TD; MAP 49.38%; Pooled]
unine [Experiment UNINEPE4; MAP 49.37%; Pooled]
opentext [Experiment OTFA09TDE; MAP 39.53%; Pooled]
qazviniau [Experiment IAUPERFA3; MAP 37.62%; Pooled]
10%
20%
30%
40%
60%
70%
80%
90%
100%
90%
80%
70%
60%
40%
30%
20%
10%
0%</p>
      <p>0%
90%
80%
70%
60%
40%
30%
20%
10%</p>
      <p>Fig. 10. Monolingual Persian
Ad−Hoc TEL Bilingual Persian Task Top 5 Participants − Standard Recall Levels vs Mean Interpolated Precision
100%</p>
      <p>qazviniau [Experiment IAUPEREN3; MAP 2.72%; Pooled]
0%
0%
10%
20%
30%
40%
We were very disappointed this year that despite the fact that 14 groups
registered for the Persian task, only four actually submitted results. And only one
of these groups was from Iran. We suspect that one of the reasons for this was
that the date for submission of results was not very convenient for the Iranian
groups. Furthermore, only one group [18] attempted the bilingual task with the
very poor results cited above. The technique they used was the same as that
adopted for their bilingual to English experiments, exploiting Wikipedia
translation links, and the reason they give for the very poor performance here is that
the coverage of Farsi in Wikipedia is still very scarce compared to that of many
other languages.</p>
      <p>In the monolingual Persian task, the top two groups had very similar
performance figures. [21] found they had best results using a light suffix-stripping
algorithm and by combining different indexing and searching strategies.
Interestingly, their results this year do not confirm their findings for the same task
last year when the use of stemming did not prove very effective. The other
group [14] tested variants of character n-gram tokenization; 4-grams, 5-grams,
and skipgrams all provided about a 10% relative gain over plain words.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In CLEF 2009 we deliberately repeated the TEL and Persian tasks offered in 2008
in order to build up our test collections. Although we have not yet had sufficient
time to assess them in depth, we are reasonably happy with the results for the
TEL task: several groups worked on tackling the particular features of the TEL
collections with varying success; evidence has been acquired on the effectiveness
of a number of different IR strategies; there is a very strong indication of the
validity of the Google Translate functionality.</p>
      <p>On the other hand, the results for the Persian task were quite disappointing:
very few groups participated; the results obtained are either in contradiction to
those obtained previously and thus need further investigation [21] or tend to be
a very straightforward repetition and confirmation of last year’s results [14].
6</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The TEL task was studied in order to provide useful input to The European
Library (TEL); we express our gratitude in particular to Jill Cousins, Programme
Director, and Sjoerd Siebinga, Technical Developer of TEL. Vivien Petras,
Humboldt University, Germany, and Nicolas Moreau, Evaluation and Language
Resources Distribution Agency, France, were responsible for the creation of the
topics and the supervision of the relevance assessment work for the ONB and
BNF data respectively. We thank them for their valuable assistance.</p>
      <p>We should also like to acknowledge the enormous contribution to the
coordination of the Persian task made by the Data Base Research group of the
University of Tehran and in particular to Abolfazl AleAhmad and Hadi Amiri.
They were responsible for the preparation of the set of topics for the Hamshahri
collection in Farsi and English and for the subsequent relevance assessments.</p>
      <p>Least but not last, we would warmly thank Giorgio Maria Di Nunzio for all
the contributions he gave in carrying out the TEL and Persian tasks.
6. G. M. Di Nunzio and N. Ferro. Appendix A: Results of the TEL@CLEF Task. In
this volume.
7. G. M. Di Nunzio and N. Ferro. Appendix B: Results of the Persian@CLEF Task.</p>
      <p>In this volume.
8. M. Sanderson and H. Joho. Forming Test Collections with No System Pooling. In
M. Sanderson, K. J¨arvelin, J. Allan, and P. Bruza, editors, Proc. 27th Annual
International ACM SIGIR Conference on Research and Development in Information
Retrieval (SIGIR 2004), pages 33–40. ACM Press, New York, USA, 2004.
9. S. Tomlinson. German, French, English and Persian Retrieval Experiments at</p>
      <p>CLEF 2009. In this volume.
10. Tomlinson, S.: Sampling Precision to Depth 10000 at CLEF 2008. Systems for
Multilingual and Multimodal Information Access: 9th Workshop of the Cross-Language
Evaluation Forum (CLEF 2008). Revised Selected Papers, Lecture Notes in
Computer Science (LNCS) 5706, Springer, Heidelberg, Germany (2009)
11. Anderka, M., Lipka, N., Stein, B.: Evaluating Cross-Language Explicit Semantic</p>
      <p>Analysis and Cross Querying at TEL@CLEF 2009. In this volume.
12. Zhou, D., Wade, V: Language Modeling and Document Re-Ranking: Trinity
Experiments at TEL@CLEF-2009. In this volume.
13. Larson, R.R.: Multilingual Query Expansion for CLEF Adhoc-TEL. In this
volume.
14. McNamee, P.: JHU Experiments in Monolingual Farsi Document Retrieval at</p>
      <p>CLEF 2009. In this volume.
15. Kuersten, J.: Chemnitz at CLEF 2009 Ad-Hoc TEL Task: Combining Different</p>
      <p>Retrieval Models and Addressing the Multilinguality. In this volume.
16. Sorg, P., Braun, M., Nicolay, D., Cimiano, P.: Cross-lingual Information Retrieval
based on Multiple Indexes. In this volume.
17. Katsiouli, P., Kalamboukis, T.: An Evaluation of Greek-English Cross Language</p>
      <p>Retrieval within the CLEF Ad-Hoc Bilingual Task. In this volume.
18. Jadidinejad, A.H., Mahmoudi, F.: Query Wikification: Mining Structured Queries
From Unstructured Information Needs using Wikipedia-based Semantic Analysis.</p>
      <p>In this volume.
19. Bosca, A., Dini, L.: CACAO Project at the TEL@CLEF 2009 Task. In this volume.
20. Leveling, J., Zhou, D., Jones, G.F., Wade, V.: TCD-DCU at TEL@CLEF 2009:</p>
      <p>Document Expansion, Query Translation and Language Modeling. In this volume.
21. Dolamic, L., Fautsch, C., Savoy, J.: UniNE at CLEF 2009: Persian Ad Hoc
Retrieval and IP. In this volume.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Agosti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Di Nunzio</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          .
          <article-title>The Importance of Scientific Data Curation for Evaluation Campaigns</article-title>
          . In C. Thanos and F. Borri, editors,
          <source>DELOS Conference 2007 Working Notes</source>
          , pages
          <fpage>185</fpage>
          -
          <lpage>193</lpage>
          . ISTI-CNR,
          <string-name>
            <surname>Gruppo</surname>
            <given-names>ALI</given-names>
          </string-name>
          , Pisa, Italy,
          <year>February 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>O.</given-names>
            <surname>Alonso</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Mizzaro</surname>
          </string-name>
          .
          <article-title>Can we get rid of TREC assessors? Using Mechanical Turk for relevance assessment</article-title>
          . In S. Geva,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Sakai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Trotman</surname>
          </string-name>
          , and E. Voorhees, editors,
          <source>Proc. SIGIR 2009 Workshop on The Future of IR Evaluation</source>
          . http://staff.science.uva.nl/~kamps/ireval/papers/paper_ 22.pdf,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          .
          <article-title>CLEF 2002 - Overview of Results</article-title>
          . In C. Peters,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          , and M. Kluck, editors,
          <source>Advances in Cross-Language Information Retrieval: Third Workshop of the Cross-Language Evaluation Forum (CLEF</source>
          <year>2002</year>
          )
          <article-title>Revised Papers</article-title>
          , pages
          <fpage>9</fpage>
          -
          <lpage>27</lpage>
          . Lecture Notes in Computer Science (LNCS) 2785, Springer, Heidelberg, Germany,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Peters</surname>
          </string-name>
          .
          <article-title>CLEF 2003 Methodology and Metrics</article-title>
          . In C. Peters,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          , and M. Kluck, editors,
          <source>Comparative Evaluation of Multilingual Information Access Systems: Fourth Workshop of the Cross-Language Evaluation Forum (CLEF 2003) Revised Selected Papers</source>
          , pages
          <fpage>7</fpage>
          -
          <lpage>20</lpage>
          . Lecture Notes in Computer Science (LNCS) 3237, Springer, Heidelberg, Germany,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>C. W.</given-names>
            <surname>Cleverdon</surname>
          </string-name>
          .
          <article-title>The Cranfield Tests on Index Languages Devices</article-title>
          . In K. Sp¨arck Jones and P. Willett, editors,
          <source>Readings in Information Retrieval</source>
          , pages
          <fpage>47</fpage>
          -
          <lpage>60</lpage>
          . Morgan Kaufmann Publisher, Inc., San Francisco, CA, USA,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>