<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Report on CLEF-2002 Experiments: Combining Multiple Sources of Evidence</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jacques Savoy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institut interfacultaire d'informatique, Université de Neuchâtel</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1996</year>
      </pub-date>
      <abstract>
        <p>For our second participation in the CLEF retrieval tasks, our first objective was to propose better and more general stopword lists for various European languages (namely, French, Italian, German, Spanish and Finnish) along with improved, simpler and efficient stemming procedures. Our second goal was to propose a combined query-translation approach that could cross language barriers and also an effective merging strategy based on logistic regression for accessing the multilingual collection. Finally, within the Amaryllis experiment, we wanted to analyze how a specialized thesaurus might improve retrieval effectiveness.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The corpora used in our experiments included newspapers such as the Los Angeles Times (1994, English) Le
Monde (1994, French), La Stampa (1994, Italian), Der Spiegel (1994/95, German) and Frankfurter Rundschau
(1994, German), NRC Handelsbald (1994/95, Dutch), Algemeen Dagblad (1995/95, Dutch) and Tidningarnas
Telegrambyrå (1994/95, Finnish). As a second source of information, we also used various articles edited by
news agencies such as EFE (1994, Spanish), and the Swiss news agency (1994, available in French, German and
Italian but without parallel translation). As shown in Table 1a and 1b, these corpora are of various sizes, with
the English, German, Spanish and Dutch collections being twice the volume of the French, Italian and Finnish
sources. On the other hand, the mean number of distinct indexing terms per document is relatively similar across
the corpora (around 120), and this number is a little bit higher for the English collection (167.33). The
Amaryllis collection contains abstracts of scientific papers written mainly in French and this corpus contains
fewer distinct indexing terms per article (70.418).</p>
      <sec id="sec-1-1">
        <title>Size (in MB)</title>
        <p># of documents
# of distinct terms
English
425 MB
113,005
330,753</p>
      </sec>
      <sec id="sec-1-2">
        <title>French</title>
        <p>243 MB
87,191
320,526
Number of distinct indexing terms / document
Mean 167.33 130.213
Standard deviation 126.315 109.151
Median 138 95
Maximum 1,812 1,622
Minimum 2 3
Max df 69,082 42,983
Number of indexing terms / document
Mean 273.846 181.559
Standard deviation 246.878 164.347
Median 212 129
Maximum 6,087 3,923
Minimum 2 3
Number of queries
Number rel. items
Mean rel./request
Standard deviation
Median
Maximum
Minimum
42
821
19.548
20.832
11.5
96 (#q:95)
1 (#q:97,98,136)</p>
        <p>50
1,383
27.66
34.293
13.5
177 (#q:95)
1 (#q:121)</p>
        <p>Italian
278 MB
108,578
503,550
When examining the number of relevant documents per request, Tables 1a and 1b show that the mean number is
always greater than the median (e.g., for the English collection, there is an average of 19.548 relevant documents
per query and the corresponding median is 11.5). These findings indicate that each collection contains numerous
queries with a rather small number of relevant items. For each collection, we encounter 50 queries except for the
Italian corpus (for which Query #120 does not have any relevant items) and the English collection (for which
Query #93, #96, #101, #110, #117, #118, #127 and #132 do not have any relevant items). The Finnish corpus
contains only 30 available requests while only 25 queries are included in the Amaryllis collection.
From the original documents and during the indexing process, we retained only the following logical sections in
our automatic runs: &lt;TITLE&gt;, &lt;HEADLINE&gt;, &lt;TEXT&gt;, &lt;LEAD&gt;, &lt;LEAD1&gt;, &lt;TX&gt;, &lt;LD&gt;, &lt;TI&gt; and &lt;ST&gt;. On
the other hand, we did conduct two experiments (indicated as manual runs), one with the French collection and
one with the German corpus, within which we retained the following tags: for the French collection: &lt;DE&gt;,
&lt;KW&gt; , &lt; TB&gt; , &lt; CHA1&gt;, &lt;SUBJECTS&gt;, &lt;NAMES&gt;, &lt;NOM1&gt;, &lt;NOTE&gt;, &lt;GENRE&gt;, &lt;PEOPLE&gt;, &lt;SU11&gt;,
&lt;SU21&gt;, &lt;GO11&gt;, &lt;GO12&gt;, &lt;GO13&gt;, &lt;GO14&gt;, &lt;GO24&gt;, &lt;TI01&gt;, &lt;TI02&gt;, &lt;TI03&gt;, &lt;TI04&gt;, &lt;TI05&gt;, &lt;TI06&gt;,
&lt;TI07&gt;, &lt;TI08&gt;, &lt;TI09&gt;, &lt;ORT1&gt;, &lt;SOT1&gt;, &lt;SYE1&gt; and &lt;SYF1&gt;; while for the German corpus and for one
experiment, we used also the following tags: &lt;KW&gt; and &lt;TB&gt;.</p>
        <p>From the topic descriptions we automatically removed certain phrases such as "Relevant document report …",
"Find documents …", "Trouver des documents qui parlent …", "Sono valide le discussioni e le decisioni …",
"Relevante Dokumente berichten …" or "Los documentos relevantes proporcionan información …".
To evaluate our approaches, we used the SMART system as a test bed for implementing the Okapi probabilistic
model [Robertson 2000] as well as other vector-space models. This year our experiments were conducted on an
Intel Pentium III/600 (memory: 1 GB, swap: 2 GB, disk: 6 x 35 GB).</p>
        <sec id="sec-1-2-1">
          <title>1.2. Stopword lists and stemming procedures</title>
          <p>In order to define general stopword lists, we used those lists already available for the English and French
languages [Fox 1990], [Savoy 1999], while for the other languages we established a general stopword list by
following the guidelines described in [Fox 1990]. These lists mainly contain the top 200 most frequent words
included in the various collections together with articles, pronouns, prepositions, conjunctions or very frequently
occurring verb forms (e.g., to be, is, has, etc.). Stopword lists used during our previous participation [Savoy
2002b] were often extended. For example for the English we used that provided by the SMART system (571
words), 431 Italian words (no change from last year), 462 French words (previously 217), 603 German words
(previously 294), 351 Spanish terms (previously 272), 1,315 Dutch terms (available at CLEF Web site) and
1,134 Finnish words (these stopword lists are available at www.unine.ch/info/clef/).</p>
          <p>After removing high frequency words, an indexing procedure uses a stemming algorithm that attempts to conflate
word variants into the same stem or root. In developing this procedure for the French, Italian, German and
Spanish languages, it is important to remember that these languages have more complex morphologies than does
the English language [Sproat 1992]. As a first approach, our intention was to remove only inflectional suffixes
such that singular and plural word forms or feminine and masculine forms conflate to the same root. More
sophisticated schemes have already been proposed for the removal of derivational suffixes (e.g., "-ize", "-ably",
"ship" in the English language), such as the stemmer developed by Lovins [1968] is based on a list of over 260
suffixes, while that of Porter [1980] looks for about 60 suffixes. Figuerola [2002] for example described two
different stemmers for the Spanish language, and the results show that removing only inflectional suffixes (88
different inflectional suffixes were defined) seemed to provide better retrieval levels than did removing both
inflectional and derivational suffixes (this extended stemmer included 230 suffixes).</p>
          <p>Our various stemming procedures can be found at www.unine.ch/info/clef/. This year we improved our
stemming algorithms for French, within which some derivational suffixes were also removed. For the Dutch
language, we use the Kraaij &amp; Pohlmann's stemmer (ruulst.let.ruu.nl:2000/uplift/ulift.html) [Kraaij 1996]. For
the Finnish language, our stemmer tries to conflate various word declinations into the same stem. Also, the
Finnish language makes a distinction between partial object and whole object (e.g., "syön leilää" or "I'm eating
bread" and "syön leivan" for "I'm eating the whole bread"). This aspect is not actually taken into consideration.
Finally, diacritic characters are usually not present in English collections (with some exceptions, such as "à la
carte" or "résumé"); and such characters are replaced by their corresponding non-accentuated letter in the Italian,
Dutch, Finnish, German and Spanish language.</p>
        </sec>
        <sec id="sec-1-2-2">
          <title>1.3. Decompounding German words</title>
          <p>Most European languages manifests other morphological characteristics that we have been considered by our
approach, with compound word constructions being just one example (e.g., handgun, worldwide). In German
compound words are widely used and this causes more difficulties than does English. For example, a life
insurance company employee would be "Lebensversicherungsgesellschaftsangestellter" (Leben + S + versicherung
+ S + gesellschaft +S + angestellter for life + insurance + company + employee). Also the morphological
marker ("S") is not always present (e.g., "Bankangestelltenlohn" built as Bank + angestellter + lohn (salary)). In
Finnish, we also encounter similar constructions as such as "rakkauskirje" (rakkaus + kirje for love + letter) or
"työviikko" (työ + viikko for work + week).</p>
          <p>schaften schaft
weisen weise
lischen lisch
lingen ling
igkeiten igkeit
lichkeit lichkeit
keiten keit
erheiten erheit
enheiten enheit
heiten heit
haften haft
halben halb
langen lang
erlichen erlich
enlichen enlich
lichen lich
baren bar
igenden igend
igungen igung
igen ig
enden end
isten ist
anten ant
ungen ung
schaft schaft
weise weise
lisch lisch
ismus ismus</p>
          <p>String sequence End of previous word Beginning of next word
. tion tion . ern er . schg
. ling ling . tät tät . schl
. igkeit igkeit . net net . schh
. lichkeit lichkeit . ens en . scht
. keit keit . ers er . dtt
. erheit erheit . ems em . dtp
. enheit enheit . ts t . dtm
. heit heit . ions ion . dtb
. lein lein . isch isch . dtw
. chen chen . rm rm . ldan
. haft haft . rw rw . ldg
. halb halb . nbr n br ldm
. lang lang . nb n b ldq
. erlich erlich . nfl n f l ldp
. enlich enlich . nfr n f r ldv
. lich lich . nf n f ldw
. bar bar . nh n h tst
. igend igend . nk n k rg
. igung igung . ntr n tr rk
. ig ig . f f f f f f rm
. end end . f f s f f rr
. ist ist . f k f k rs
. ant ant . f m f m rt
. tum tum . f p f p rw
. age age . f v f v rz
. ung ung . f w f w f p
. enden end . schb sch b f s f
. eren er . schf sch f gss
According to Monz &amp; de Rijke [2002] or [Chen 2002], including both compounds and their composite parts
(only noun-noun decompositions in [Monz 2002]) in queries and documents can result in better performance
while according to Molina-Salgado [2002], the decomposition of German words seems to reduce average
precision.</p>
          <p>Our approach seeks to break up those words having an initial length greater than or equal to eight characters.
Moreover, decomposition cannot take place before an initial sequence [V]C, meaning that a word might begin
with a series of vowels that must be followed by at least one consonant. The algorithm then seeks the
occurrence of one of the models described in Table 2. For example, the last model "gss g s" indicates that when
we encounter the character string "gss" the computer is allowed to cut the compound term, ending the first word
with "g" and beginning the second with "s". All the models depicted in Table 2 often include letters sequences
impossible to find in a simple German word such as "dtt," "fff," or "ldm". Once it has detected this pattern, the
computer makes sure that the right part consists of at least four characters, potentially beginning with a series of
vowels (criterion noted as [V]), followed by a CV sequence. If decomposition proves to be possible, the
algorithm begins working on the right part of the decomposed word.</p>
          <p>As an example, take the compound word "Betreuungsstelle" (meaning "care center" and made up "Betreuung"
(care) and "Stelle" (center, place)). This word is definitely more than seven characters long. Once this has been
verified, the computer begins searching for substitution models for the third character. The computer will find a
match with the last model described in Table 2, and form the words "Betreuung" and "Stelle." This break is
validated because the second word has a length greater than four characters. This term also meets criterion [V]CV
and finally, given that the term "Stelle" has less than eight letters, the computer will not attempt to continue
decomposing this term.</p>
        </sec>
        <sec id="sec-1-2-3">
          <title>1.4. Indexing and searching strategy</title>
          <p>In order to obtain a broader view of the relative merit of various retrieval models, we first adopted a binary
indexing scheme within which each document (or request) is represented by a set of keywords, without any
weight. To measure the similarity between documents and requests, we count the number of common terms,
computed according to the inner product (retrieval model denoted "doc=bnn, query=bnn" or "bnn-bnn"). For
document and query indexing however binary logical restrictions however are often too limiting. In order to
weight the presence of each indexing term in a document surrogate (or in a query), we may take account of the
term occurrence frequency which allows for better term distinction and increases indexing flexibility (retrieval
model notation: "doc=nnn, query=nnn" or "nnn-nnn").
nnn
atn
npn
Lnu
ntc
wij =
wij =
wij = tfij
wij = idfj . [0.5+ 0.5.tfij / max tfi.]</p>
          <p>⎡(n - df j )
wij = tf ij ⋅ ln⎢
⎢
⎣
⎤
⎥
df j ⎥</p>
          <p>⎦
⎛1 + ln(tfij )
⎜
⎝</p>
          <p>⎞
1+pivot⎠⎟
(1- slope) ⋅ pivot + slope ⋅ nti</p>
          <p>tfij ⋅ idf j
t
∑ (tfik ⋅ idfk )
k =1</p>
          <p>2
t
∑ ((ln(ln(tf ik ) +1) + 1)⋅ idf k )
k =1</p>
          <p>2
t
∑ ((ln(tfik ) + 1) ⋅ idfk )</p>
          <p>2
k=1
bnn
ltn
nfn
Okapi
lnc
dtc
ltc
dtu
wij = 1
wij = (ln(tfij ) + 1) . idfj</p>
          <p>⎡
wij = ln ⎢ n
⎣</p>
          <p>⎤
df j ⎦⎥
((k1 + 1) ⋅ tf ij )
wij =
wij =</p>
          <p>ln(tf ij ) + 1
t
∑ (ln(tf ik ) +1)
k =1</p>
          <p>2
wij =
(K + tf ij )
(ln(ln(tf ij) + 1) + 1)⋅ idf j
wij =</p>
          <p>(ln(tfij ) + 1)⋅ idf j
wij =
⎛ (1 + l n(1 + ln(tf ij))) ⋅idf j
⎜
⎜
⎝
⎞
⎟
1+pivot⎟</p>
          <p>⎠
(1- slope) ⋅ pivot + slope ⋅ nti
Those terms however that do occur very frequently in the collection are not considered very helpful in
distinguishing between relevant and non-relevant items. Thus we might count their frequency in the collection,
or more precisely the inverse document frequency (denoted by idf), resulting in more weight for sparse words and
less weight for more frequent ones. Moreover, a cosine normalization could prove beneficial and each indexing
weight could vary within the range of 0 to 1 (retrieval model notation: "ntc-ntc", Table 3 depicts the exact
weighting formulation).</p>
          <p>Other variants may also be created, especially if we consider the occurrence of a given term in a document is a
rare event. Thus, it may be a good practice to give more importance to the first occurrence of this word as
compared to any successive or repeating occurrences. Therefore, the tf component may be computed as 0.5 + 0.5
· [tf / max tf in a document] (retrieval model denoted "doc=atn").</p>
          <p>Finally, we should consider that a term's presence in a shorter document provides stronger evidence than it does
in a longer document. To account for this, we integrate document length within the weighting formula, leading
to more complex IR models; for example, the IR model denoted by "doc=Lnu" [Buckley 1996], "doc=dtu"
[Singhal 1999]. Finally for CLEF-2002, we also conducted various experiments using the Okapi probabilistic
model [Robertson 2000] within with K = k1 · [ ( 1 - b ) + b · ( l i / avdl)], representing the ratio between the
length of Di measured by li (sum of tfij ) and the collection mean noted by advl.</p>
          <p>In our experiments, the constants b, k1, advl, pivot and slope are fixed according to values listed in Table 4. To
evaluate the retrieval performance of these various IR models, we adopted the non-interpolated average precision
(computed on the basis of 1,000 retrieved items per request by the TREC-EVAL program), allowing for both
precision and recall using a single number.
For the German language, we determined that 5-gram indexing, decompounded indexing and word-based document
representation methods to be distinct and independent sources of evidence for German language document content.
We therefore decided to combine these three indexing schemes and to do so we normalized similarity values
obtained by each document extracted from these three separate retrieval models, according to Equation 1 (see
advl
Given that French, Italian and Spanish morphology is comparable to that of English, we decided to index French,
Italian and Spanish documents based on word stems. For the German, Dutch and Finnish languages and their
more complex compounding morphology, we decided to use a 5-gram approach [McNamee 2002]. However,
contrary to [McNamee 2002], our generation of 5-gram indexing terms does not span word boundaries. This
value of 5 was chosen because it performed better with the CLEF-2000 corpora [Savoy 2001a]. Using this
indexing scheme, the compound «das Hausdach» (the roof of the house) will generate the following indexing
terms: «das», «hausd», «ausda», «usdac» and «sdach».</p>
          <p>Our evaluation results as reported in Tables 5 show that the Okapi probabilistic model performs best with the
use of five different languages. In the second position, we usually find the vector-space model "doc=Lnu,
query=ltc" and in the third "doc=dtu, query=dtc". Finally, the traditional tf-idf weighting scheme ("doc=ntc,
query=ntc") does not exhibit very satisfactory results, and the simple term-frequency weighting scheme
("doc=nnn, query=nnn") or the simple coordinate match ("doc=bnn, query=bnn") results in poor retrieval
performance.</p>
        </sec>
      </sec>
      <sec id="sec-1-3">
        <title>Average precision</title>
        <p>Query T-D
Model
Section 3). The resulting average precision for these four approaches is shown in Table 5b, thus demonstrating
how the combined model usually results in better retrieval performance.
It was observed that pseudo-relevance feedback (blind-query expansion) seems to be a useful technique for
enhancing retrieval effectiveness. In this study, we adopted Rocchio's approach [Buckley 1996] with α = 0.75,
β = 0.75 whereby the system was allowed to add m terms extracted from the n best ranked documents from the
original query. To evaluate this proposition, we used the Okapi probabilistic model and we enlarged the query by
10 to 20 terms provided by the 5 or 10 best-retrieved articles. The results depicted in Table 6a and 6b indicate
that the optimal parameter setting seems to be collection-dependant. Moreover, performance improvement seems
also to be collection dependant (or language dependant) with no improvement for the English corpus yet an
increase of 8.55% for the Spanish corpus (from an average precision of 51.71 to 56.13), 9.85% for the French
corpus (from 48.41 to 53.18), 12.91% for the Italian language (41.05 to 46.35) and 13.26% for the German
collection (from 41.25 to 46.72, combined model, Table 6b).</p>
      </sec>
      <sec id="sec-1-4">
        <title>Average precision German 5-gram 50 queries</title>
        <p>This year, we also participated in the Dutch and Finnish monolingual tasks, the results of which are depicted in
Table 7, and the average precision of the Okapi model using blind-query expansion is given in Table 8. For
these two languages, we also applied or combined an indexing model based on 5-gram indexing and word-based
document representations. While for the Dutch language, our combined model seems to enhance the retrieval
effectiveness, for the Finnish language it does not. This however was a first trial for our proposed stemmer and
it seemed to improve the average precision over a baseline trial without stemming procedure (Okapi model,
unstemmed 23.04, with stemming 30.45, an improvement of +32.16%).
In the monolingual track, we submitted six runs along with their corresponding descriptions, as listed in
Table 9. Four of them were fully automatic using the request's Title and Descriptive logical sections, while the
last three used more other document sections, based on the request's Title, Descriptive and Narrative sections. In
these last three runs, two were labeled "manual" because we used logical sections containing manually assigned
index terms. For all other runs however we did not use any manual intervention during the indexing and retrieval
procedures.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Bilingual information retrieval</title>
      <p>In order to overcome language barriers, we based our approach on free and readily available translation resources
that automatically translate queries into the desired target language. More precisely, the original queries were
written in English and we used no parallel or aligned corpora to derive statistically or semantically related words
in the target language. Section 2.1 describes our combined strategy for cross-lingual retrieval while Section 2.2
provides some examples of translation errors.</p>
      <p>This year, we used five machine translation systems, namely SYSTRAN™ (babel.altavista.com/translate.dyn),
GOOGLE.COM (www.google.com/language_tools), FREETRANSLATION.COM (www.freetranslation.com),
INTERTRAN (www.tranexp.com:2000/InterTran) and REVERSO ONLINE (translation2.paralink.com). As
bilingual dictionary we used the BABYLON system (www.babylon.com).</p>
      <sec id="sec-2-1">
        <title>2.1. Query automatic translation</title>
        <p>In order to develop a fully automatically approach, we chose to translate the requests using five different machine
translation (MT) systems. We also translated query terms word-by-word using the BABYLON bilingual
dictionary, provides not only one but several terms for the translation for each word submitted. In our
experiments, we decided to pick the first translation available (labeled "baby1"), the first two terms (labeled
"baby2") or the first three available translations (labeled "baby3").
The first part of Table 10 lists the average precision for each translation devices used along the performance
achieved by manually translated requests. For German, we also reported the retrieval effectiveness achieved by
the three difference approach, namely using words as indexing terms, decompounding the German words
according to our approach and the 5-grams model. While the REVERSO system seems to be the better choice for
German and Spanish, FREETRANSLATION is the best choice for Italian and BABYLON 1 the best for French.
In order to improve search performance, we tried combining different machine translation systems with the
bilingual dictionary approach. In this case, we formed the translated query by concatenating the different
translations provided by the various approaches. Thus the column header "Comb 1", we combined one machine
translation system with the bilingual dictionary ("baby1"). Similarly, under columns "Comb 2" or "Comb 2b,"
we listed the results of two machine translation approaches and three machine translation systems under column
headings "Comb 3", "Comb 3b" or "Comb 3b2". With the exception of the performance under "Comb 3b2,"
we also included terms provided by the "baby1" dictionary look-up in the translated requests. In columns
"MT 2" and "MT 3," we evaluated the combination of two and three machine translation systems respectively.
Finally, we could also combine all translation sources (under heading "All") or all machine translation
approaches under the heading "MT all."
Since the performance of each translation device depends on the target language, in the lower part of Table 10 we
included the exact specification for each of the combined runs. For the German language, for each of the three
indexing models, we used the same combination of translation resources. From an examination of the retrieval
effectiveness of our various combined approaches listed in the middle part of Table 10, a clear recommendation
cannot be made. Overall, it seems better to combine two or three machine translation systems with the bilingual
dictionary approach ("baby1"). However, combining the five machine translation systems (heading "MT all") or
all translation tools (heading "All") does not result in a very effective performance.</p>
        <sec id="sec-2-1-1">
          <title>Query T-D</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>Combined</title>
          <p>Expand # docs / # terms
Corrected
Official
Query T-D</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>Combined Expand # docs / # terms Corrected Official</title>
          <p>In order to obtain a preliminary picture of the automatic translation approach's underlying difficulties, we
analyzed some queries through comparing translations produced by our six machine-based tools with the request
formulation written by a human being (examples are given in Table 12). As a first example, the title of Query
#113 is "European Cup". In this case, the term "cup" was analyzed as a teacup by all automatic translation
tools, resulting in the French translations "tasse" or "verre" (or "tazza" in Italian, "Schale" in German ("Pokal"
can be viewed as a correct translation alternative) and "taza" or "Jícara" (small teacup) in Spanish).
In Query #118 ("Finland's first EU Commissioner"), the machine translation systems failed to give the
appropriate Spanish term "comisario" for "Commissioner" but returned "comisión" (commission) or
"Comisionado" (adjective relative to commission). For this same request number, the manually translated query
seemed to contain a spelling error in Italian ("commis ario" instead of "commis s ario"). For the same request, the
translation given in German "Beauftragter" (delegate) does not correspond to the appropriate term "Kommissar"
(more the missing "-" in the translation "EUBEAUFTRAGTER").</p>
          <p>Other examples: for Query #94 ("Return of Solzhenitsyn") which is translated manually in German ("Rückkehr
Solschenizyns"), our automatic translation systems fail to translate the proper noun (returning "Solzhenitsyn"
instead of "Solschenizyns"). Query #109 ("Computer Security") is translated manually Spanish as "Seguridad
Informática" and our various translations devices return different terms for "Computer" (e.g., "Computadora",
"Computador", or "ordenador") but not the word "Informática".</p>
          <p>&lt;num&gt; C113 (query translations failed in French, Italian, German and Spanish)
&lt;EN-title&gt; European Cup
&lt;FR-title manually translated&gt; Coupe d'Europe de football
&lt;FR-title FREETRANSLATION&gt; Tasse européenne
&lt;FR-title BYBYLON 1&gt; Européen verre
&lt;FR-title BYBYLON 2&gt; Européen résident de verre tasse
&lt;FR-title BYBYLON 3&gt; Européen résident de l'Europe verre tasse coupe
&lt;IT-title manually translated&gt; Campionati europei
&lt;IT-title SYSTRAN&gt; Tazza Europea
&lt;IT-title GOOGLE&gt; Tazza Europea
&lt;GE-title manually translated&gt; Fussballeuropameisterschaft
&lt;GE-title SYSTRAN&gt; Europäische Schale
&lt;GE-title R EVERSO&gt; Europäischer Pokal
&lt;ES-title manually translated&gt; Eurocopa
&lt;ES-title I NTERTRAN&gt; Europea Jícara
&lt;ES-title REVERSO&gt; Taza europea
&lt;num&gt; C118 (query translations failed in Italian, German and Spanish)
&lt;EN-title&gt; Finland's first EU Commissioner.
&lt;IT-title manually translated&gt; Primo commisario europeo per la Finlandia
&lt;IT-title GOOGLE&gt; Primo commissario dell'Eu della Finlandia.
&lt;IT-title F REETRANSLATION&gt; Finlandia primo Commissario di EU.
&lt;GE-title manually translated&gt; Erster EU-Kommissar aus Finnland
&lt;GE-title GOOGLE&gt; Finnlands erster EUBEAUFTRAGTER.
&lt;GE-title R EVERSO&gt; Finlands erster EG-Beauftragter
&lt;ES-title manually translated&gt; Primer comisario finlandés de la UE
&lt;ES-title GOOGLE&gt; Primera comisión del EU de Finlandia.
&lt;ES-title REVERSO&gt; El primer Comisionado de Unión Europea de Finlandia.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Multilingual information retrieval</title>
      <p>Using our combined approach to automatically translate a query, we were able to search a document collection for
a request written in English. This stage however represents only the first step in a proposal for multi-language
information retrieval systems. We also need to investigate situations where users write a request in English in
order to retrieve pertinent documents in English, French, Italian, German and Spanish. To deal with this
multilanguage barrier, we divided our document sources according to language and thus formed five different
collections. After searching in these corpora and obtaining five results lists, we needed to merge them in order to
provide users with a single list of retrieved articles.</p>
      <p>Recent works have suggested various solutions to merging the separate result list obtained from different
collections or distributed information services. As a first approach, we will assume that each collection contains
approximately the same number of pertinent items and that the distribution of the relevant documents is similar
across the result lists. Based solely on the rank of the retrieved records, we can interleave the results in a
roundrobin fashion. According to previous studies [Voorhees 1995], the retrieval effectiveness of such an interleaving
scheme is around 40% below that achieved from a single retrieval scheme working with a single huge collection,
representing the entire set of documents.</p>
      <p>To take account of the document score computed for each retrieved item (or the similarity value between the
retrieved record and the request, denoted score rsvj), we might formulate the hypothesis that each collection is
searched by the same or a very similar search engine and that the similarity values are therefore directly
comparable [Kwok 1995]. Such a strategy, called raw-score merging, produces a final list sorted by the
document score computed by each collection. However, collection-dependent statistics in document or query
weights may vary widely among collections, and therefore this phenomenon may invalidate the raw-score
merging hypothesis.</p>
      <p>To account for this fact, we might normalize the document scores within each collection by dividing them by the
maximum score (i.e. the document score of the retrieved record in the first position). As a variant of this
normalized score merging scheme, Powell et al. [2000] suggest normalizing the document score rsvj according to
the following formula:
in which rsvj is the original retrieval status value (or document score), and rsvmax and rsvmin are the maximum and
minimum document score values that a collection could achieve for the current request. In this study, the rsvmax
is given by the document score achieved by the first retrieved item and the retrieval status value obtained by the
1000th retrieved record gives the value of rsvmin.</p>
      <p>As a fourth strategy, we might use the logistic regression [Flury 1997, Chapter 7] to predict the probability of a
binary outcome variable, according to a set of explanatory variables. Based on this statistical approach, Le Calvé
and Savoy [2000] and Savoy [2002a] described how to predict the probability of relevance of those documents
retrieved by different retrieval schemes or collections. The resulting estimated probabilities would be predicted
according to both the original document score rsvi and the logarithm of the ranki attributed to the corresponding
document Di. Based on these estimated relevance probabilities, we sorted the records retrieved from separate
collections in order to obtain a single ranked list. However, in order to estimate the underlying parameters, this
approach requires a training set, in this case the CLEF-2001 topics and their relevance assessments.</p>
      <p>Prob [Di is rel | rank i , rsv i ] =</p>
      <p>eα+β1⋅ln(ranki) +β2 ⋅rsvi
1 + eα+β1⋅ln(ranki )+β 2⋅rsvi
within which ranki denotes the rank of the retrieved document Di, ln() is the natural logarithm, and rsvi is the
retrieval status value (or document score) of the document Di. In this equation, the coefficients α, β1 and β2 are
unknown parameters that are estimated according the method of the maximum likelihood (the required
computations have been done with the S language).</p>
      <p>rsv′ j = (rsvj - rsv min )</p>
      <p>(rsvmax - rsv min )
Our official and corrected results are shown in Table 14 while some statistics about the number of documents
provided by each collection are given in Table 15. From this data, we can see that the normalized score merging
(UniNEm1) extracts more documents for the English corpus (in mean 24.94 items) than the logistic regression
When searching in multi-lingual corpora using Okapi, the round-robin scheme or the raw-score merging strategy
provide very similar retrieval performances (see Table 13). The normalized score merging based on Equation 1
shows an enhancement over the round-robin approach (36.62 vs. 34.27, an improvement of +6.86% in our first
experiment, and 36.90 vs. 33.97, +8.63% in our second run). Using our logistic model with only the rank as
explanatory variable (or more precisely the ln(ranki), performance depicted under the label "Log ln(ranki)"), the
resulting average precision is lower than the normalized score merging. When merging the result lists based on
the logistic regression approach (using both the rank and the document score as explanatory variables) presents
the best average precision.
model (UniNEm2 where in mean 11.44 documents are coming from the English collection). Moreover, the
logistic regression scheme takes more documents from the Spanish and German collections Finally, we can see
that the percentage of relevant items is relatively similar when comparing CLEF01 and CLEF02 test-collections.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Amaryllis experiments</title>
      <p>For the Amaryllis experiments, we wanted to determine whether a specialized thesaurus might improve the
retrieval effectiveness over a baseline, ignoring term relationships. From the original documents and during the
indexing process, we retained only the following logical sections in our runs: &lt;text&gt;, &lt;ti&gt;, &lt;ab&gt;, &lt;mc&gt;, &lt;kw&gt;.
&lt;RECORD&gt;
&lt;TERMFR&gt; Analyse de poste
&lt;TRADENG&gt; Station Analysis
…
&lt;RECORD&gt;
&lt;TERMFR&gt; Bureau poste
&lt;TRADENG&gt; Post offices
&lt;RECORD&gt;
&lt;TERMFR&gt; Bureau poste
&lt;TRADENG&gt; Post office
…
&lt;RECORD&gt;
&lt;TERMFR&gt; Isolation poste électrique
&lt;TRADENG&gt; Substation insulation
…
&lt;RECORD&gt;
&lt;TERMFR&gt; Caserne pompier
&lt;TRADENG&gt; Fire houses
&lt;SYNOFRE1&gt; Poste incendie
…
&lt;RECORD&gt;
&lt;TERMFR&gt; Habitacle aéronef
&lt;TRADENG&gt; Cockpits (aircraft)
&lt;SYNOFRE1&gt; Poste pilotage
…
&lt;RECORD&gt;
&lt;TERMFR&gt; La Poste
&lt;TRADENG&gt; Postal services
…
&lt;RECORD&gt;
&lt;TERMFR&gt; Poste conduite
&lt;TRADENG&gt; Operation platform
&lt;SYNOFRE1&gt; Cabine conduite
…
&lt;RECORD&gt;
&lt;TERMFR&gt; POSTE DE TRAVAIL
&lt;TRADENG&gt; WORK STATION
&lt;RECORD&gt;
&lt;TERMFR&gt; Poste de travail
&lt;TRADENG&gt; Work Station
&lt;RECORD&gt;
&lt;TERMFR&gt; Poste de travail
&lt;TRADENG&gt; Work station
&lt;RECORD&gt;
&lt;TERMFR&gt; Poste de travail
&lt;TRADENG&gt; workstations
&lt;SYNOFRE1&gt; Poste travail
…
From the given thesaurus, we have extracted 126,902 terms having a relationship with one or more terms (the
thesaurus owns 173,946 entries delimited by the tags &lt;RECORD&gt; … &lt;/RECORD&gt;, however only 149,207 entries
have at least one relationship with another term. From these 149,207 entries, we found 22,305 multiple entries
(that are removed, as for example, the term "Poste de travail" or "Bureau poste" in Table 16). In building our
thesaurus, we removed the accents, wrote all terms in lowercase, and ignored numbers and terms given between
parenthesis. For example, the word "poste" appears in 49 records (usually as part of a compound entry in the
&lt;TERMFR&gt; field).</p>
      <p>From our 126,902 entries, we counted 107,038 TRADEENG relationships, 14,590 SYNOFRE1, 26,772 AUTOP1
relationships and 1,071 VAUSSI1 relationships (see examples given in Table 16). In a first set of experiments,
we did not use this thesaurus and we used the Title and Descriptive logical sections of the requests (second
column of Table 17a) or the Title, Descriptive and Narrative parts of the queries (last column of Table 17a). In
a second set of experiments, we included all related words that could be found in the thesaurus using only the
search keywords (average precision depicted under the label "Qthes"). In a third experiment, we enlarged only
document representatives using our thesaurus (performance shown under column heading "Dthes"). In a last
experiment, we take account for related words found in the thesaurus only for document surrogates and under the
additional condition that such relationship can be found with at least three terms (e.g. "moteur à combustion" is a
valid candidate but not single term like "moteur"). On the other hand, we also included in the query all
relationships that can be found using the search keywords (performance shown under the column heading
"Dthes3Qthes").
From the achieved average precision depicted in Tables 17a and 17b, we cannot infer that the available thesaurus
is really helpful in improving retrieval effectiveness, at least as implemented in this study.</p>
      <sec id="sec-4-1">
        <title>UniNEama1</title>
        <p>UniNEama2
UniNEama3
UniNEama4
UniNEamaN1
T-D
T-D
T-D
T-D
T-D-N</p>
      </sec>
      <sec id="sec-4-2">
        <title>Form</title>
        <p>automatic
automatic
automatic
automatic
automatic</p>
      </sec>
      <sec id="sec-4-3">
        <title>Model</title>
      </sec>
      <sec id="sec-4-4">
        <title>Okapi</title>
        <p>Okapi
Okapi
Okapi
Okapi</p>
      </sec>
      <sec id="sec-4-5">
        <title>Thesaurus</title>
        <p>no
with query terms
with documents
both query &amp; doc
no
25 docs / 50 terms
25 docs / 25 terms
25 docs / 50 terms
10 docs / 15 terms
25 docs / 50 terms</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>For our second participation in CLEF retrieval tasks, we suggested a general stopword list and stemming
procedure for the French, Italian, German, Spanish and Finnish languages. We also suggested a simple
decompounding approach for the German language. For the Dutch, Finnish and German languages we were to
consider 5-gram indexing and word-based (and decompounding-based) document representations to be distinct and
independent sources of evidence on document content, and it would be a good practice to combine these two (or
three) indexing schemes.</p>
      <p>To improve bilingual information retrieval, we suggest using not only one but two or three different translation
sources to translate the query into the target languages. Such a combination seems to improve the retrieval
effectiveness. In the multilingual environment, we demonstrated that a learning scheme such as logistic
regression could perform effectively. As a second best solution, we suggested using a simple normalization
procedure based on the document score.</p>
      <p>Finally, in the Amaryllis experiments, we studied various possible ways we could use a specialized thesaurus to
improve average precision. However, the various strategies used in this paper do not demonstrate clear
enhancement over a baseline that ignores the term relationships stored in the thesaurus.</p>
      <p>Acknowledgments
The author would like to thank C. Buckley from SabIR for giving us the opportunity to use the SMART
system, without which this study could not have been conducted. This research was supported in part by the
SNSF (Swiss National Science Foundation) under grants 21-58 813.99 and 21-66 742.01.
[Chen 2002]
[Figuerola 2002]
[Flury 1997]
[Fox 1990]
[Kraaij 1996]</p>
      <p>Buckley, C., Singhal, A., Mitra, M. &amp; Salton, G. (1996). New retrieval approaches using
SMART. In Proceedings of TREC'4, (pp. 25-48). Gaithersburg: NIST Publication
#500236.</p>
      <p>Chen, A. (2002). Multilingual information retrieval using English and Chinese queries. In
C. Peters, M. Braschler, J. Gonzalo &amp; M. Kluck (Eds.), Evaluation of cross-language
information retrieval systems. Lecture Notes in Computer Science #2409. Berlin:
Springer-Verlag.</p>
      <p>Figuerola, C.G., Gómez, R. &amp; Zazo Rodríguez, A.F. (2002). Stemming in Spanish: A
first approach to its impact on information retrieval. In C. Peters, M. Braschler, J.
Gonzalo &amp; M. Kluck (Eds.), Evaluation of cross-language information retrieval systems.
Lecture Notes in Computer Science #2409. Berlin: Springer-Verlag.</p>
      <p>Flury, B. (1997). A first course in multivariate statistics . New York: Springer.</p>
      <p>Fox, C. (1990). A stop list for general text. ACM-SIGIR Forum, 24, 19-35.</p>
      <p>Kraaij, W. &amp; Pohlmann, R. (1996). Viewing stemming as recall enhancement. In
Proceedings of the 19th International Conference of the ACM-SIGIR'96, (pp. 40-48). New
York: The ACM Press.</p>
      <p>Kwok, K.L., Grunfeld, L. &amp; Lewis, D.D. (1995). TREC-3 ad-hoc, routing retrieval and
thresholding experiments using PIRCS. In Proceedings of TREC'3, (pp. 247-255).
Gaithersburg: NIST Publication #500-225.</p>
      <p>Le Calvé, A., Savoy, J. (2000). Database merging strategy based on logistic regression.
Information Processing &amp; Management, 36(3), 341-359.</p>
      <p>Lovins, J. B. (1968). Development of a stemming algorithm. Mechanical Translation and
Computational Linguistics, 11(1), 22-31.
[McNamee 2002] McNamee, P. &amp; Mayfield, J. (2002). JHU/APL Experiments at CLEF: Translation
Resources and Score Normalization. In C. Peters, M. Braschler, J. Gonzalo &amp; M. Kluck
(Eds.), Evaluation of Cross-Language Information Retrieval Systems. Lecture Notes in
Computer Science #2409. Berlin: Springer-Verlag.
[Molina-Salgado 2002] Molina-Salgado, H., Moulinier, I., Knutson, M., Lund, E. &amp; Sekhon, K. (2002).</p>
      <p>Thomson legal and regulatory at CLEF 2001: Monolingual and bilingual experiments. In
C. Peters, M. Braschler, J. Gonzalo &amp; M. Kluck (Eds.), Evaluation of cross-language
information retrieval systems. Lecture Notes in Computer Science #2409. Berlin:
Springer-Verlag.
[Monz 2002] Monz, C. &amp; de Rijke, M. (2002). The University of Amsterdam at CLEF 2001. In C.</p>
      <p>Peters, M. Braschler, J. Gonzalo &amp; M. Kluck (Eds.), Evaluation of cross-language
information retrieval systems. Lecture Notes in Computer Science #2409. Berlin:
Springer-Verlag.
[Porter 1980] Porter, M.F. (1980). An algorithm for suffix stripping. Program, 14, 130-137.
[Powell 2000] Powell, A.L., French, J. C., Callan, J., Connell, M. &amp; Viles, C.L. (2000). The impact of
database selection on distributed searching. In Proceedings of the 23rd International
Conference of the ACM-SIGIR'2000, (pp. 232-239). New York: The ACM Press.
[Robertson 2000] Robertson, S.E., Walker, S. &amp; Beaulieu, M. (2000). Experimentation as a way of life:</p>
      <p>Okapi at TREC. Information Processing &amp; Management, 36(1), 95-108.
[Savoy 1999] Savoy, J. (1999). A stemming procedure and stopword list for general French corpora.</p>
      <p>Journal of the American Society for Information Science, 50(10), 944-952.
[Savoy 2002a] Savoy, J. (2002). Cross-language information retrieval: Experiments based on CLEF-2000
corpora. Information Processing &amp; Management, to appear.
[Savoy 2002b] Savoy, J. (2002). Report on CLEF-2001 Experiments: Effective Combined
QueryTranslation Approach. In C. Peters, M. Braschler, J. Gonzalo &amp; M. Kluck (Eds.),
Evaluation of cross-language information retrieval systems. Lecture Notes in Computer
Science #2409. Berlin: Springer-Verlag.
[Savoy 2002c] Savoy, J. (2002). Recherche d'informations dans des corpus en langue française :
Utilisation du référentiel Amaryllis. TSI, Technique et Science Informatiques, 21(3),
345373.
[Singhal 1999] Singhal, A., Choi, J., Hindle, D., Lewis, D.D. &amp; Pereira, F. (1999). AT&amp;T at TREC-7. In</p>
      <p>Proceedings TREC-7, (pp. 239-251). Gaithersburg: NIST Publication #500-242.
[Sproat 1992] Sproat, R. (1992). Morphology and computation. Cambridge: The MIT Press.
[Voorhees 1995] Voorhees, E.M., Gupta, N.K. &amp; Johnson-Laird, B. (1995). The collection fusion problem.</p>
      <p>In Proceedings of TREC'3, (pp. 95-104). Gaithersburg: NIST Publication #500-225.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>