<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Selection and Merging Strategies for Multilingual Information Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jacques Savoy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierre-Yves Berger</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jacques.Savoy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Université de Neuchâtel</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>For our fourth participation in the CLEF evaluation campaigns, our objective was to verify whether our combined query translation approach would work well with new requests and new languages (Russian and Portuguese in this case). As a second objective, we suggested a selecting procedure that could extract a smaller number of documents from collections that for the current request seem to contain no or only few relevant items. We also applied different merging strategies in order to obtain more evidence on the respective relative merits.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>1. Bilingual Information Retrieval
SYSTRAN
GOOGLE
FREETRANSLATION
INTERTRAN
REVERSO
WORLDLINGO
BABELFISH
PROMPT
ONLINE
BABYLON
www.systranlinks.com/
www.google.com/language_tools
www.freetranslation.com/web.htm
intertran.tranexp.com/
www.reverso.fr/url_translation.asp
www.worldlingo.com/
babelfish.altavista.com/
webtranslation.paralink.com/
www.online-translator.com/srvurl.asp?lang=en
www.babylon.com</p>
      <p>When using the Babylon bilingual dictionary to translate an English request word-by-word, usually more
than one translation is provided, in an unspecified order. We decided to pick only the first translation available
(labeled “Babylon 1”), the first two terms (labeled “Babylon 2”) or the first three available translations (labeled
“Babylon 3”).</p>
      <p>
        Table 1 shows the resulting mean average precision using the various translation tools and the Okapi
probabilistic model (see Savoy (2004c) for implementation details). Of course, not all tools can be used for
each language, and thus as shown in Table 1 various entries are missing (indicated with the label “N/A”). From
this data, we can see that the results from the FreeTranslation MT system usually obtain satisfactory retrieval
performances (around 82% of the MAP of the corresponding monolingual search). As another good translation
system, we may mention Reverso or BabelFish for the French, Prompt for the Russian or Online for both the
Russian and Portuguese languages. For the Finnish language we found only two translation tools, but
unfortunately their overall performance levels were not very good (a similar low level performance was also
found when translating English topics into various Asian languages
        <xref ref-type="bibr" rid="ref11 ref12 ref13">(Savoy 2004b)</xref>
        ). Not surprisingly, we
found there was a relationship between the various translation tools. For example, the Systran, BabelFish, and
WorldLingo MT systems appeared to be nearly identical MT systems.
      </p>
      <sec id="sec-1-1">
        <title>Manual</title>
        <p>Systran
Google
FreeTrans
InterTrans
Reverso
WorldLingo
BabelFish
Prompt
Online
Babylon 1
Babylon 2
Babylon 3</p>
        <p>French</p>
        <p>Okapi
49 queries</p>
        <p>0.4685
0.3729 (79.6%)
0.3680 (78.5%)
0.3845 (82.1%)
0.2664 (56.9%)
0.3830 (81.8%)
0.3728 (79.6%)
0.3729 (79.6%)</p>
        <p>N/A</p>
        <p>
          N/A
0.3706 (79.1%)
0.3356 (71.6%)
0.3378 (72.1%)
It is known that while a given translation tool may produce acceptable translations for a given set of
requests, it may perform poorly for other queries
          <xref ref-type="bibr" rid="ref10">(Savoy 2003; 2004a)</xref>
          . To date we have not been able to detect
very precisely when a given translation will produce satisfactory retrieval performance and when it will fail. In
this vein,
          <xref ref-type="bibr" rid="ref5">Kishida et al. (2004)</xref>
          suggest using a linear regression model to predict the average precision of the
current query, based on both manual evaluations of translation quality for the current query and the underlying
topic difficulty. In this study, before carrying out the retrievals, we chose to concatenate two or more
translations before submitting a query for translation.
        </p>
        <p>Language
Combination
Comb!1
Comb!2
Comb!3
Comb!4
Comb!5
Best single
Comb!1
Comb!2
Comb!3
Comb!4
Comb!5</p>
        <p>French</p>
        <p>Okapi
49 queries</p>
        <p>Bab2+Free
Bab2+Reverso
Reverso+Systran</p>
        <p>Free+Rev
Bab2+Free+</p>
        <p>Reverso
0.3845
0.3784
0.3857
0.3858
0 . 4 0 6 6
0.3962</p>
        <p>Finnish
Okapi
45 queries
Bab1+Inter</p>
        <p>Mean average precision</p>
        <p>Finnish
Okapi
45 queries</p>
        <p>Bab1+Inter
0.2290
0 . 2 5 2 9
0.2653
0 . 3 0 4 2</p>
        <p>Russian
Okapi
34 queries
Bab1+Free
Free+Prompt
Prompt+Online</p>
        <p>Free+Online
Bab1+Free+</p>
        <p>Online
0.3067
0 . 3 8 8 8
0.3032
0.2964
0.3043
0.3324</p>
        <p>Portuguese</p>
        <p>Okapi
46 queries
Free+Online
Bab1+Systran
Bab1+Free+Onl
Bab1+Free+Sys</p>
        <p>Bab1+Free+
Online+Systran
0.4057
0.4072
0.3713
0 . 4 2 0 4
0.3996
0.4070</p>
        <p>
          The resulting retrieval performances shown in Table 2 are sometimes better than the best single translation
scheme indicated in the row labeled “Best single” (e.g., the strategies “Comb 4” or “Comb 5” for French, or
“Comb 1” for Russian, and “Comb 3” for the Portuguese language). Of course, the main difficulty in this
bilingual search was the translation of English topics into Finnish, due to limited number of free translation
tools. When handling those languages less-often speaking around the world, it seems it would be worthwhile
considering other translation alternatives, such as probabilistic translation based on parallel corpora
          <xref ref-type="bibr" rid="ref9">(Nie et al.
1999)</xref>
          ,
          <xref ref-type="bibr" rid="ref8">(MacNamee &amp; Mayfield 2003)</xref>
          .
        </p>
        <p>For monolingual searches, as described in Savoy (2004c), we used a data fusion search strategy that
combined the Okapi and Prosit probabilistic models (see details in Section 2). The data shown in Table 3
indicates that our data fusion approaches may result in better retrieval effectiveness (except for the Finnish
4gram indexing scheme or the Russian corpus). Of course before combining the result lists we could also
automatically expand the translated queries, using a pseudo-relevance feedback method (Rocchio’s approach in
the present case). The resulting mean average precision as shown in Table 4 did not improve the retrieval
effectiveness when compared to the best single approach. In Tables!3 and 4, under the heading “Z-scoreW”, we
attached a weight of 1.5 to the Prosit model, and 1 to the Okapi model. Finally, Table 5 depicts the parameters
used for our official bilingual runs.
2. Multilingual Information Retrieval</p>
        <p>
          Our multilingual information retrieval system is based on the use of a query translation strategy instead of
either translating all documents into a common language (e.g., English), combining both query and document
translations
          <xref ref-type="bibr" rid="ref3 ref4">(Chen &amp; Gey 2003)</xref>
          or ignoring the translation phase
          <xref ref-type="bibr" rid="ref2">(Buckley et a l . 1998)</xref>
          ,
          <xref ref-type="bibr" rid="ref8">(MacNamee &amp;
Mayfield 2003)</xref>
          ; for a general overview of these questions, see
          <xref ref-type="bibr" rid="ref1">(Braschler &amp; Peters 2004)</xref>
          . In our approach,
when a request was received (in English in this study), we automatically translated it into the desired target
languages and then searched for pertinent items within each of the four corpora (English, French, Finnish and
Russian). After receiving a result list from each search engine, we needed to introduce a merging procedure to
provide a unique ranked result list. As a first approach to this problem, we considered the round-robin approach
whereby we took one document in turn from each individual list
          <xref ref-type="bibr" rid="ref14">(Voorhees et al. 1995)</xref>
          .
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>Russian</title>
        <p>34 queries
Prosit (3/15)
Okapi (3/10)</p>
        <p>Round-robin
Pro-Free-Reverso
0.2962
UniNEBru2</p>
      </sec>
      <sec id="sec-1-3">
        <title>Portuguese</title>
        <p>46 queries
Prosit (10/20)
Okapi (5/15)
Norm RSV
Onl-Free-Bab1
0.4704
UniNEBpt1</p>
      </sec>
      <sec id="sec-1-4">
        <title>Portuguese</title>
        <p>46 queries
Okapi (0/0)
Prosit (0/0)</p>
        <p>Norm RSV
Onl-Free-Sys-Bab1</p>
        <p>0.4491
UniNEBpt2</p>
        <p>
          To account for the document score computed for each retrieved item (denoted RSVk for document Dk), we
might formulate the hypothesis that each collection is searched by the same or a very similar search engine and
that the similarity values are therefore directly comparable
          <xref ref-type="bibr" rid="ref6">(Kwok et al. 1995)</xref>
          . Such a strategy is called
rawscore merging and produces a final list sorted by the document score computed by each collection. When using
the same IR model (with the same or very similar parameter settings) to search into all collections, such a
merging strategy may produce good retrieval performance (e.g., with a logistic regression IR model in
          <xref ref-type="bibr" rid="ref3 ref4">(Chen
2003)</xref>
          ).
        </p>
        <p>Unfortunately the document scores cannot always be directly compared, thus as a third merging strategy we
normalized the document scores within each collection by dividing them by the maximum score (i.e. the
document score of the retrieved record in the first position) and denoted them “Norm Max”. As a variant of this
normalized score merging scheme (denoted “Norm RSV”), we could normalize the document RSVk scores within
the ith result list, according to the following formula:</p>
        <p>Norm RSVk = ((RSVk - MinRSVi) / (MaxRSVi - MinRSVi))</p>
        <p>
          As a fifth merging strategy, we might use logistic regression to predict the probability of a binary outcome
variable, according to a set of explanatory variables
          <xref ref-type="bibr" rid="ref7">(Le Calvé &amp; Savoy 2000)</xref>
          . In our current case, we predicted
the probability of relevance of document Dk given both the logarithm of its rank (indicated by ln(rankk)) and the
original document score RSVk as indicated in Equation 2. Based on these estimated relevance probabilities
(computed independently for each language using the S+ software), we sorted the records retrieved from separate
collections in order to obtain a single ranked list. However, in order to estimate the underlying parameters, this
approach requires that a training set is available. To achieve this, we used the CLEF-2003 topics and their
relevance assessments in our evaluations.
        </p>
        <p>Pr ob [Dk is rel | rank k , rsvk ] =</p>
        <p>ea+b1⋅ln(rank k )+b2 ⋅rsv k
1 + ea+b1⋅ln(rank k )+b2 ⋅rsv k
(1)
(2)</p>
      </sec>
      <sec id="sec-1-5">
        <title>TD Queries</title>
        <p>†</p>
      </sec>
      <sec id="sec-1-6">
        <title>Condition A</title>
        <p>IR model 1 (#docs/#terms)
IR model 2 (#docs/#terms)
Data fusion operator
Translation tools
Mean average precision</p>
      </sec>
      <sec id="sec-1-7">
        <title>Condition B</title>
        <p>IR model 1 (#docs/#terms)
IR model 2 (#docs/#terms)
Data fusion operator
Translation tools
Mean average precision</p>
      </sec>
      <sec id="sec-1-8">
        <title>Condition C IR model (#docs/#terms) Translation tools Mean average precision</title>
        <p>Parameters of each single run according to each language
English French Finnish (4-gram) Russian (word)
42 queries 49 queries 45 queries 34 queries
Okapi (3/15)
Prosit (3/10)</p>
        <p>Z-score
0.5580
Okapi (3/15)
Prosit (3/10)</p>
        <p>
          Z-score
0.5580
Prosit (3/10)
Finally, we suggest merging the retrieved documents according to the Z-score, taken from their document
scores
          <xref ref-type="bibr" rid="ref10">(Savoy 2003)</xref>
          . Within this scheme, we need to compute, for the ith result list, the average of the RSVk
(denoted MeanRSVi) and the standard deviation (denoted StdevRSVi). Based on these values, we can normalize
the retrieval status value of each document Dk provided by the ith result list, by computing the following
formula:
        </p>
        <p>Z-Score RSVk = ai . [((RSVk-MeanRSVi) / StdevRSVi ) + d i] with di = ((MeanRSVi- MinRSVi)/StdevRSVi) (3)
within which the value of di is used to generate only positive values, and ai (usually fixed at 1) is used to
reflect the retrieval performance of the underlying retrieval model and to account for the fact that pertinent items
are not uniformly distributed across all collections.</p>
        <p>
          Table 6 depicts the exact parameters used to search in the four different collections. For the Russian
collection, we only considered the word-based indexing strategy while for the Finnish language we only used
the 4-gram indexing scheme. In the top part of Table 6, it can be seen that we used a combined query
translation strategy for French, Finnish and Russian languages. As described in our monolingual experiments
          <xref ref-type="bibr" rid="ref1 ref11 ref12 ref13">(Savoy 2004c)</xref>
          , we might also apply a data fusion phase before merging the result lists. Thus when searching
into the English or French corpus, we combined the Okapi and Prosit result lists (both with blind query
expansion). In a second multilingual experiment (denoted Condition B), we have applied a data fusion approach
for all bilingual searches (descriptions given in the middle part of Table 6). Finally, we decided to search
through all corpora using the same retrieval model, Prosit in this case, as shown in the bottom part of Table 6
(and corresponding to Condition C).
        </p>
        <p>Table 7 depicts the retrieval effectiveness of various merging strategies using three different bilingual search
parameter settings. In this table, the round-robin scheme will be used as a baseline. On the one hand, when
different search engines are merged (Condition A and Condition B), the raw-score merging strategy results in
very poor mean average precision. On the other hand, when the same search engine is used (Condition C), the
resulting performance is better, but this is not the best one we should be able to achieve. The normalized score
merging based on Equation 1 shows degradation over the simple round-robin approach when using parameter
setting Condition B (0.1042 vs. 0.2340, or -4.9% in relative performance). Applying our logistic model using
both the rank and the document score as explanatory variables, the resulting mean average precision is clearly
better than the round-robin merging strategy and than other merging approaches (under Condition A or C).
Under Condition B, the difference between our logistic model and the Z-score merging strategy is rather small
(0.3111 vs. 0.3019, or 3.1% in relative performance).</p>
        <p>As a simple alternative, we also suggest a biased round-robin approach which extracts not one document per
collection per round but one document for the Russian corpus and two from the English, French and Finnish
collection (because the last three represent larger corpora). This merging strategy results in good retrieval
performance, better that the simple round-robin approach. Finally, the Z-score merging approach seems to
provide generally satisfactory performance. Moreover, we may multiply the Z-score by an a value (performance
under the label “ai = 1.5” with the ai values set as follows: EN: 1.5, FR: 1.5, FI: 1.0, and RU: 1.0).</p>
      </sec>
      <sec id="sec-1-9">
        <title>Parameters setting Merging Strategy Round-robin (baseline) Raw-score</title>
        <p>Norm Max
Norm RSV (Eq. 1)
Logistic reg. (ln(rank), RSV)
Biased round-robin
Z-score (Eq. 3)
Z-score (Eq. 3) ai = 1.5
Logistic reg. &amp; Selection (0)
Logistic reg. &amp; Selection (3)
Logistic reg. &amp; Selection (10)
Logistic reg. &amp; Selection (20)
Logistic reg. &amp; Selection (50)
Logistic reg. &amp; OptimalSelect</p>
      </sec>
      <sec id="sec-1-10">
        <title>Mean average precision (% change)</title>
        <p>Condition A Condition B Condition C
50 queries 50 queries 50 queries</p>
        <p>0.2386 0.2430 0.2358
0.0642 (-73.1%) 0.0650 (-73.2%) 0.3067 (+30.1%)
0.2552 (+7.0%) 0.1044 (-57.0%) 0.2484 (+5.3%)
0.2899 (+21.5%) 0.1042 (-57.1%) 0.2646 (+12.2%)
0.3090 (+29.5%) 0.3111 (+28.0%) 0.3393 (+43.9%)
0.2639 (+10.6%) 0.2683 (+10.4%) 0.2613 (+10.8%)
0.2677 (+12.2%) 0.2903 (+19.5%) 0.2555 (+8.4%)
0.2669 (+11.9%) 0.3019 (+24.2%) 0.2867 (+21.6%)
0.2957 (+23.9%) 0.2959 (+21.8%) 0.3405 (+44.4%)
0.2953 (+23.8%) 0.2982 (+22.7%) 0.3378 (+43.3%)
0.2990 (+25.3%) 0.3008 (+23.8%) 0.3381 (+43.4%)
0.3010 (+26.1%) 0.3029 (+24.7%) 0.3384 (+43.5%)
0.3044 (+27.6%) 0.3064 (+26.1%) 0.3388 (+43.7%)
0.3234 (+35.5%) 0.3261 (+34.2%) 0.3558 (+50.9%)</p>
        <p>It cannot be expected however that each result list would always contain pertinent items in response to a
given request. In fact, a given corpus may contain no relevant information regarding the submitted request or
the pertinent articles cannot be found by the search engine. In a cross-lingual environment we have found an
additional problem: important facets of the original request were translated with inappropriate words or
expressions. In all these cases, it is not useful to include items provided by such collections (or such search
engines) in the final result list. In addition, the number of pertinent documents is usually not uniformly
distributed across all four collections. For a given request (e.g., related to a regional or a national event), only
one or two collections may contain relevant documents describing this particular event.</p>
        <p>To take into account these phenomena, we have designed a selection procedure which works as follows.
First, for each result list we normalize the document score according to our logistic regression method (given in
Equation 2). After this step, each document score represents the probability that the underlying article is
relevant (with respect to the submitted query and the collection). In the second step, for each result list (or
language) we sum the document scores of the first 15 top-ranked documents. If this sum exceeds a given
threshold (depending on the collection or search engine), we can thus consider that the corresponding collection
contains many pertinent documents. Otherwise, we might only include the m best ranking retrieved items from
the corpus (with a relatively small m value). We may thus limit the number of items extracted from a given
corpus while also taking account of the fact that each collection usually contains few pertinent items. Table 7
lists the mean average precision achieved using this selection strategy under the label “Logistic reg. &amp; Selection
(m),” where the value m indicates that we always include the m best retrieved items from each corpus in our final
result list. Of course, when we set m = 0, the system will not extract any documents from a collection having a
poor overall score. Finally under the label “Logistic reg. &amp; OptimalSelect“, we have computed the mean
average precision that can be achieved when the selection is done without any error (with m = 0). When using
such an ideal selection system, the mean average precision is clearly better than all other merging strategies (e.g,
under Condition C, the MAP is 0.3558 vs. 0.3393 with the logistic regression without selection).</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Conclusion</title>
      <p>Query lang.</p>
      <p>Query type
English
English
English
English
English</p>
      <p>TD
TD
TD
TD
TD</p>
      <p>Type
automatic
automatic
automatic
automatic
automatic</p>
      <p>Merging
logistic
Z-scoreW
raw-score
logistic &amp; select</p>
      <p>Z-scoreW</p>
      <p>Parameters</p>
      <p>Condition A
Cond.!A, ai!=!1.5</p>
      <p>Condition C
Cond. A, m = 20</p>
      <p>Condition B</p>
      <p>MAP</p>
      <p>In multilingual tasks, searching documents written in different languages represents a real challenge. In this
case we propose a new simple selecting strategy which will avoid extracting a relatively large number of
documents from collections when these documents are of little interest with respect to the current request (see
Table 7). In this multilingual task, it is also interesting to mention that combining the result lists provided by
same search engine (Condition C in Table 7) may sometimes produce good retrieval effectiveness compared to
combining different search models (Condition A in Table 7).</p>
      <p>Acknowledgments</p>
      <p>The authors would like to thank the CLEF-2004 task organizers for their efforts in developing various
European languages test-collections. The authors would also like to thank C. Buckley from SabIR for giving us
the opportunity to use the SMART system, together with Samir Abdou for his help in translating the English
topics. This research was supported by the Swiss National Science Foundation under Grant #21-66 742.01.</p>
      <p>C201
C202
C203
C204
C205
C206
C207
C208
C209
C210
C211
C212
C213
C214
C215
C216
C217
C218
C219
C220
C221
C222
C223
C224
C225</p>
      <sec id="sec-2-1">
        <title>Domestic Fires</title>
        <p>Nick Leeson's Arrest
East Timor Guerrillas
Victims of Avalanches
Tamil Suicide Attacks
G7 Summit in Halifax
Fireworks Injuries
“Sophie's World”
Tour de France Winner
Nobel Peace Prize Candidates
Peru-Ecuador Border Conflict
Sportswomen and Doping
Papal Travels
Multi-billionaires
Re-election of Peru's President
Glue-sniffing Youngsters
AIDS in Africa
Andreotti and the Mafia
EU Commissioner Candidates
European Cars in Russia
2002 Olympic Winter Games
Presidential elections in France
Chernobyl Disaster outside ex-USSR
Woman solos Everest
Nuclear Power Plant of Sosnovyi Bor
C226
C227
C228
C229
C230
C231
C232
C233
C234
C235
C236
C237
C238
C239
C240
C241
C242
C243
C244
C245
C246
C247
C248
C249
C250</p>
      </sec>
      <sec id="sec-2-2">
        <title>Sex-change Operations Altai Ice Maiden Prehistorical Art Dam Building</title>
        <p>Atlantis-Mir Docking
New Portuguese Prime Minister
Pension Schemes in Europe
Greenhouse Effect
Deaf and Society
Seal-hunting
A typhoon in the Philippines
Panchen Lama
Lady Diana
Mental Health of the Young
Sioux Ghost Shirt
New political parties
Record Permanence in Space
Films of Kieslowski
Footballer of the Year 1994
Christopher Reeve
Castro visits UN
Alexander the Great's Tomb
Macedonia Name Dispute
Women's Ten Thousand Metres Champion
Rabies in Humans
Title of the queries of the CLEF-2004 test-collection</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Braschler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Cross-language evaluation forum: Objectives, results and achievements</article-title>
          .
          <source>IR Journal</source>
          ,
          <volume>7</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>7</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Buckley</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitra</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waltz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Cardie</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>Using clustering and superconcepts within SMART</article-title>
          .
          <source>In Proceedings of TREC-6</source>
          , (pp.
          <fpage>107</fpage>
          -
          <lpage>124</lpage>
          ).
          <source>Gaithersburg: NIST Special Publication</source>
          <volume>500</volume>
          -240.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Gey</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Combining query translation and document translation in cross-language retrieval</article-title>
          .
          <source>In Proceedings CLEF-2003</source>
          , (pp.
          <fpage>39</fpage>
          -
          <lpage>48</lpage>
          ). Trondheim.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Cross-language retrieval experiments at CLEF 2002</article-title>
          . In C. Peters,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          , &amp; M. Kluck, (Eds),
          <source>Advances in Cross-Language Information Retrieval</source>
          , (pp.
          <fpage>28</fpage>
          -
          <lpage>48</lpage>
          ), Springer-Verlag, Berlin, LNCS #
          <fpage>2785</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Kishida</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuriyama</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Eguchi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>Prediction of performance on cross-lingual information retrieval by regression models</article-title>
          .
          <source>In Proceedings NTCIR-4</source>
          , (pp.
          <fpage>219</fpage>
          -
          <lpage>224</lpage>
          ). Tokyo: NII.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Kwok</surname>
            ,
            <given-names>K.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grunfeld</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>D.D.</given-names>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>TREC-3 ad-hoc, routing retrieval and thresholding experiments using PIRCS</article-title>
          .
          <source>In Proceedings of TREC'3</source>
          , (pp.
          <fpage>247</fpage>
          -
          <lpage>255</lpage>
          ). Gaithersburg: NIST Publication #
          <fpage>500</fpage>
          -
          <lpage>225</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Le</given-names>
            <surname>Calvé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            &amp;
            <surname>Savoy</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>Database merging strategy based on logistic regression</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>36</volume>
          (
          <issue>3</issue>
          ),
          <fpage>341</fpage>
          -
          <lpage>359</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>MacNamee</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Mayfield</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>JHU/APL experiments in tokenization and non-word translation</article-title>
          .
          <source>In Proceedings CLEF-2003</source>
          , (pp.
          <fpage>19</fpage>
          -
          <lpage>28</lpage>
          ). Trondheim.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Nie</surname>
            ,
            <given-names>J. Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isabelle</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Durand</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>1999</year>
          ).
          <article-title>Cross-language information retrieval based on parallel texts and automatic mining of parallel texts from the Web</article-title>
          .
          <source>In Proceedings of the ACM-SIGIR'99</source>
          , (pp.
          <fpage>74</fpage>
          -
          <lpage>81</lpage>
          ). New York: The ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Savoy J.</surname>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Report on CLEF-2003 multilingual tracks</article-title>
          .
          <source>In Proceedings CLEF-2003</source>
          , (pp.
          <fpage>7</fpage>
          -
          <lpage>12</lpage>
          ). Trondheim.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2004a</year>
          ).
          <article-title>Combining multiple strategies for effective monolingual and cross-lingual retrieval</article-title>
          .
          <source>IR Journal</source>
          ,
          <volume>7</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>121</fpage>
          -
          <lpage>148</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2004b</year>
          ).
          <article-title>Report on CLIR task for the NTCIR-4 evaluation campaign</article-title>
          .
          <source>In Proceedings NTCIR-4</source>
          , (pp.
          <fpage>178</fpage>
          -
          <lpage>185</lpage>
          ). Tokyo: NII.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2004c</year>
          ).
          <article-title>Report on CLEF-2004 monolingual tracks</article-title>
          .
          <source>In Proceedings CLEF-2004 (this volume)</source>
          .
          <source>Bath.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>N.K.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Johnson-Laird</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>The collection fusion problem</article-title>
          .
          <source>In Proceedings of TREC'3</source>
          , (pp.
          <fpage>95</fpage>
          -
          <lpage>104</lpage>
          ). Gaithersburg: NIST Publication #
          <fpage>500</fpage>
          -
          <lpage>225</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>