<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Merging Mechanisms in Multilingual Information Retrieval</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Wen-Cheng Lin and Hsin-Hsi Chen Department of Computer Science and Information Engineering National Taiwan University Taipei</institution>
          ,
          <country country="TW">TAIWAN, R.O.C</country>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>7</lpage>
      <abstract>
        <p>National Taiwan University (NTU) Natural Language Processing Laboratory (NLPL) participated in MLIR task in CLEF 2002. We submitted five official multilingual runs. In this paper, we try to resolve the collection fusion problem. We experimented with several merging strategies that merge the results of several intermediate runs. Multilingual Information Retrieval [4] uses a query in one language to retrieve documents in different languages. A multilingual data collection is a set of documents that are written in different languages. There are two types of multilingual data collection. The first one contains several monolingual document collections. The second one consists of multilingual-documents. A multilingual-document is written in more than two languages. Some multilingual-documents have a major language, i.e. most part of the document is written in the same language. For example, a document is written in Chinese, but the abstract is in English. Therefore this document is a multilingual-document and Chinese is its major language. The significances of different languages in a multilingual-document may be different. For example, the English translation of a Chinese proper noun is a useful clue when using English queries to retrieve Chinese documents. Therefore, the English translation should have higher weight. Figure 1 shows these two types of multilingual data collections. In Multilingual Information Retrieval, queries and documents are in different languages. We can translate queries, or translate documents, or translate both queries and documents into an intermediate language to unify the languages of queries and documents. Figure 2 shows some MLIR architectures when query translation is adopted. The front-end controller processes queries, translates queries, submits translated queries to the monolingual IR systems, collects the relevant document lists reported by IR systems and merges them. Figure 3 shows document translation architectures.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>…
L1
L2
. . .</p>
      <p>L2</p>
      <p>L1
L2
L1
L2
…
…
L1
L2
L1-1</p>
      <p>L1-m</p>
      <p>L2-1</p>
      <p>L2-n
…</p>
      <p>Lt-1
…</p>
      <p>Lt-k
(a) A set of monolingual document collections
. . .
(b) Some types of multilingual-documents
Query</p>
      <sec id="sec-1-1">
        <title>Result</title>
      </sec>
      <sec id="sec-1-2">
        <title>Query</title>
      </sec>
      <sec id="sec-1-3">
        <title>Result</title>
      </sec>
      <sec id="sec-1-4">
        <title>MLIR front-end controller</title>
      </sec>
      <sec id="sec-1-5">
        <title>MLIR front-end controller</title>
      </sec>
      <sec id="sec-1-6">
        <title>Query</title>
      </sec>
      <sec id="sec-1-7">
        <title>Result</title>
      </sec>
      <sec id="sec-1-8">
        <title>MLIR front-end controller</title>
        <p>Fig 2. MLIR Architectures when Query Translation is Adopted</p>
      </sec>
      <sec id="sec-1-9">
        <title>Query</title>
      </sec>
      <sec id="sec-1-10">
        <title>Result</title>
      </sec>
      <sec id="sec-1-11">
        <title>Query</title>
      </sec>
      <sec id="sec-1-12">
        <title>Result</title>
        <p>MLIR front-end
controller</p>
      </sec>
      <sec id="sec-1-13">
        <title>Mono IR</title>
      </sec>
      <sec id="sec-1-14">
        <title>Mono IR</title>
      </sec>
      <sec id="sec-1-15">
        <title>Mono IR</title>
      </sec>
      <sec id="sec-1-16">
        <title>Mono IR D L1 D L2’</title>
        <p>D L3’</p>
        <p>D L1
D L2’
D L3’</p>
        <p>DT
L21
DT
L31</p>
        <p>DT
L21
DT
L31
Fig 3. Document Translation Architectures</p>
      </sec>
      <sec id="sec-1-17">
        <title>Mono IR</title>
      </sec>
      <sec id="sec-1-18">
        <title>Mono IR</title>
      </sec>
      <sec id="sec-1-19">
        <title>Mono IR</title>
      </sec>
      <sec id="sec-1-20">
        <title>Mono IR 1</title>
      </sec>
      <sec id="sec-1-21">
        <title>Mono IR 2</title>
      </sec>
      <sec id="sec-1-22">
        <title>Mono IR 3 D_L1 D_L2</title>
        <p>D_L3
D_L2
D_L3
D_L1
D_L2
D_L3</p>
        <p>D L2</p>
        <p>D L3
D L2
D L3</p>
        <p>In addition to language barrier issue, how to conduct a ranked list that contains documents in different
languages from several text collections is also critical. There are two architectures in MLIR: centralized and
distributed. The first two architectures in Figure 2 and first architecture in Figure 3 are distributed architectures;
the remaining architectures in Figures 2 and 3 are centralized architectures. In a centralized architecture, a huge
collection that contains documents in different languages is used. In a distributed architecture, documents in
different languages are indexed and retrieved separately. The results of each run are merged into a multilingual
ranked list. Several merging strategies have been proposed. Raw-score merging selects documents based on
their original similarity scores. Normalized-score merging normalizes the similarity score of each document and
sorts all documents by the normalized score. For each topic, the similarity score of each document is divided by
the maximum score in this topic. Round-robin merging interleaves the results in the intermediate runs. In this
experiment, we adopted distributed architecture and proposed a merging strategy to merge the result lists.</p>
        <p>The rest of this paper is organized as follows. Section 2 describes the indexing method. Section 3 shows the
query translation process. Section 4 describes our merging strategies. Section 5 shows the experiment results.
Section 6 concludes the remarks.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Indexing</title>
      <p>The document set used in CLEF2002 MLIR task consists of English, French, German, Spanish and Italian. The
numbers of documents in English, French, German, Spanish and Italian document sets are 113,005, 87,191,
225,371, 215,738 and 108,578 respectively.</p>
      <p>The IR model we used is the basic vector space model. Documents and queries are represented as term
vectors, and cosine vector similarity formula is used to measure the similarity of a query and a document. The
term weighting function is tf*idf. Appropriate terms are extracted from each document in indexing stage. In the
experiment, the &lt;HEADLINE&gt; and &lt;TEXT&gt; sections in English documents were used for indexing. For
Spanish documents, the &lt;TITLE&gt; and &lt;TEXT&gt; sections were used. When indexing French, German and Italian
documents, the &lt;TITLE&gt;, &lt;TEXT&gt;, &lt;TI&gt;, &lt;LD&gt; and &lt;TX&gt; sections were used. The words in these sections
were stemmed, and stopwords were removed. The stopword lists and stemmers were developed by University of
Neuchatel (The stopword lists and stemmers are available at http://www.unine.ch/info/clef/) [5]</p>
    </sec>
    <sec id="sec-3">
      <title>Query translation</title>
      <p>In the experiment, the English queries were used as source queries and translated into target languages, i.e.
French, German, Spanish and Italian. In our previous experiments [1, 3], we used CO model to translate queries.
CO model uses word co-occurrence information trained from a target language text collection to disambiguate
the translations of query terms. In this experiment, we didn’t have enough time to train the word co-occurrence
information for the languages used in CLEF 2002 MLIR task. Thus, we used a simple method to translate the
queries. We adopted a dictionary-based approach to translate the queries. For each English query term, we
found its translation equivalents by looking up a dictionary and selected the first two translation equivalents to
be the target language query terms. The dictionaries we used are Ergane English-French, English-German,
English-Spanish and English-Italian dictionaries (The Ergane dictionaries are available at
http://www.travlang.com/Ergane). There are 8,839, 9,046, 16,936 and 4,595 terms in Ergane English-French,
English-German, English-Spanish and English-Italian dictionaries, respectively.</p>
    </sec>
    <sec id="sec-4">
      <title>Merging strategies</title>
      <p>There are two architectures in MLIR, i.e., centralized and distributed. In a centralized architecture, document
collections in different languages are viewed as a single document collection and are indexed in one huge index
file. The advantage of a centralized architecture is that it avoids the merging problem. It needs only one
retrieving phase to produce a result list that contains documents in different languages. One of problems of
centralized architecture is that the weights of index terms are over weighting. The total number of documents
increases but the number of occurrences of a term does not. Thus, the idf of a term is increased and the weight is
over-weighting. This phenomenon is acuter in small text collection. For example, the N in idf formula is 87,191
when French document is used. However, this value is increased to 749,883, i.e. about 8.6 times larger, if the
three document collections are merged together. Comparatively, the weights of German index terms are
increased 3.33 times due to the size of N. The increments of weights are unbalance for document collections in
different size. Thus, it makes retrieval result preferring documents in small document collection.</p>
      <p>The second architecture is a distributed MLIR. Documents in different languages are indexed and retrieved
separately. The ranked lists of all monolingual and cross-lingual runs are merged into one multilingual ranked
list. How to merge result lists is a problem. Recent works have proposed various approaches to deal with
merging problem. A simple merging method is raw-score merging that sorts all results by their original
similarity scores and then selects the top ranked documents. Raw-score merging is based on the assumption that
the similarity scores across collections are comparable. However, the collection-dependent statistics in
document or query weights invalidates this assumption [2, 6]. Another approach, round-robin merging,
interleaves the results based on the rank. This approach assumes that each collection has approximately the
same number of relevant documents and the distribution of relevant documents is similar across the result lists.
Actually, different collections do not contain equal numbers of relevant documents. Thus the performance of
round-robin merging may be poor. The third approach is normalized-score merging. For each topic, the
similarity score of each document is divided by the maximum score in this topic. After adjusting scores, all
results are put into a pool and sorted by the normalized score. This approach maps the similarity scores of
different result lists into the same range, from 0 to 1, and makes the scores more comparable. But it has a
problem. If the maximum score is much higher than the second one in a result list, the normalized-score of
document at rank 2 would be low even if its original score is high. Thus, the final rank of this document would
be lower than that of the top ranked documents with very low but similar original scores in another result list.</p>
      <p>Similarity score reflects the degree of similarity between a document and a query. A document with higher
similarity score seems more relevant to the desired query. But, if the query is not formulated well, e.g.,
unappropriate translation of a query, a document with high score still does not meet the user’s information need.
When merging results, such documents that have high scores should not be included in the final result list. Thus,
we have to consider the effectiveness of each individual run in the merging stage. The basic idea of our merging
strategy is that adjusting the similarity scores of documents in each result list to make them more comparable
and to reflect their confidence. The similarity scores are adjusted by the following formula.</p>
      <p>1</p>
      <p>
        Sˆij = Sij × Sk ×Wi (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
where Sij is the original similarity score of the document at rank j in the ranked list of topic i,
ˆij is the adjusted similarity score of the document at rank j in the ranked list of topic i,
S
Sk is the average similarity score of top k documents, and
      </p>
      <p>Wi is the weight of query i in a cross lingual run.</p>
      <p>We divide the weight adjusting process into two steps. First, we use a modified score normalization method
to normalize the similarity scores. The original score of each document is divided by the average score of top k
documents instead of the maximum score. We call this normalized-by-top-k. Second, the normalized score
multiplies a weight that reflects the retrieval effectiveness of the desired topic in each text collection. However,
due to not knowing the retrieval performance in advance, we have to guess the performance of each run. For
each language pair, the queries are translated into target language and then the target language documents are
retrieved. A good translation should have better performance. We can predict the retrieval performance based
on the translation performance. There are two factors affecting the translation performance, i.e., the degree of
translation ambiguity and the number of unknown words. For each query, we compute the average number of
translation equivalents of query terms and the number of unknown words in each language pair, and use them to
compute the weights of each cross lingual run. The weight can be determined by the following formula:
  1    </p>
      <p>
        Wi = c1 + c2 ×  Ti   + c3 × 1 − Unii  (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
where Wi is the weight of query i in a cross lingual run,
      </p>
      <p>Ti is the average number of translation equivalents of query terms in query i,
Ui is the number of unknown words in query i,
ni is the number of query terms in query i, and
c1, c2 and c3 are tunable parameters, and c1+c2+c3=1.</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>We submitted five multilingual runs. All runs use title and description fields. The five multilingual runs use
English topics as source queries. The English topics were translated into French, German, Spanish and Italian.
The source English topics and translated French, German, Spanish and Italian topics were used to retrieve the
corresponding document collections. Then, we merged these five result lists. We used different merging
strategies for the five multilingual runs:
1. NTUmulti01</p>
      <p>The result lists were merged by normalized-score merging strategy. The maximum similarity score was
used for normalization. After normalization, all results were put in a pool and were sorted by the adjusted
score. The top1000 documents were selected as the final results.
2. NTUmulti02</p>
      <p>
        In this run, we used the modified normalized-score merging method. The average similarity score of top
100 documents were used for normalization. We did not consider the performance drop down caused by
query translation. That is, the weight Wi in formula (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) was 1 for every sub run.
3. NTUmulti03
      </p>
      <p>
        First, the similarity scores of each document were normalized. The maximum similarity score was used for
normalization. Then we assigned a weight Wi to each intermediate run. The weight was determined by
formula (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ). The values of c1, c2 and c3 were 0, 0.4 and 0.6, respectively.
4. NTUmulti04
      </p>
      <p>
        We used formula (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) to adjust the similarity score of each document. We used the average similarity score
of top 100 documents for normalization. The weight Wi was determined by formula (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ). The values of c1,
c2 and c3 were 0, 0.4 and 0.6, respectively.
5. NTUmulti05
      </p>
      <p>In this run, the merging strategy is similar to run NTUmulti04. The difference was that each intermediate
run was assigned a constant weight. The weights assigned to English-English, English-French,
EnglishGerman, English-Italian and English-Spanish intermediate runs were 1, 0.7, 0.4, 0.6 and 0.6 respectively.</p>
      <p>The results of our official runs are shown in Table 1. The performance of normalized-score merging is bad.
The average precision of run NTUmulti01 is 0.0173. When using our modified normalized-score merging
strategy, the performance is better. The average precision is increased to 0.0266. Run NTUmulti03 and
NTUmulti04 have considered the performance drop down caused by query translation. Table 2 shows the
unofficial evaluation of intermediate monolingual and cross-lingual runs. The performance of English
monolingual run is much better than cross-lingual runs. Therefore, the cross-lingual runs should have lower
weights when merging results. The results show that the performances are improved by decreasing the
importance of un-effective cross-lingual runs. The average precisions of runs NTUmulti03 and NTUmulti04 are
0.0336 and 0.0373, which are better than run NTUmulti01 and NTUmulti02. Run NTUmulti05 assigned constant
weights to each intermediate runs. Its performance is slightly worse than run NTUmulti04. All our official runs
didn’t perform well. This is because that the performances of cross-lingual runs are very bad. In the
experiments, we did not disambiguate the senses of query terms when translating queries and the numbers of
words contained in the bilingual dictionaries we used are too few, the queries were not translated well, thus the
performances of cross-lingual runs are very poor. If we use larger dictionaries and disambiguate the word senses,
the performance should be better.</p>
      <p>In order to compare the effectiveness of different merging strategies, we also conducted several unofficial
runs:
1. ntu-multi-raw-score</p>
      <p>We used raw-score merging to merge result lists.
2. ntu-multi-round-robin</p>
      <p>We used round-robin merging to merge result lists.
3. ntu-multi-centralized</p>
      <p>This run adopted centralized architecture. All document collections were indexed in one index file. The
topics contained source English query terms, and other translated query terms.
# Topic</p>
      <sec id="sec-5-1">
        <title>Average Precision 0.0381 0.0224 0.0398</title>
        <p>The results of unofficial runs are shown in Table 3. The performance of raw-score merging is good. This is
probably because we use the same IR model and term weighting scheme for all text collections. When using
round-robin merging strategy, the performance is bad. The best run is ntu-multi-centralized which indexes all
documents in different languages together. In this run, most top ranked documents are in English, French and
Italian in most topics. The performance of English monolingual retrieval is much better than other runs. The
average precisions of English-German and English-Spanish cross-lingual runs are quite low. Therefore, the
result list should not contain too many German and Spanish documents. The over-weighting phenomenon in
centralized architecture makes the scores of French, Italian and English documents increase. Thus, more French,
Italian and English documents are included in the result list of run ntu-multi-centralized. This makes the
performance better. If we use the German or Spanish queries as source queries, the performance of centralized
architecture may be not so good.
6.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Concluding Remarks</title>
      <p>In the experiment, we proposed some merging strategies to integrate the result lists of collections in different
languages. The results showed that the performance of our merging strategies was similar to that of raw-score
merging and was better than normalized-score and round-robin merging. The performance of run NTUmulti05
was good. The weights of English-German and English-Spanish runs were not low enough. If we decrease the
weights of these two runs, the performance should be better. The results showed that we could gain better
performance by adjusting the similarity score appropriately. The similar results also appeared in NTCIR
multilingual IR task. How to determine appropriate weights is an important issue. Considering the degree of
ambiguity, i.e. lowering the weights of more ambiguous query terms, improves some performance. The
centralized approach performed well. We will do more experiments to find out if the centralized approach still
works when using German or Spanish queries as source queries.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bian</surname>
            ,
            <given-names>G.W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>W.C.</given-names>
          </string-name>
          ,
          <year>1999</year>
          .
          <article-title>Resolving translation ambiguity and target polysemy in cross-language information retrieval</article-title>
          .
          <source>In Proceedings of 37th Annual Meeting of the Association for Computational Linguistics</source>
          , Maryland,
          <year>June 1999</year>
          . Association for Computational Linguistics,
          <fpage>215</fpage>
          -
          <lpage>222</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <year>1992</year>
          .
          <article-title>LSI meets TREC: A Status Report</article-title>
          .
          <source>In Proceedings of the First Text REtrieval Conference (TREC1)</source>
          , Gaithersburg, Maryland, November,
          <year>1992</year>
          . NIST Publication,
          <volume>137</volume>
          -
          <fpage>152</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>W.C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H.H.</given-names>
          </string-name>
          ,
          <year>2002</year>
          .
          <article-title>NTU at NTCIR3 MLIR Task</article-title>
          .
          <source>In Working Notes for NTCIR3 workshop</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Oard</surname>
            ,
            <given-names>D.W.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Dorr</surname>
            ,
            <given-names>B.J.</given-names>
          </string-name>
          ,
          <year>1996</year>
          .
          <article-title>A Survey of Multilingual Text Retrieval</article-title>
          .
          <source>Technical Report UMIACS-TR-96-19</source>
          , University of Maryland, Institute for Advanced Computer Studies.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <year>2001</year>
          . Report on CLEF-2001 Experiments:
          <article-title>Effective Combined Query-Translation Approach</article-title>
          .
          <source>In Evaluation of Cross-Language Information Retrieval Systems, Lecture Notes in Computer Science</source>
          , Vol.
          <volume>2406</volume>
          , Darmstadt, Germany, September,
          <year>2001</year>
          . Springer,
          <fpage>27</fpage>
          -
          <lpage>43</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>N.K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Johnson-Laird</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>1995</year>
          .
          <article-title>The Collection Fusion Problem</article-title>
          .
          <source>In proceedings of the Third Text REtrieval Conference (TREC-3)</source>
          , Gaithersburg, Maryland, November,
          <year>1994</year>
          . NIST Publication,
          <volume>95</volume>
          -
          <fpage>104</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>