<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Introduction
September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <pub-date>
        <year>2000</year>
      </pub-date>
      <volume>5</volume>
      <issue>2000</issue>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>and Der Spiegel). For indexing the collection, we used a stopwordlist that contained also capitalized versions
retrieval.
The original query topics in German were searched against the German collection (Frankfurter Rundschau
this collection, we used a stopwordlist, the French-to-lower normalizer and the French stemmer (from association
query depending on familiarity with the context and meaningfulness of the returned documents and top ranked
in German were extended with additional query terms obtained by searching the German CLEF collection
for these original queries (with the help of Aitao Chen’s Cross-language Text Retrieval System Web-interface).
were searched against the Italian collection (La Stampa). For indexing this collection, we used a stopwordlist,
(Frankfurter Rundschau and Der Spiegel) with the original German query topics and looking at the results
BKMOFFA2 (Berkeley Monolingual French against French Automatic Run 2)
words and the German stemmer (from association dictionary) build by Aitao Chen. We didn’t use a normalizer
for this collection.
4. BKMOGGM1 (Berkeley Monolingual German against German Manual Run 1) The original query topics
of words and the German stemmer (from association dictionary) described in Section 3.4. We did not use a
BKMOIIA3 (Berkeley Monolingual Italian against Italian Automatic Run 3) The original query topics in Italian
the Italian-to-lower normalizer and the Italian stemmer (from association dictionary) described in Section 4.
document terms. For indexing the collection, we used a stopwordlist that contained also capitalized versions of
dictionary) described in Section 4.
normalizer for this collection because all nouns in German are capitalized and hence this clue might be used in
The additional query terms were obtained by either directly looking at the documents or looking at the top
The original query topics in French were searched against the French collection (Le Monde). For indexing
ranked document terms for the original query text. The searcher spent about 10 to 25 minutes per topic or
BKMOGGA1 (Berkeley Monolingual German against German Automatic Run 1)</p>
    </sec>
    <sec id="sec-2">
      <title>Manual</title>
      <p>Automatic
Automatic
Automatic
evaluation we chose to create a stemmer mechanistically from common leading substring analysis of the entire
Several interesting questions have arisen in recent research on CLIR. First, is CLIR merely a matter of a marriage
corpus. The impact of the stemmer on performance will be discussed at the end of the ocial results discussion.
CLEF work we made use of two widely available machine translation packages, the SYSTRAN system found at
of TREC have convinced us that some form of stemming will always improve performance. For this particular
in French, German, and Italian. The two translated les for eac h languages were pooled together and then put
from multiple sources, having found that dieren t packages made dieren t mistakes on particular topics. Second,
The original query topics in English were translated once with the Systran system
1. BKMUEAA1 (Berkeley Multilingual English against all Automatic Run 1)
Third, is performance improved by creating a multilingual index by pooling all documents together in one index
what is the role of language specic stemming in impro ved performance? Our experience with the Spanish tracks
fusion which combines the individual rankings into a unied ranking independen t of language? This was one of
the major focuses of our experiments at CLEF.
(http://babel.altavista.com/translate.dyn) and with L&amp;H Powertranslator. The English topics were translated
made comparisions to Power Translator. For CLEF multilingual we combined translations and dictionary lookup
the AltaVista site, and the Lernout and Hauspie Power Translator Pro Version 7.0. For the GIRT retrieval we
of convenience between machine translation combined with ordinary (monolingual) information retrieval? In our
or by creating separate language indexes and doing monolingual retrieval for each language followed by data
together in one query le (the English original query topics w ere multiplied by 2 to gain the same frequency of
3
version, one in French with the Systran and Powertranslator version, and one in Italian accordingly). The original
German topics were translated into English, French, and Italian. The 2 translated versions for each language
collection, and the pooled Italian topics le was searched against the Italian collection. The frequency of the
description of the collection see BKMUEAA1.
4. BKMUGAM1 (Berkeley Multilingual German against all Manual Run 1)
result les with the 1000 top rank ed records for each topic. These 4 result les w ere then pooled together and
search terms was divided by 2 to avoid over-emphasis of equally translated search terms. This resulted in 4
le w as searched against the German collection, the pooled French topics le w as searched against the French
3. BKMUGAA2 (Berkeley Multilingual German against all Automatic Run 2)
of the collection see BKMUEAA1.
The original query topics in English were translated once with Systran and with L&amp;H PowerTranslator. The
Powertranslator into English, French and Italian. These translations were pooled together with the German
pooled together in one query le (resulting in 3 topics les, one in German with the Systran and P owertranslator
originals in one le. This topics le w as searched against the whole collection including all 4 languages. For a
The original query topics in German were translated once with Systran and with Powertranslator. The
2. BKMUEAA2 (Berkeley Multilingual English against all Automatic Run 2)
English topics le w as searched against the English collection (Los Angeles Times). The pooled German topics
sorted by weight (rank) for each record and topic. The pooling method is described below. For a description of
the collections see BKMOGGM1, BKMOFFA2, BKMOIIA3, BKMUEAA1.
were pooled together in one query le. The original German topics le was multiplied by 2 to gain the same
frequency of query terms in the query le searc hed. The nal topics le con tained 2 German (original), English,
English topics were translated into French, German, and Italian. The 2 translated versions for each language were
French, and Italian versions (one Powertranslator and one Systran) for each topic. During the search, we divided
the frequency of the search terms by 2 to avoid over-emphasis of equally translated search terms. For a description
The manually extended German query topics (see description from BKMOGGM1) were now translated with
4
Table 3: Results of four ocial CLEF m ultilingual runs.
its English equivalent became ’de-nationalization’ a very uncommon synonym for ’privatization,’ and one which
stands out. Query 40 about the privatization of the German national railway was one which seems to have given
of our machine translation softwares. The German version of the topic was not much better { in translation
we were particularly vexed by the use of the English spelling ’privatisation’ which couldn’t be recognized by either
everyone problems (the median precision over all CLEF runs was 0.0537 for this query). As an American group,
A query-by-query analysis can be done to identify problems. We have not had time to do this, but one query
the base form, nouns to the singular form, and adjectives to the positive form. All the French words which have
the same English translations after normalization were grouped together to form a class. A member from each
The German stemmer and Italian stemmer were created alike.
are shown in column 3 of table 4. Column 4 shows the overall precision with the French, German, and Italian
We submitted four monolingual runs and four multilingual runs. These eight runs were repeated without
French collection into English using SYSTRAN. The English translations were normalized by reducing verbs to
A stemmer for the French collection was created by rst translating all the distinct F rench words found in the
class representative in indexing.
class is selected to represent the whole class in indexing. All the words in the same class were replaced by the
the French, German, and Italian stemmers. The overall precision for each of the eight runs without stemming
stemmers. Column 5 shows the improvement in precision which can be attributed to the stemmers.
have been assigned individual classication iden tiers b y human indexers. These classication iden tiers come
TREC-8 [5]
are expended on developing these classication on tologies and applying them to index documents, it seems only
tion is managed and indexed the GESIS organization (http://www.social-science-gesis.de). GIRT is an excellent
has been done in Europe with the GEMET (General European Multilingual Environmental Thesaurus) eort
Russian. We worked extensively with a previous version of the GIRT collection in our cross-language work for
thesauri can be found in [8].
The GIRT collection consists of reports and papers (grey literature) in the social science domain. The
collecnatural to attempt to exploit the resources previously expended to the fullest extent possible to improve retrieval.
example of a collection indexed by a multilingual thesaurus, originally German-English, recently translated into
A special emphasis of our current funding has focussed upon retrieval of specialized domain documents which
In some cases such thesauri are developed with identiers translated (or pro vided) in multiple languages. This
and with the OECD General Thesaurus (available in English, French, and Spanish). A review of multilingual
from what we call "domain ontologies", of which thesauri are a particular case. Since many millions of dollars
percent of English query words untranslated. After examining the untranslated English query words carefully, we
from the GIRT thesaurus and used it to translate the English queries to German. This approach left about 50
the English query. One problem with thesaurus lookup is how to match the phrasal items in a thesaurus. We
by looking up the thesaurus:
have taken a simple approach to deal with this problem: use POS tagger to identify noun phrases.
contains English items and their corresponding German translations. This "vocabulary discovery" approach
b. Use the part-of-speech tagger LT-POS developed by University of Edinburgh
has a corresponding English translation. We took the following steps to translate the English query to German
likely to occur in a domain-specic thesaurus lik e GIRT. Examples are "country", "car", "foreign", "industry",
a. Create an English-German transfer dictionary from the Social Science Thesaurus. This transfer dictionary
For last year’s GIRT task at the TREC-8 evaluation, we extracted an English-German transfer dictionary
(http://www.ltg.ed.ac.uk/software/pos/index.html) to tag the English query and identify noun phrases in
The GIRT social science Thesaurus is a German-English bilingual thesaurus. Each German item in this thesaurus
was taken by Eichmann, Ruiz and Srinivasan for medical information cross-language retrieval using the UMLS
found that most of them fell into the following two categories: one category contains general terms that are not
Metathesaurus[9].
6</p>
    </sec>
    <sec id="sec-3">
      <title>There are 76128 German documents in GIRT subtask collection. Of them, about 54275 (72 percent) have English</title>
      <p>TREC-8.
In our experiments, we indexed only the TITLE and TEXT sections in each document (not the E-TITLE or
our CLEF runs this year we added a German stemmer similar to the Porter stemmer for the German language.
TITLE sections. 5317 documents (7 percent) have also English TEXT sections. Almost all the documents contain
manually assigned thesaurus terms. On average, there are about 10 thesaurus terms assigned to each document.
Using this stemmer led to a 15 percent increase in average precision when tested using the GIRT-1 collection of
E-TEXT). The CLEF rules specied that indexing an y other eld w ould need to be declared a manual run. For</p>
    </sec>
    <sec id="sec-4">
      <title>Query translation is almost always selected for practical reasons of eciency , and because translation errors in</title>
      <p>documents can propagate without discovery since the maintainers of a text archive rarely read every document.
In CLIR, essentially either queries or documents or both need to be translated from one language to another.
collection.
We applied the following three methods to translate the English queries to German: Thesaurus lookup, Entry
Vocabulary Module (EVM), machine translation (MT). The resulted German queries were run against the GIRT
For the CLEF GIRT task, our focus has been to compare the performance of dieren t translation strategies.
entry vocabulary method to map from query term to thesaurus term, the top ranked thesaurus term and its
Our GIRT results are summarized in Table 5. The runs can be described as follows: BKGREGA4 used our
translation was used to create the German query. BKGREGA3 used the results of machine translation by the
7</p>
    </sec>
    <sec id="sec-5">
      <title>English queries. More details about this work can be found in [5].</title>
      <p>in English Title and text sections to German thesaurus terms. This mapping can then be used to translate the
In the GIRT collection, about 72 percent of the documents have both German titles and English titles. 7 percent
have also English text sections. This feature allows us to build a EVM which maps the English words appearing
for example,
Fuzzy matching also found related terms for some query terms which do not appear in the thesaurus at all,
from dieren t runs are commensurable. So, for each document retrieved, we used the sum of its probability from
translation methods may lead to better performance than of any one of the methods. Since we use the same
GIRT. When analyzing the experimental training results, we noticed that dieren t translation methods retrieved
the dieren t runs as its nal probabilit y to create the ranking for the merged results.
While our CLEF Multilingual strategy focussed on merging monolingual results run independently on dieren t
subcollections, one per language, all our GIRT runs were done on a single subcollection, the German text part of
retrieval algorithm and data collection for all the runs, the probability that a document is relevant to a query
sets of documents that contain dieren t relevant documents. This implies that merging the results from dieren t
at 0.50
at 0.70
at 0.60
at 1.00
Run ID
at 0.00
at 0.30
Med. Prec.
at 0.40
at 0.20
at 0.80
Retrieved
Relevant
Brk. Prec.
at 0.10
Rel. Ret.
at 0.90</p>
    </sec>
    <sec id="sec-6">
      <title>References</title>
      <p>8
5 Summary and Acknowlegments
Table 5: Results of four ocial GIR T English-German runs.
0.1611
827
0.2004
0.2465
0.2299
0.1252
0.3583
0.3292
0.1477
0.1938
0.6139
0.4482
BKGREGA4
0.0003
0.0612
1193
23000
[3] A Chen J He L Xu F Gey and J Meggs. Chinese text retrieval without using a dictionary. In A.
DeACM SIGIR Conference on Research and Development in Information Retrieval, Philadelphia, pages 42{49,
1997.
sai Narasimhalu Nicholas J. Belkin and Peter Willett, editors, Proceedings of the 20th Annual International</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>