<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Language-Independent Approach to European Text Retrieval</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Paul McNamee and James Mayfield Johns Hopkins University Applied Physics Lab</institution>
          <addr-line>11100 Johns Hopkins Road Laurel, MD 20723-6099</addr-line>
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2000</year>
      </pub-date>
      <abstract>
        <p>We present an approach to multilingual information retrieval that does not depend on the existence of specific linguistic resources such as stemmers or thesaurii. Using the HAIRCUT system we participated in the monolingual, bilingual, and multilingual tasks of the CLEF-2000 evaluation. Our method, based on combining the benefits of words and character n-grams, was effective for both language-independent monolingual retrieval as well as for cross-language retrieval with translated queries. After describing our monolingual retrieval approach we compare a translation method using aligned parallel corpora to commercial machine translation software.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Background</title>
    </sec>
    <sec id="sec-2">
      <title>Overview</title>
      <p>The Hopkins Automated Information Retriever for Combing Unstructured Text (HAIRCUT) is a
research retrieval system developed at the Johns Hopkins University Applied Physics Lab (APL). One of
the research areas that we want to investigate with HAIRCUT is the relative merit of different
tokenization schemes. In particular we use both character n-grams and words as indexing terms. Our
experiences in the TREC evaluations have led us to believe that while n-grams and words are
comparable in retrieval performance, a combination of both techniques outperforms the use of a single
approach. Through the CLEF-2000 evaluation we demonstrate that unsophisticated,
languageindependent techniques can form a credible approach to multilingual retrieval. We also compare query
translation methods based on parallel corpora with automated machine translation.</p>
      <p>We participated in the monolingual, bilingual, and multilingual tasks. For all three tasks we used the
same 8 indices, a word and an n-gram based index in each of the four languages. Information about each
index is provided in Table 1. In all of our experiments documents were indexed in their native language
because we prefer query translation over document translation for reasons of efficiency.</p>
      <p># docs</p>
      <sec id="sec-2-1">
        <title>English 110,282</title>
      </sec>
      <sec id="sec-2-2">
        <title>French 44,013</title>
      </sec>
      <sec id="sec-2-3">
        <title>German</title>
        <p>153,694
Italian
58,051
collection size
(MB gzipped)
163
62
153
78
name
# terms</p>
        <p>index size (MB)
enw
en6
frw
fr6
gew
ge6
itw
it6
We used two methods of translation in the bilingual and multilingual tasks. We used the Systran®
translator to convert French and Spanish queries to English for our bilingual experiments and to convert
English topics to French, German and Italian in the multilingual task. For the bilingual task we also used
a method based on extracting translation equivalents from parallel corpora. Parallel English/French
documents were most readily available to us, so we only applied this method when translating French to
English.
boundaries are not recorded. Thus n-grams with leading, central, or trailing spaces are formed at word
boundaries. We used 6-grams with success in the TREC-8 CLIR task [6] and decided to do the same
thing this year. As can be seen from Table 1, the use of 6-grams as indexing terms increases both the size
of the inverted file and the dictionary.</p>
        <p>
          Query Processing
HAIRCUT performs rudimentary preprocessing on queries to remove stop structure, e.g., affixes such as
“… would be relevant” or “relevant documents should….” A list of about 1000 such English phrases was
translated into French, German, and Italian using both Systran and the FreeTranslation.com translator.
Other than this preprocessing, queries are parsed in the same fashion as documents in the collection.
The HAIRCUT HMM is a simple two-state model that captures both document and collection statistics
[
          <xref ref-type="bibr" rid="ref5">7</xref>
          ]. After the query is parsed each term is weighted by the query term frequency and an initial retrieval is
performed followed by a single round of relevance feedback. To perform relevance feedback we first
retrieve the top 1000 documents. We use the top 20 documents for positive feedback and the bottom 75
documents for negative feedback, however we check to see that no duplicate or neo-duplicate documents
are included in these sets. We then select terms for the expanded query based on three factors, a term’s
initial query term frequency (if any), the cube root of the (α=3, β=2, γ=2) Rocchio score, and a third term
selection metric that incorporates an idf component. The top-scoring terms are then used as the revised
query. After retrieval using this expanded and reweighted query, we have found a slight improvement by
penalizing document scores for documents missing many query terms. We multiply document scores by
a penalty factor:
        </p>
        <p>⎛ # of missing terms ⎞
PF = 1.0 − ⎜⎜⎝ total number of terms in query ⎠⎟
⎟
1.25
We use only about one-fifth of the terms of the expanded query for this penalty function
words
6-grams
# Top Terms
60
400
# Penalty terms
12
75
We conducted our work on a 4-node Sun Microsystems Ultra Enterprise 450 server. The workstation
had 2 GB of physical memory and access to 50 GB of dedicated hard disk space.</p>
        <p>The HAIRCUT system comprises approximately 25,000 lines of Java code.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Monolingual Experiments</title>
      <p>Our approach to monolingual retrieval was to focus on language independent methods. We refrained
from using language specific resources such as stoplists, lists of phrases, morphological stemmers,
dictionaries, thesauri, decompounders, or semantic lexicons (e.g. Euro WordNet). We emphasize that this
decision was made, not from a belief that these resources are ineffective, but because they are not
universally available (or affordable) and not available in a standard format. Our processing for each
language was identical in every regard and was based on a combination of evidence from word-based
and 6-gram based runs. We elected to use all of the topic sections for our queries.
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0
0%
10%
20%
30%
40%
60%
70%
80%</p>
      <p>The retrieval effectiveness of our monolingual runs is fairly similar for each of the four languages as
evidenced by Figure 1. We expected to do somewhat worse on the Italian topics since the use of
diacritical marks differed between the topic statements and the document collection; consistent with our
‘language-independent’ approach we did not correct for this. Given the generally high level of
performance and the number of ‘best’ and ‘above median’ topics for the monolingual tasks (see Table 2),
we believe that language independent techniques can be quite effective.</p>
      <p>avg prec
recall
# topics
One of our objectives was to compare the performance of the constituent word and n-gram runs that were
combined for our official submissions. Figure 2 shows the precision-recall curves for the base and
combined runs for each of the four languages. Our experience in the TREC-8 CLIR track led us to
believe that n-grams and words are comparable, however each seems do perform slightly better in
different languages. In particular, n-grams performed appreciably better on translated German queries,
something we attribute to a lack of decompounding in our word-based runs. This trend was continued
this year, with 6-grams performing just slightly better in Italian and French, somewhat better in German,
but dramatically worse in our unofficial runs of English queries against the bilingual relevance
judgments. We are stymied by the disparity between n-grams and words in English and have never seen
such a dramatic difference in other test collectiions. Nonetheless, the general trend seems to indicate that
combination of these two schemes has a positive effect as measured by average precision. Our method of
combining two runs is to normalize scores for each topic in a run and then to merge multiple runs by the
normalized scores.
M o no ling ual  E ng lis h  P erf o rmanc e  ( U no f f ic ial)
M o no ling ual  F renc h  P erf o rmanc e
0%
10%
20%
30%
40%
60%
70%
80%
90%
100%
0%
10%
20%
30%
40%
60%
70%
80%
90%
100%
aplmoge( 0.4501)
  w ords ( 0.3920)
6-­‐grams ( 0.4371)
aplmoit (0.4187)
  w ords ( 0.3915)
6-­‐grams ( 0.4018)</p>
    </sec>
    <sec id="sec-4">
      <title>Bilingual Experiments</title>
      <p>Our goal for the bilingual task was to evaluate two methods for translating queries, commercial machine
translation software and a method based on aligned parallel corpora. While high quality MT products are
available only for certain languages, the languages used most commonly in Western Europe are well
represented. We used the Systran product which supports bi-directional conversion between English and
the French, German, Italian, Spanish, and Portuguese languages. We did not use any of the domain
specific dictionaries that are provided with the product.</p>
      <p>The run, aplbifrc, was created by converting the French topic statements to English using Systran and
searching the LA Times collection. As with the monolingual task both 6-grams and words were used
separately and the independent results were combined. Our other official run using Systran was aplbispa
that was based on the Spanish topic statements.</p>
      <p>
        We only had access to large aligned parallel texts in English and French. We were therefore unable to
conduct experiments in corpora-based translation in other languages. Our English / French dataset
included text from the Hansard Set-A[
        <xref ref-type="bibr" rid="ref4">5</xref>
        ], Hansard Set-C[
        <xref ref-type="bibr" rid="ref4">5</xref>
        ], United Nations[
        <xref ref-type="bibr" rid="ref4">5</xref>
        ], RALI[
        <xref ref-type="bibr" rid="ref6">8</xref>
        ], and JOC[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
corpora. The Hansard data accounts for the vast majority of the collection.
      </p>
      <p>Description
Hansard Set-A 2.9 million aligned sentences
Hansard Set-C aligned documents, converted to ~400,000 aligned sentences
United Nations 25,000 aligned documents
RALI 18,000 aligned documents</p>
      <p>
        JOC 10,000 aligned sentences
Table 3. Description of the parallel collection used for aplbifrb
The process that we used for translating an individual topic is shown in Figure 3. First we perform a
pretranslation expansion on a topic by running that topic in its source language on a contemporaneous
expansion collection and extracting terms from top ranked documents. Thus for our French to English
run we use the Le Monde collection to expand the original topic which is then represented as a weighted
list of 60 words. Each of these words is then translated to the target language (English) using the
statistics of the aligned parallel collection. We selected a single ‘best’ translation for each word and the
translated word retained the weight assigned during topic expansion. Our method of producing
translations is based on a term similarity measure similar to mutual information [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]; we do not use any
dimension reduction techniques such as CL-LSI [4]. An example is shown for Topic C003 in Table 4.
Finally we ran the translated query on the target collection in four ways, using both 6-grams and words
and by using and not using relevance feedback.
&lt;F-narr&gt; Les documents pertinents exposent la réglementation et les décisions du gouvernement néerlandais
concernant la vente et la consommation de drogues douces et dures.
      </p>
      <sec id="sec-4-1">
        <title>Official English Query</title>
        <p>&lt;E-title&gt;
Drugs in Holland
&lt;E-desc&gt;
What is the drugs policy in the Netherlands?
&lt;E-narr&gt; Relevant documents report regulations and decisions made by the Dutch government regarding the sale
and consumption of hard and soft drugs.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Systran translation of French query</title>
        <p>
          &lt;F-title&gt;
Drug in Holland
&lt;F-desc&gt;
Which is the policy of the Netherlands as regards drug?
&lt;F-narr&gt; The relevant documents expose the regulation and the decisions of Dutch government concerning the sale
and the consumption of soft and hard drugs.
1.00
0.90
0.80
0.70
0.60
0.50
0.40
0.30
0.20
0.10
0.00
fra ( 0.3212)
frb ( 0.2223)
frc ( 0.3358)
frb2 ( 0.2595)
We obtained superior results using translation software instead of our corpora-based translation. The
precision-recall graph in Figure 5 shows a clear separation between the Systran-only run (aplbifrc) with
average precision 0.3358 and the corpora-only run (aplbifrb) with average precision of 0.2223. We do
not interpret this difference as a condemnation of our approach to corpus-based translation. Instead we
agree with Brachler et al. that “MT cannot be the only solution to CLIR [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].” Both translation systems
and corpus-based methods have their weaknesses. A translation system is particularly susceptible to
named entities not being found in its dictionary. Perhaps as few as 3 out of the 40 topics in the test set
mention obscure names: topics 2, 8, and 12. Topics 2 and 8 have no relevant English documents, so it is
difficult to assess whether the corpora-based approach would outperform the use of dictionaries or
translation tools on these topics. The run aplbifra is simply a combination of aplbifrb and aplbifrc that
we had expected to outperform the individual runs.
        </p>
        <p>There are several reasons why our translation scheme might be prone to error. First of all, the collection
is largely based on the Hansard data, which are transcripts of Canadian parliamentary proceedings. The
fact that the domain of discourse in the parallel collection is narrow compared to the queries could
account for some difficulties. And the English recorded in the Hansard data is formal, spoken, and uses
Canadian spellings whereas the English document collection in the tasks is informal, written, and
published in the United States. It should be also noted that generating 6-grams from a list of words
rather than from prose leaves out any n-grams that span word boundaries; such n-grams might capture
phrasal information and be of particular value. Finally we had no opportunity to test our approach prior
to submitting our results; we are confident that this technique can be improved.</p>
        <p>With some post-hoc analysis we found one way to improve the quality of our corpus-based runs. We had
run the translated queries both with, and without the use of relevance feedback. It appears that the
relevance feedback runs perform worse than those without this normally beneficial technique. The
dashed curve in Figure 5 labeled ‘frb2’ is the curve produced when relevance feedback is not used with
the corpora-translated query. Perhaps the use of both pre-translation and post-translation expansions
introduces too much ambiguity about the query.</p>
        <p>Below are our results for the bilingual task. There were no relevant English documents for topics 2, 6, 8,
23, 25, 27, and 35, leaving just 33 topics in the task.</p>
        <p>avg prec
recall
Multilingual Experiments
We did not focuse our efforts on the multilingual task. We selected English as the topic language for the
task and used Systran to produce translations in French, German, and Italian. We performed retrieval
using 6-grams and words and then performed a multi-way merge using two different approaches,
merging normalized scores and merging runs by rank.</p>
        <p>avg prec</p>
        <p>recall
The large number of topics with no relevant documents in the collections of various languages suggests
that the workshop organizers were successful in selecting challenging queries for merging. It seems
clear that more sophisticated methods of multilingual merging are required to avoid a large drop in
precision from the monolingual and bilingual tasks.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>The CLEF workshop provides an excellent opportunity to explore the practical issues involved in
crosslanguage information retrieval. We approached the monolingual task believing that it is possible to
achieve good retrieval performance using language-independent methods. This methodology appears to
have borne out based on our results using a combination of words and n-grams. For the bilingual task we
kept our philosophy of simple methods, but also used a high-powered machine translation product.
While our initial efforts using parallel corpora were not as effective as those with machine translated
queries, the results were still quite credible and we are confident this technique can be improved further.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Brachler</surname>
          </string-name>
          ,
          <string-name>
            <surname>M-Y. Kan</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Schauble</surname>
          </string-name>
          , '
          <article-title>The SPIDER Retrieval System and the TREC-8 CrossLanguage Track</article-title>
          .' In E. M. Voorhees and
          <string-name>
            <surname>D. K</surname>
          </string-name>
          . Harman, eds.,
          <source>Proceedings of the Eighth Text REtrieval Conference (TREC-8)</source>
          . To appear.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K. W.</given-names>
            <surname>Church</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Hanks</surname>
          </string-name>
          , 'Word Association Norms, Mutual Information, and Lexicography.' In Computational Linguistics,
          <volume>6</volume>
          (
          <issue>1</issue>
          ),
          <fpage>22</fpage>
          -
          <lpage>29</lpage>
          ,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>European</given-names>
            <surname>Language Resource Association</surname>
          </string-name>
          (ELRA), http://www.icp.grenet.fr/ELRA/home.html [4]
          <string-name>
            <given-names>T. K.</given-names>
            <surname>Landauer</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Littman</surname>
          </string-name>
          , '
          <article-title>Fully automated cross-language document retrieval using latent semantic indexing</article-title>
          .
          <source>' In the Proceedings of the Sixth Annual Conference of the UW Centre for the New Oxford English Dictionary and Text Research</source>
          .
          <volume>31</volume>
          -
          <fpage>38</fpage>
          ,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Linguistic</given-names>
            <surname>Data</surname>
          </string-name>
          <article-title>Consortium (LDC)</article-title>
          , http://www.ldc.upenn.edu/ [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Mayfield</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>McNamee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Piatko</surname>
          </string-name>
          , '
          <article-title>The JHU/APL HAIRCUT System at TREC-8</article-title>
          .' In E. M. Voorhees and
          <string-name>
            <surname>D. K</surname>
          </string-name>
          . Harman, eds.,
          <source>Proceedings of the Eighth Text REtrieval Conference (TREC-8)</source>
          . To appear.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D. R. H.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Leek</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Schwartz</surname>
          </string-name>
          , '
          <string-name>
            <given-names>A Hidden</given-names>
            <surname>Markov Model Information Retrieval System</surname>
          </string-name>
          .'
          <source>In the Proceedings of the 22nd International Conference on Research and Development in Information Retrieval (SIGIR-99)</source>
          , pp.
          <fpage>214</fpage>
          -
          <lpage>221</lpage>
          ,
          <year>August 1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Recherche</given-names>
            <surname>Appliquée en Linguistic Informatique</surname>
          </string-name>
          (RALI), http://www-rali.iro.umontreal.ca/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>