<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A poor man's approach to CLEF</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arjen P. de Vries</string-name>
          <email>arjen@acm.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amsterdam</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>The Netherlands</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Twente</institution>
          ,
          <addr-line>Enschede</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Mirror DBMS [dV99] aims speci cally at supporting both data management and content management in a single system. Its design separates the retrieval model from the speci c techniques used for implementation, thus allowing more exibility to experiment with a variety of retrieval models. Its design based on database techniques intends to support this exibility without causing a major penalty on the e ciency and scalability of the system. The support for information retrieval in our system is presented in detail in [dVH99], [dV98], and [dVW99]. The primary goal of our participation in CLEF is to acquire experience with supporting Dutch users. Also, we want to investigate whether we can obtain a reasonable performance without requiring expensive (but high quality) resources. We do not expect to obtain impressive results with our system, but hope to obtain a baseline from which we can develop our system further. We decided to submit runs for all four target languages, but our main interest is in the bilingual Dutch to English runs. We have used only `o -the-shelf' tools for stopping, stemming, compound-splitting (only for Dutch) and translation. All our tools are available for free, without usage restrictions for research purposes.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Goals</title>
      <p>The Ergane translation dictionaries2 were made available by Gerard van Wilgen. To avoid the
necessity of a bilingual wordlist for every possible language combination, Ergane uses the arti cial
language Esperanto as an interlingua. Ergane supports translation from and to no less than 57
languages, although some languages are only covered by a few hundred words. The number of
entries in the dictionaries used are summarized in Table 2.</p>
      <p>Because of synonyms, the size of bilinugal dictionaries might actually be bigger than the size
of the smallest word-list of a language pair. After removal of multiword expressions, the number
of Dutch entries in the bilingual translation lexicons are presented in Table 3.</p>
      <p>Note that these dictionary sizes are really small compared to dictionaries used in other
crosslanguage retrieval experiments. For instance, Hiemstra and Kraaij have used professional
dictionaries that are about 15 times as large [HK99].</p>
      <sec id="sec-1-1">
        <title>Compound-splitting</title>
        <p>Compound-splitting was only used for the Dutch queries. We applied a simple compound-splitter
developed at the University of Twente. The algorithm tries to split any word that is not in
the bilingual dictionary using the full word-list of about 50,000 Dutch words from Ergane. The
algorithm tries to split the word in as little parts as possible. It encodes a morphological rule to
handle a property known as `tussen-s', but it does not use part-of-speech information to search
for linguistically plausible compounds.</p>
        <p>Because the Dutch word-list used for splitting was much larger than the number of entries in the
bilingual dictionaries, compound-splitting might result in words that are only partially translated.
For example, the Dutch word `wereldbevolkingsconferentie' (topic 13, English: `World Population
Conference') was correctly splitted in three parts: `wereld', `bevolking' and `conferentie' of which
only the rst two words have entries in the Dutch-to-French dictionary.3
3</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>System</title>
      <p>For a detailed description of our retrieval system, we refer the interested user to [dVH99]. The
underlying retrieval model is best explained in our technical report4 [HdV00]. It supplements the
2http://www.travlang.com/Ergane/
3This example also illustrates the `tussen-s' rule: the `s' between `bevolking' and `conferentie' has been correctly
removed.</p>
      <p>4http://wwwhome.cs.utwente.nl/~hiemstra/papers/index.html#ctit</p>
      <p>English
French
German
Italian
Bi-lingual
Multi-lingual
theoretical basis of the model with a series of experiments, comparing this model with other, more
common retrieval models.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>This section discusses the results obtained with our system. We discuss the retrieval results
expressed in average precision, and, the coverage of our translations. After discussing the o cial
runs, we present some tests performed with pre-processing Dutch topics.
4.1</p>
      <p>O</p>
      <p>cial results
All experiments were done using the title and description elds of the topics. The average query
length for Dutch was 10.5 after stopping (which is of course rather long compared to the average
query size people enter in e.g. web search engines).</p>
      <p>Table 6 summarizes our results. The second column shows the number of queries with hits
in the monolingual runs; the third and fourth columns show the mean average precision5. The
monolingual results for English have been based on the bilingual qrels. The last column summarizes
the drop in average precision that can be attributed to the translation process.</p>
      <p>We hypothesize from the relatively low average precision (0.3134) on the monolingual German
task that we really have to perform compound-splitting of this corpus. Another possible cause of
the lower score for German is that we had to merge the runs from the two subcollections, which
were handled separately. But, our experiments on TREC-8 showed that this cannot really explain
such a performance drop.</p>
      <p>We attribute the large drop in performance for e.g. the bilingual Italian task (only 24% of
the average precision from on the monolingual task) to the small coverage of our translation
dictionaries. The coverage of the topic translations produced has been summarized in table 7.</p>
      <p>5The mean average precision for the bilingual runs as given by trec eval, normalized for the number of queries
with hits in the monolingual case.</p>
      <p>Together, the inferior results on German and Italian explain the disappointing average precision
obtained on the multilingual retrieval task (0.0864).
4.2</p>
      <sec id="sec-3-1">
        <title>Morphological normalisation and compound-splitting</title>
        <p>Our primary goal with CLEF participation is to test whether we could provide a Dutch interface
to our retrieval systems. To con rm our intuition about stemming and compound-splitting, we
performed some test runs to analyze the e ects of morphological normalisation and
compoundsplitting for Dutch. We either performed stemming or not, and performed compound-splitting or
not, resulting in four variants of the system:
nlen1:
nlen2:
nlen3:
nlen4:
base-line translation using full-form dictionary
translation using Dutch stemmer and a dictionary with stemmed entries
translation using compound-splitter for Dutch and full-form dictionary
translation using compound-splitter and dictionary with stemmed entries</p>
        <p>The results of these runs are summarized in Table 8. We conclude that compound-splitting is
very important, and stemming seems a useful pre-processing step.</p>
        <p>To support these conclusions, Table 9 summarizes the coverage of the various translations used
in the Dutch runs. Compound-splitting and morphological stemming of Dutch words nearly triples
the relative coverage of the translation dictionaries. The total of 92 untranslated Dutch terms in
the English queries include about 13 proper names like `Weinberg', `Salam' and `Glashow' (topic
2) and a view terms that were left untranslated in the Dutch topics like `Academie Francaise'
(topic 15) and `Deutsche Bundesbahn' (topic 40).
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and future work</title>
      <p>Summarizing our experiments, we may conclude that our retrieval models works well for all
monolingual runs, except for German. Future experiments will have to con rm whether a process like
compound-splitting will indeed bring our monolingual results to a level comparable to the other
languages. The in uence of compound-splitting of Dutch topics on the bilingual results raises our
expectations on this end.</p>
      <p>We were not at all unhappy with our bilingual results. But, from the coverage of the
translations, we still have to conclude that a poor man's approach should not expect to result in rich
men's retrieval results. But, we cannot blame it all on the dictionaries. The current version of
our retrieval system does not use query expansion techniques to improve mediocre translations; it
remains to be seen if better statistical techniques can bring us closer to the results obtained with
`proper' linguistic tools.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>Without Djoerd Hiemstra's help, I would never have obtained CLEF results: he helped me with
most of the pre-processing. More importantly, our discussions about his and competing retrieval
models has improved signi cantly my understanding Information Retrieval. I also like to thank
Gerard van Wilgen for making available the Ergane dictionaries.
[dV98]
[HK99]</p>
      <p>
        D. Hiemstra and W. Kraaij. Twenty-One at TREC-7: Ad-hoc and cross-language track.
In E.M. Voorhees and D.K. Harman, editors, Proceedings of the Seventh Text Retrieval
Conference TREC-7, number 500-242 in NIST Speci
        <xref ref-type="bibr" rid="ref2">al publications, 1999</xref>
        .
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [dV99]
          <string-name>
            <surname>A.P. de Vries</surname>
          </string-name>
          . Mirror:
          <article-title>Multimedia query processing in extensible databases</article-title>
          .
          <source>In Proceedings of the fourteenth Twente workshop on language technology (TWLT14): Language Technology in Multimedia Information Retrieval</source>
          , pages
          <volume>37</volume>
          {
          <fpage>48</fpage>
          ,
          <string-name>
            <surname>Enschede</surname>
          </string-name>
          , The Netherlands,
          <year>December 1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>A.P. de Vries</surname>
          </string-name>
          .
          <article-title>Content and multimedia database management systems</article-title>
          .
          <source>PhD thesis</source>
          , University of Twente, Enschede, The Netherlands,
          <year>December 1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [dVH99]
          <string-name>
            <surname>A.P. de Vries</surname>
            and
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Hiemstra</surname>
          </string-name>
          .
          <article-title>The Mirror DBMS at TREC</article-title>
          .
          <source>In Proceedings of the Seventh Text Retrieval Conference TREC-8</source>
          , Gaithersburg, Maryland,
          <year>November 1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [dVW99]
          <string-name>
            <surname>A.P. de Vries</surname>
            and
            <given-names>A.N.</given-names>
          </string-name>
          <string-name>
            <surname>Wilschut</surname>
          </string-name>
          .
          <article-title>On the integration of IR and databases</article-title>
          .
          <source>In Database issues in multimedia; short paper proceedings, international conference on database semantics (DS-8)</source>
          , pages
          <fpage>16</fpage>
          {
          <fpage>31</fpage>
          ,
          <string-name>
            <surname>Rotorua</surname>
          </string-name>
          , New Zealand,
          <year>January 1999</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>