<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The University of Amsterdam at WebCLEF 2005</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jaap Kamps</string-name>
          <email>S@1</email>
          <email>S@10</email>
          <email>S@5</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maarten de Rijke</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bo¨rkur Sigurboj¨rnsson</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Archives and Information Studies, University of Amsterdam</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Informatics Institute, University of Amsterdam</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Measurement</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>We describe the University of Amsterdam's participation in the WebCLEF track at CLEF 2005. We submitted runs for both the mixed monolingual task and the multilingual task.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>7 Digital Libraries</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>In the CLEF 2005 WebCLEF track, we took part in two of the retrieval tasks. We took part
in the WebCLEF mixed monolingual task. Our participation here was aimed at evaluating the
effectiveness of standard ad hoc retrieval settings for a stream of topics in various languages. Our
assumption was that this would shed new light on the robustness of modern information retrieval
techniques.</p>
      <p>
        We also took part in the WebCLEF multilingual task. Our participation here was aimed at
evaluating the effectiveness of straightforwardly combining runs using a number of translations of
the original English queries. Such methods have previously been used successfully at the CLEF
multilingual ad hoc retrieval task [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ].
      </p>
      <p>This paper is structured as follows. In Section 2 we describe our retrieval system as well as
the approaches used for the two WebCLEF tasks in which we participate. Section 3 describes our
official retrieval runs for WebCLEF 2005, and Section 4 discusses the results we have obtained.
Finally, in Section 5, we offer some conclusions regarding our multilingual web retrieval efforts.</p>
    </sec>
    <sec id="sec-2">
      <title>Description</title>
      <p>
        Our retrieval system is based on the Lucene engine with a number of home-grown extensions [
        <xref ref-type="bibr" rid="ref3 ref8">3, 8</xref>
        ].
2.1
      </p>
      <sec id="sec-2-1">
        <title>Retrieval Approach</title>
        <p>
          For our ranking, we used the default similarity measure in Lucene [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], i.e., for a collection D,
document d and query q containing terms ti:
sim(q, d) =
t∈q
X tft,q · idft tft,d · idft
normq · normd
· coordq,d · weightt ,
where
tft,X
idft
normd
coordq,d
normq
=
=
=
=
=
pfreq(t, X)
        </p>
        <p>|D|
freq(t, D)
1 + log
p d</p>
        <p>| |
|q ∩ d|
|q|
t∈q
sX tft,q · idft 2
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Tokenization</title>
        <p>We indexed the whole collection by simply extracting the full text from the documents. We did
not apply any stemming nor did we use a stopword list. We applied case-folding and normalized
marked characters to their unmarked counterparts, i.e., mappingo¨ to o, ae to ae, ˆıto i, etc. The
only language specific processing we did was a transformation of the multiple Russian encodings
into an ASCII transliteration.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Translation</title>
        <p>
          We used the WorldLingo machine translation [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] for translating the English topic statements
into eight languages: Dutch, French, German, Greek, Italian, Portuguese, Russian, and Spanish.
Combined with the English source topic statements, this gave us short topic statements in nine
European languages.
2.4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Combination</title>
        <p>
          We combined various ‘base’ runs using the unweighted CombSUM function of Fox and Shaw [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
The runs were combined after normalizing the retrieval status values (RSVs) to the interval [
          <xref ref-type="bibr" rid="ref1">0,1</xref>
          ]
as suggested in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
3
3.1
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Runs</title>
      <sec id="sec-3-1">
        <title>Mixed-monolingual task</title>
        <p>We submitted one run to the mixed-monolingual task. The run uses the short topic statement
in the htitlei efild of the WebCLEF 2005 topics. Our run uses Lucene’s standard ranking formula
applied on our full-text index (as discussed in Section 2 above).
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Multilingual task</title>
        <p>We submitted four runs to the multilingual task. All runs use the English version of the short
topic statement in the htranslation language=”EN”i field of the WebCLEF 2005 topics, and the
translations mentioned in Section 2.3.</p>
        <p>We experimented along two dimensions. The rfist dimension is the number of topic languages:
All topics
Home pages
Named pages</p>
        <p>Recall that WebCLEF provides a stream of topics, with topics from arbitrary languages. For the
multilingual task, we use the English short topic statement. The downside of this is, of course,
that finding the targeted page in the source language becomes a formidable problem. The upside
is that, at least, the topic language is known, and the same holds for the translations we obtained.
The second dimension we experiment with is trying to exploit this knowledge:
All results Topics in one language may likely retrieve pages in other languages as well. A case
in point is WebCLEF topic WC0014, whose English topic statement (“Chancellery at the
Spreebogen”) could still allow us to retrieve German pages targeted by the German topic
statement (“Bundeskanzleramt am Spreebogen”). Hence, we may simply use all pages
retrieved by a topic of a particular, known language.</p>
        <p>Language restricted Since we know the language of the topic in each of the translations, and
the intention of the translated topic is to retrieve pages in that language, we may decide to
restrict the pages returned by our retrieval system. We do this by restricting retrieved pages
to the dominant domains. For example, for a run with the topics translated to Dutch, we
restrict pages to come from either the .nl or the .eu.int domain.</p>
        <p>Combining the two dimensions naturally suggests the four following cases:
1. using nine topic languages without restriction;
2. using nine topic languages and restricting pages to dominant domains;
3. using five topic languages without restriction; and
4. using five topic languages and restricting pages to dominant domains.</p>
        <p>For each of the cases, we obtain vfie to nine different runs, which we combine using unweighted
CombSUM. This results in the four runs submitted to WebCLEF 2005.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <sec id="sec-4-1">
        <title>Mixed-monolingual task</title>
        <p>We submitted a single no-thrills run for the mixed monolingual task, using standard ad hoc
document retrieval setting (as discussed in Section 3). Table 1 reports the result of the mixed
monolingual run. A number of observations present themselves. First, we see that, on average,
the desired page is found in the top three. That is a reassuring result for the mixed monolingual
All topics
Nine languages
Nine languages, restricted
Five languages
Five languages, restricted</p>
        <sec id="sec-4-1-1">
          <title>Home pages</title>
          <p>Nine languages
Nine languages, restricted
Five languages
Five languages, restricted</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>Named pages</title>
          <p>Nine languages
Nine languages, restricted
Five languages
Five languages, restricted
MRR
0.0072
0.0157
0.0084
0.0163
MRR
0.0109
0.0158
0.0129
0.0168
S@1
0.0041
0.0124
0.0041
0.0124</p>
          <p>S@1
0.0066
0.0066
0.0066
0.0066
S@5
0.0083
0.0165
0.0083
0.0165</p>
          <p>
            S@5
0.0066
0.0230
0.0098
0.0230
S@10
0.0124
0.0165
0.0124
0.0207
S@10
0.0197
0.0262
0.0197
0.0262
task. Somewhat worrying is the success rate at rank 10, with no relevant page found for over
40% of the topics. Second, named page topics score somewhat higher than home page topics, on
all measures. This is well-known from other web retrieval tasks [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ], which also suggests that the
scores for home page nfiding can be substantially improved using specicfi web centric techniques
such as various document representations and non-content priors [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ].
4.2
          </p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Multilingual task</title>
        <p>We submitted four runs for the multilingual task (as discussed in Section 3). We will rfist look
at the overall results, and then focus on the effectiveness for each of the languages in which we
translated the English topics.</p>
        <p>Table 2 reports the result of the multilingual runs. Again, we make a number of observations.
First, we see that scores are substantially lower than for the mixed monolingual task. The
complexity of the multilingual task can hardly be overestimated: given an English query we have to
guess what page in any language has to be returned to the user. Obvious ways of limiting this
wealth of options are the use of topic meta-fields, or of sophisticated techniques to extract target
language cues. Second, our experiment with the number of translations to use, points conclusively
to the smaller set of vfie language used frequently in the topic set. It is a reassuring fact that
the improvement is moderate, and the extended set of translations is far from detrimental to the
performance. Note that the extended set includes, for example, Italian, which is not used in any
of the topics. Third, our experiment with restricting our intention to pages in the language of the
topic translation is clearly successfull. It leads to substantial improvement of the score.</p>
        <p>We now zoom in on the effectiveness of the individual translations. Table 3 lists the results
of the translated queries, both evaluated against the whole topic set, as well as against all topics
targeting a page in the language at hand. We see the following. First, when looking at the
restricted topic sets, effectiveness varies from total failure (Greek) to perfection (French). The
score for the five frequent languages is reasonable compared to those of the mixed monolingual
task. Hence, one may conclude that the automatic topic translations are effective. Second, when
looking at all topics, the scores are generally unimpressive and mirroring the frequency with which
a topic of the given language appears in the topic set. This comes as no surprise, given that the
topic set covers eleven languages, and each of the topic translations will dominantly target only
one of them. Third, the translated topics pick up relevant pages in languages other than the
target language. In particular, the Italian topics do pick up a relevant page for 35 of the topics.
Fourth, the single topic language runs are still much more effective than the combined multilingual
runs in Table 2. This is a disappointing result, and clearly indicates that the straightforward
run combination is ineffective. On a more positive note, however, the results for the individual
translations strongly suggest that more sensible methods are possible.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>This paper documents the University of Amsterdam’s participation in the CLEF 2005 WebCLEF
track. The EuroGOV collection used at WebCLEF is based on a crawl of governmental information
from a range of sites. Such a collection of web data is much noisier than traditional collections
of newswire and newspaper data originating from a single source. Moreover, the linguistic variety
in the collection makes it harder to apply language-specific processing methods such as stemming
algorithms. Hence, we simply indexed the collection by extracting the full text from the documents.</p>
      <p>For the mixed monolingual task, we submitted a single, standard ad hoc retrieval run. Our
main finding is that such a straightforward approach is relatively effective, that uses no web specific
settings. Considering the fact that we are dealing with a stream of topics in eleven languages, and
with an even greater number of languages in the collection, this sheds new light on the robustness
of modern information retrieval techniques.</p>
      <p>For the multilingual task, we experimented with different numbers of translations of the English
queries, and with restricting the returned pages to the now-known language of the query at hand.
Our experiments show benecfiial effects for restricting the number of translation to those occurring
frequently in the topic set, as well as for limiting query translations to return only pages in the
language of the query. In general, however, the combined results for the multilingual task are
unimpressive. The individual query translations, however, seem relatively successful in targeting
their share of relevant pages. This casts considerable doubt on the effectiveness of standard
combination methods for this particular task.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We want to thank Valentin Jijkoun for help with the Russian collection. Jaap Kamps was
supported by a grant from the Netherlands Organization for Scienticfi Research (NWO) under project
numbers 612.066.302 and 640.001.501. Maarten de Rijke was supported by grants from NWO
under project numbers 017.001.190, 220-80-001, 264-70-050, 354-20-005, 612-13-001, 612.000.106,
612.000.207, 612.066.302, and 612.069.006.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Craswell</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Hawking</surname>
          </string-name>
          .
          <article-title>Overview of the TREC-2004 Web Track</article-title>
          .
          <source>In Proceedings TREC</source>
          <year>2004</year>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E.A.</given-names>
            <surname>Fox</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.A.</given-names>
            <surname>Shaw</surname>
          </string-name>
          .
          <article-title>Combination of multiple searches</article-title>
          .
          <source>In The Second Text REtrieval Conference (TREC-2)</source>
          , pages
          <fpage>243</fpage>
          -
          <lpage>252</lpage>
          .
          <article-title>National Institute for Standards and Technology</article-title>
          .
          <source>NIST Special Publication 500-215</source>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>ILPS.</surname>
          </string-name>
          <article-title>The ILPS extension of the Lucene search engine</article-title>
          ,
          <year>2005</year>
          . http://ilps.science.uva. nl/Resources/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          .
          <article-title>Web-centric language models</article-title>
          .
          <source>In Proceedings of the Fourteenth ACM Conference on Information and Knowledge Management (CIKM</source>
          <year>2005</year>
          ). ACM Press, New York NY, USA,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Fissaha</given-names>
            <surname>Adafre</surname>
          </string-name>
          , and M. de Rijke.
          <article-title>Effective translation, tokenization and combination for cross-lingual retrieval</article-title>
          . In C. Peters,
          <string-name>
            <given-names>P. D.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kluck</surname>
          </string-name>
          , and B. Magnini, editors,
          <source>Multilingual Information Access for Text</source>
          ,
          <article-title>Speech and Images: Results of the Fifth CLEF Evaluation Campaign</article-title>
          , volume
          <volume>3491</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>123</fpage>
          -
          <lpage>134</lpage>
          . Springer Verlag, Heidelberg,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Monz</surname>
          </string-name>
          , M. de Rijke, and
          <string-name>
            <given-names>B.</given-names>
            <surname>Sigurboj</surname>
          </string-name>
          <article-title>¨rnsson. Language-dependent and languageindependent approaches to cross-lingual text retrieval</article-title>
          . In Carol Peters, Julio Gonzalo,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Braschler</surname>
          </string-name>
          , and Michael Kluck, editors,
          <source>Comparative Evaluation of Multilingual Information Access Systems, CLEF</source>
          <year>2003</year>
          , volume
          <volume>3237</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>152</fpage>
          -
          <lpage>165</lpage>
          . Springer,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.H.</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>Combining multiple evidence from different properties of weighting schemes</article-title>
          .
          <source>In Proceedings of the 18th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , pages
          <fpage>180</fpage>
          -
          <lpage>188</lpage>
          . ACM Press, New York NY, USA,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Lucene</surname>
          </string-name>
          .
          <source>The Lucene search engine</source>
          ,
          <year>2005</year>
          . http://jakarta.apache.org/lucene/.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Worldlingo</surname>
          </string-name>
          . Online translator,
          <year>2005</year>
          . http://www.worldlingo.com/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>