<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Web Retrieval Experiments with the EuroGOV Corpus at the University of Hildesheim</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Niels Jensen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>René Hackl</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Mandl</string-name>
          <email>mandl@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robert Strötgen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Science, University of Hildesheim</institution>
          ,
          <addr-line>Marienburger Platz 22 D-31141 Hildesheim</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the CLEF 2005 initiative, multlingual web retrieval was integrated as a task for the first time. This paper describes experiments based on one multilingual index carried out at the University of Hildesheim. Several indexing strategies based on a multi-lingual index have been tested with the EuroGOV corpus. Boosting topic fields with higher weight led to best results during post submission runs. The experiments also led to experiences in working with large test collections and the challenges associated with them.</p>
      </abstract>
      <kwd-group>
        <kwd>Web Retrieval</kwd>
        <kwd>Multilingual Information Retrieval</kwd>
        <kwd>N-gram Indexing</kwd>
        <kwd>Evaluation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Web search engines has become a part of every day life for many people. The development of information
retrieval systems for the web is faced with many challenges
        <xref ref-type="bibr" rid="ref1">(Arasu et al. 2001)</xref>
        . Systems give different answers
to these challenges and it is difficult to judge the effect of decisions during the design of search enigne. As a
consequence, there is a great need for evaluation in web retrieval
        <xref ref-type="bibr" rid="ref3">(Hawking 2000)</xref>
        . The web is also a natural
source for multilingual documents.
      </p>
      <p>
        Within the Cross Language Evaluation Forum (CLEF) the web track has been created (
        <xref ref-type="bibr" rid="ref10 ref9">Sigurbjörnsson
et al. 2005</xref>
        b). A large multilingual corpus has been collected and distributed (
        <xref ref-type="bibr" rid="ref10 ref9">Sigurbjörnsson et al. 2005</xref>
        a). In our
first participation, we intended to tune our system to the challanges of a large web corpus. For the experiments,
language resources in all languages were not available from ad-hoc retrieval. As a consequence, we considered
n-gram indexing for the web retrieval task
        <xref ref-type="bibr" rid="ref8">(McNamee &amp; Mayfield 2004)</xref>
        .
      </p>
    </sec>
    <sec id="sec-2">
      <title>2 Data Pre-Processing</title>
      <p>
        Since the files of the EuroGOV corpus were not released in well formed XML, substantial effort for data
preprocessing was necessary. A corpus in well formed XML would allow us to use the System implemented during
the CLEF 2004 campaign for multilingual ad-hoc tasks
        <xref ref-type="bibr" rid="ref2">(Hackl et al. 2005)</xref>
        . The two main items in the EuroGOV
files that needed replacing were predeclared entities. This step was required for ampersand with the associated
entity reference in the URL fields of the individual documents and all nested CDATA tags. The first attempt to
reformat the files has been carried out by a Perl-script. At the first view, it seemed that the Perl-scirpt would
work perfectly for our needs. Unfortunately, we realized that during the process of indexing the corpus, the XML
parser would frequently report “parser exceptions” that we traced back to the fact that the XML files still
contained a couple of not adjusted predeclared entities. Having this in mind, a Java program was developed that
worked through the whole corpus perfectly. It seems that Perl is not able to process EuroGOV files bigger than
250 MB since we successfully tested the Perl-script with the small-size files (22 MB, 59 MB &amp; 220 MB) of the
corpus.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3 Submitted Retrieval Experiments with EuroGOV</title>
      <p>
        As mentioned in the introduction, one multilingual index was created. In order to generate a slim index we
assembled a multilingual stopwordlist. The bases for this list were the stopwordlists supplied by the University
of Neuchatel1 and a list developed specifically for the Czech language
        <xref ref-type="bibr" rid="ref4">(Hofman Miquel 2005)</xref>
        . All lists were
combined and revised into one file. This multilingual stopwordlist covers twelve languages and was used for the
indexing process of the corpus.
      </p>
      <p>For our retrieval experiments, we created three different multilingual indexes. Two were created with
the Lucene StandardAnalyzer2, which does not implement any linguistic processing apart from word
segmentation. The first index covered the whole corpus whereas the second index cut off the indexing process
after a maximum of 200 characters for each individual document. Due to this approach, the sizes of the indexes
varies from 5 GB to 700 MB.</p>
      <p>
        The third index was created with a NGram Analyzer also applied to multilingual ad-hoc retrieval before
        <xref ref-type="bibr" rid="ref2">(Hackl et al. 2005)</xref>
        . Because of performance and time restrictions the trigram approach was only applied to the
title field of the individual documents in the corpus files. As a result the size of the index is down to 300 MB
which led to a very quick and stable performance at retrieval time. These three indexes are the foundation for our
experiments. As a main retrieval engine, we used Lucene 1.43. Some of the basic code for retrieval and n-gram
analysis was adopted from previous CLEF ad-hoc experiments
        <xref ref-type="bibr" rid="ref2">(Hackl et al. 2005)</xref>
        . Six different baseline runs
were submitted. We did not use any of the metadata that was supplied by the topics due to time and resource
constraints. Our monolingual queries were created with the title field of the topic whereas the multilingual
queries were based on the monolingual title field and the translation language English field. Both types of
queries were sent to one multilingual index. Results are shown in table 1.
1 Stopwordlists: http://www.unine.ch/Info/clef/ verified August 11th 2005
2 Lucene StandardAnalyzer: http://lucene.apache.org verified on August 11th 2005
3 Lucene: http://lucene.apache.org verified August 11th 2005
compensated or even improved. As table 2 shows quite obviously, even a more complete index was not able to
improve the MRR of the trigram runs. The results declined by approx. 50 %.
In the second part of our post experiments we took the four indexes we had generated, and modified the weights
of the query fields. The ratio for the two query fields were 10 to 1 and vice versa. The results that are shown in
table 3 and 4 show that by boosting the title field of the query the results improve by 0.0144 MRR points on
average. Applying this procedure, the performance of the multilingual run based on the StandardAnalyzer Index
results in higher MRR values. The boosted multilingual run has a better result than any monolingual run and is
the best run of all our experiments.
For the first web track at CLEF we intended to tune our system to be able to cope with a large amount of data.
We suceeded in returning valid results for several runs.
      </p>
      <p>
        In future experiments, we intend to step beyond the baseline runs and try to involve the metadata that is
being provided by the WebCLEF topics. We also want to include advanced quality measures into consideration.
Link based quality measures seem to be integral part of commercial search engines. They have been evaluated at
the web track at TREC
        <xref ref-type="bibr" rid="ref3">(Hawking 2000)</xref>
        . Advanced quality measures take more features into account, especially
information and design aspects
        <xref ref-type="bibr" rid="ref2 ref7">(Mandl 2005)</xref>
        .
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Arasu</surname>
          </string-name>
          , Arvind; Cho, Junghoo; Garcia-Molina, Hector; Paepcke, Andreas; Raghavan,
          <string-name>
            <surname>Sriram</surname>
          </string-name>
          (
          <year>2001</year>
          )
          <article-title>: Searching the Web</article-title>
          .
          <source>In: ACM Transactions on Internet Technology</source>
          <volume>1</volume>
          (
          <issue>1</issue>
          ) pp.
          <fpage>2</fpage>
          -
          <lpage>43</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Hackl</surname>
          </string-name>
          , René; Mandl, Thomas; Womser-Hacker,
          <source>Christa</source>
          (
          <year>2005</year>
          )
          <article-title>: Mono-</article-title>
          and
          <string-name>
            <surname>Cross-lingual Retrieval</surname>
          </string-name>
          Experiments at the University of Hildesheim. In: Peters, Carol; Clough, Paul; Gonzalo, Julio; Kluck, Michael; Jones, Gareth; Magnini, Bernard (eds):
          <article-title>Multilingual Information Access for Text, Speech and Images: Results of the Fifth CLEF Evaluation Campaign</article-title>
          . Berlin et al.:
          <source>Springer [Lecture Notes in Computer Science</source>
          <volume>3491</volume>
          ] pp.
          <fpage>165</fpage>
          -
          <lpage>169</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Hawking</surname>
          </string-name>
          ,
          <string-name>
            <surname>David</surname>
          </string-name>
          (
          <year>2000</year>
          )
          <article-title>: Overview of the TREC-9 Web Track</article-title>
          .
          <source>In: The Ninth Text Retrieval Conference (TREC-9)</source>
          .
          <source>NIST Special Publication 500-249. National Institute of Standards and Technology. Gaithersburg, Maryland. November</source>
          <year>2000</year>
          . http://trec.nist.gov/pubs/trec9/t9_proceedings.html
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Hofman</given-names>
            <surname>Miquel</surname>
          </string-name>
          ,
          <string-name>
            <surname>Laura</surname>
          </string-name>
          (
          <year>2005</year>
          )
          <article-title>Informationslinguistische Ressourcen für das Information Retrieval in der tschechischen Sprache im Rahmen des Cross Language Evaluation Forums (CLEF)</article-title>
          .
          <source>Master Thesis Information Science</source>
          , University of Hildesheim.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Jensen</surname>
          </string-name>
          ,
          <string-name>
            <surname>Niels</surname>
          </string-name>
          (2005a)
          <article-title>Web Information Retrieval am Beispiel des WEB-GOV Korpus</article-title>
          .
          <source>Master Thesis Information Science</source>
          , University of Hildesheim.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Jensen</surname>
          </string-name>
          ,
          <string-name>
            <surname>Niels</surname>
          </string-name>
          (2005b)
          <article-title>Mehrsprachiges Information Retrieval mit einem WEB-Korpus</article-title>
          . In: Mandl, Thomas; Womser-Hacker, Christa (Eds.):
          <source>Proceedings Vierter Hildesheimer Information Retrieval und Evaluierungsworkshop (HIER 2005) Hildesheim, 20.7</source>
          .
          <year>2005</year>
          . Universitätsverlag Konstanz [Schriften zur Informationswissenschaft] to appear.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Mandl</surname>
          </string-name>
          , Thomas (
          <year>2005</year>
          ):
          <article-title>The quest for the best pages on the web</article-title>
          .
          <source>In: Information Service &amp; Use</source>
          . To appear
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>McNamee</surname>
          </string-name>
          , Paul; Mayfield,
          <string-name>
            <surname>James</surname>
          </string-name>
          (
          <year>2004</year>
          )
          <article-title>: Character N-Gram Tokenization for European Language Text Retrieval</article-title>
          .
          <source>In: Information Retrieval</source>
          , vol.
          <volume>7</volume>
          (
          <issue>1</issue>
          /2). pp.
          <fpage>73</fpage>
          -
          <lpage>98</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Sigurbjörnsson</surname>
          </string-name>
          , Börkur; Kamps, Jaap; de Rijke,
          <article-title>Maarten (2005a): Blueprint of a Cross-Lingual Web Retrieval Collection</article-title>
          .
          <source>In: Journal of Digital Information Management</source>
          , vol.
          <volume>3</volume>
          (
          <issue>1</issue>
          ) pp.
          <fpage>9</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Sigurbjörnsson</surname>
          </string-name>
          , Börkur; Kamps, Jaap; de Rijke,
          <article-title>Maarten (2005b): Overview of WebCLEF 2005</article-title>
          . In this volume.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>