<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mono- and Crosslingual Retrieval Experiments at the University of Hildesheim</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>René Hackl</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Mandl</string-name>
          <email>mandl@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christa Womser-Hacker</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Hildesheim, Information Science</institution>
          ,
          <addr-line>Marienburger Platz 22 D-31141 Hildesheim</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this year's participation we continued to evaluate open source information retrieval software. We used mainly the system Lucene and experimented with some of the most effective optimization strategies applied in CLEF. The effectiveness of open source and other free tools can be enhanced by these optimization strategies. For most languages, blind relevance feedback leads to considerable improvement. Indexing strategies with n-grams have not led to improvements within Lucene. In the CLEF 2004 campaign, we tested an adaptive fusion system based on the MIMOR model with several mono- and multi-lingual tasks. As a basic retrieval system we employed the open source system Lucene. Our main goal is to measure the quality of open source product in comparison to the best systems at CLEF. We exploit some of the most promising optimization techniques applied at CLEF in order to observe the potential for improvement of standard IR systems like Lucene. This work contributes to the practical application of the results from CLEF. Lucene has proved to be very efficient in CLEF as well as in other projects (e.g. cf. Hackl, Mandl &amp; Schwantner 2004) and is becoming increasingly popular. We expect Lucene to be employed in many more contexts. Therefore, we intend to continue testing its effectiveness within the CLEF campaign. Our basic fusion approach MIMOR is described in more detail in Womser-Hacker 1997 and Mandl &amp; WomserHacker 2004a. MIMOR has already been applied to CLEF experiments (Hackl, Kölle et al. 2004).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The tools we employed this year include Lucene 1.4-final1 and JavaTM-based snowball2 language analyzers. Last
year we had also evaluated the MySQL’s full text indexing and search module, but due to their poor
performance they were excluded this year. For this year’s participation we focussed on different indexing
methods such as different stemmers and n-gram-techniques.</p>
      <p>We took part in the monolingual tracks for Russian and Finnish, the bilingual track English to Russian and the
multilingual track.</p>
      <p>Firstly, we ran some preliminary monolingual experiments on the collections from 2003 without query
expansion (Table 1). Note that we did not index the LA Times 1994, as well as the 1994 volumes of the French
collections as they were not needed for this year.</p>
      <p>Secondly, we had planned to try out every combination of indexing methods for a language, to find out whether
there is some fusion potential. This can be seen in the table for Finnish (“Finnish all”), where all result lists from
the different indices were merged into a single one.</p>
      <p>Russian character handling in Java led to problems which caused a very low performance.</p>
      <p>For evaluation of the test runs we used our beta stage Java clone of the official trec_eval program.</p>
      <sec id="sec-1-1">
        <title>1 Lucene: http://jakarta.apache.org/lucene/docs/index.html 2 Snowball: http://jakarta.apache.org/lucene/docs/lucene-sandbox/snowball/</title>
        <p>For the submitted runs we used the title and descriptor topic fields, which were also mandatory. We applied
pseudo-relevance feedback for all tasks. For runs involving Russian, we also created one run without BRF. To
translate the queries we used the internet service freetranslation.com3 which provided some surprisingly good
translations from English to Russian. The results can be seen in Table 2.</p>
        <p>Runs
UHImlt1
UHImlt2
UHIenru1
UHIenru2
UHIenru3
UHIru1
UHIru2
UHIru3
UHIfi1
UHIfi2
For Finnish, the performance is quite high. The snowball stemmer works very well.</p>
        <p>For Russian, our results are very bad due to some encoding problems. BRF still worked well for Russian under
these circumstances. Also, the test runs had indicated that the Lucene stemmer seems very capable of dealing
with Russian and it held up to that expectation.</p>
        <p>The multilingual runs suffered severely from the obstacles that led to the bad results for Russian. We do also
have only limited insight into the usefulness of intertran.com as a translation tool for Finnish.</p>
      </sec>
      <sec id="sec-1-2">
        <title>3 http://www.freetranslation.com</title>
        <p>This year’s Russian tracks posed some challenges we could not easily overcome. Despite working with Java
only and unicode-based character sets, the Russian stopwords could not be eliminated. We did not have the
resources to work on a more sophisticated approach.
4</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Outlook</title>
      <p>Our system is far from well adapted the task. It has about 30 weighting parameters. This year, we could only
experiment with a few ones.</p>
      <p>
        Furthermore, in future years we intend to exploit the observed relation between the number of named entities in
topics and retrieval performance
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">(cf. Mandl &amp; Womser-Hacker 2004b)</xref>
        .
      </p>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgements References</title>
      <p>We would like to thank the Jakarta and Apache projects’ teams for sharing Lucene with a wide community as
well as the providers of Snowball. Furthermore, we acknowledge the work of several students from the
University of Hildesheim who implemented MIMOR as part of their course work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Hackl</surname>
          </string-name>
          ,
          <string-name>
            <surname>René</surname>
          </string-name>
          (
          <year>2004</year>
          ):
          <article-title>Multilinguales Information Retrieval im Rahmen von CLEF 2003</article-title>
          .
          <source>Master Thesis</source>
          , University of Hildesheim, Information Science.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Hackl</surname>
          </string-name>
          , René; Kölle, Ralph; Mandl, Thomas; Ploedt, Alexandra; Scheufen,
          <string-name>
            <surname>Jan-Hendrik;</surname>
          </string-name>
          Womser-Hacker,
          <source>Christa</source>
          (
          <year>2004</year>
          )
          <article-title>: Multilingual Retrieval Experiments with MIMOR at the University of Hildesheim</article-title>
          . To appear in: Peters, Carol; Braschler, Martin; Gonzalo, Julio; Kluck, Michael (eds.):
          <article-title>Evaluation of CrossLanguage Information Retrieval Systems</article-title>
          .
          <source>Proceedings of the CLEF 2003 Workshop</source>
          . Berlin et al.: Springer [Lecture Notes in Computer Science]
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Hackl</surname>
          </string-name>
          , René; Mandl, Thomas; Schwantner,
          <string-name>
            <surname>Michael</surname>
          </string-name>
          (
          <year>2004</year>
          ):
          <article-title>Evaluierung und Einsatz des open source VolltextRetrievalsystems Lucene am FIZ Karlsruhe</article-title>
          . In: Ockenfeld, Marlies (ed.):
          <source>Information Professional</source>
          <year>2011</year>
          : Strategien - Allianzen
          <source>- Netzwerke. Proceedings 26. DGI Online-Tagung. Frankfurt a.M. 15</source>
          .-
          <fpage>17</fpage>
          . Juni. pp.
          <fpage>147</fpage>
          -
          <lpage>153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Mandl</surname>
          </string-name>
          , Thomas; Womser-Hacker,
          <article-title>Christa (2004a): A Framework for long-term Learning of Topical User Preferences in Information Retrieval</article-title>
          . In: New Library World. vol.
          <volume>105</volume>
          (
          <issue>5</issue>
          /6). pp.
          <fpage>184</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Mandl</surname>
          </string-name>
          , Thomas; Womser-Hacker,
          <article-title>Christa (2004b): Analysis of Topic Features in Cross-Language Information Retrieval Evaluation</article-title>
          .
          <source>In: 4th International Conference on Language Resources</source>
          and
          <string-name>
            <surname>Evaluation (LREC) Lisbon</surname>
          </string-name>
          , Portugal, May
          <volume>24</volume>
          -30. Workshop Lessons Learned from Evaluation:
          <article-title>Towards Transparency and Integration in Cross-Lingual Information Retrieval (LECLIQ)</article-title>
          . pp.
          <fpage>17</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Womser-Hacker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>1997</year>
          )
          <article-title>: Das MIMOR-Modell</article-title>
          .
          <article-title>Mehrfachindexierung zur dynamischen Methoden-ObjektRelationierung im Information Retrieval</article-title>
          . Habilitationsschrift. Universität Regensburg, Informationswissenschaft.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>