<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Robust Retrieval Experiments at the University of Hildesheim</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ben Heuwing</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Mandl</string-name>
          <email>mandl@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Multilingual Retrieval</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robust Retrieval</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Evaluation Measures</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Measurement</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Hildesheim</institution>
          ,
          <addr-line>Information Science Marienburger Platz 22 D-31141 Hildesheim</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper reports on experiments submitted for the robust task at CLEF 2007. We applied a system previously tested for ad-hoc retrieval. Experiments were focused on the effect of blind relevance feedback and named entities. Experiments for monolingual English and French are presented. Categories and Subject Descriptors H.3 [Information Storage and Retrieval]: H.3.1 Content Analysis and Indexing; H.3.3 Information Search and Retrieval; H.3.4 Systems and Software System Setup Five runs for the English and three for the French monolingual data were submitted. The results for both test and training topics are shown in table 1 and 2, respectively. Optimization of the Blind Feedback parameters on the English training topics of 2006 showed the best results when the query was expanded with 30 Terms from the top10 documents and the query-expansion was given a relative weight of 0.05 compared to the rest of the query. The same improvements (compared to the base run) can be seen on a smaller scale for the submitted runs. The use of Named Entities from the English documents did not have an effect on the retrieval quality. For the French runs the use of a heavy-weighted (equal to the rest of the query) query expansion with 50 terms from the best 5 documents came out as the best Blind Relevance parameters - even though for the training topics the base run performed better.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Introduction
We intended to provide a base line for the robust task at CLEF 2007. Our basic system was used at CLEF
campaigns previously (Hackl et al. 2005).</p>
      <p>For the base line experiments, we optimized blind relevance feedback (BRF) parameters. The underlying basic
retrieval engine of the system is the open source search engine Apache Lucene.</p>
    </sec>
    <sec id="sec-2">
      <title>HiMoEnBase</title>
      <p>HiMoEnBrf1
HiMoEnBrf2
HiMoEnBrfNe
HiMoEnNe
HiMoFrBase
HiMoFrBrf
HiMoFrBrf2</p>
      <sec id="sec-2-1">
        <title>Language</title>
      </sec>
      <sec id="sec-2-2">
        <title>Stemming NE MAP</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>English</title>
      <p>English
English
English
English
French
French
French
snowball
snowball
snowball
snowball
snowball
lucene
lucene
lucene</p>
      <p>BRF
(weight-docs-terms)</p>
      <p>1.0-10-30
0.05-10-30
0.05-10-30
0.5-5-25
1.0-5-50
1.0
2.0
snowball
snowball
snowball
snowball
snowball
lucene
lucene
lucene</p>
      <p>BRF
(weight-docs-terms)</p>
      <p>1.0-10-30
0.05-10-30
0.05-10-30
0.5-5-25
1.0-5-50
0.1634
0.1489
0.1801
0.1801
0.1634
0.2081
0.2173
0.2351
Only the runs for French have reached a competitive level of above 0.2 MAP. The results for the geometric
average for the English topics are worse, because low performance for several topics leads to a sharp drop in the
performance according to this measure.
3</p>
      <p>
        Future Work
For future experiments, we intend to exploit the knowledge on the impact of named entities on the retrieval
process
        <xref ref-type="bibr" rid="ref2">(Mandl &amp; Womser-Hacker 2005)</xref>
        as well as selective relevance feedback strategies in order to improve
robustness.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Mandl</surname>
          </string-name>
          , Thomas; Hackl, René; Womser-Hacker,
          <source>Christa</source>
          (
          <year>2007</year>
          )
          <article-title>: Robust Ad-hoc Retrieval Experiments with French</article-title>
          and English at the University of Hildesheim. In: Peters, Carol et al. (Eds.).
          <source>7th Workshop of the Cross-Language Evaluation Forum</source>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2006</year>
          , Alicante, Spain,
          <source>Revised Selected Papers</source>
          . Berlin et al.:
          <source>Springer [Lecture Notes in Computer Science</source>
          <volume>4730</volume>
          ] pp.
          <fpage>127</fpage>
          -
          <lpage>128</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Mandl</surname>
          </string-name>
          , Thomas; Womser-Hacker,
          <source>Christa</source>
          (
          <year>2005</year>
          ):
          <article-title>The Effect of Named Entities on Effectiveness in Cross-Language Information Retrieval Evaluation</article-title>
          .
          <source>In: Applied Computing 2005: Proc. ACM SAC Symposium on Applied Computing (SAC)</source>
          .
          <article-title>Information Access and Retrieval (IAR) Track</article-title>
          . Santa Fe, New Mexico, USA. March 13.-
          <fpage>17</fpage>
          .
          <year>2005</year>
          . pp.
          <fpage>1059</fpage>
          -
          <lpage>1064</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Peters</surname>
          </string-name>
          , Carol et al. (
          <year>2007</year>
          )
          <article-title>: Overview of the ad-hoc Track at CLEF</article-title>
          . In this volume.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>