<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>MediaLab BV</institution>
          ,
          <addr-line>Schellinkhout</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Peter van der Weerd</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This report describes the participation of MediaLab BV in the CLEF-2003 evaluations. This year we participated in the monolingual Dutch task, experimenting with a keyword disambiguation tool. MediaLab developed this tool to exploit human assigned keywords in the search engine in a better way than just blind searching with the keywords themselves. Although this tool was not planned to be used for CLEF-like applications it was fun to check if it could help boosting the search quality.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Disambiguation</title>
    </sec>
    <sec id="sec-2">
      <title>2. Approach</title>
      <p>The CLEF data-collection contains some extra fields with keyword information (HTR) and with geographical
information (GEO). We used both extra fields to feed the disambiguation tool. The searching process is done in
the following way:
• doing a “normal” search giving result R1
• determine the top5 of disambiguation items and search all these items giving result R2
• than recomputed the weights in R1 by adding a fraction of the weights in R2
• the modified R1 is used as the submission
We submitted a base run, and for each field (HTR and GEO) we submitted 5 runs with different relative weights
of the second result.</p>
      <p>HTR-5 means: normal result combined with a 50% weight of the HTR-result.</p>
      <p>HTR-2 means: normal result combined with a 20% weight of the HTR-result.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>The following table summarizes some measures of the runs.</p>
      <p>Rel_ret
It is clear that blind boosting the results with HTR or GEO data helps a little bit to retrieve more relevant
documents. In case of boosting the results by the HTR data the optimum is about 10% to 20%, increasing the
average precision with 3%. However, the effect is rather small.</p>
      <p>The profit of boosting by the GEO data is less convincing, probably caused by the quality of the GEO-data.
Looking at the data made us already hesitate about using it.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
    </sec>
    <sec id="sec-5">
      <title>5. References</title>
      <p>
        Although MediaLab’s disambiguation tool was not intended for blind boosting search results, it might be used
for it. Probably better results are achieved by using these fields in a normal blind relevance feedback procedure
as used McNamee and Mayfield and others [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>APL</given-names>
            <surname>Experiments at</surname>
          </string-name>
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          :
          <article-title>Translation Resources and Score Normalization, Paul McNamee</article-title>
          and
          <string-name>
            <given-names>James</given-names>
            <surname>Mayfield</surname>
          </string-name>
          , Johns Hopkins University, USA, in [1]
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>