<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CLEF 2005: Domain-Specific Track Overview</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michael Kluck</string-name>
          <email>michael.kluck@swp-berlin.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maximilian Stempfhuber</string-name>
          <email>stempfhuber@iz-soz.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Informationszentrum Sozialwissenschaften (IZ).</institution>
          <addr-line>Bonn</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Stiftung Wissenschaft und Politik (SWP), German Institute for International and Security Affairs</institution>
          ,
          <addr-line>Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Sub-task Multi-lingual Bilingual X → DE Bilingual X → EN Bilingual X → RU Monolingual DE Monolingual EN Monolingual RU Sum</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The domain-specific track aims at mono- and cross-language information retrieval on structured scientific data. This track studies retrieval in a domain-specific context using two social science databases: The German Indexing and Retrieval Testdatabase (GIRT) (forth version GIRT-4: German/English pseudo-parallel corpus with identical documents) with 302,638 documents in total, and the Russian Social Science Corpus (RSSC) with 94,581 documents. Different sub-tasks have been available: 1. Monolingual task: German topics against German data GIRT4-DE, English topics against English data GIRT4-EN, Russian topics against Russian data RSSC; 2. Bilingual task: German topics against English data GIRT4-EN, German topics against Russian data RSSC, English topics against German data GIRT4-DE, English topics against Russian data RSSC, Russian topics against German data GIRT4-DE, Russian topics against English data GIRT4-EN; 3. Multilingual task: German topics against all data GIRT4-DE, GIRT4-EN, RSSC, English topics against all data GIRT4-DE, GIRT4-EN, RSSC, Russian topics against all data GIRT4-DE, GIRT4-EN, RSSC. The domain-specific task attracted 8 participating groups (three of them from Berkeley), which delivered a total of 76 runs: 40 monolingual runs, 33 bilingual runs and 3 multilingual runs. For detailed figures see the following table: Remarks and figures on the assessment process will conclude the overview of the domain-specific task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p># Participants
1
5
4
3
6
6
5
8
# Runs
3
15
13
5
17
15
8
76</p>
      <p>Topic Language
DE 1; EN 1; RU 1
EN 14; RU 1
DE 7; RU 6
DE 2; EN 3
The retrieval systems explicitly mentioned by the participants are based on logistic regression or OKAPI
formula. For translation purposes several MT systems have been used: L+H, SYSTRAN, PROMT, WorldLingo,
IMTranslator, FreeTranslation, Eurodictautom. Some groups concentrated on data fusion aspects. Most groups
used the thesaurus information provided with GIRT to translate queries or to produce a translation vocabulary.
Linguistic treatment reached from the use of stemmers, POS, de-compounding to the extraction of semantically
related concepts or WordNet concepts. Some groups did experiments with blind relevance feedback.
Concerning the results some groups were emphasizing the importance of robustness of the used methodology
and of high-quality results on a per query basis rather than high average precision computed over all queries.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>