<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Domain-Specific Track CLEF 2006: Overview of Results and Approaches, Remarks on the Assessment Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maximilian Stempfhuber</string-name>
          <email>stempfhuber@iz-soz.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Baerisch</string-name>
          <email>baerisch@iz-soz.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Informationszentrum Sozialwissenschaften (IZ)</institution>
          ,
          <addr-line>Lennéstrasse 30, 53113 Bonn</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The CLEF domain-specific track uses databases in different languages from the social science domain as the basis for retrieval of documents relevant to a user's query. Predefined topics simulate the user's information needs and are used by the research groups participating in this track to generate actual queries. The documents are structured into individual elements so that they can be used for optimizing the search strategies. One type of the retrieval task was on crosslanguage retrieval, the finding of information using queries in a different language than the actual language of the documents. English, German and Russian were used as languages for queries and documents. Queries therefore had to be translated from one of the languages to one or more of the other languages. Besides the cross-language tasks also monolingual retrieval tasks were carried out were query and document are in the same language. The focus here was to map the query onto the language and internal structure of the documents. This paper gives an overview of the domain-specific track and reports on noteworthy trends in approaches and results as well as on the topic creation and assessment process.</p>
      </abstract>
      <kwd-group>
        <kwd>Information Retrieval</kwd>
        <kwd>Evaluation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The CLEF domain-specific task is focused on cross-language retrieval in English, German and Russian.
Queries as well as documents are taken from the social science domain and are provided in all three languages.
This section presents the domain-specific task with the focus on the corpora used. It gives an overview on the
2006 track. The following sections are on the methods used and results achieved in the retrieval and on the
processes of topic-creation and assessment. The paper concludes with an outlook on the future development of
the domain-specific track.
•</p>
      <sec id="sec-1-1">
        <title>Sub-task</title>
      </sec>
      <sec id="sec-1-2">
        <title>Multi-lingual</title>
      </sec>
      <sec id="sec-1-3">
        <title>Bilingual X → DE</title>
      </sec>
      <sec id="sec-1-4">
        <title>Bilingual X → EN</title>
      </sec>
      <sec id="sec-1-5">
        <title>Bilingual X → RU</title>
      </sec>
      <sec id="sec-1-6">
        <title>Monolingual DE</title>
      </sec>
      <sec id="sec-1-7">
        <title>Monolingual EN</title>
      </sec>
      <sec id="sec-1-8">
        <title>Monolingual RU</title>
      </sec>
      <sec id="sec-1-9">
        <title>Total</title>
        <p>Queries in the CLEF are provided as topics. Each topic is the representation of an information need form the
social science domain. Topics as well as documents are structured with the topic including a title, a short
description of the information need and a longer narrative for further clarification. The documents, among other
information, include fields such as author, title and year of publication, terms from a thesaurus and an abstract:</p>
        <p>The participating groups submits one or more runs for each subtask. A run includes the 1000 highest ranked
document for each of the 25 topics. The runs from all groups for a given subtask are pooled, the highest ranking
document are then assessed by domain experts in order to obtain a reference on the relevance of the documents.
Overview of the 2006 Domain Specific Track
2.3</p>
        <sec id="sec-1-9-1">
          <title>Methods and Results Overview</title>
          <p>Details on the results and employed methods of all groups are given in the corresponding chapters of this
volume. This chapter gives only a brief overview of the methods employed. The University of Hagen used a
reranking approach based on antagonistic terms, improving the mean average precision. The Chemnitz Technical
University used Apache Lucene as the basis for the retrieval and achieved the best results with a combination of
suffix stripping, stemming, and decompounding. The University of California, Berkeley, did not use methods for
thesaurus-based query expansion and de-compounding as they were employed before. The University of
Neuchatel used DFR GL2 and Okapi with several combinations of the fields included in a topic against the
different fields of GIRT4 documents.
2.4</p>
        </sec>
        <sec id="sec-1-9-2">
          <title>Topics</title>
          <p>The team responsible for the topic-creation process changed in 2006. This has lead to changes in the
topiccreation process itself as well. In order to produce a large number of potential information needs covering a wide
area of the social sciences, a number of domain experts familiar with the GIRT corpus were asked for
contributions to the topics. 25 topics were selected from 42 submissions. These topics were compared to the
topics from the years 2001-2005 in order to remove topics too similar to topics already used. Removed topics
were replaced from the pool of previously unused topics. The definition of topics by domain experts with
knowledge about the corpus resulted in many topics of interest but also caused more effort in the selection of the
actual topic-set. This was due to the fact that not all submitters of topics had detailed knowledge of the CLEF
topic-creation rules and corrections had to be made afterwards to provide consistent structure and wording. For
future CLEF-campaigns, this process will be further optimized.</p>
          <p>The definition of topics suitable for the GIRT corpus as well as the Russian INION corpus remains
challenging. Lack of Russian language skills in the topic-creation team increases the difficulty of creating a set
of topics well suited for both corpora. An additional factor in 2006 was lack of experience with the INION
corpus. Also in this regard, we hop to improve the topic cretation process in the next campaign as the
multilingual corpus is now stable again.
2.5</p>
        </sec>
        <sec id="sec-1-9-3">
          <title>Assessment Process</title>
          <p>As in 2005 the assessment for English and German were done using a Java Swing program while our partner
from the Research Computing Center of the M.V.Lomonosov Moscow State University used the CLEF
assessment tool for the assessment of the Russian documents. The decision to use this software for English and
German was made based on the fact that the assessments of English and German documents from the
pseudoparallel corpus should be compared to each other. An interesting finding was the statement of both assessors that
they liked the ability to choose between keyboard-shortcuts and a mouse driven interface since it allowed for
different working styles in the otherwise monotonous assessment task. Figures 1a and 1b show the number of
documents in the pool for each topic and the percentage of relevant documents. It should be noted that, while the
number of documents found for each topic is relatively stable, the percentage of relevant documents shows
outliers with the topics number 167 and 169. Investigation on this outliers and discussion with the assessors
showed that the outliers were caused by complex exclusion criteria in case of topic 167 and the use of the
complex concepts of gender specific education and the specifics of the German primary education system in case
of topic 169.</p>
          <p>Assessment Results: Number of Documents per Topic for GIRT4
1000
900
800
700
600
500
400
300
200
100
0</p>
          <p>917
533</p>
          <p>506
344
568
428
500
333
681
584
854
450
495
420</p>
          <p>DE
591
486
620</p>
          <p>585
466</p>
          <p>409
151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175
Figure 1a: Number of documents per topic
80
70
60
50
40
30
20
10
0</p>
          <p>Assessment Results: Percentage of relevant documents
151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Results</title>
      <p>
        The detailed information on recall and precision for the participants and their runs can be found in the groups
chapters in this volume [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2,3,4,5</xref>
        ] and shall not be repeated here in full. The Recall-Precision graph for the tracks
which attracted the most participants can be found in the figures 2, 3 and 4.
      </p>
      <p>Top 4 ParticipantsDS Monolingual German Recall vsPrecision</p>
      <p>Neuchatel-DS-MONO-DECLEF2006_unine_uninede3
Hagen-DS-MONO-DECLEF2006_hagen_fuhggyynbfl500r
Chemnitz-DS-MONO-DECLEF2006_tuchemniz_tucmigirtde4
Berkeley-DS-MONO-DE</p>
      <p>CLEF2006_berkeley_berk_mo_de_t2fb
120
100
80
60
40
20
0
0
10
20
30
40
50
60
70
80
90
100</p>
    </sec>
    <sec id="sec-3">
      <title>Outlook</title>
      <p>Several methods are being investigated to improve the quality of the topic-creation and assessment process.
In Order to create topics covering a broader scope of the social sciences as well as realistic information needs it
is planned to involve more domain experts into the topic creation process. These experts will be asked for topic
proposals in the form of narratives. The most suitable topics from the resulting pool will then be selected and
produced as topics conforming to the topic creation guidelines. In order to better verify the topics against the
INION and GIRT4 corpora, a new retrieval system will be used which will contain both corpora (deployment
planned for autumn of 2006) and will allow also our Russian partner from the Research Computing Center of the</p>
      <sec id="sec-3-1">
        <title>M.V.Lomonosov Moscow State University to better participate in the topic creation process.</title>
        <p>Concerning the assessment process several methods are evaluated in order to allow the assessors to better
coordinate their assessment criteria. Proposed methods include breaking the assessment und reassessment
processes into smaller batches of 5 topics or the notification of the assessors in case of notable differences in the
results for a topics.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>We greatly acknowledge the partial funding by the European Commission through the DELOS Network of
Excellence on Digital Libraries. Natalia Lookachevitch from the Research Computing Center of the
M.V.Lomonosov Moscow State University helped in acquiring the Russion corpus and organized the translation
of topics into Russian and the assessment process for the Russian corpus. Claudia Henning and Jeof Spiro did
the German and English assessments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kluck</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The GIRT Data in the Evaluation of CLIR Systems - from 1997 until 2003</article-title>
          . In: Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Kluck</surname>
          </string-name>
          , M. (eds.):
          <source>Comparative Evaluation of Multilingual Information Access Systems. 4th Workshop of the Cross-Language Evaluation Forum</source>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2003</year>
          , Trondheim, Norway,
          <source>August 21-22</source>
          ,
          <year>2003</year>
          ,
          <source>Revised Selected Papers. Lecture Notes in Computer Science</source>
          , Vol.
          <volume>3237</volume>
          . Springer-Verlag Berlin Heidelberg New York,
          <fpage>379</fpage>
          -
          <lpage>393</lpage>
          , (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kluck</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The Domain-Specific track in CLEF 2004: Overview of the Results and Remarks on the Assessment Process</article-title>
          . In: Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.J.F.</given-names>
            ,
            <surname>Kluck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Magnini</surname>
          </string-name>
          ,
          <string-name>
            <surname>B</surname>
          </string-name>
          . (eds.):
          <article-title>Multilingual Information Access for Text, Speech and Images. 5th Workshop of the Cross-Language Evaluation Forum</article-title>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2004</year>
          ,
          <article-title>Bath</article-title>
          , UK,
          <year>September 2004</year>
          ,
          <source>Revised Selected Papers. Lecture Notes in Computer Science</source>
          , Vol.
          <volume>3491</volume>
          . Springer-Verlag Berlin Heidelberg New York,
          <fpage>260</fpage>
          -
          <lpage>270</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Jens</given-names>
            <surname>Kürsten</surname>
          </string-name>
          , Maximilian Eibl:
          <article-title>Monolingual Retrieval Experiments with a Domain Specific Document Corpus at</article-title>
          the Chemnitz Technical University
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ray R. Larson</surname>
          </string-name>
          <article-title>: Domain Specic Retrieval: Back to Basics (this volume) (</article-title>
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Johannes Leveling: University of Hagen at CLEF2006:
          <article-title>Reranking documents for the domain-specific task (this volume) (</article-title>
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Jacques</given-names>
            <surname>Savoy</surname>
          </string-name>
          , Samir Abdou: UniNE at CLEF 2006:
          <article-title>Experiments with Monolingual, Bilingual, Domain-Specific and Robust Retrieval (this volume) (</article-title>
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>