<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UniNE at Domain-Specific IR - CLEF 2008: Scientific Data Retrieval: Various Query Expansion Approaches</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claire Fautsch</string-name>
          <email>Claire.Fautsch@unine.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ljiljana Dolamic</string-name>
          <email>Ljiljana.Dolamic@unine.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jacques Savoy</string-name>
          <email>Jacques.Savoy@unine.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Experimentation</institution>
          ,
          <addr-line>Performance, Measurement, Algorithms</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Natural Language Processing with European Languages</institution>
          ,
          <addr-line>Digital Libraries, German Language, Russian Language</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Neuchatel</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Our first objective in participating in this domain-specific evaluation campaign is to propose and evaluate various indexing and search strategies for the German, English and Russian languages, in an effort to obtain better retrieval effectiveness than that of the language-independent approach (n-gram). To do so we evaluate the GIRT-4 test-collection using the Okapi, various IR models derived from the Divergence from Randomness (DFR) paradigm, the statistical language model (LM) together with the classical tf.idf vector-processing scheme.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Domain-specific retrieval is an interesting task, one in which we access bibliographic notices (usually
composed of a title and an abstract) extracted from two German social science sources and one Russian corpus. The
records in these notices also contain manually assigned keywords extracted from a controlled vocabulary by
librarians who are knowledgeable of the discipline to which the indexed articles belong. These descriptors should
be helpful in improving document surrogates and consequently the extraction of more pertinent information, while
also discarding irrelevant abstracts. Access to the underlying thesaurus would also improve retrieval performance.</p>
      <p>The rest of this paper is organized as follows: Section 2 describes the main characteristics of the GIRT-4
(written in the German and English languages) and ISISS (Russian) test-collections. Section 3 outlines the main
aspects of our stopword lists and light stemming procedures, along with the IR models used in our experiments.
Section 4 explains different blind query expansion approaches and evaluates their use with the available corpora.
Section 5 provides our official runs and results.</p>
      <p>
        In the domain-specific retrieval task, the two available corpora are composed of bibliographic records extracted
from various sources in the social sciences domain. Typical records (see Figure 1 for a German example) in this
corpus consist of a title (tag &lt;TITLE-DE&gt;), author name (tag &lt;AUTHOR&gt;), document language (tag
&lt;LANGUAGE-CODE&gt;), publication date (tag &lt;PUBLICATION-YEAR&gt;) and abstract (tag &lt;ABSTRACT-DE&gt;).
Manually assigned descriptors and classifiers are provided for all documents. An inspection of this German corpus
reveals that all bibliographic notices consist of a title and 96.4% of them include an abstract. In addition to this
information provided by the author, a typical record contains on average 10.15 descriptors
(“&lt;CONTROLLED-TERM-DE&gt;”), 2.02 classification terms (“&lt;CLASSIFICATION-TEXT-DE&gt;”), and 2.42
methodological terms (“&lt;METHOD-TEXT-DE&gt;“ or “&lt;METHOD-TERM-DE&gt;“). The manually assigned descriptors
are extracted from the controlled list known as the “Thesaurus for the Social Sciences”. Finally, associated with
each record is a unique identifier (“&lt;DOCNO&gt;”).
        <xref ref-type="bibr" rid="ref8">Kluck (2004)</xref>
        provides a more complete description of this
corpus.
      </p>
      <p>&lt;DOC&gt;
&lt;DOCNO&gt; GIRT-DE19909343
&lt;TITLE-DE&gt; Die sozioökonomische Transformation einer Region : Das Bergische Land von 1930 bis 1960
&lt;AUTHOR&gt; Henne, Franz J.
&lt;AUTHOR&gt; Geyer, Michael
&lt;PUBLICATION-YEAR&gt; 1990
&lt;LANGUAGE-CODE&gt; DE
&lt;CONTROLLED-TERM-DE&gt; Rheinland
&lt;CONTROLLED-TERM-DE&gt; historische Entwicklung
&lt;CONTROLLED-TERM-DE&gt; regionale Entwicklung
&lt;CONTROLLED-TERM-DE&gt; sozioökonomische Faktoren
&lt;METHOD-TERM-DE&gt; historisch
&lt;METHOD-TERM-DE&gt; Aktenanalyse
&lt;CLASSIFICATION-TEXT-DE&gt; Sozialgeschichte
&lt;ABSTRACT-DE&gt; Die Arbeit hat das Ziel, anhand einer regionalen Studie die Entstehung des "modernen"
fordistischen Wirtschaftssystems und des sozialen Systems im Zeitraum zwischen 1930 und 1960 zu
beleuchten; dabei geht es auch um das Studium des "Sozial-imaginären", der Veränderung von Bewußtsein und
Selbst-Verständnis von Arbeitern durch das Erlebnis und die Erfahrung der Depression, des
Nationalsozialismus und der Nachkriegszeit, welches sich in den 1950er Jahren gemeinsam mit der
wirtschaftlichen Veränderung zu einem neuen "System" zusammenfügt.
&lt;DOC&gt; …</p>
      <p>The above-mentioned German collection was translated into British English, mainly by professional translators
whose native language was English. Included in all English records is a translated title (listed under “&lt;TITLE-EN&gt;”
in Figure 2), manually assigned descriptors (“&lt;CONTROLLED-TERM-EN&gt;”), classification terms
(“&lt;CLASSIFICATION-TEXT-EN&gt;”) and methodological terms (“&lt;METHOD-TERM-EN&gt;”). Abstracts however
were not always translated (in fact they are available for only around 15% of the English records).</p>
      <p>In addition to this bilingual corpus, we may also access the GIRT thesaurus, containing 10,623 entries (all
including both the &lt;GERMAN&gt; and &lt;GERMAN-CAPS&gt;) tags together with 9,705 English translations. It also
contains 2,947 &lt;BROADER-TERM&gt; relationships and 2,853 &lt;NARROWER-TERM&gt; links. The synonym
relationship between terms is expressed through &lt;USE-INSTEAD&gt; (2,153) links, &lt;RELATED-TERM&gt; (1,528) or
&lt;USE-COMBINATION&gt; (3,263).</p>
      <p>As a third language, we access bibliographic records written in the Russian language composed of the ISISS
(Russian Economic and Social Science) bibliographic data collection (see Figure 3 for an example of a record
extracted from the Russian collection). Using a pattern similar to that of the other two corpora, records include a
title (“&lt;TITLE-RU&gt;” in Figure 3), sometimes an abstract (“&lt;ABSTRACT-RU&gt;”), and certain manually assigned
descriptors (“&lt;KEYWORDS-RU&gt;”).</p>
      <p>During the indexing process, we retained all pertinent sections in order to build document representatives.
Additional information such as author name, publication date and the language in which the bibliographic notice
was written are of less importance, particularly from an IR perspective, and thus they will be ignored in our
experiments.</p>
      <p>As shown in Appendix 2, the available topics cover various subjects (e.g., Topic #206: “Environmental
justice,” Topic #209: “Doping and sports,” Topic #221: “Violence in schools,” or Topic #211: “Shrinking cities”),
and some of them may cover a relative large domain (e.g. Topic #212: “Labor market and migration”).</p>
      <sec id="sec-1-1">
        <title>3.1 Indexing and IR Models</title>
        <p>
          For the English, German and Russian language, we used the same stopword lists and stemmers that we selected
for our previous CLEF participation
          <xref ref-type="bibr" rid="ref5">(Fautsch et al., 2008)</xref>
          . Thus for English it was the SMART stemmer and
stopword list (containing 571 items), while for the German we apply our light stemmer (available at
http://www.unine.ch/info/clef/) and stopword list (603 words). For all our German experiments we also apply our
decompounding algorithm
          <xref ref-type="bibr" rid="ref12">(Savoy, 2004)</xref>
          . For the Russian language, the stopword list contains 430 words and we
apply our light stemming procedure (based on 53 rules to remove the final suffix representing gender (masculine,
feminine, and neutral), number (singular, plural) and the six Russian grammatical cases (nominative, accusative,
genitive, dative, instrumental, and locative)).
        </p>
        <p>In order to obtain a broader view of the relative merit of various retrieval models, we may first adopt the
classical tf idf indexing scheme. In this case, the weight attached to each indexing term in a document surrogate (or
in a query) combines the term's occurrence frequency (denoted tfij for indexing term tj in document Di) and also the
inverse document frequency (denoted idfj).</p>
        <p>
          In addition to this vector-processing model, we may also consider probabilistic models such as the Okapi model
(or BM25)
          <xref ref-type="bibr" rid="ref10">(Robertson et al., 2000)</xref>
          . As a second probabilistic approach, we may implement four variants of the
DFR (Divergence from Randomness) family suggested by
          <xref ref-type="bibr" rid="ref2">Amati &amp; van Rijsbergen (2002</xref>
          ). In this framework, the
indexing weight wij attached to term tj in document Di combines two information measures as follows.
        </p>
        <p>wij = Inf1ij · Inf2ij = –log2[Prob1 ij(tf)] · (1 – Prob2ij(tf))</p>
        <sec id="sec-1-1-1">
          <title>The first model PB2 is based on the following equations:</title>
          <p>Prob1ij = (e-λj · λtfij) / tfij!</p>
          <p>with λj = tcj / n
Prob2ij = 1- [(tcj+1) / (df j · (tfnij+1))]</p>
          <p>with tfnij = tfij · log2[1 + ((c · mean dl) / li)
where tcj represents the number of occurrences of term tj in the collection, dfj the number of documents in which the
term tj appears, and n the number of documents in the corpus. Moreover, c and mean dl (average document length)
are constants whose values are given in the Appendix 1.</p>
        </sec>
        <sec id="sec-1-1-2">
          <title>The second model GL2 is defined as:</title>
          <p>Prob1ij = [1 / (1+λj)] · [λj / (1+λj)]tfnij
Prob2ij = tfnij / (tfnij + 1)</p>
          <p>Inf1ij = tfnij · log2[(n+1) / (dfj+0.5)]</p>
          <p>For the third model I(n)B2, we still use Equation 2 to compute Prob2ij but the implementation of Inf1ij is
modified as:</p>
          <p>For the fourth model I(ne)C2 the initial value of Prob2ij is obtained from Equation 2 and for the value Inf1ij we
use:</p>
          <p>Inf1ij = tfnij · log2[(n+1) / (ne+0.5)]</p>
          <p>with ne = n · [1 - [(n-1) / n]tcj]</p>
          <p>
            Finally, we also consider an approach based on a statistical language model (LM)
            <xref ref-type="bibr" rid="ref6">(Hiemstra 2000; 2002)</xref>
            ,
known as a non-parametric probabilistic model (both Okapi and DFR are viewed as parametric models). Thus, the
probability estimates would not be based on any known distribution (as in Equations 1, or 3), but rather be
estimated directly based on the occurrence frequencies in document D or corpus C. Within this language model
(LM) paradigm, various implementations and smoothing methods might be considered, and in this study we adopt
a model proposed by
            <xref ref-type="bibr" rid="ref7">Hiemstra (2002)</xref>
            as described in Equation 7, which combines an estimate based on document
(P[tj | Di]) and on corpus (P[tj | C]) (Jelinek-Mercer smoothing method).
          </p>
          <p>P[Di | Q] = P[Di] . ∏tj∈Q [λj . P[tj | Di] + (1-λj) . P[tj | C]]
with P[tj | Di] = tfij/li and P[tj | C] = dfj/lc
with lc = ∑k dfk
where λj is a smoothing factor (constant for all indexing terms tj, and usually fixed at 0.35) and lc an estimate of the
size of the corpus C.</p>
        </sec>
      </sec>
      <sec id="sec-1-2">
        <title>3.2 Overall Evaluation</title>
        <p>To measure the retrieval performance, we adopted the mean average precision (MAP) (computed on the basis of
1,000 retrieved items per request by the new TREC-EVAL program). In the following tables, the best performances
under the given conditions (with the same indexing scheme and the same collection) are listed in bold type.</p>
        <p>From this table, we can see that when using word-based indexing, the DFR I(ne)B2 or the LM models tend to
perform the best. With the 4-gram indexing approach, the LM model always presents the best performing schemes.
The short query formulation (T) tends to produce a better retrieval performance than medium (TD) topic
formulation. As shown in the last line, when comparing the word-based and 4-gram indexing systems, the relative
difference is seen to be rather short (around 4.6%) and favors the 4-gram approach.</p>
        <p>Using our evaluation approach, evaluation differences occur when comparing with values computed according
to the official measure (the latter always takes 25 queries into account).</p>
        <sec id="sec-1-2-1">
          <title>Query type</title>
          <p>Indexing / stemmer
IR Model
DFR GL2
DFR I(ne)B2
LM (λ=0.35)
Okapi
tf idf
Mean
% change over T
over stemming</p>
          <p>
            To provide a better match between user information needs and documents, various query expansion techniques
have been suggested. The general principle is to expand the query using words or phrases having similar meanings
to, or related to those appearing in the original request. To achieve this, query expansion approaches consider
various relationships between these words, along with term selection mechanisms and term weighting schemes.
Specific answers regarding the best technique may vary, thus leading to a variety of query expansion approaches
            <xref ref-type="bibr" rid="ref4">(Efthimiadis, 1996)</xref>
            .
          </p>
          <p>
            In our first attempt to find related search terms, we might ask the user to select additional terms to be included in
an expanded query. This could be handled interactively through displaying a ranked list of retrieved items returned
by the first query. As a second strategy,
            <xref ref-type="bibr" rid="ref11">Rocchio (1971)</xref>
            proposed taking the relevance or non-relevance of
top-ranked documents into account, as indicated manually by the user. In this case, a new query would then be built
automatically in the form of a linear combination of the term included in the previous query and terms automatically
extracted from both relevant (with a positive weight) and non-relevant documents (with a negative weight).
Empirical studies have demonstrated that such an approach is usually quite effective.
          </p>
          <p>
            As a third technique,
            <xref ref-type="bibr" rid="ref3">Buckley et al. (1996)</xref>
            suggested that even without looking at them or asking the user, it
could be assumed that the top-k ranked documents would be relevant. This method, denoted as the
pseudo-relevance feedback or blind-query expansion approach does not require user intervention. Moreover, using
the MAP as performance measure is a strategy that usually tends to enhance performance measures.
          </p>
          <p>In the current context, we used Rocchio’s formulation (denoted “Rocchio”) with α = 0.75, β = 0.75, whereby
the system was allowed to add m terms extracted from the k best ranked documents from the original query. For the
German corpus (Table 4, third column), such a search technique does not seem to enhance the MAP. For the
English collection (Table 5, second and third column), Rocchio’s blind query expansion may improve the MAP
from +9.3% (DFR PB2, 0.3101 vs. 0.3392) or hurt the retrieval performance -8.72% (Okapi model, 0.3039 vs.
0.2774). For the Russian language (Table 6, second and forth column), blind query expansion improves the MAP
(e.g., +28.98% with the Okapi model, 0.1740 vs. 0.1349 or +2.3% with the DFR I(ne)B2 model, 0.1503 vs. 0.1468).</p>
        </sec>
        <sec id="sec-1-2-2">
          <title>Mean average precision</title>
          <p>German German</p>
          <p>Rocchio idf
DFR I(n)B2 0.4179 DFR I(n)B2 0.4179
5/70 0.3965 5/70 0.4120
10/100 0.3965 10/100 0.4025
10/200 0.3992 10/200 0.4104</p>
        </sec>
        <sec id="sec-1-2-3">
          <title>Mean average precision</title>
          <p>English English</p>
          <p>Rocchio idf
DFR PB2 0.3101 DFR PB2 0.3101
Rocchio's query expansion approach however does not always significantly improve the MAP. Such a query
expansion approach is based on term co-occurrence data and tends to include additional terms that occur very
frequently in the documents. In such cases, these additional search terms will not always be effective in
discriminating between relevant and non-relevant documents, and the final effect on retrieval performance could be
negative.</p>
          <p>
            As another pseudo-relevance feedback technique we may apply an idf-based approach (denoted “idf” in
following tables)
            <xref ref-type="bibr" rid="ref1">(Abdou &amp; Savoy, 2008)</xref>
            . In this query expansion scheme, the inclusion of new search terms is
based on their idf values, tending to enlarge the query with more infrequent terms. Overall this idf-based term
selection performs rather well and usually its retrieval performance is more robust.
          </p>
          <p>For example, with the Russian language (Table 6, third and fifth column), this idf-based blind query expansion
may also improve the MAP (e.g., +19.5% with the Okapi model, 0.1612) but, on the other hand, with the DFR
I(ne)B2 model, the MAP is slightly reduced (-2.3% from 0.1468 to 0.1433).</p>
          <p>However, the idf-based query expansion tends to include rare terms, without considering the context. Thus
among the top-k retrieved documents such a scheme may add terms appearing far away from where the search terms
occurred. The single selection criterion is based only on idf values, not the position of those additional terms in the
top-ranked documents. This year we investigated retrieval effectiveness when including a second criterion in the
selection of terms to be included in the new expanded query. We considered it to be important to expand the query
using terms appearing close to a search term (fixed at 10 indexing terms in the current experiments). This short
window includes 10 terms to the right and 10 terms to the left of each query term. This type of query expansion
method is denoted as “idf-window” in Table 7.</p>
          <p>Finally, to find words or expressions related to the current request, we considered using commercial search
engines (e.g., Google) or online encyclopedia (e.g., Wikipedia). In this case, we submitted a query containing the
short topic formulation (T or title-only) to each information service. When using Google, we fetched the first two
text snippets and added them as additional terms to the original topic formulation, forming a new expanded query.
When using Wikipedia, we fetched the first returned article and added the ten most frequent terms (tf) contained in
the extracted article.
The retrieval effectiveness of our two new query expansion approaches is depicted in Table 7 (German
collection) and is compared to two other query expansion techniques. Compared to the performance before query
expansion (0.4096), Rocchio's and the idf-based blind query expansion cannot improve the MAP. On the other
hand, the variant “idf-window” presents a better retrieval performance (+4.9%, from 0.4069 to 0.4247). Using the
first two text snippets returned by Google, we may also enhance slightly the MAP (from 0.4096 to 0.4196, or
+2.4%). The MAP variation varied according to approaches and parameter settings, while the largest enhancement
could be found using the idf+window technique (forth column in Table 7). Finally, using Google to find related
terms or phrases implied that we required more processing time.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>5 Official Results</title>
      <p>As a complementary search technique, we used two stemmers when defining the official run UniNEDSde3. In
this case we first applied our light stemming approach and then a more aggressive one. If the same term was
produced by the two stemmers, we only kept one occurrence. On the other hand, if the returned stem differed, we
added the two forms to the query formulation.
Run name Language Query
UniNEDSde1 German TD</p>
      <p>TD</p>
      <p>TD
UniNEDSde2 German TD</p>
      <p>TD</p>
      <p>TD
UniNEDSde3 German T
special TD</p>
      <p>TD
UniNEDSde4 German TD</p>
      <p>TD</p>
      <p>TD
UniNEDSen1 English TD</p>
      <p>TD</p>
      <p>TD
UniNEru1 Russian TD</p>
      <p>TD</p>
      <sec id="sec-2-1">
        <title>UniNEru2 Russian</title>
      </sec>
      <sec id="sec-2-2">
        <title>UniNEru3 Russian</title>
      </sec>
      <sec id="sec-2-3">
        <title>UniNEru4 Russian Index dec</title>
      </sec>
      <sec id="sec-2-4">
        <title>Z-score 0.4399</title>
      </sec>
      <sec id="sec-2-5">
        <title>Z-score 0.4251</title>
      </sec>
      <sec id="sec-2-6">
        <title>Z-score 0.4343</title>
      </sec>
      <sec id="sec-2-7">
        <title>Z-score</title>
        <p>0.3770</p>
      </sec>
      <sec id="sec-2-8">
        <title>Z-score</title>
        <p>0.1594
(0.1531)
Z-score
0.1628
(0.1563)</p>
      </sec>
      <sec id="sec-2-9">
        <title>Z-score</title>
        <p>0.1655
(0.1589)
Z-score
0.1890
(0.1815)</p>
        <p>This year we suggest two new query expansion techniques. The first, denoted "idf-window", is based on
co-occurrence of relatively rare terms in a close context (within 10 terms from the occurrence of a search term in a
retrieved document). As a second approach, we add the first two text snippets found by Google to expand the query.
Compared to the performance before query expansion (e.g., with Okapi the MAP is 0.4096), Rocchio's and the
idf-based blind query expansion cannot improve this retrieval performance. On the other hand, the variant
“idf-window” presents a better retrieval performance (+4.9%, from 0.4069 to 0.4247). Using the first two text
snippets returned by Google, we may also enhance slightly the MAP (from 0.4096 to 0.4196, or +2.4%).</p>
        <p>Acknowledgments</p>
        <p>The authors would like to also thank the GIRT - CLEF-2008 task organizers for their efforts in developing
domain-specific test-collections. This research was supported in part by the Swiss National Science Foundation
under Grant #200021-113273.</p>
        <sec id="sec-2-9-1">
          <title>Appendix 1: Parameter Settings</title>
        </sec>
      </sec>
      <sec id="sec-2-10">
        <title>Language</title>
        <p>German GIRT
English GIRT
Russian word
Russian 4-gram
b
C201
C202
C203
C204
C205
C206
C207
C208
C209
C210
C211
C212</p>
      </sec>
      <sec id="sec-2-11">
        <title>Migrant organizations</title>
        <p>Violence in old age
Tobacco advertising
Islamist parallel societies in Western Europe
Poverty and social exclusion
Generational differences on the Internet
(Intellectually) Gifted</p>
      </sec>
      <sec id="sec-2-12">
        <title>Healthcare for prostitutes Violence in schools Commuting and labor mobility</title>
      </sec>
      <sec id="sec-2-13">
        <title>Media in the preschool age</title>
        <p>Employment service
Chronic illnesses</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Abdou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Searching in Medline: Stemming, query expansion, and manual indexing evaluation</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>44</volume>
          (
          <issue>2</issue>
          ), p.
          <fpage>781</fpage>
          -
          <lpage>789</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Amati</surname>
            ,
            <given-names>G</given-names>
          </string-name>
          . &amp; van
          <string-name>
            <surname>Rijsbergen</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Probabilistic models of information retrieval based on measuring the divergence from randomness</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          ,
          <volume>20</volume>
          (
          <issue>4</issue>
          ), p.
          <fpage>357</fpage>
          -
          <lpage>389</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Buckley</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singhal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitra</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Salton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>1996</year>
          ).
          <article-title>New retrieval approaches using SMART</article-title>
          .
          <source>In Proceedings of TREC-4</source>
          , Gaithersburg: NIST Publication #
          <fpage>500</fpage>
          -
          <lpage>236</lpage>
          , p.
          <fpage>25</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Efthimiadis</surname>
            ,
            <given-names>E.N.</given-names>
          </string-name>
          (
          <year>1996</year>
          ).
          <article-title>Query expansion</article-title>
          .
          <source>Annual Review of Information Science and Technology</source>
          ,
          <volume>31</volume>
          , p.
          <fpage>121</fpage>
          -
          <lpage>187</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Fautsch</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dolamic</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , (
          <year>2008</year>
          ).
          <article-title>Domain-Specific IR for German, English and Russian Languages</article-title>
          . In C. Peters,
          <string-name>
            <given-names>P.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.C.</given-names>
            <surname>Gey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlgen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Magini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.W.</given-names>
            <surname>Oard</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. de Rijke &amp; M. Stempfhuber</surname>
          </string-name>
          (Eds.),
          <source>8th Workshop of the Cross-Language Evaluation Forum. LNCS #5152</source>
          , Springer-Verlag, Berlin, p.
          <fpage>196</fpage>
          -
          <lpage>199</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Hiemstra</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>Using language models for information retrieval</article-title>
          .
          <source>CTIT Ph.D. Thesis.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Hiemstra</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Term-specific smoothing for the language modeling approach to information retrieval</article-title>
          .
          <source>In Proceedings of the ACM-SIGIR</source>
          , The ACM Press, Tempere, p.
          <fpage>35</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Kluck</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>The GIRT data in the evaluation of CLIR systems - from 1997 until 2003</article-title>
          . In C. Peters,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          , M. Kluck (Eds.),
          <source>Comparative Evaluation of Multilingual Information Access Systems. LNCS #3237</source>
          . Springer-Verlag, Berlin,
          <year>2004</year>
          , p.
          <fpage>376</fpage>
          -
          <lpage>390</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>McNamee</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Mayfield</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Character n-gram tokenization for European language text retrieval</article-title>
          .
          <source>IR Journal</source>
          ,
          <volume>7</volume>
          (
          <issue>1-2</issue>
          ), p.
          <fpage>73</fpage>
          -
          <lpage>97</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Beaulieu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>Experimentation as a way of life: Okapi at TREC</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>36</volume>
          (
          <issue>1</issue>
          ), p.
          <fpage>95</fpage>
          -
          <lpage>108</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Rocchio</surname>
            ,
            <given-names>J.J.</given-names>
          </string-name>
          <string-name>
            <surname>Jr.</surname>
          </string-name>
          (
          <year>1971</year>
          ).
          <article-title>Relevance feedback in information retrieval</article-title>
          . In G. Salton (Ed.):
          <article-title>The SMART Retrieval System</article-title>
          .
          <article-title>Prentice-Hall Inc</article-title>
          ., Englewood Cliffs (NJ), p.
          <fpage>313</fpage>
          -
          <lpage>323</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Report on CLEF-2003 monolingual tracks: Fusion of probabilistic models for effective monolingual retrieval</article-title>
          . In C. Peters,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          , M. Kluck (Eds.),
          <source>Comparative Evaluation of Multilingual Information Access Systems. LNCS #3237</source>
          . Springer-Verlag, Berlin,
          <year>2004</year>
          , p.
          <fpage>322</fpage>
          -
          <lpage>336</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Berger</surname>
          </string-name>
          , P.-Y. (
          <year>2005</year>
          )
          <article-title>: Selection and merging strategies for multilingual information retrieval</article-title>
          . In: Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.J.F.</given-names>
            ,
            <surname>Kluck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Magnini</surname>
          </string-name>
          ,
          <string-name>
            <surname>B</surname>
          </string-name>
          . (Eds.):
          <article-title>Multilingual Information Access for text</article-title>
          ,
          <source>Speech and Images. Lecture Notes in Computer Science</source>
          : Vol.
          <volume>3491</volume>
          . Springer, Heidelberg, p.
          <fpage>27</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>