<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>D. A. Hull. Stemming Algorithms: A Case Study for Detailed Evaluation. In Journal of The
American Society For Information Science</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>West Group at CLEF2000: Non-English Monolingual Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Isabelle Moulinier</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. Andrew McCulloh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elizabeth Lund West Group</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Opperman Drive Eagan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>USA Isabelle.Moulinier@westgroup.com</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>[TTYF95] P. Thompson, H. Turtle</institution>
          ,
          <addr-line>B. Yang and J. Flood</addr-line>
          ,
          <institution>"TREC-3 Ad Hoc Retrieval and Routing Experiments using the WIN System," in Overview of the 3rd Text Retrieval Conference (TREC-3)</institution>
          ,
          <addr-line>NIST Special Publication 500-225, Gaithersburg, MD</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>[Tur90] H. Turtle. Inference Networks for Document Retrieval. PhD Thesis, Computer Science Department, University of Massassuchets</institution>
          ,
          <addr-line>Amherst, 1990</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1994</year>
      </pub-date>
      <volume>47</volume>
      <issue>1</issue>
      <abstract>
        <p>West Group participated in the non-English monolingual retrieval task for French and German. Our primary interest was to investigate whether retrieval of German or French documents was any different from the retrieval of English documents. We focused on two aspects: stemming for both languages and compound breaking for German, and studied several query formulations to take advantage of compounds. Our results suggest that German retrieval is indeed different from English or French retrieval, inasmuch as breaking compounds can significantly improve performance.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>•
•</p>
      <p>German:
•
•
•
•
French</p>
      <p>We used a third-party stemmer based on a morphological analyzer. One of the features was
compound decomposition. Forcing decomposition or not was a parameter in our experiments.
We indexed both German collections as one single retrieval collection. We did not investigate
merging retrieved sets.</p>
      <sec id="sec-1-1">
        <title>We added a tokenization rule to handle elision.</title>
        <p>We used two kinds of stemmers: a third-party stemmer based on a morphological analyzer, and a
rule-based stemmer (a la Porter) from the Muscat project.</p>
        <p>A WIN query consists of concepts extracted from natural language text. Normal WIN query processing
eliminates stopwords, noise phrases (or introductory phrases) and recognizes phrases or other important
concepts for special handling. Many of the concepts ordinarily recognized by WIN are specific to both
English documents and the legal domain. To perform these tasks, WIN relies on various resources: a
stopword list, a list of introductory phrases (“Find cases about…”, “A relevant document describes…”) , a
dictionary of (legal) phrases.</p>
        <p>Query processing for French was similar to English query processing. We used a stopword list of 1745
terms (highly frequent terms, and noise terms like adverbs). Using the TREC-6, 7 and 8 topics, we refined
the list of introductory patterns we created for Amaryllis-2. In the end, there were 160 patterns (a pattern
is a regular expression that handles case variants and some spelling errors). We did not use phrase
identification for lack of a general French phrase dictionary.</p>
        <p>We investigated several options to structuring German queries, decomposing or not decomposing
compounds. This specific processing is described below. We used a stopword list of 333 terms. Using the
TREC-6, 7 and 8 topics, we derived a set of introductory patterns for German. There were 11 regular
expressions, summarizing over 200 noise phrases. We did not perform phrase identification through a
dictionary. However German compounds have been treated as “natural phrases” in some of our runs.
Finally, we extracted concepts from the full topics. However, we gave more weight to concepts appearing
in the Title or Description fields than concepts extracted from the Narrative field. Following West’s
participation at TREC3 [TTYF95], we assigned a weight of 4 to concepts extracted from the Title field,
while concepts originating from the Description and Narrative fields were given a weight of 2 and 1,
respectively.</p>
        <p>German monolingual retrieval experiments and results
Our experiments with monolingual German retrieval focused on query processing and compound
decomposition. Our submitted runs rely on decomposing compounds, but we also experimented with no
decomposition, and no stemming at all. Indexing followed the choice made for query processing. For
instance, when no decomposition was performed for query terms, parts of compounds were not indexed.
When dealing with breaking compound terms, we faced the choice of considering a compound term as a
single concept in our WIN query, or treating the compound as several concepts (as many concepts as
there were parts in the compound). The submitted run WESTgg1 considers that a compound corresponds
to several concepts; the run WESTgg2 handles a compound as a single concept.</p>
        <p>When faced with a compound Energiequellen, the structured query in WESTgg1 introduces 2 concepts,
Energie and Quelle; the structured query in WESTgg2 introduces 1 concept, #PHRASE(Energie Quelle).
The #PHRASE operator is a soft phrase, i.e. the component terms must appear with 3 words of one
another. The score of the #PHRASE concept in our experiment was set to be the maximum score of the
soft phrase itself or of its components.</p>
        <p>R-Prec.</p>
      </sec>
      <sec id="sec-1-2">
        <title>Best</title>
      </sec>
      <sec id="sec-1-3">
        <title>Above</title>
      </sec>
      <sec id="sec-1-4">
        <title>Median</title>
      </sec>
      <sec id="sec-1-5">
        <title>Below</title>
        <p>Worst
0.3706
0.3628
0.3080
0.3141
3
3
0
0
21
18
15
18
3
6
1
1
9
9
19
15
1
1
2
3</p>
        <p>We expected a greater difference between our two submitted runs. WESTgg1 allows compound terms to
contribute more to the score of a document, while WESTgg2 gives the same contribution to compound
and non-compound terms. The contribution of a compound term in WESTgg1 is weighted by the number
of parts in the compound, so one would expect its occurrence in a document to significantly alter a
document score.</p>
        <p>After reviewing the individual queries, we noticed the following behavior. First, for those queries where
both the compounds and their parts had an average frequency, neither particularly common nor
particularly rare, the two runs behaved similarly. Then, the parts helped locate documents, but did not add
to or draw away from the document relevance score. Second, for those queries where the compound itself
is above average, but the individual parts are average, or even fairly common, then the weighted
contribution provided in WESTgg1 performed better. Third, for those queries where at least one part of a
compound was very common, the high occurrence of that part degraded the weighting scheme of
WESTgg1, thus the single concept construct of WESTgg2 provided a more representative score.
Also, compound handling in WESTgg1 as well as WESTgg2 is only as influential as there are compounds
in the query. In the 40 German topics, roughly 16% of the query terms are compound terms. In addition, it
should noted that we indexed the individual parts of compounds. As a result, a simple query term may
also match the part of a compound in a document.</p>
        <p>French monolingual retrieval experiments and results
The goal of our experiments with French document retrieval was to assess the difference between
stemming algorithms. Our motivation was to further investigate the particularity of French compared to
English. [Hul96] reported results on various kinds of stemmers for English document retrieval. So far, we
have studied two types of stemmers (out of the 5 types in [Hul96]) as well as no stemming at all:
•
•
a stemmer based on an inflectional morphological analyzer, e.g. it conflates verb forms to the
infinitive of the verb, noun forms to the singular noun, adjectives to the masculine singular form.
This stemmer is based on a lexicon.
a rule-based stemmer “a la Porter” that approximates mainly inflectional rules, but also provides a
limited set of derivational rules based on suffix stripping, e.g. it strips suffixes like –able or –isme.
Our runs also took advantage of the multiple TEXT elements in a document. We considered those
elements to mark paragraph boundaries and used this information for document scoring and ranking.
Our submitted run, WESTff, used the inflectional stemmer. Table 2 summarizes the performance of runs
using the inflectional stemmer, the Porter stemmer and no stemmer at all. We also ran experiments when
a document was either scored as a whole or as its best paragraph. Those runs are not reported here, as they
did not perform as well as the combined score.</p>
      </sec>
      <sec id="sec-1-6">
        <title>Performance of individual queries Avg. Prec. R-Prec.</title>
      </sec>
      <sec id="sec-1-7">
        <title>Above</title>
      </sec>
      <sec id="sec-1-8">
        <title>Median</title>
      </sec>
      <sec id="sec-1-9">
        <title>Below</title>
      </sec>
      <sec id="sec-1-10">
        <title>Worst Run</title>
      </sec>
      <sec id="sec-1-11">
        <title>WESTff</title>
      </sec>
      <sec id="sec-1-12">
        <title>Porter</title>
        <p>While we usually consider not stemming as a baseline, our tests showed that no stemming performed
better on several topics. In those instances, we found that the Porter stemmer was too aggressive and
stemmed important query terms to very common forms. For instance, parti was stemmed to part, directive
to direct and français to franc. The inflectional stemmer did exactly what it was supposed to do, e.g. stem
française to français. However, certain stems were very common, while their raw form was less common.
Phrase identification, e.g. académie française, and monnaie européenne, may likely improve performance,
as it has proven to be beneficial for the English version of the WIN search engine.</p>
        <p>In addition, the inflectional stemmer is only as good as its lexicon. We found a couple of queries where
the Porter stemmer performed better because important query terms were not in the lexicon.
While our analysis is only partial at this time, it appears that our French stemming results follow the
patterns exhibited by [Hul96] for English stemming, except that inflectional stemming seems slightly
superior. We do not know yet whether this is a particularity of the French language or of this particular
collection and set of topics.
1 For some queries, our runs achieved an average precision that was better than the best average precision
reported at CLEF.</p>
        <p>0.9
0.8
0.7
0.6
0.3
0.2
0.1
0
WESTff
WESTgg1
WESTgg2</p>
        <p>1</p>
        <p>Recall</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Summary References</title>
      <p>The WIN retrieval system achieved good performance for both German and French document retrieval
without any major modification being made to its retrieval engine. On the one hand, we showed that
German document retrieval required specific handling because of the use of compound words in the
language. Our results showed that decomposing compounds during indexing and query processing
enhanced the capabilities of our system. Our French experiments, on the other hand, did not uncover any
striking difference between French and English retrieval, except a preference towards the use of an
inflectional stemmer.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>