<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>A MacFarlane, Centre for Interactive Systems Research, Department of Information Science, City University</institution>
          ,
          <addr-line>Northampton Square, LONDON EC1V 0HB</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2002</year>
      </pub-date>
      <abstract>
        <p>We test the utility of European language stemmers created using the Snowball language [1]. This allows us to experiment with PLIERS in languages other than English. We also report on some BM25 tuning constant experiments conducted in order to find the best settings for our searches. In this paper we briefly describe our experiments at CLEF 2002. We address a number of issues as follows. The snowball language was recently created by Porter [2] in order to provide generic mechanism for creating stemmers. The main purpose of the experiments was to investigate the utility of these stemmers and as to whether a reasonable level of retrieval effectiveness could be achieved by using Snowball in an information retrieval system. This allowed us to port the PLIERS system [3] for use on languages other than English. We also investigate the variation of tuning constants for the BM25 weighting function for the chosen European languages. The languages used for our experiments are as follows: German, French, Dutch, Italian, German and Finnish. All experiments were conducted on a Pentium 4 machine with 256 MB of memory and 240 GB of disk space. The operating system used was Red Hat Linux 7.2. All search runs were done using the Robertson/Sparck Jones Probabilistic model - the BM25 weighting model was used. All our runs are in the monolingual track. All queries derived from topics are automatic. The paper is organised as follows. Section two describes our motivation for doing this research. In section three we describe our indexing methodology and results. In section 4 we describe some preparatory results using CLEF 2001 data. Section 5 describes our CLEF 2002 results, and a conclusion is given at the end. We have several different strands in our research agenda. The most significant of these is the issue of using Snowball stemmers in information retrieval, both in terms of retrieval effectiveness and retrieval efficiency. In terms of retrieval efficiency we want to quantify the cost of stemming both for indexing collections and servicing queries over them. How expensive is stemming in terms of time when processing words? Our hypothesis for search would be that stemming will increase inverted list size and that this would in turn lead to slower response times for queries. Stemming will slow down indexing, but by how much? Due to time constraints we restrict our discussion on search efficiency to CLEF 2002 runs. With respect to retrieval effectiveness we hope to demonstrate that using Snowball stemmers leads to an increase in retrieval effectiveness for all languages, and any deterioration in results should not be significantly worse. Our hypothesis is therefore: “stemming using Snowball will lead to an increase in retrieval effectiveness”. A further issue that we wish to address is that of tuning constants for the BM25 weighting model. Our previous research on this issue has been done on collections of English [3], but we want to investigate the issue on other types of European languages. A hypothesis is formulated in section 4 where the issue is discussed in more detail.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Motivation for the Research</title>
    </sec>
    <sec id="sec-3">
      <title>3. Indexing methodology and results</title>
      <sec id="sec-3-1">
        <title>3.1 Indexing methodology</title>
        <p>We used a simple and straightforward methodology for indexing: parsing, remove stop words, stemming in the
given language. The PLIERS HTML/SGML parser needed to be altered to detect non-ascii characters such as
those with umlauts, accents, circumflexes etc. The stemmers were easily incorporated into the PLIERS library.
We used various stop word lists for the language, gathered from the internet. The official runs for Finnish did not
use stemming, as no stemmer was available for that language while experiments were being conducted.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Indexing results</title>
        <p>The key results here are that stemmed builds take up slightly less space for most languages (it is a significant
saving for French and Italian) and that builds with stemming take significantly longer than builds with no
stemming. Builds with no stemming index text at a rate of 3.7 to 4.5 GB per hour compared with 0.48 to 0.89 GB
per hour for stemmed builds. The results for stemmed builds are acceptable however.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Preparatory experiments: working with CLEF 2001 data</title>
      <p>In order to find the best tuning constants for our CLEF 2002 runs we conducted tuning constant variation
experiments on the CLEF 2001 data for the following languages: French, German, Dutch, Spanish and Italian.
We were unable to conduct experiments with Finnish data as this track was not run in 2001: we arbitrarily chose
K1=1.5 and B=0.8 for our Finnish runs.</p>
      <p>
        We give a brief description of the BM25 tuning constants being discussed here [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The K1 constant alters the
influence of term frequency in the BM25 function, while the B constant alters the influence of normalised
average document length. Values of K1 can range from 0 to infinity, whereas the values of B are with the range 1
(document lengths used unaltered) to 0 (document length data not used at all).
      </p>
      <p>We used the following strategy for tuning constant variation. For K1 we start with a value of 0.25 with
increments of 0.25 to a maximum of 3.0, stopping when it was obvious that no further useful data could be
gathered. For B we used a range of 0.1 to 0.9 with increments of 0.1. A maximum of 135 experiments were
therefore conducted for each language.
Due to lack of time, our aim in these experiments was to achieve a better than baseline retrieval effectiveness for
our preparatory CLEF 2001 experiments. For the most part we succeeded in doing this for both types of build.
However we can separate our results into three main groups:
•</p>
      <p>For Dutch and German we were able to better six systems with our runs.</p>
      <p>For French and Italian, we were unable to better more than one run, but our effectiveness is
considerably better than the official baseline runs.</p>
      <p>For Spanish we were unable to show much of an improvement over the baseline run.</p>
      <p>Having said that we have a long way to go before our runs achieve the levels of performance of groups such as
the University of Neuchatel particularly for languages such as French and Spanish.</p>
      <p>We were unable to investigate the reason for the levels of performance achieved because of time constraints, but
it is believed that the automatic query generator used to select terms for queries is simplistic and needs to be
replaced with a more sophisticated mechanism. Our reason for believing that this might be the problem is that
our results for Title only queries on the Spanish run are superior to those queries that were derived from Title and
Description: this is counter to what we would expect.</p>
      <p>When comparing experiments on those builds which used stemming and those that did not, we can separate our
runs into three main groups:
•
•
•</p>
      <p>Stemming is an advantage: We were able to demonstrate that stemming was a positive advantage for
both Italian and Spanish runs.</p>
      <p>Stemming makes no difference: Using stemming on French made very little difference either way.</p>
      <p>Stemming is a disadvantage: Stemming proved to be problematic for both Dutch and German.
We need to investigate the reason for these results. We are surprised that stemming is a disadvantage for any
language
5</p>
    </sec>
    <sec id="sec-5">
      <title>CLEF 2002 results</title>
      <sec id="sec-5-1">
        <title>5.1 Retrieval efficiency results</title>
        <sec id="sec-5-1-1">
          <title>Language</title>
          <p>German
French
Dutch
Italian
Spanish
Finnish</p>
          <p>In general this is what we would expect as inverted lists on stemmed builds tend to be larger than those of builds
with no stemming. It is interesting to examine the exceptions and outliers, however. The reason German (TD)
runs are faster on stemmed builds, is that the average query size is slightly larger by about half a term (more
inverted lists are being processed on average). The Dutch no stem run is significantly faster on average than
those of stemmed runs, but this is largely due to query size: queries with no stems on the Dutch collection are
have 1.7 less terms on average than those of stemmed queries (fewer inverted lists are being processed). An
interesting result with title only Finnish runs is that queries with no stems are nearly 3.5 times smaller than
stemmed queries, but runs times are virtually identical: execution speeds are so small here it is difficult to
separate them. It should be noted when we compared the size of queries with stems to those without stems, there
is no clear pattern. We would suggest therefore that a simple hypothesis which suggested that runs on builds
without stemming is faster on average than runs on stemmed builds cannot be supported with the evidence given
here. It is clear that the number of inverted lists processes is an important factor as well as the size of the inverted
lists.</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>Language</title>
          <p>German
French
Dutch
Italian
Spanish
Finnish</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>5.2 Retrieval effectiveness results</title>
        <p>The status for German and Dutch is unchanged from our CLEF 2001 results (stemming runs produced worse
results), and also for Italian where the stemmer runs produced better results. Our results for French have
improved comparatively, but for Spanish the results have deteriorated. The runs on the Finnish collection are
particularly disappointing: the results for average precision on builds with stemming being about 40% worse for
title only queries and nearly 60% worse on title/description queries than experiments on builds without
stemming. The reason for this loss in performance could be because of the morphological complexity of Finnish
and merits significant further investigation. It is also interesting that the initial version of the Finnish stemmer
did slightly better in terms of average precision than the final version: this is offset with a slight loss in precision
at 5 and 10 documents retrieved. The reduced effectiveness found on title/description queries compared with title
only queries in our Spanish CLEF 2001 experiments was not repeated in our CLEF 2002 runs. However there is
a slight loss in performance on the second version of the Finnish stemmer when comparing title/description
queries to title only queries: this loss is consistent across all shown precision measures.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>A number of research questions have been raised during this investigation of the effectiveness of stemming
utilitising Snowball. They are as follows:
•
•
•
•</p>
      <p>Why are results on Dutch and German consistently worse using Snowball stemmers?
Why are the results using the Finnish snowball stemmer significantly worse?
Why are the results on the Spanish collection inconsistent, both with the Snowball stemmer and varying
the query type (title only and title/description)?
Why are the runs on Italian collections consistent, and how do we use evidence from these runs to
improve the results of other romance languages such as Spanish and French?
Our hypothesis that suggested that the use of Snowball stemmers is always beneficial has not been confirmed. It
may be possible to investigate this hypothesis further when we have addressed the research questions given
above. We also have some conclusions with regard to retrieval efficiency and stemming:
•</p>
      <p>Stemming using the Snowball stemmers is costly when indexing, but does not slow down the process of
inverted file generation to an unacceptable level.</p>
      <p>We have confirmed that inverted file size is not the only factor in search speed, query size that requires
processing of more inverted lists play an important part as well.</p>
      <p>We are unable to comment on our tuning constant experiments due to time constraints, and hence our hypothesis
for this workshop version of the paper, but will have a poster which shows the results graphically and the
information will be included in the final version of the paper.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements References</title>
      <p>The author is grateful to Martin Porter for his efforts to produce a snowball stemmer for Finnish.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] Snowball web site [http://snowball.sourceforge.net] -
          <source>visited 19th July</source>
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Porter. M.</surname>
          </string-name>
          ,
          <string-name>
            <surname>Snowball:</surname>
          </string-name>
          <article-title>A language for stemming algorithms</article-title>
          , [http://snowball.sourceforge.net/texts/introduction.html] -
          <source>visited 19th July</source>
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>MacFarlane</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.E..</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCann</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <source>PLIERS AT TREC8</source>
          , In: Voorhees,
          <string-name>
            <given-names>E.M.</given-names>
            , and
            <surname>Harman</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.K.</surname>
          </string-name>
          , (eds),
          <source>The Eighth Text Retrieval Conference (TREC-8)</source>
          , NIST Special Publication 500-246, NIST: Gaithersburg,
          <year>2000</year>
          ,
          <fpage>p241</fpage>
          -
          <lpage>252</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          , and
          <string-name>
            <given-names>Sparck</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Simple</surname>
          </string-name>
          , proven approaches to text retrieval,
          <source>University of Cambridge Technical report, May</source>
          <year>1997</year>
          , TR356 , [http://www.cl.cam.ac.uk/Research/Reports/TR356-ksj
          <article-title>-approaches-to-text-retrieval</article-title>
          .
          <source>html] - visited 22nd July</source>
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Frakes</surname>
            ,
            <given-names>W.B.</given-names>
          </string-name>
          ,
          <article-title>Introduction to information storage and retrieval systems</article-title>
          . In: Frakes,
          <string-name>
            <given-names>W.B.</given-names>
            and
            <surname>Baeza-Yates</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds),
          <article-title>Information retrieval; data structures and algorithms</article-title>
          , Prentice Hall,
          <year>1992</year>
          ,
          <fpage>p1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>