<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>ITC-irst - Centro per la Ricerca Scientifica e Tecnologica I-38050 Povo</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2000</year>
      </pub-date>
      <fpage>2</fpage>
      <lpage>7</lpage>
      <abstract>
        <p>This paper presents work on document retrieval for Italian carried out at ITC-irst. Two different approaches to information retrieval were investigated, one based on the Okapi weighting formula and one based on a statistical model. Development experiments were carried out using the Italian sample of the TREC-8 CLIR track. Performance evaluation was done on the Cross Language Evaluation Forum (CLEF) 2000 Italian monolingual track.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
    </sec>
    <sec id="sec-2">
      <title>TEXT PREPROCESSING</title>
      <p>Document and query preprocessing implies
several stages: tokenization, morphological analysis
of words, part-of-speech (POS) tagging of text,
base form extraction, stemming, and stop-terms
removal.</p>
      <p>Tokenization. Tokenization of text is performed
in order to isolate words from punctuation marks,
recognize abbreviations and acronyms, correct
possible word splits across lines, and discriminate
between accents and quotation marks.</p>
      <sec id="sec-2-1">
        <title>Morphological analysis. A morphological ana</title>
        <p>lyzer decomposes each Italian inflected word into
its morphemes, and suggests all possible POSs and
base forms of each valid decomposition. By base
forms we mean the usual not inflected entries of a
dictionary.</p>
        <p>
          POS tagging. POS tagging is based on a Viterbi
decoder that computes the best text-POS alignment
on the basis of a bigram POS language model and a
discrete observation model
          <xref ref-type="bibr" rid="ref4">(Merialdo, 1994)</xref>
          . The
employed tagger works with 57 tag classes and has
an accuracy around 96%.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Base form extraction. Once the POS and the</title>
        <p>morphological analysis of each word in the text
is computed, a base form can be assigned to each
word.</p>
        <p>Stemming. Word stemming is applied at the level
of tagged base forms. POS specific rules were
developed that remove suffixes from verbs, nouns,
and adjectives.</p>
        <p>Stop-terms removal. Words in the collection that
are considered non relevant for the purpose of
information retrieval are discarded in order to save index
space. Words are filtered out on the basis either of
their POS or their inverted document frequency. In
particular, punctuation is eliminated together with
articles, determiners, quantifiers, auxiliary verbs,
prepositions, conjunctions, interjections, and
pronouns. Among the remaining terms, those with a
low inverted document frequency, i.e. that occur in
many different documents, are eliminated.</p>
        <p>An example of text preprocessing is presented in
Table 8.</p>
        <sec id="sec-2-2-1">
          <title>Terms</title>
          <p>text
base forms
stems
base forms
stems
Stop
no
no
no
yes
yes</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. INFORMATION RETRIEVAL MODELS</title>
      <p>w number of documents containing
w frequency of in the collection
d length of document
w q frequency of in query
w d frequency of word in document
d vocabulary size of document
length of the collection
mean document length
number of documents
average document vocabulary size
vocabulary size of the collection
V
160K
126K
101K
125K
100K
d evance of a document versus a query q. In this</p>
      <sec id="sec-3-1">
        <title>3.1. Okapi Model</title>
        <p>
          Okapi
          <xref ref-type="bibr" rid="ref8">(Robertson et al., 1994)</xref>
          is the name of
a retrieval system project that developed a family
of weighting functions in order to evaluate the
relwork, the following Okapi weighting function was
applied:
P (q j d) where represents the likelihood of q, given
P (d) d, represents the a-priori probability of d, and
P (q) is a normalization term. By assuming no
ad q the probability of generating and assume an
priori knowledge about the documents, and
disregarding the normalization factor, documents can be
ranked, with respect to q, just by the likelihood
term. If we interpret the likelihood function as
order-free multinomial model, the following
logprobability score can be derived:
w The probability that a term is generated by
d can be estimated by applying statistical language
modeling techniques. Previous work on statistical
information retrieval
          <xref ref-type="bibr" rid="ref5 ref7">(Miller et al., 1998; Ng, 1999)</xref>
          proposed to interpolate relative frequencies of each
document with those of the whole collection, with
interpolation weights empirically estimated from
the data.
        </p>
        <p>
          In this work we use an interpolation formula
which applies the smoothing method proposed by
          <xref ref-type="bibr" rid="ref10">(Witten and Bell, 1991)</xref>
          . This method linearly
smoothes word frequencies of a document and the
amount of probability assigned to never observed
terms is proportional to the number of different
words contained in the document. Hence, the
following probability estimate is applied:
Vd
134
129
126
80
77
(3)
(1)
(2)
where:
k1(1
fd(w)
        </p>
        <p>Vd
mAvPr</p>
        <sec id="sec-3-1-1">
          <title>Data Set (topic #s’)</title>
          <p>CLIR (54-81)
title
description
narrative
CLEF (1-40)
title
description
narrative
Min
41
3
8
25
31
3
7
14</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>EXPERIMENTS</title>
      <p>
        (8)
(9)
B ing all the terms of the top documents according
T documents are taken and the most relevant terms
in them are added to the query. Hence, the retrieval
phase is repeated with the augmented query. In
this work, new search terms are extracted by
sortto
        <xref ref-type="bibr" rid="ref3">(Johnson et al., 1999)</xref>
        :
r 1 : : : provided against a given query q, let rk
q The AvPr for is defined as the average of the
pre
      </p>
      <p>This section presents work done to develop and
test the presented models. Development and
testing were done on two different Italian document
retrieval tasks. Performance was measured in terms
of Average Precision (AvPr) and mean Average
Precision (mAvPr). Given the document ranking
be the ranks of the retrieved relevant documents.
cision values achieved at all recall points, i.e.:</p>
      <p>The mAvPr of a set of queries corresponds to the
mean of the corresponding query AvPr values.</p>
      <sec id="sec-4-1">
        <title>4.1. Development</title>
        <p>For the purpose of parameter tuning,
development material made available by CLEF was used.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.3. Blind Relevance Feedback Tuning</title>
        <p>1CLIR topics without Italian relevant documents are
60, 63, 76, and 80.</p>
        <p>B
0.7 0</p>
        <p>5
0.6
0.5
T = 15 were chosen, whose combination gives
B = 5 uments and the number of relevant terms
ters is shown. Finally, the number of relevant
doca mAvPr of 49.2%, corresponding to a 6.8%
improvement over the first step.</p>
        <p>Further work was done to optimize the
performance of the first retrieval step. Indeed,
performance of the BRF procedure is determined by the
precision achieved, by the first retrieval phase, on
the very top ranking documents. In particular, an
higher resolution for documents and queries was
considered by using base forms instead of stems.</p>
        <p>In Table 6 mAvPr values are shown by considering
different combinations of text preprocessing before
and after BRF. In particular, we considered using
base forms before and after BRF, using word stems
before and after BRF, and using base forms before
BRF and stems after BRF. The last combination
achieved the largest improvement (8.6%) and was
adopted for the final system.
of the topics do not have corresponding documents
in the collection they are not taken into account 2.</p>
        <p>More details about the CLEF collection and topics
are in Tables 3, 4, and 5.</p>
        <p>Official results of the Okapi and statistical
models are reported in Figure 3 with the names irst1
and irst2, respectively. Figure 3 shows the
difference in AvPr between each run and the median
reference provided by the CLEF organization. As a
further reference, performance differences between
the best result of CLEF and the median are also
plotted. The mAvPr of irst1 and irst2 are 49.0%
and 47.5%, respectively. Both methods score above
the median reference mAvPr, which is 44.5%. The
mAvPr of the median reference was computed by
taking the average over the median AvPr scores.</p>
        <p>I
st
ba
ba</p>
        <p>II
st
ba
st</p>
        <p>5
46.4
46.2
46.7</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.4. Official Evaluation</title>
        <p>The two presented models were evaluated on
the CLEF 2000 Italian monolingual track. The test
collection consists of newspaper articles published
by La Stampa, during 1994, and 40 topics. As six
5. IMPROVEMENTS
r(B(d))]2
(10)
2CLEF topics without Italian relevant documents are
3, 6, 14, 27, 28, and 40.
35
merge
best
0.7 0
0.6</p>
        <p>S By taking the average of over all the queries 3,
S 1=(N and B, has mean 0 and variance 1). On
[0; 1] of each model were normalized in the range
A Under the hypothesis of independence between
the contrary, in case of perfect correlation the S
statistics has value 1.
a rank correlation of 0.4 resulted between the irst1
and irst2 runs.</p>
        <p>This results confirms some degree of
independence between the two information retrieval
models. Hence, a combination of the two models was
implemented by just taking the sum of scores.
Actually, in order to adjust scale differences, scores
before summation. By using the official relevance
assessments of CLEF, a mAvPr of 50.0% was
achieved by the combined model.</p>
        <p>In Figure 4 and Figure 5 detailed results of
the combined model (merge) are provided for each
query, respectively, against the CLEF references
and the irst1 and irst2 runs. It results that the
combined model performs better than the median
reference on 24 topics of 34, while irst1 and irst2
improved the median AvPr 16 e 17 times,
respectively. Finally, the combined model improves the
best reference on two topics (20 and 36).
30 35
merge vs. irst1
merge vs. irst2
-0.4 1 3 5 7 9 11 13 15 17 19 21 23 25 27 29 31 33 35 37 39</p>
        <p>Topic Number</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Federico</surname>
          </string-name>
          , Marcello,
          <year>2000</year>
          .
          <article-title>A system for the retrieval of italian broadcast news</article-title>
          .
          <source>Speech Communication</source>
          ,
          <volume>33</volume>
          (
          <issue>1</issue>
          -
          <fpage>2</fpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Federico</surname>
          </string-name>
          , Marcello and Renato De Mori,
          <year>1998</year>
          .
          <article-title>Language modelling</article-title>
          . In Renato De Mori (ed.),
          <article-title>Spoken Dialogues with Computers, chapter 7</article-title>
          . London, UK: Academy Press.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Johnson</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Jourlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Spark</given-names>
            <surname>Jones</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.C.</given-names>
            <surname>Woodland</surname>
          </string-name>
          ,
          <year>1999</year>
          .
          <article-title>Spoken document retrieval for TREC-</article-title>
          8 at Cambridge University.
          <source>In Proceedings of the 8th Text REtrieval Conference</source>
          . Gaithersburg, MD.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Merialdo</surname>
          </string-name>
          , Bernard,
          <year>1994</year>
          .
          <article-title>Tagging English text with a probabilistic model</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>20</volume>
          (
          <issue>2</issue>
          ):
          <fpage>155</fpage>
          -
          <lpage>172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>David R. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tim</surname>
            <given-names>Leek</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Richard</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Schwartz</surname>
          </string-name>
          ,
          <year>1998</year>
          .
          <article-title>BBN at TREC-7: Using hidden Markov models for information retrieval</article-title>
          .
          <source>In Proceedings of the 7th Text REtrieval Conference</source>
          . Gaithersburg, MD.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Mood</surname>
          </string-name>
          ,
          <string-name>
            <surname>Alexander</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franklin</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Graybill</surname>
          </string-name>
          , and
          <string-name>
            <surname>Duane</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Boes</surname>
          </string-name>
          ,
          <year>1974</year>
          . Introduction to the Theory of Statistics. Singapore:
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Ng</surname>
          </string-name>
          , Kenney,
          <year>1999</year>
          .
          <article-title>A maximum likelihood ratio information retrieval model</article-title>
          .
          <source>In Proceedings of the 8th Text REtrieval Conference</source>
          . Gaithersburg, MD.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.</given-names>
            E., S.
          </string-name>
          <string-name>
            <surname>Walker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>M. M.</given-names>
          </string-name>
          <string-name>
            <surname>Hancock-Beaulieu</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Gatford</surname>
          </string-name>
          ,
          <year>1994</year>
          .
          <article-title>Okapi at TREC-3</article-title>
          .
          <source>In Proceedings of the 3rd Text REtrieval Conference</source>
          . Gaithersburg, MD.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Sparck</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <source>Karen and Peter Willett (eds.)</source>
          ,
          <year>1997</year>
          . Readings in Information Retrieval. San Francisco, CA: Morgan Kaufmann.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Witten</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ian H</surname>
          </string-name>
          . and
          <string-name>
            <surname>Timothy</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Bell</surname>
          </string-name>
          ,
          <year>1991</year>
          .
          <article-title>The zero-frequency problem: Estimating the probabilities of novel events in adaptive text compression</article-title>
          .
          <source>IEEE Trans. Inform. Theory, IT37</source>
          (4):
          <fpage>1085</fpage>
          -
          <lpage>1094</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>