<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Attempts to Search Czech Spontaneous Spoken Interviews - the University of West Bohemia at CLEF 2007 CL-SR track</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pavel Ircing</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ludek Muller</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of West Bohemia</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The paper presents an overview of the system build and experiments performed for the CLEF 2007 CL-SR track by the University of West Bohemia. We have concentrated on the monolingual experiments using the Czech collection only. The approach that was successfully employed by our team in the last year's campaign (simple tf.idf model with blind relevance feedback, accompanied with solid linguistic preprocessing) was used again but the set of performed experiments was broadened.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Experimentation</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        The Czech subtask of the CL-SR track, which was rst introduced at CLEF 2006 campaign,
is enormously challenging | let us repeat once again that the goal is to identify appropriate
replay points (that is, the moments where the discussion about the queried topics starts) in a
continuous stream of text generated by automatic transcription of spontaneous speech. Therefore,
it is neither the standard document retrieval task (as there are no true documents de ned) nor the
fully- edged speech retrieval (since the participants do not have the speech data nor the lattices,
so they can't explore alternative hypotheses and must rely on one-best transcription). However,
in order to lower the barrier of entry for teams pro cient at classic document retrieval (or, for
that matter, even total IR beginners), the last year's organisers prepared a so called Quickstart
collection with arti cially de ned \documents" that were created by sliding 3-minute window over
the stream of transcriptions with a 2-minute step (i.e., the consecutive documents have a one
minute overlap).1 The last year's Quickstart collection was further equipped with both manually
1It turned out later that the actual timing was di erent due to some faulty assumptions during the Quickstart
collection design, but since the principle of the document creation remains the same, we will still use the \intended"
time gures instead of the actual ones, just for the sake of readability.
and automatically generated keywords (see [
        <xref ref-type="bibr" rid="ref2">5</xref>
        ] for details) but they have shown itself to be of no
bene t for IR performance [3](the former for the timing problems, the latter for the problems with
their assignment that yet remain to be identi ed) and thus have been dropped from this year's
data. The scripts for generating such Quickstart collection with variable window and overlap times
were also included in the data release.
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>System description</title>
      <p>Our current system largely builds upon the one that was successful in the last year's campaign [3],
with only minor modi cations and larger set of tested settings.
2.1</p>
      <sec id="sec-3-1">
        <title>Linguistic preprocessing</title>
        <p>Stemming (or lemmatization) is considered to be vital for good IR performance even in the case
of weakly in ected languages such as English; thus it is probably even more crucial for Czech as
the representative of the richly in ectional language family. This assumption was experimentally
proven by our group in the last year's CLEF CL-SR track [3]. Thus we have used the same method
of linguistic preprocessing, that is, the serial combination of Czech morphological analyser and
tagger [2], which provides both the lemma and stem for each input word form, together with a
detailed morphological tag. This tag (namely it's rst position) is used for stop-word removal |
we removed from indexing all the words that were tagged as prepositions, conjunctions, particles
and interjections.
2.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Retrieval</title>
        <p>All our retrieval experiments were performed using the Lemur toolkit [1], which o ers a variety of
retrieval models. We have decided to stick to the tf.idf model where both documents and queries
are represented as weighted term vectors d~i = (wi;1; wi;2; ; wi;n) and ~qk = (wk;1; wk;2; ; wk;n),
respectively (n denotes the total number of distinct terms in the collection). The inner-product
of such weighted term vectors then determines the similarity between individual documents and
queries. There are many di erent formulas for computation of the weights wi;j , we have tested
two of them, varying in the tf component:
Raw term frequency
wi;j = tfi;j log
d
dfj
where tfi;j denotes the number of occurrences of the term tj in the document di (term frequency), d
is the total number of documents in the collection and nally dfj denotes the number of documents
that contain tj .</p>
        <p>BM25 term frequency
wi;j =</p>
        <p>k1 tfi;j
tfi;j + k1(1
b + b llCd )
log
d
dfj
where tfi;j , d and dfj have the same meaning as in (1), ld denotes the length of the document, lC
the average length of a document in the collection and nally k1 and b are the parameters to be
set.</p>
        <p>The tf components for queries are de ned analogously, except for the average length of a
query, which obviously cannot be determined as the system is not aware of the full query set and
processes one query at a time. The Lemur documentation is however not clear about the exact
way of handling the lC value for queries.
(1)
(2)</p>
        <p>
          The values of k1 and b were set according to the suggestions made by [
          <xref ref-type="bibr" rid="ref4">7</xref>
          ] and [
          <xref ref-type="bibr" rid="ref3">6</xref>
          ], that is k1 = 1:2
and b = 0:75 for computing document weights and k1 = 1 and b = 02 for query weights.
        </p>
        <p>
          We have also tested the in uence of the blind relevance feedback. The simpli ed version of the
Rocchio's relevance feedback implemented in Lemur [
          <xref ref-type="bibr" rid="ref4">7</xref>
          ] was used for this purposes. The original
Rocchio's algorithm is de ned by the formula
~qnew = ~qold +
~
dR
~
dR
where R and R denote the set of relevant and non-relevant documents, respectively, and d~R
and d~R denote the corresponding centroid vectors of those sets. In other words, the basic idea
behind this algorithm is to move the query vector closer to the relevant documents and away from
the non-relevant ones. In the case of blind feedback, the top M documents from the rst-pass run
are simply considered to be relevant. The Lemur modi cation of this algorithm sets the = 0
and keeps only the K top-weighted terms in d~R.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Evaluation</title>
      <p>We have created 3 di erent indices from the collection | using original data and their lemmatized
and stemmed version. There were 29 training topics and 42 evaluation topics de ned by the
organisers. We have rst run the set of experiments for the training topics (see Table 1), comparing:
Results obtained for the queries constructed by concatenating the tokens (either words,
lemmas or stems) from the &lt;title&gt; and &lt;desc&gt; elds of the topics (TD - upper section of
the table) with results for queries made from all three topic elds, i.e. &lt;title&gt;, &lt;desc&gt; and
&lt;narr&gt; (TDN - lower section).</p>
      <p>Results achieved on the \original" Quickstart collection (i.e. 3-minute window with 1-minute
overlap - Segments 3-1) with results computed using the collection created by using 2-minute
window with 1-minute overlap (Segments 2.1).</p>
      <p>
        In all cases the performance of raw term frequency (Raw TF) and BM25 term frequency (BM25
TF) is tested, both with (BRF) and without (no FB) application of the blind relevance feedback.
The mean Generalized Average Precision (mGAP) is used as the evaluation metric | the details
about this measure can be found in [
        <xref ref-type="bibr" rid="ref1">4</xref>
        ].
2This is actually not a choice, as the value of b is hard-set to 0 for queries in Lemur.
TDN
      </p>
      <p>Segments 3-1</p>
      <p>Raw TF BM25 TF
no FB BRF no FB BRF
words 0.0105 0.0121 0.0088 0.0121
lemmas 0.0168 0.0189 0.0126 0.0126
stems 0.0188 0.0205 0.0132 0.0161
words 0.0113 0.0142 0.0089 0.0108
lemmas 0.0205 0.0226 0.0114 0.0150
stems 0.0215 0.0215 0.0092 0.0107</p>
      <p>Two minute \documents" seem to perform better than the three minute ones | probably
the three minute segmentation is too coarse.</p>
      <p>The simplest raw term frequency weighting scheme generally outperforms the more
sophisticated BM25 | one possible explanation is that in a standard document retrieval setup
the BM25 scheme pro ts mostly from its length normalization component that is completely
unnecessary in our case (remember that our documents all have approximately identical
length by design).</p>
      <p>The fact that both stemming and lemmatization boost the performance by about the same
margin was already observed in the last year's experiments.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In the CLEF 2007 CL-SR task, we have made just a little step further towards successful searching
of Czech spontaneous speech. In order to make a bigger progress, we would need to really take the
speech part of the task into account | that is, to use the speech recognizer lattices when searching
for the desired information, or even to modify the ASR components so that it will be more likely
to produce output useful for IR (for example, enrich the language model with rare named entities
that are currently often being misrecognized).</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments References</title>
      <p>This work was supported by the Grant Agency of the Czech Academy of Sciences project No.
1ET101470416 and the Ministry of Education of the Czech Republic project No. LC536.
[1] Carnegie Mellon University and the University of Massachusetts. The Lemur Toolkit for
Language Modeling and Information Retrieval. (http://www.lemurproject.org/), 2006.
[2] Jan Hajic. Disambiguation of Rich In ection. (Computational Morphology of Czech).</p>
      <p>Karolinum, Prague, 2004.
[3] Pavel Ircing and Ludek Muller. Bene t of Proper Language Processing for Czech Speech
Retrieval in the CL-SR Task at CLEF 2006. In C. Peters, P. Clough, F. Gey, J. Karlgren,
B. Magnini, D. Oard, M. de Rijke, and M. Stempfhuber, editors, Evaluation of Multilingual and
Multi-modal Information Retrieval - 7th Workshop of the Cross-Language Evaluation Forum,
CLEF 2006, Lecture Notes in Computer Science, Alicante, Spain, 2007.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Baolong</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Douglas</given-names>
            <surname>Oard</surname>
          </string-name>
          .
          <article-title>One-Sided Measures for Evaluating Ranked Retrieval Effectiveness with Spontaneous Conversational Speech</article-title>
          .
          <source>In Proceedings of SIGIR 2006</source>
          , pages
          <fpage>673</fpage>
          {
          <fpage>674</fpage>
          , Seattle, Washington, USA,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Douglas</given-names>
            <surname>Oard</surname>
          </string-name>
          , Jianqiang Wang, Gareth Jones, Ryen White, Pavel Pecina, Dagobert Soergel, Xiaoli Huang, and
          <string-name>
            <given-names>Izhak</given-names>
            <surname>Shafran</surname>
          </string-name>
          .
          <article-title>Overview of the CLEF-2006 Cross-Language Speech Retrieval Track</article-title>
          . In C. Peters,
          <string-name>
            <given-names>P.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Magnini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Oard</surname>
          </string-name>
          , M. de Rijke, and M. Stempfhuber, editors,
          <source>Evaluation of Multilingual and Multi-modal Information Retrieval - 7th Workshop of the Cross-Language Evaluation Forum, CLEF 2006, Lecture Notes in Computer Science</source>
          , Alicante, Spain,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Robertson</surname>
          </string-name>
          and
          <string-name>
            <given-names>Steve</given-names>
            <surname>Walker</surname>
          </string-name>
          .
          <source>Okapi/Keenbow at TREC-8. In The Eight Text REtrieval Conference (TREC-8)</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Chengxiang</given-names>
            <surname>Zhai</surname>
          </string-name>
          .
          <source>Notes on the Lemur TFIDF model. Note with Lemur 1.9 documentation</source>
          , School of CS, CMU,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>