<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Clairvoyance CLEF-2003 Experiments</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yan Qu</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Greg Grefenstette</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David A. Evans Clairvoyance Corporation</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Baum Boulevard</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Suite</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pittsburgh</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>grefen</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>dae}@clairvoyancecorp.com</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>CLARIT Cross-Language Information Retrieval</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>In CLEF 2003, Clairvoyance participated in the bilingual retrieval track with the German and Italian language pair. As we did not have any German-to-Italian translation resources, we used the Babel Fish translation service provided by Altavista.com for translating German topics into Italian, with English as a pivot language. Then the translated Italian topics were used for retrieving Italian documents from the Italian document collection. The translated Italian topics and the document collections were indexed using three different kinds of units: (1) linguistically meaningful units, (2) character 6-grams, and (3) a combination of 1 and 2. We submitted three automatic runs with the three indexing units. Clairvoyance participated in the CLEF 2003 bilingual retrieval track using the German and Italian language pair. As we did not have German-to-Italian translation resources, we used the free Babel Fish translation service provided by Altavista.com for translating German topics into Italian, with English as a pivot language. The resulting translated Italian topics were used for retrieving Italian documents from the Italian document collection. The translated Italian topics and the document collections were indexed using three different kinds of features: (1) linguistically meaningful units (e.g., words and NPs), (2) character 6-grams, and (3) a combination of 1 and 2. We submitted three automatic runs, each based on one of the three indexing units. In the following sections, we describe the details of our submission and present the performance results. In CLEF 2003, we adopted query translation as the means for bridging the language gap between the query language and the document language for cross-language information retrieval. For German-to-Italian information retrieval, first, a German query string was translated into Italian via machine translation; then the translated Italian topics were used for retrieving Italian documents from the Italian document collection. For query and document processing, we used the CLARIT system [1], in particular, those components encompassing a newly developed Italian NLP module (for extracting Italian phrases), indexing (term weighting and phrasal decomposition), retrieval, and “thesaurus extraction” (for extracting terms to support pseudo relevance feedback).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2.1 Query Translation via a Pivot Language</title>
      <p>The Babel Fish translation service (altavista.com) provides translation between selected language pairs including
German-to-English and English-to-Italian. It does not provide translation service between German and Italian
directly. So we used English as a pivot language, first translating the German topics to English and then
translating the English topics into Italian. As an illustration of typical results for this process, Figure 1 provides
the translations from Babel Fish for Topic 141.</p>
      <p>Even though there was increased degradation in query quality after translation, we felt that, except for translation
of proper names, the quality of the translation from German to English and from English to Italian by Babel Fish
was adequate for the purpose of cross-language information retrieval. We quantitatively evaluate this impression
in Section 3.</p>
    </sec>
    <sec id="sec-3">
      <title>2.2 Italian Topic Processing</title>
      <p>
        Once the topics were translated into Italian, we extracted two types of terms from the topics: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) linguistically
meaningful units or character n-grams.
      </p>
      <p>
        To extract linguistically meaningful units, we used CLARIT Italian NLP. This NLP module makes use of a
lexicon and finite-state grammar for extracting phrases such as NPs. The lexicon is based on the Multext Italian
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) Original German topic from CLEF-2003:
      </p>
      <sec id="sec-3-1">
        <title>Briefbombe für Kiesbauer .</title>
      </sec>
      <sec id="sec-3-2">
        <title>Finde Informationen über die Explosion einer Briefbombe im Studio der Moderatorin Arabella Kiesbauer beim</title>
      </sec>
      <sec id="sec-3-3">
        <title>Fernsehsender PRO7.</title>
        <p>
          (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) English translation of (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) by Babel Fish:
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Letter bomb for gravel farmer. Find information about the explosion of a letter bomb in the studio of the host</title>
      </sec>
      <sec id="sec-3-5">
        <title>Arabella gravel farmer with the television station PRO7.</title>
        <p>
          (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) Ideal English topic from CLEF-2003:
        </p>
      </sec>
      <sec id="sec-3-6">
        <title>Letter Bomb for Kiesbauer</title>
      </sec>
      <sec id="sec-3-7">
        <title>Find information on the explosion of a letter bomb in the studio of the TV channel PRO7 presenter Arabella</title>
      </sec>
      <sec id="sec-3-8">
        <title>Kiesbauer.</title>
        <p>
          (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) Italian translation of (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) by Babel Fish:
        </p>
      </sec>
      <sec id="sec-3-9">
        <title>Bomba della lettera per il coltivatore della ghiaia. Trovi le informazioni sull'esplosione di una bomba della lettera</title>
        <p>nell'studio del coltivatore della ghiaia di Arabella ospite con la stazione PRO7 della televisione.
(5) Ideal Italian topic from CLEF-2003:</p>
      </sec>
      <sec id="sec-3-10">
        <title>Lettera Bomba per Kiesbauer</title>
      </sec>
      <sec id="sec-3-11">
        <title>Recupera le informazioni relative all'esplosione di una lettera bomba nello studio della presentatrice della rete televisiva PRO7.</title>
        <p>lexicon1, which was expanded by adding punctuations and special characters. In addition, entries with accented
vowels were duplicated by substituting the accented vowels with their corresponding unaccented vowels
followed by an apostrophe (“’”). The final lexicon contained about 135,000 entries. An Italian stop word list2,
which contained 433 entries, was used to filter out stop words. The grammar specified the rules for constructing
phrases, especially NPs, and morphological normalization rules for normalizing morphological variants to their
root forms, e.g., “previsto” to “prever”. In CLEF 2003 experiments, we extracted Adjectives, Verbs, and NPs as
indexing terms.</p>
        <p>Another way to construct terms is to use overlapping character n-grams. We have observed that our
lexiconbased term extraction did not have complete coverage for morphological normalization. The n-gram approach
we adopted was aimed at mitigating such an effect. For the submissions, we have used overlapping 6-grams, as
it was previously reported to be effective [2]. Spaces and punctuations were included in the character 6-grams.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>2.3 CLARIT Indexing and Retrieval</title>
      <p>CLARIT indexing involves statistical analysis of a text corpus and construction of an inverted index, with each
index entry specifying the index word and a list of texts. CLARIT allows the index to be built upon full
documents or variable-length subdocuments. We used subdocuments as the basis for indexing and document
scoring in our experiments. The size of a subdocument was in the range of 8 sentences to 12 sentences.
CLARIT retrieval is based on the vector space retrieval model. Various similarity measures are supported in the
model. For CLEF 2003, we used the dot product function for computing similarities between a query and a document:
sim ( P , D ) =</p>
      <p>W P ( t ) ⋅ W D ( t ).</p>
      <p>t ∈ P ∩ D
where WP(t) is the weight associated with the query term t and WD(t) is the weight associated with the term t in
the document D. The two weights were computed as follows:
1 http://www.lpl.univ-aix.fr/projects/multext/LEX/LEX.SmpIt.html
2 Obtained from http://www.unine.ch/Info/clef/</p>
      <p>W D ( t ) = TF D ( t ) ⋅ IDF ( t ).</p>
      <p>W P ( t ) = C ( t ) ⋅ TF P ( t ) ⋅ IDF ( t )
where IDF and TF are standard inverse document frequency and term frequency statistics, respectively. IDF(t)
was computed with the target corpus for retrieval. The coefficient C(t) is an “importance coefficient”, which can
be modified either manually by the user or automatically by the system (e.g., updated during feedback).</p>
    </sec>
    <sec id="sec-5">
      <title>2.4 Post-Translation Query Expansion</title>
      <p>Query expansion through (pseudo) relevance feedback has proved to be effective for improving IR performance
[3]. We used pseudo relevance feedback for augmenting the queries. After retrieving some documents for a
given topic from the target corpus, we took a set of top ranked documents, regarding them as relevant documents
to the query, and extracted terms from the these documents. The terms were ranked based on the following
formula:</p>
      <p>Prob2(t)
= log(R t + 1) x log(</p>
      <p>N - R + 2 - 1) - log(
N t - R t + 1</p>
      <p>R + 1</p>
      <p>R t
- 1)
where N is the number of documents in the target corpus, Nt is the number of documents in the corpus that
contain term t, R is the number of documents for feedback that are (presumed to be) relevant to the topic, and Rt
is the number of documents that are (presumed to be) relevant to the topic and contain term t.
3</p>
    </sec>
    <sec id="sec-6">
      <title>Experiments</title>
      <p>We submitted three automatic runs to CLEF 2003. All the queries used the title and description fields
(Ttitle+Description) of the topics provided by CLEF 2003. The results presented below are based on relevance
judgments of 42 topics, which have relevant documents in the Italian corpus. The three runs were:
•
•
•
ccwrd: with linguistically meaningful units as indexing terms
ccngm: with character 6-grams as indexing terms
ccmix: a combination of linguistic units and character 6-grams as indexing units
With the ccmix run, the combinations were constructed through a simple concatenation of the terms nominated
by ccwrd and ccngm. We ran Italian monolingual experiments to obtain the baseline with ideal translations
after obtaining the relevance judgments from CLEF 2003.</p>
      <p>
        All the experiments were run with post-translation pseudo relevance feedback. The feedback-related parameters
were based on training over CLEF 2002 topics. The settings for German-to-Italian retrieval were: extracting
T=80 terms from the top N=25 retrieved documents with the Prob2 method. For the n-gram based indexing and
the mixed model, an additional term cutoff percentage set to P=0.01. For the word-based indexing, the
percentage cutoff is set to P=0.25. For Italian monolingual retrieval with words as indexing terms: T=50, N=50,
P=0.1. For Italian monolingual retrieval with n-grams and the mixed model as indexing terms: T=80, N=25,
P=0.05.
ccwrd2002
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) Translated English (from German) to Italian
Performance change compared with (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
Performance change compared with (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) Ideal English to Italian
Performance change compared with (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) Ideal Italian
4
      </p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions References</title>
      <p>Due to the lack of resources, our participation in CLEF 2003 was limited. We succeeded in submitting three
runs for German-to-Italian retrieval, examining word based indexing and n-gram based indexing. Our results
with CLEF 2002 and CLEF 2003 did not provide firm evidence of which indexing method is better. Future
analysis is required in this direction.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>R.G.</given-names>
            <surname>Lefferts. CLARIT-TREC Experiments</surname>
          </string-name>
          .
          <source>Information Processing and Management</source>
          , Vol.
          <volume>31</volume>
          , No.
          <issue>3</issue>
          , pp.
          <fpage>385</fpage>
          -
          <lpage>395</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>McNamee</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Mayfield</surname>
          </string-name>
          .
          <article-title>Scalable Multilingual Information Access</article-title>
          . In C. Peters, editor,
          <source>Working Notes for the CLEF 2002 Workshop</source>
          , pp.
          <fpage>133</fpage>
          -
          <lpage>140</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Ballesteros</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>Statistical Methods for Cross-Language Information Retrieval</article-title>
          . In G. Grefenstette, editor,
          <source>Cross-Language Information Retrieval, Chapter</source>
          <volume>3</volume>
          . Kluwer Academic Publishers, Boston, pp.
          <fpage>23</fpage>
          -
          <lpage>40</lpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <article-title>The Philosophy of Information Retrieval Evaluation</article-title>
          . In C. Peters,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          , and M. Kluck, editors,
          <source>Evaluation of Cross-Language Information Retrieval Systems: Proceedings of the CLEF 2001 Workshop, Lecture Notes in Computer Science 2406</source>
          , Springer, pp.
          <fpage>355</fpage>
          -
          <lpage>370</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>