<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LIC2M experiments at CLEF 2004</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Romaric Besancon</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olivier Ferret</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Fluhr CEA-LIST/LIC</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M fromaric.besancon</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>olivier.ferret</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>christian. uhrg@cea.fr</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>For its second participation in the CLEF campaign, the LIC2M participated in the multilingual task. Our challenge for this participation was to improve the results obtained for French and English and integrate two new languages in the system, Russian and Finnish. Our results are not good on Russian and Finnish, which shows that our system strongly depends on a correct linguistic analysis on the documents.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The LIC2M cross-language retrieval system is a weighted boolean search engine based on a
linguistic analysis of the query and the documents. This system has been used in the small multilingual
task of the previous CLEF 2003 campaign [BdCF+03].</p>
      <sec id="sec-1-1">
        <title>Document processing</title>
        <p>The documents are processed to extract informative linguistic elements from the text parts. The
processing includes a part-of-speech tagging of the words, their lemmatization and the extraction
of compounds and named entities. This linguistic processing requires the de nition of a set of
resources for each language:
a full form dictionary, containing for each word form its possible part-of-speech tags and
linguistic features (gender, number, etc);
a set of trigrams and bigrams of part-of-speech categories that are used for part-of-speech
tagging (these trigrams and bigrams are learned from a corpus);
a set of rules for the shallow parsing of sentences. This parsing identi es syntactic relations
that are used to extract compounds from the sentences.
a set of rules for the identi cation of named entities: these rules are composed of gazetteers
and of some contextual rules that uses special triggers to identify named entities and their
type.</p>
        <p>The introduction of Russian and Finnish in the multilingual task raised a di culty concerning
this linguistic processing. For Russian, we used a language dictionary that allowed us to simply
associate the words with their possible part-of-speech. We had no time to train a part-of-speech
tagger nor to develop sets of rules for syntactic analysis or named entities. The processing of
Russian has then been quite straightforward since we only used the words and their categories.</p>
        <p>For Finnish, since we did not have a full form dictionary, we used a simple stemmer (Porter
Snowball stemmer [Por02]) and no part-of-speech. We also apply the stoplist provided by Jacques
Savoy [Sav].
2.2</p>
      </sec>
      <sec id="sec-1-2">
        <title>Query processing</title>
        <p>All query processing is automatic. Each query is rst processed through the linguistic analyzer
corresponding to the query language. For two out of the three submitted runs (see section 3), the
three elds of the query, title (T), description (D) and narrative (N) were kept for this analysis.
For the third run, only title and description were taken.</p>
        <p>When using the narrative eld in the query processing, a stoplist containing meta-words was
used to lter out non-relevant words (words used in the narrative to describe what are relevant
documents, such as : \document", \relevant" etc.). These meta-words stoplists were built on
the basis of CLEF 2002 topics, from a rst selection using frequency information, and revised
manually.</p>
        <p>The result is a query composed of a list of the linguistic elements extracted from the analysis,
possibly ltered by the meta-words stoplists. These elements are called the concepts of the query.
Each concept is reformulated into search terms in the language of the considered index, either
using bilingual dictionaries or, in the case of monolingual search, using monolingual reformulation
dictionaries (adding synonyms and related words) and/or a topical expansion, based on a network
of lexical cooccurrences, as described in [BdCF+03].</p>
        <p>For translation, we had bilingual dictionaries for French-English and English-Russian pairs.
The dictionary we used for the reformulation into Finnish language was the FreeLang bilingual
English-Finnish dictionary [HK]. Other translations (French-Russian, French-Finnish) were
performed through a multi-step translation (using English as a pivot language).
2.3</p>
      </sec>
      <sec id="sec-1-3">
        <title>Search and Merge Strategy</title>
        <p>The search and merging techniques are the same as the ones used in previous CLEF 2003 campaign
and are explained in details in [BdCF+03]. They are brie y described in this section.</p>
        <p>The original topic is associated, during the query processing, to four di erent sets of search
terms, one for each language. Each search term set is used as an independent query against the
index of the corresponding language. N documents are retrieved for each language. The 4 N
retrieved documents from the four corpora are then merged and sorted by their relevance to the
topic. Only the rst 1000 are kept (in the submitted runs, we took N = 1000).</p>
        <p>For each language, our system retrieves, for each search term, the documents containing the
term (until N documents are retrieved). A concept pro le is associated with each document, each
component of which indicates the presence or absence of a query concept in the document (a
concept is present in a document if one of its reformulated search term is present). Retrieved
documents sharing the same concept pro le are clustered together. This clustering allows a
straightforward merging strategy that takes into account the original query concepts and the way they
have been reformulated: since the concepts are in the original query language, the concept pro les
associated with the clusters formed for di erent target languages are comparable, and the clusters
having the same pro le are simply merged.</p>
        <p>To compute the relevance weight of each cluster, we rst compute a cross-lingual pseudo-idf
weight of each concept, using only the corpus composed of the 4 N documents kept as the
N
result of the search. This weight is computed by the formula idf (c) = log d4f (c) , where df (c) is the
number of documents containing the concept c. The weight associated with a cluster is then the
sum of the weights of the concepts present in its concept pro le.</p>
        <p>The clusters are then sorted by their weights: all documents in a cluster are given the weight of
the cluster (the documents are not sorted inside the clusters). The list of the rst 1000 documents
from the best clusters is then built and used for the evaluation.
3</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Results</title>
      <p>We submitted three runs to the multilingual task, described in Table 1. The rst two use English
topics (one using the title, description and narrative elds, the other using only title and description
elds), the third one uses French topics (using title, description and narrative elds for the query
processing and topical expansion of the query).</p>
      <p>lic2men1
lic2men2
lic2mfr1
query language</p>
      <p>English
English
French
query elds</p>
      <p>T+D
T+D+N
T+D+N</p>
      <p>query expansion
dictionary reformulation
dictionary reformulation
topical expansion + dictionary reformulation</p>
      <p>n
o
ii
s
c
e
r
P
1
0.8
0.6
0.4
0.2
used for translation are based on lemmas and parts-of-speech. We should integrate in our system
some default processing for the di erent steps of linguistic processing that would not require the
complete de nition of linguistic resources but relies on basic schemas and training data. This would
allow to better integrate new languages in the existing design of our system1. Another possible
improvement is to enrich the reformulation by techniques such as transliteration or approximate
matching (for proper names in particular), or use reformulation data automatically learned from
aligned corpora.</p>
      <p>The results presented in Table 2 also show that our system seems to work better when using
all information available in the query (title, description and narrative). The narrative seems to
introduce some relevant information by giving di erent formulations of the topic and without
adding much noise after the basic ltering of meta-words by a specialized stop-list. A more precise
analysis of the results should be performed to also study the e ect of the negative formulations in
the narrative (\documents that contain ... are not relevant ").</p>
      <p>For French and English, the results are better than for Russian and Finnish but are not as
good as we could expect. A rst analysis suggests several possible adjustments, that have been
tested in a new run:
monolingual reformulation introduces too many rare synonyms (or synonyms of too rare
senses of the words) that cause non-relevant documents to be retrieved. For the new test,
we simply deactivated this monolingual reformulation (in the future, the monolingual
reformulation dictionaries will be checked to improve the relevance of added terms).
the importance of named entities was neglected in the runs we submitted. Giving a special
importance to named entities, relatively to other words, improves the results. For the new
test, we set a double weight for named entities, relatively to other words.
the value of N (number of documents retrieved for one language) is also important. Indeed,
the documents are retrieved until the number of documents N is reached: if this number is
too small, all search terms may not be exploited. For the new test, we set this number at
5000.</p>
      <p>With these three changes in the system con guration, the results obtained (for English topics,
using the T+D+N elds, and only on French and English corpora) are given in Table 3, and
show a signi cant improvement: 90% of the relevant documents are retrieved. Some tuning of our
system remains to be done to improve the ordering of the documents.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>These experiments in the multilingual track of CLEF 2004 show some improved results of our
system, relatively to last year, on French and English corpora. On the other hand, the poor
1Notice that this would not solve problems speci c to certain languages such as the decompounding of Finnish
words.</p>
      <p>avg p
relret
fre/eng eng fre
0.243 0.44 0.238
1168 (90.5%) 362 (96.5%) 806 (88.1%)
results obtained for Russian and Finnish show that the introduction of new languages in our
system with simpli ed linguistic processing or stemming/stoplist approaches do not perform well.
This integration should be made easier either by making the system more exible (de ning for
instance robust default processing for some steps of linguistic analysis) or by allowing the search
system to take as input the result of a completely di erent approach for new languages (for
instance, simple linguistic analysis combined with a reformulation based on statistical translation
lexicons learned from aligned corpora). In this case, we would have to tackle the di culty of
merging the results obtained with di erent processings. Further experiments in these directions
will be undertaken.
[HK]
[Sav]</p>
      <p>Kimmo Hamalainen and Toivo Kivirinta.
http://www.kasvua.org/ kphamala/dict.html.</p>
      <p>Freelang</p>
      <p>Jacques Savoy. A stopword list for nnish. http://www.unine.ch/info/clef/.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [BdCF+03]
          <string-name>
            <surname>Romaric</surname>
            <given-names>Besancon</given-names>
          </string-name>
          , Gael de Chalendar, Olivier Ferret, Christian Fluhr, Olivier Mesnard, and
          <string-name>
            <given-names>Hubert</given-names>
            <surname>Naets</surname>
          </string-name>
          .
          <article-title>The LIC2M's CLEF 2003 system</article-title>
          .
          <source>In Working Notes for the CLEF 2003 Workshop</source>
          , Trondheim, Norway,
          <fpage>21</fpage>
          -
          <lpage>22</lpage>
          August
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Por02]
          <article-title>nnish-english dictionary</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Martin</given-names>
            <surname>Porter</surname>
          </string-name>
          .
          <article-title>Finnish snowball stemmer</article-title>
          . http://snowball.tartarus.org/ nnish/stemmer.html,
          <year>September 2002</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>