<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ontology-Based Multilingual Information Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jacques Guyot</string-name>
          <email>Jacques.Guyot@rolex.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saïd Radhouani</string-name>
          <email>Said.Radhouani@cui.unige.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gilles Falquet</string-name>
          <email>Gilles.Falquet@cui.unige.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre universitaire d'informatique 24</institution>
          ,
          <addr-line>rue Général-Dufour, CH-1211 Genève 4</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1993</year>
      </pub-date>
      <abstract>
        <p>For our first participation in the CLEF evaluation campaign, our aim is to explore a translation-free technique for multilingual information retrieval. This technique is based on an ontological representation of documents and queries. We use a multilingual ontology for documents/queries representation. For each language, we use the multilingual ontology to map a term to its corresponding concept. The same mapping is applied to each document and each query. Then, we use a classic vector space model for the indexing and the querying. The main advantages of our approach are: no merging phase is required, no dependency on automatic translators between all pairs of languages exists, and adding a new language only requires a new mapping dictionary to the multilingual ontology.</p>
      </abstract>
      <kwd-group>
        <kwd>Multilingual Ontology</kwd>
        <kwd>Conceptual Indexing</kwd>
        <kwd>Multilingual Information Retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>1. Ontology based Multilingual Information Retrieval</p>
      <sec id="sec-1-1">
        <title>1.1 Multilingual ontology</title>
        <p>A Multilingual ontology is defined by one ontology and a set of dictionary (one dictionary for each language). An
ontology is a formal, explicit specification of a shared conceptualisation [Gruber 1993]. It contains a set of distinct
and identified concepts C related by a set of relations R. In our approach, we only need to use the set of concepts.
Here, we present two examples of concepts extracted from our ontology:
•
•
•
•</p>
        <p>DEN(28845)= {thumb}.</p>
        <p>SFR(pouce)= {8612, 28845 }.
28845: the thick short innermost digit of the forelimb.</p>
        <p>A dictionary DL is an association of ontology concepts C with a terms set TL pertaining to a language L. We
denote: DL : C  TL. Indeed, the concept c is labelled by a set of terms t1,t2,..,tn in the language L. We denote
DL(c)={t1,t2, …,tn}. We also define the reciprocal relation SL : TL  C by SL (t)={c∈ C | t ∈ DL(c)}. Actually, the
term t indicates the concepts c1,c2,…, cm. We also denote SL(t)={ c1,c2, …,cm}. Here, we present two examples of
associations between terms and concepts:</p>
      </sec>
      <sec id="sec-1-2">
        <title>1.2 Ontology based multilingual information retrieval</title>
        <p>In our approach, for each document in the whole collection, we use the multilingual ontology to map each term to
its corresponding concept. We apply the same process on the queries.</p>
        <p>The document dL = &lt; t1, t2,…, tn&gt; is a sequence of terms from the set TL of the language L. To carry out the
termconcept mapping, we apply the function SL on each term ti of the document dL: SL(t)={c∈ C | t ∈ DL(c)}. So we
obtain the conceptual representation of the document that we denote: CR(dL)= &lt; S(t1), S(t2), …, S(tn)&gt;. Finally,
CR(dL) is a sequence of sets of concepts.</p>
        <p>We did not introduce any treatment for the term ambiguity. In fact, if the term is ambiguous, we replace it by all
its corresponding concepts.</p>
        <p>Before the term mapping step, we use a “stop word list” for each language, and a dedicated stemming system. We
have used Snowball, a small string processing language designed for creating stemming algorithms in Information
Retrieval [Snow 2005].</p>
        <p>We did not introduce any morpho-syntactic or processing (like n-grams) to break composite words in Dutch,
German, or Finnish.</p>
        <p>For indexing and querying, we use the vector space model [Salton et al. 83].</p>
        <p>TERMS</p>
        <p>CONCEPTS
Finnish terms
corpus terms
Spanish
corpus
French
corpus</p>
        <p>...</p>
        <p>Spanish
queries
French
queries
...</p>
        <p>Finnish
queries
stem
ming
indexing process
conceptual
document
Spanish dictionary
French dictionary</p>
        <p>...</p>
        <p>Finnish dictionary</p>
        <p>ontology
quering process
conceptual
queries
indexing
conceptual
index
of
corpus
search
engine</p>
        <p>results</p>
      </sec>
      <sec id="sec-1-3">
        <title>1.3 Official runs description</title>
        <p>In our approach, each query Q is composed by two fields: a topic field and a body field. We denote
Q=&lt;Topic,Body&gt;. The content of each field depends on the runs. Each field is composed by a list of terms extracted
from the original text query.</p>
        <p>As the queries are precise, we use the topic field to query the whole collection. As a result, we obtain a set of
documents containing the topic concepts. Then we use the body field to rank this set of documents.</p>
        <p>Here we present an example of a query composed by a topic field (text between topic tags) and a body field (text
between all the other tags).
&lt;top&gt;
&lt;num&gt; C182 &lt;/num&gt;
&lt;topic&gt; Normandië Landing &lt;/topic&gt;
&lt;NL-title&gt; 50e Herdenkingsdag van de Landing in Normandië &lt;/NL-title&gt;
&lt;NL-desc&gt; Zoek verslagen over de dropping van veteranen boven Sainte-Mère-Église tijdens de viering van de 50e
herdenkingsdag van de landing in Normandië. &lt;/NL-desc&gt;</p>
        <p>&lt;NL-narr&gt; Ongeveer veertig veteranen sprongen tijdens de viering van de 50e herdenkingsdag van de landing in Normandië
met een parachute boven Sainte-Mère-Église, net zoals ze vijftig jaar eerder op D-day hadden gedaan. Alle informatie over het
programma of over de gebeurtenis zelf worden als relevant beschouwd. &lt;/NL-narr&gt;
&lt;/top&gt;</p>
        <p>Now we present our official runs. In the following three runs, we use English when submitting queries:
1.</p>
        <p>AUTOEN: the topic field is composed by the terms of the title of the original query (text between the
title tags). The body field is composed by the text of the original query.</p>
        <p>ADJUSTEN: the topic field is composed by the modified title by adding and/or removing terms. The
adding terms are extracted from the original query text. The body field is composed by the text of the
original query.
3. FEEDBCKEN: the topic field is composed by the modified title as in ADJUSTEN. The body field is
composed by the original text query and the first relevant document (if it exists) in the first 30
documents found by the previous ADJUSTEN run.</p>
        <p>In order to compare the system results using different languages when submitting queries, we carried out three
more runs: ADJUSTDU, ADJUSTFR, and ADJUSTSP. For each run, we use respectively Dutch, French, and
Spanish when submitting queries. In these runs, topic field is composed by modified title like in ADJUSTEN and
body field is composed by the original query text.</p>
        <p>For all the four runs, we obtain almost the same mean average precision: 13.90% for ADJUSTDU, 13.47% for
ADJUSTFR, 13.80% for ADJUSTSP and 16.85 % for ADJSUTEN. Our system is not dependent of the query
language. It gives nearly the same results when submitting queries in four different languages. It’s difficult to explain
difference because the coverage and the quality of ontological dictionaries are important.</p>
        <p>Run name
AUTO-EN
ADJSUT-EN
ADJUST-DU
ADJUST-FR
ADJUST-SP
FEEDBCK-EN</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Conclusion</title>
      <p>In this preliminary work, we tried only to prove the feasibility of our approach. We tried also to prove that our
system is independent of the query language. We still have some limits in our system because we did not introduce
any morpho-syntactic processing to break composite words in Dutch, German, or Finnish. Moreover, our ontology is
incomplete and dirty (we have imported many errors with automatic translation).</p>
      <p>We have also used the same approach in the bi-text alignment field. We have used other language like Chinese,
Arabic and Russian [Guyot 2005].</p>
      <sec id="sec-2-1">
        <title>Acknowledgments</title>
        <p>We would like to thank the CLEF-2005 organisers for their efforts. We would also like to thank Metaread for
giving us the possibility to use the “idxvli” information retrieval system (A fast indexer for big corpora).
[Chen at al. 2003] Chen, A. and Gey, F. Combining query translation and document translation in cross-language
retrieval. In proceedings CLEF-2003, pp. 39.48. Trondheim.
[ERG 2005] Ergane: http//download.travlang.com/, see also
http://www.majstro.com/[Guyot 2005] GUYOT, J. yaaa: yet another alignment algorithm - Alignement ontologique bi-texte pour un corpus
multilingue. Cahier du CUI 2005.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>