<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Ontology-Based Multilingual Information Retrieval</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Jacques</forename><surname>Guyot</surname></persName>
							<email>jacques.guyot@rolex.com</email>
							<affiliation key="aff0">
								<orgName type="department">Centre universitaire d&apos;informatique</orgName>
								<address>
									<addrLine>24, rue Général-Dufour</addrLine>
									<postCode>CH-1211</postCode>
									<settlement>Genève 4</settlement>
									<country key="CH">Switzerland</country>
								</address>
							</affiliation>
							<affiliation key="aff1">
								<orgName type="laboratory">Laboratoire CLIPS-IMAG</orgName>
								<address>
									<postBox>P. 53</postBox>
									<postCode>38041, cedex 9</postCode>
									<settlement>Grenoble</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Saïd</forename><surname>Radhouani</surname></persName>
							<email>said.radhouani@cui.unige.ch</email>
							<affiliation key="aff0">
								<orgName type="department">Centre universitaire d&apos;informatique</orgName>
								<address>
									<addrLine>24, rue Général-Dufour</addrLine>
									<postCode>CH-1211</postCode>
									<settlement>Genève 4</settlement>
									<country key="CH">Switzerland</country>
								</address>
							</affiliation>
							<affiliation key="aff1">
								<orgName type="laboratory">Laboratoire CLIPS-IMAG</orgName>
								<address>
									<postBox>P. 53</postBox>
									<postCode>38041, cedex 9</postCode>
									<settlement>Grenoble</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Gilles</forename><surname>Falquet</surname></persName>
							<email>gilles.falquet@cui.unige.ch</email>
							<affiliation key="aff0">
								<orgName type="department">Centre universitaire d&apos;informatique</orgName>
								<address>
									<addrLine>24, rue Général-Dufour</addrLine>
									<postCode>CH-1211</postCode>
									<settlement>Genève 4</settlement>
									<country key="CH">Switzerland</country>
								</address>
							</affiliation>
							<affiliation key="aff1">
								<orgName type="laboratory">Laboratoire CLIPS-IMAG</orgName>
								<address>
									<postBox>P. 53</postBox>
									<postCode>38041, cedex 9</postCode>
									<settlement>Grenoble</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">Ontology-Based Multilingual Information Retrieval</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">6368747B929E19F5ED74B665F8F35670</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-25T00:37+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>Multilingual Ontology</term>
					<term>Conceptual Indexing</term>
					<term>Multilingual Information Retrieval</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>For our first participation in the CLEF evaluation campaign, our aim is to explore a translation-free technique for multilingual information retrieval. This technique is based on an ontological representation of documents and queries. We use a multilingual ontology for documents/queries representation. For each language, we use the multilingual ontology to map a term to its corresponding concept. The same mapping is applied to each document and each query. Then, we use a classic vector space model for the indexing and the querying. The main advantages of our approach are: no merging phase is required, no dependency on automatic translators between all pairs of languages exists, and adding a new language only requires a new mapping dictionary to the multilingual ontology.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>The existing approaches use either translation of all documents into a common language, either automatic translation of the queries, or combination of both query and document translations <ref type="bibr" target="#b0">[Chen at al. 2003</ref>]. In these cases, we need automatic translators between all pairs of languages. If we translate queries, after receiving a result list from each search engine, we need to use a merging procedure to provide a unique ranked result list. Moreover adding a new language (query or document) requires as much translators as existing languages.</p><p>In our approach, we tried to "dissolve" these problems by using of a multilingual ontology. Based on this ontology, we conducted different experiments involving multilingual test-collection. We retrieve documents written in Dutch, English, Finnish, French, German, Italian, Spanish and Swedish, independently of the query language. First, we tried to prove the feasibility of our approach using English when submitting queries. We also tried to prove that our system is independent of the query language. Thus, we have used Dutch, French, and Spanish when submitting queries, In the next section, we describe our approach and present our official runs.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Ontology based Multilingual Information Retrieval</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.1">Multilingual ontology</head><p>A Multilingual ontology is defined by one ontology and a set of dictionary (one dictionary for each language). An ontology is a formal, explicit specification of a shared conceptualisation <ref type="bibr" target="#b0">[Gruber 1993</ref>]. It contains a set of distinct and identified concepts C related by a set of relations R. In our approach, we only need to use the set of concepts. Here, we present two examples of concepts extracted from our ontology: A dictionary D L is an association of ontology concepts C with a terms set T L pertaining to a language L. We denote: D L : C  T L . Indeed, the concept c is labelled by a set of terms t 1 ,t 2 ,..,t n in the language L. We denote D L (c)={t 1 ,t 2 , …,t n }. We also define the reciprocal relation S L : T L  C by S L (t)={c∈ C | t ∈ D L (c)}. Actually, the term t indicates the concepts c 1 ,c 2 ,…, c m . We also denote S L (t)={ c 1 ,c 2 , …,c m }. Here, we present two examples of associations between terms and concepts:</p><formula xml:id="formula_0">• D EN (28845)= {thumb}.</formula><p>• S FR (pouce)= {8612, 28845 }.  To bootstrap and build dictionaries, we have used Esperanto dictionaries found on the Web (principally from Ergane <ref type="bibr">[ERG 2005]</ref>). We have also used automatic translations to complete some of them. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.2">Ontology based multilingual information retrieval</head><p>In our approach, for each document in the whole collection, we use the multilingual ontology to map each term to its corresponding concept. We apply the same process on the queries.</p><p>The document d L = &lt; t 1 , t 2 ,…, t n &gt; is a sequence of terms from the set T L of the language L. To carry out the termconcept mapping, we apply the function S L on each term t i of the document d L : S L (t)={c∈ C | t ∈ D L (c)}. So we obtain the conceptual representation of the document that we denote: CR(d L )= &lt; S(t 1 ), S(t 2 ), …, S(t n )&gt;. Finally, CR(d L ) is a sequence of sets of concepts.</p><p>We did not introduce any treatment for the term ambiguity. In fact, if the term is ambiguous, we replace it by all its corresponding concepts.</p><p>Before the term mapping step, we use a "stop word list" for each language, and a dedicated stemming system. We have used Snowball, a small string processing language designed for creating stemming algorithms in Information Retrieval <ref type="bibr">[Snow 2005</ref>].</p><p>We did not introduce any morpho-syntactic or processing (like n-grams) to break composite words in Dutch, German, or Finnish.</p><p>For indexing and querying, we use the vector space model <ref type="bibr">[Salton et al. 83]</ref>. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.3">Official runs description</head><p>In our approach, each query Q is composed by two fields: a topic field and a body field. We denote Q=&lt;Topic,Body&gt;. The content of each field depends on the runs. Each field is composed by a list of terms extracted from the original text query.</p><p>As the queries are precise, we use the topic field to query the whole collection. As a result, we obtain a set of documents containing the topic concepts. Then we use the body field to rank this set of documents.</p><p>Here we present an example of a query composed by a topic field (text between topic tags) and a body field (text between all the other tags). &lt;top&gt; &lt;num&gt; C182 &lt;/num&gt; &lt;topic&gt; Normandië Landing &lt;/topic&gt; &lt;NL-title&gt; 50e Herdenkingsdag van de Landing in Normandië &lt;/NL-title&gt; &lt;NL-desc&gt; Zoek verslagen over de dropping van veteranen boven Sainte-Mère-Église tijdens de viering van de 50e herdenkingsdag van de landing in Normandië. &lt;/NL-desc&gt; &lt;NL-narr&gt; Ongeveer veertig veteranen sprongen tijdens de viering van de 50e herdenkingsdag van de landing in Normandië met een parachute boven Sainte-Mère-Église, net zoals ze vijftig jaar eerder op D-day hadden gedaan. Alle informatie over het programma of over de gebeurtenis zelf worden als relevant beschouwd. &lt;/NL-narr&gt; &lt;/top&gt; Now we present our official runs. In the following three runs, we use English when submitting queries:</p><p>1. AUTOEN: the topic field is composed by the terms of the title of the original query (text between the title tags). The body field is composed by the text of the original query.</p><p>2. ADJUSTEN: the topic field is composed by the modified title by adding and/or removing terms. The adding terms are extracted from the original query text. The body field is composed by the text of the original query.</p><p>3. FEEDBCKEN: the topic field is composed by the modified title as in ADJUSTEN. The body field is composed by the original text query and the first relevant document (if it exists) in the first 30 documents found by the previous ADJUSTEN run.</p><p>Table <ref type="table">1</ref> shows the result of each run. Of course it's difficult to have a good result for the AUTOEN run while the topic contains all the concepts corresponding to the query title terms. We succeeded in improving the result of 63.11% by using the adjusted topic. Finally, by using the relevance feedback, we improved the result of 24.74%. This improvement is due to the vector space model which gives better results when the documents/queries vectors are long.  In order to compare the system results using different languages when submitting queries, we carried out three more runs: ADJUSTDU, ADJUSTFR, and ADJUSTSP. For each run, we use respectively Dutch, French, and Spanish when submitting queries. In these runs, topic field is composed by modified title like in ADJUSTEN and body field is composed by the original query text.</p><p>For all the four runs, we obtain almost the same mean average precision: 13.90% for ADJUSTDU, 13.47% for ADJUSTFR, 13.80% for ADJUSTSP and 16.85 % for ADJSUTEN. Our system is not dependent of the query language. It gives nearly the same results when submitting queries in four different languages. It's difficult to explain difference because the coverage and the quality of ontological dictionaries are important. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Conclusion</head><p>In this CLEF evaluation campaign, we evaluated a multilingual ontology-based approach for multilingual information retrieval. We did not use any translation either for documents or for queries. We carried out a common document/query representation based on multilingual ontology. Then, we used the vector space model for indexing and querying. Compared with the existing approaches, our approach has several advantages. Indeed, there is no dependency on automatic translators between all pairs of languages. When we add a new language, we only add, in the ontology, a new mapping dictionary. Also, we do not need any merging technique to rank the list of retrieved documents.</p><p>In this preliminary work, we tried only to prove the feasibility of our approach. We tried also to prove that our system is independent of the query language. We still have some limits in our system because we did not introduce any morpho-syntactic processing to break composite words in Dutch, German, or Finnish. Moreover, our ontology is incomplete and dirty (we have imported many errors with automatic translation).</p><p>We have also used the same approach in the bi-text alignment field. We have used other language like Chinese, Arabic and Russian <ref type="bibr" target="#b0">[Guyot 2005</ref>].</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>•</head><label></label><figDesc>8612 : a unit of length (in United States and Britain) equal to one twelfth of a foot• 28845: the thick short innermost digit of the forelimb.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head>Figure 1 .</head><label>1</label><figDesc>Figure 1. Example of concept definition in UNL[UNL].Figure 2. Example of term-concept association.</figDesc><graphic coords="2,73.20,218.68,209.28,217.92" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_2"><head>Figure 2 .</head><label>2</label><figDesc>Figure 1. Example of concept definition in UNL[UNL].Figure 2. Example of term-concept association.</figDesc><graphic coords="2,296.40,235.00,333.36,335.04" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_3"><head></head><label></label><figDesc>Description of the linguistic used resources.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_4"><head>Figure 3 .</head><label>3</label><figDesc>Figure 3. Indexing and querying process</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_5"><head>Figure 4 .Figure 5 .</head><label>45</label><figDesc>Figure 4. Comparison of the system result using three strategies</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_1"><head>Table 2 .</head><label>2</label><figDesc>Description and Mean Average Precision (MAP) of our official multilingual runs</figDesc><table><row><cell>Run name</cell><cell>Query language</cell><cell>Type</cell><cell>MAP</cell></row><row><cell>AUTO-EN</cell><cell>English</cell><cell>Automatic</cell><cell>10.33 %</cell></row><row><cell>ADJSUT-EN</cell><cell>English</cell><cell>Adjusted Topic</cell><cell>16.85 %</cell></row><row><cell>ADJUST-DU</cell><cell>Dutch</cell><cell>Adjusted Topic</cell><cell>13.90 %</cell></row><row><cell>ADJUST-FR</cell><cell>French</cell><cell>Adjusted Topic</cell><cell>13.47 %</cell></row><row><cell>ADJUST-SP</cell><cell>Spanish</cell><cell>Adjusted Topic</cell><cell>13.80 %</cell></row><row><cell>FEEDBCK-EN</cell><cell>English</cell><cell>Manual</cell><cell>21.02 %</cell></row></table></figure>
		</body>
		<back>

			<div type="acknowledgement">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Acknowledgments</head><p>We would like to thank the CLEF-2005 organisers for their efforts. We would also like to thank Metaread for giving us the possibility to use the "idxvli" information retrieval system (A fast indexer for big corpora).</p></div>
			</div>

			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">Combining query translation and document translation in cross-language retrieval</title>
		<author>
			<persName><forename type="first">Chen</forename><surname>Chen</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Gey</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Guyot</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Guyot</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Gruber</surname></persName>
		</author>
		<ptr target="2005" />
	</analytic>
	<monogr>
		<title level="m">R. A translation Approach to Portable Ontology Specifications, Knowledge Acquisition</title>
				<meeting><address><addrLine>Trondheim</addrLine></address></meeting>
		<imprint>
			<date type="published" when="1993">2003. 2005. 1993. 1993</date>
			<biblScope unit="volume">5</biblScope>
			<biblScope unit="page" from="199" to="220" />
		</imprint>
	</monogr>
	<note>yet another alignment algorithm -Alignement ontologique bi-</note>
</biblStruct>

<biblStruct xml:id="b1">
	<monogr>
		<ptr target="http://cui.unige.ch/isi/unl/&amp;http://www.undl.org/" />
		<title level="m">Universal Networking Language</title>
				<imprint/>
		<respStmt>
			<orgName>UNL</orgName>
		</respStmt>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<analytic>
		<title level="a" type="main">Introduction to the modern Information Retrieval</title>
		<author>
			<persName><surname>Salton</surname></persName>
		</author>
		<ptr target="http://snowball.tartarus.org/" />
	</analytic>
	<monogr>
		<title level="m">Snow 2005</title>
				<imprint>
			<publisher>McGraw-Hill</publisher>
			<date type="published" when="1983">1983</date>
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
