<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">ElhuyarIXA: semantic relatedness and crosslingual passage retrieval</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Eneko</forename><surname>Agirre</surname></persName>
							<email>e.agirre@ehu.es</email>
							<affiliation key="aff0">
								<orgName type="laboratory">IXA NLP Group</orgName>
								<orgName type="institution">University of the Basque Country. Donostia</orgName>
								<address>
									<country>Basque Country</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Olatz</forename><surname>Ansa</surname></persName>
							<email>olatz.ansa@ehu.es</email>
							<affiliation key="aff0">
								<orgName type="laboratory">IXA NLP Group</orgName>
								<orgName type="institution">University of the Basque Country. Donostia</orgName>
								<address>
									<country>Basque Country</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Xabier</forename><surname>Arregi</surname></persName>
							<email>xabier.arregi@ehu.es</email>
							<affiliation key="aff0">
								<orgName type="laboratory">IXA NLP Group</orgName>
								<orgName type="institution">University of the Basque Country. Donostia</orgName>
								<address>
									<country>Basque Country</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Maddalen</forename><surname>Lopez De Lacalle</surname></persName>
							<email>maddalen@elhuyar.com</email>
							<affiliation key="aff1">
								<orgName type="department">R&amp;D</orgName>
								<orgName type="institution">Elhuyar Foundation. Usurbil</orgName>
								<address>
									<country>Basque Country</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Arantxa</forename><surname>Otegi</surname></persName>
							<email>arantza.otegi@ehu.es</email>
							<affiliation key="aff0">
								<orgName type="laboratory">IXA NLP Group</orgName>
								<orgName type="institution">University of the Basque Country. Donostia</orgName>
								<address>
									<country>Basque Country</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Xabier</forename><surname>Saralegi</surname></persName>
							<email>xabiers@elhuyar.com</email>
							<affiliation key="aff1">
								<orgName type="department">R&amp;D</orgName>
								<orgName type="institution">Elhuyar Foundation. Usurbil</orgName>
								<address>
									<country>Basque Country</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Hugo</forename><surname>Zaragoza</surname></persName>
							<email>hugoz@yahooinc.com</email>
							<affiliation key="aff2">
								<orgName type="institution">Yahoo! Research. Barcelona</orgName>
								<address>
									<country key="ES">Spain</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">ElhuyarIXA: semantic relatedness and crosslingual passage retrieval</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">1E6CFFEA65026016D47BFA34BE6208EA</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-24T23:00+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>This article describes the participation of the joint ElhuyarIXA group in the RespubliQA exercise at QA&amp;CLEF. We put together tools developed separately and we combined and shared knowledge and technology between the two groups. In particular, we participated in the English-English monolingual task and in the Basque-English crosslingual one. Our focus has been threefold: (1) to check to what extent IR can achieve good results in passage retrieval without question analysis and answer validation, (2) to check Machine Readable Dictionary techniques for the Basque to English retrieval when faced with the lack of parallel corpora for Basque in this domain, and (3) to check the contribution of semantic relatedness based on WordNet to expand the passages to related words. Our results show that IR provides good results in the monolingual task, that our crosslingual system lowers the performance compared to the monolingual runs, and that semantic relatedness improves the results in both tasks (by 6 and 2 points, respectively).</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">Introduction</head><p>The joint team was formed by two different groups, on the one hand the Elhuyar Foundation, and on the other hand the IXA NLP group. The Elhuyar Foundation is a nonprofit making organization located in the Basque Country. The Elhuyar Foundation's mission is to popularize science and technology and promote the development of the Basque language. To achieve this, it offers the Basque public at large quality services, tools and resources that are a reference, with innovation being pivotal and with a commitment to operate in the area of education. Related to these objectives it deals with many activities, such as R&amp;D in Natural Language Processing and Information Retrieval fields. The IXA NLP group of the University of the Basque Country has previously participated at CLEF, specifically, in the CLEF 2008 Basque to Basque monolingual QA task <ref type="bibr" target="#b2">[3]</ref> and in the CLEF 2008 RobustWSD task. IXA has participated in the CLEF 2009 RobustWSD Task this year too.</p><p>Both Elhuyar and IXA considered that it would be interesting to share experience and knowledge on QA oriented (CL)IR. We decided to form a single team for participating in the ResPubliQA track. This collaboration allowed us to tackle the English-English monolingual task and the Basque-English crosslingual one.</p><p>Question answering systems typically rely on a passage retrieval system. Given that passages are shorter than documents, vocabulary mismatch problems are more important than in full document retrieval. Most of the previous work on expansion techniques has focused on pseudorelevance feedback and other query expansion techniques. In particular, WordNet has been used previously to expand the terms in the query with little success <ref type="bibr" target="#b8">[9,</ref><ref type="bibr" target="#b9">10,</ref><ref type="bibr" target="#b10">11,</ref><ref type="bibr" target="#b12">13]</ref>. The main problem is ambiguity, and the limited context available to disambiguate the word in the query effectively. As an alternative, we felt that passages would provide sufficient context to disambiguate and expand the terms in the passage. In fact, we do not do explicit WSD, but rather apply a stateoftheart semantic relatedness method <ref type="bibr" target="#b0">[1]</ref> in order to select the best terms to expand the documents.</p><p>With respect to the BasqueEnglish task, we met the challenge of retrieving English passages for Basque questions. We tackled this problem by translating the lexical units of the questions into English. The main setback is that no parallel corpus is available for Basque, given that there is no Basque version of the JRCAcquis collection. So we have explored a corpus parallel free approach for translating queries which could also be interesting for other less resourced languages. Even so, we regarded the crosslingual exercise as interesting. In our opinion, bearing in mind the idiosyncrasy of the European Union, it is worthwhile dealing with the search of passages that answer questions formulated in unofficial languages.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">System overview 2.1 Question preprocessing</head><p>We analysed the Basque questions by reusing the linguistic processors of the Ihardetsi questionanswering system <ref type="bibr" target="#b2">[3]</ref>. This module uses two general linguistic processors: the lemmatizer/tagger named Morfeus <ref type="bibr" target="#b6">[7]</ref>, and the Named Entity Recognition and Classification (NERC) processor called Eihera <ref type="bibr" target="#b1">[2]</ref>. The use of the lemmatizer/tagger is particularly suited to Basque, as it is an agglutinative language. It returns only one lemma and one part of speech for each lexical unit, which includes single word terms and multiword terms (MWT) (those included in the MRD introduced in the next subsection). The NERC processor, Eihera, captures entities such as person, organization and location. The numerical and temporal expressions are captured by the lemmatizer/tagger. The questions thus analyzed are passed to the translation module.</p><p>English queries were tokenized without further analysis.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2">Translation of the query terms (BasqueEnglish runs)</head><p>Once the questions had been linguistically processed, we translated them into English. Among the main strategies and methods proposed in the literature to deal with language barriers in IR problems we adopted a Machine Readable Dictionary (MRD)based method. Due to the scarcity of parallel corpora for a small language or even for big languages in certain domains we have explored a MRDbased method. However, MRDbased approaches have inherent problems, such as the presence of ambiguous translations and outofvocabulary (OOV) words. To tackle these problems, both translation ambiguity and OOV words, some techniques have been proposed such as structured querybased techniques <ref type="bibr" target="#b5">[6,</ref><ref type="bibr" target="#b13">14]</ref> and concurrencesbased techniques <ref type="bibr" target="#b3">[4,</ref><ref type="bibr" target="#b7">8,</ref><ref type="bibr" target="#b11">12]</ref>. These approaches have been compared for Basque by obtaining best MAP results with structured queries <ref type="bibr" target="#b14">[15]</ref>. However, structured queries were not supported in the retrieval algorithm used (see Section 2.3), so we adopted a concurrencesbased translation selection strategy.</p><p>The translation process designed comprises two steps and takes the keywords (Name Entities, MWT and singles words tagged as noun, adjective or verb) of the question as source words.</p><p>In the first step the translation candidates of each source word are obtained. The translation candidates for the lemmas of the source words are taken from a bilingual euen MRD composed from the BasqueEnglish Morris dictionary<ref type="foot" target="#foot_0">1</ref> , and the Euskalterm terminology bank<ref type="foot" target="#foot_1">2</ref> which includes 38,184 MWTs. After that, OOV words and ambiguous translations are dealt with. The number of OOV words quantified out of a total of 421 keywords for the 77 questions of the development set was 42 (10%). These 77 questions were translated by hand from English to Basque in order to carry out the development phase. Nevertheless, it must be said that many of these OOV words were wrongly tagged lemmas and entities. We deal with OOV words by searching for their cognates in the target collection. The cognate detection is done in two phases. Firstly, we apply several transliteration rules to the source word. Then we calculate the Longest Common Subsequence Ratio (LCSR) among words with a similar length (+10%) from the target collection (see Figure <ref type="figure" target="#fig_0">1</ref>). The ones which reach a previously established threshold (0.9) are selected as translation candidates. The attempt to select the best translation candidate will be held in the translation selection phase. The MWT terms that are not found in the dictionary are translated word by word, as we realized that most of the MWT could be translated correctly in that way, exactly 91% of the total MWTs identified by hand in the 77 development questions. In the second step of the translation process we perform a translation selection step. In the translation selection step, we select the best translation of each source keyword according to an algorithm based on target collection concurrences. This algorithm sets out to obtain the translation candidate combination that maximizes their global association degree. We take the algorithm proposed by Monz and Dorr <ref type="bibr" target="#b11">[12]</ref>.</p><p>Initially, all the translation candidates are equally likely. Assuming that t is a translation candidate of the set of all candidates tr s i  for a query term s i given by the MRD, then:</p><p>Initialization step:</p><formula xml:id="formula_0">w T 0  t|s i  = 1 |tr  s i  |</formula><p>In the iteration step, each translation candidate is iteratively updated using the weights of the rest of the candidates and the weight of the link connecting them.</p><p>Iteration step:</p><formula xml:id="formula_1">w T n  t|s i =w T n−1  t|s i   ∑ t' ∈inlink t  w L  t,t' •w T  t'|s i </formula><p>where inlink  t  is the set of translation candidates that are linked to t, and w L t ,t '  is the association degree between t and t' on the target passages measured by Loglikelihood ratio. These concurrences were calculated by taking the passages of the documents of the target collection as window.</p><p>After recomputing each term weight they are normalized. Normalization step:</p><formula xml:id="formula_2">w L n  t|s i  = w L n  t |s i ∑ m=1 |tr s i | w L n  t i,m |s i </formula><p>The iteration stops when the variations of the term weights become smaller than a predefined threshold.</p><p>We have modified the iteration step by adding a factor w T t,t'  to increase the association degree w L t ,t '  between translation candidates t and t' whose source words w T t,t'  are near each other (distance dis in words is low) in the source query Q, and whose source words so t , so t '  belong to the same MultiWord Unit (MWU) Z smw t ,t '  . As the global association degree between translation candidates is estimated from the association degree of pairs of candidates, we score positively these two characteristics when the association degree for a pair of candidates is calculated. Thus, the modified association degree w ' L t , t '  between t and t' will be calculated in this way in the iteration step:</p><formula xml:id="formula_3">w ' L t , t ' =w L t , t '  * w T t , t '  w T t , t ' = max s i , s j ∈Q  dis s i , s j  dis so t  , so t '  * 2 smwt , t' </formula><p>smw t ,t ' = { 1 {so t , so t ' }⊆Z where Z ∈MWU 0</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3">Passage retrieval</head><p>The purpose of the passage retrieval module is to retrieve passages from the document collection which are likely to contain an answer. The main feature of this module is that the passages are expanded based on their related concepts, as explained in the following sections.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3.1">Document preprocessing and application of semantic relatedness</head><p>Given that the system needs to return paragraphs, we first split the document collection into paragraphs, which are delimited by the mark &lt;p&gt; in the documents. Then we lemmatized and POS tagged those passages using the OpenNLP open source software <ref type="foot" target="#foot_2">3</ref> .</p><p>After preprocessing the documents, we expanded the passages based on semantic relatedness. To this end, we used UKB <ref type="foot" target="#foot_3">4</ref> , a collection of programs for performing graphbased Word Sense Disambiguation and lexical similarity/relatedness using a preexisting knowledge base <ref type="bibr" target="#b0">[1]</ref>, in this case WordNet 3.0.</p><p>Given a passage, UKB returns a vector of scores for concepts in WordNet. Each of these concepts has a score, and the higher the score, the more related the concept is to the given passage, where we represent the passage using the lemmas of all nouns, verbs, adjectives and adverbs in the passage.</p><p>Given the list of related concepts, we took the highestscoring 100 concepts and expanded them to all variants (words that lexicalize the concepts) in WordNet. An example of a document expansion is shown in Figure <ref type="figure">2</ref>.</p><p>The variants for those expanded concepts were included in a new field of the passage representation. This way, we were able to use the original words only, or alternatively, to also include the variants for the most related 100 concepts, as we will be explaining in Section 2.3.2 and Section 2.3.3.</p><p>We applied the expansion strategy only to passages which had more than 10 words (half of the passages), for two reasons: the first one is that most of these passages were found not to contain relevant information for the task (e.g. "Article 2", "Having regard to the proposal from the Commission" or "HAS ADOPTED THIS REGULATION"), and the second is that we thus saved some computation time.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3.2">Indexing</head><p>We indexed the new expanded documents using the MG4J searchengine <ref type="bibr" target="#b4">[5]</ref>. MG4J makes it possible to combine several indices over the same document collection. We created one index for each field: one for the original words and one for the expanded words. Porter stemmer was used as per usual.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3.3">Retrieval</head><p>We used the BM25 ranking function with the following parameters: 1.0 for k1 and 0.6 for b. We did not tune these parameters. MG4J allows multiindex queries, where one can specify which of the indices one wants to search in, and assign different weights to each index. We conducted different experiments, by using the original words alone (the index made of original words) and also by using the index with the expansion of concepts, giving different weights to the original words and the expanded concepts. The weight of the index which was created using the original words from the passages was 1.00 for all the runs. 1.00 was also the weight of the index that included the expanded words for the monolingual run, but it was 1.78 for the bilingual run. These weights were fixed following a training phase with the English development questions provided by the organization, and after the Basque questions had been translated by hand (as no development Basque data was released).The submitted runs are described in the next section.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3">Description of runs</head><p>We participated in the EnglishEnglish monolingual task and the BasqueEnglish crosslingual task, with two runs per language pair. We did not analyze the English queries for the monolingual run, and we just removed the stopwords.</p><p>For the bilingual runs, we first analyzed the questions (see Section 2.1), then we translated the question terms from Basque to English (see Section 2.2), and, finally, we retrieved the relevant passages for the translated query terms (see Section 2.3).</p><p>As we were interested in the performance of passage retrieval on its own, we did not carry out any answer validation, and we just chose the first passage returned by the passage retrieval module as the response. We did not leave any question unanswered.</p><p>For both tasks, the only difference between the submitted two runs is the use (or not) of the expansion in the passage retrieval module. That is, in the first run ("run 1" in Table <ref type="table" target="#tab_0">1</ref>), during the retrieval we only used the original words that were in the passage. In the second run ("run 2" in Table <ref type="table" target="#tab_0">1</ref>), apart from the original words, we also used the expanded words.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">Results</head><p>Table <ref type="table" target="#tab_0">1</ref> summarizes the results of our submitted runs, explained in Section 3. The results show that the use of the expanded words (run 2) was effective for both tasks, improving the final result by 6 % in the monolingual task. Figure <ref type="figure">2</ref> shows an example of a document expansion which was effective for answering the English question number 32: "Into which plant may genes be introduced and not raise any doubts about unfavourable consequences for people's health?" doc_id: jrc31998D0293en.xml p_id: 17 original passage: Whereas the Commission, having examined each of the objections raised in the light of Directive 90/220/EEC, the information submitted in the dossier and the opinion of the Scientific Committee on Plants, has reached the conclusion that there is no reason to believe that there will be any adverse effects on human health or the environment from the introduction into maize of the gene coding for phosphinotricine acetyltransferase and the truncated gene coding for betalactamase; some expanded words: cistron factor gene coding cryptography secret_writing ... acetyl acetyl_group acetyl_radical ethanoyl_group ethanoyl_radical beta_lactamase penicillinase common_market ec eec eu europe european_community european_economic_community european_union ... directive directing directional guiding citizens_committee committee environment environs surround surroundings corn indian_corn maize zea_mays health wellness health adverse contrary homo human human_being man adverse inauspicious untoward gamboge lemon lemon_yellow ... unfavorable unfavourable ... set_up expostulation objection remonstrance remonstration dissent protest believe light lightly belief feeling impression notion opinion ... reason reason_out argue jurisprudence law consequence effect event issue outcome result upshot ...</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>#answered</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Figure 2: Example of a document expansion</head><p>In the last part of the example we can see some words that we obtained after applying the expansion process explained in Section 2.3.1 to the original passage showed in the example too. As we can see, there are some new words among the expanded words that are not in the original passage, such as unfavourable or consequence. Those two words were in the question we mentioned before (number 32). That could be why we answered that question correctly when using the expanded words (in run 2), but not when using the original words only (without using the expanded words, in run 1).</p><p>As expected, the best results were obtained in the monolingual task. With the intention of finding reasons to explain the significant performance drop in the bilingual run, we analyzed manually 100 query translations obtained in the query translation process of the 500 test queries, and detected several types of errors arising from both the question analysis process and from the query translation process. In the question analysis process, some lemmas were not correctly identified by the lemmatizer/tagger, and in other cases some entities were not returned by the lemmatizer/tagger causing us to lose important information for the subsequent translation and retrieval processes. In the query translation process, leaving aside the incorrect translation selections, the words appearing in the source questions were not exactly the ones that figured in many queries that had been correctly translated. In most cases this happened because the English source query word was not a translation candidate in the MRD. If we assume that the answers contain words that appear in the questions and therefore in the passage that we must return, this will negatively affect the final retrieval process.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5">Conclusions</head><p>The joint ElhuyarIxa team has presented a system which works on passage retrieval alone, without any question analysis and answer validation steps. Our EnglishEnglish results show that good results can be achieved by means of this simple strategy. We experimented with applying semantic relatedness in order to expand passages prior to indexing, and the results are highly positive, especially for EnglishEnglish. The performance drop in the BasqueEnglish bilingual runs is significant, and is caused by the accumulation of errors in the analysis and translation of the query mentioned.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>Figure 1 :</head><label>1</label><figDesc>Figure 1: Example of cognate detection</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_0"><head>Table 1 :</head><label>1</label><figDesc>Results for submitted runs</figDesc><table><row><cell></cell><cell></cell><cell>correctly</cell><cell>#answered incorrectly</cell><cell>c@1</cell></row><row><cell>English English</cell><cell>run 1 run 2</cell><cell>211 240</cell><cell>289 260</cell><cell>0.42 0.48</cell></row><row><cell>Basque English</cell><cell>run 1 run 2</cell><cell>78 91</cell><cell>422 409</cell><cell>0.16 0.18</cell></row></table></figure>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_0">English/Basque dictionary including 67,000 entries and 120,000 senses.</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2" xml:id="foot_1">Terminological dictionary including 100,000 terms in Basque with equivalences in Spanish, French, English and Latin.</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="3" xml:id="foot_2">http://opennlp.sourceforge.net/</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="4" xml:id="foot_3">The algorithm is publicly available at http://ixa2.si.ehu.es/ukb/</note>
		</body>
		<back>

			<div type="acknowledgement">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Acknowledgments</head><p>This work has been supported by KNOW (TIN200615049C0301), imFUTOURnet (IE08233) and KYOTO (ICT2007211423). Arantxa Otegi's work is funded by a PhD grant from the Basque Government.</p></div>
			</div>

			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">A study on similarity and relatedness using distributional and WordNetbased approaches</title>
		<author>
			<persName><forename type="first">E</forename><surname>Agirre</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Soroa</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Alfonseca</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Hall</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Kravalova</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Pasca</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of annual meeting of the North American Chapter of the Association of Computational Linguistics (NAACL)</title>
				<meeting>annual meeting of the North American Chapter of the Association of Computational Linguistics (NAACL)<address><addrLine>Boulder, USA</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2009-06">June 2009</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<analytic>
		<title level="a" type="main">Development of a Named Entity Recognizer for an Agglutinative Language</title>
		<author>
			<persName><forename type="first">I</forename><surname>Alegria</surname></persName>
		</author>
		<author>
			<persName><forename type="first">O</forename><surname>Arregi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Balza</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Ezeiza</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Fernandez</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Urizar</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">IJCNLP</title>
				<imprint>
			<date type="published" when="2004">2004</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<analytic>
		<title level="a" type="main">Ihardetsi question answering system at QA@CLEF</title>
		<author>
			<persName><forename type="first">O</forename><surname>Ansa</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><surname>Arregi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Otegi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Soraluze</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Working Notes of the CrossLingual Evaluation Forum</title>
				<meeting><address><addrLine>Aarhus, Denmark</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2008">2008. 2008</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">Resolving Ambiguity for Crosslanguage Retrieval</title>
		<author>
			<persName><forename type="first">L</forename><surname>Ballesteros</surname></persName>
		</author>
		<author>
			<persName><forename type="first">W</forename><forename type="middle">Bruce</forename><surname>Croft</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval</title>
				<meeting>the 21st annual international ACM SIGIR conference on Research and development in information retrieval</meeting>
		<imprint>
			<date type="published" when="2008">2008</date>
			<biblScope unit="page" from="64" to="71" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<analytic>
		<title level="a" type="main">MG4J at TREC</title>
		<author>
			<persName><forename type="first">P</forename><surname>Boldi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Vigna</surname></persName>
		</author>
		<ptr target="http://mg4j.dsi.unimi.it/" />
	</analytic>
	<monogr>
		<title level="m">The Fourteenth Text REtrieval Conference (TREC 2005) Proceedings, number SP 500266 in Special Publications</title>
				<editor>
			<persName><forename type="first">Ellen</forename><forename type="middle">M</forename><surname>Voorhees</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">Lori</forename><forename type="middle">P</forename><surname>Buckland</surname></persName>
		</editor>
		<imprint>
			<publisher>NIST</publisher>
			<date type="published" when="2005">2005. 2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">Probabilistic structured Query Methods</title>
		<author>
			<persName><forename type="first">K</forename><surname>Darwish</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">W</forename><surname>Oard</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 26th annual international ACM SIGIR conference on Research and development in information retrieval</title>
				<meeting>the 26th annual international ACM SIGIR conference on Research and development in information retrieval</meeting>
		<imprint>
			<date type="published" when="2003">2003</date>
			<biblScope unit="page" from="338" to="344" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b6">
	<analytic>
		<title level="a" type="main">Combining Stochastic and RuleBased Methods for Disambiguation in Agglutinative Languages</title>
		<author>
			<persName><forename type="first">N</forename><surname>Ezeiza</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Aduriz</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Alegria</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">M</forename><surname>Arriola</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Urizar</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">COLINGACL</title>
				<imprint>
			<date type="published" when="1998">1998</date>
			<biblScope unit="page" from="380" to="384" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b7">
	<analytic>
		<title level="a" type="main">Improving Query Translation for Crosslanguage Information Retrieval using Statistcal Models</title>
		<author>
			<persName><forename type="first">J</forename><surname>Gao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">Y</forename><surname>Nie</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Xun</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Zhang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Zhou</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Huang</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 24th annual international ACM SIGIR conference on Research an development in information retrieval</title>
				<meeting>the 24th annual international ACM SIGIR conference on Research an development in information retrieval</meeting>
		<imprint>
			<date type="published" when="2001">2001</date>
			<biblScope unit="page">96104</biblScope>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b8">
	<analytic>
		<title level="a" type="main">Information retrieval using word senses: Root sense tagging approach</title>
		<author>
			<persName><forename type="first">S</forename><surname>Kim</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Seo</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Rim</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of SIGIR</title>
				<meeting>SIGIR</meeting>
		<imprint>
			<date type="published" when="2004">2004</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b9">
	<analytic>
		<title level="a" type="main">An effective approach to document retrieval via utilizing WordNet and recognizing phrases</title>
		<author>
			<persName><forename type="first">S</forename><surname>Liu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Liu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Yu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">W</forename><surname>Meng</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of SIGIR</title>
				<meeting>SIGIR</meeting>
		<imprint>
			<date type="published" when="2004">2004</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b10">
	<analytic>
		<title level="a" type="main">Word Sense Disambiguation in Queries</title>
		<author>
			<persName><forename type="first">S</forename><surname>Liu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Yu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">W</forename><surname>Meng</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of ACM Conference on Information and Knowledge Managment</title>
				<meeting>ACM Conference on Information and Knowledge Managment</meeting>
		<imprint>
			<publisher>CIKM</publisher>
			<date type="published" when="2005">2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b11">
	<analytic>
		<title level="a" type="main">Iterative translation disambiguation for crosslanguage Information Retrieval</title>
		<author>
			<persName><forename type="first">C</forename><surname>Monz</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><forename type="middle">J</forename><surname>Dorr</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 28th annual international ACM SIGIR conference on Research and development in information retrieval</title>
				<meeting>the 28th annual international ACM SIGIR conference on Research and development in information retrieval</meeting>
		<imprint>
			<date type="published" when="2005">2005</date>
			<biblScope unit="page">520527</biblScope>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b12">
	<analytic>
		<title level="a" type="main">UCMY!R at CLEF 2008 Robust and WSD tasks</title>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">R</forename><surname>Pérezagüera</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Zaragoza</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Working Notes of the CrossLingual Evaluation Forum</title>
				<meeting><address><addrLine>Aarhus, Denmark</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2008">2008</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b13">
	<analytic>
		<title level="a" type="main">The effects of query structure and dictionary setups in dictionarybased crosslanguage information retrieval</title>
		<author>
			<persName><forename type="first">A</forename><surname>Pirkola</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval</title>
				<meeting>the 21st annual international ACM SIGIR conference on Research and development in information retrieval</meeting>
		<imprint>
			<date type="published" when="1998">1998</date>
			<biblScope unit="page">5563</biblScope>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b14">
	<analytic>
		<title level="a" type="main">Comparing different approaches to treat Translation Ambiguity in CLIR: Structured Queries v. Target Cooccurrence Based Selection</title>
		<author>
			<persName><forename type="first">X</forename><surname>Saralegi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>López De Lacalle</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">6th TIR workshop</title>
				<imprint>
			<date type="published" when="2009">2009</date>
		</imprint>
	</monogr>
	<note>To appear</note>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
