<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Web Retrieval Experiments with the EuroGOV Corpus at the University of Hildesheim</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Niels</forename><surname>Jensen</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">Information Science</orgName>
								<orgName type="institution">University of Hildesheim</orgName>
								<address>
									<addrLine>Marienburger Platz 22</addrLine>
									<postCode>D-31141</postCode>
									<settlement>Hildesheim</settlement>
									<country key="DE">Germany</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">René</forename><surname>Hackl</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">Information Science</orgName>
								<orgName type="institution">University of Hildesheim</orgName>
								<address>
									<addrLine>Marienburger Platz 22</addrLine>
									<postCode>D-31141</postCode>
									<settlement>Hildesheim</settlement>
									<country key="DE">Germany</country>
								</address>
							</affiliation>
						</author>
						<author role="corresp">
							<persName><forename type="first">Thomas</forename><surname>Mandl</surname></persName>
							<email>mandl@uni-hildesheim.de</email>
							<affiliation key="aff0">
								<orgName type="department">Information Science</orgName>
								<orgName type="institution">University of Hildesheim</orgName>
								<address>
									<addrLine>Marienburger Platz 22</addrLine>
									<postCode>D-31141</postCode>
									<settlement>Hildesheim</settlement>
									<country key="DE">Germany</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Robert</forename><surname>Strötgen</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">Information Science</orgName>
								<orgName type="institution">University of Hildesheim</orgName>
								<address>
									<addrLine>Marienburger Platz 22</addrLine>
									<postCode>D-31141</postCode>
									<settlement>Hildesheim</settlement>
									<country key="DE">Germany</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">Web Retrieval Experiments with the EuroGOV Corpus at the University of Hildesheim</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">07A9AC91AF0CC80744063C2C6ED6E20E</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-25T00:40+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>H.3 [Information Storage and Retrieval]: H.3.1 Content Analysis and Indexing</term>
					<term>H.3.3 Information Search and Retrieval</term>
					<term>H.3.4 Systems and Software Measurement, Performance, Experimentation Web Retrieval, Multilingual Information Retrieval, N-gram Indexing, Evaluation</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>In the CLEF 2005 initiative, multlingual web retrieval was integrated as a task for the first time. This paper describes experiments based on one multilingual index carried out at the University of Hildesheim. Several indexing strategies based on a multi-lingual index have been tested with the EuroGOV corpus. Boosting topic fields with higher weight led to best results during post submission runs. The experiments also led to experiences in working with large test collections and the challenges associated with them.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">Introduction</head><p>Web search engines has become a part of every day life for many people. The development of information retrieval systems for the web is faced with many challenges <ref type="bibr">(Arasu et al. 2001)</ref>. Systems give different answers to these challenges and it is difficult to judge the effect of decisions during the design of search enigne. As a consequence, there is a great need for evaluation in web retrieval <ref type="bibr" target="#b6">(Hawking 2000)</ref>. The web is also a natural source for multilingual documents.</p><p>Within the Cross Language Evaluation Forum (CLEF) the web track has been created <ref type="bibr">(Sigurbjörnsson et al. 2005b)</ref>. A large multilingual corpus has been collected and distributed <ref type="bibr">(Sigurbjörnsson et al. 2005a</ref>). In our first participation, we intended to tune our system to the challanges of a large web corpus. For the experiments, language resources in all languages were not available from ad-hoc retrieval. As a consequence, we considered n-gram indexing for the web retrieval task <ref type="bibr">(McNamee &amp; Mayfield 2004</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">Data Pre-Processing</head><p>Since the files of the EuroGOV corpus were not released in well formed XML, substantial effort for data preprocessing was necessary. A corpus in well formed XML would allow us to use the System implemented during the CLEF 2004 campaign for multilingual ad-hoc tasks <ref type="bibr">(Hackl et al. 2005)</ref>. The two main items in the EuroGOV files that needed replacing were predeclared entities. This step was required for ampersand with the associated entity reference in the URL fields of the individual documents and all nested CDATA tags. The first attempt to reformat the files has been carried out by a Perl-script. At the first view, it seemed that the Perl-scirpt would work perfectly for our needs. Unfortunately, we realized that during the process of indexing the corpus, the XML parser would frequently report "parser exceptions" that we traced back to the fact that the XML files still contained a couple of not adjusted predeclared entities. Having this in mind, a Java program was developed that worked through the whole corpus perfectly. It seems that Perl is not able to process EuroGOV files bigger than 250 MB since we successfully tested the Perl-script with the small-size files (22 MB, 59 MB &amp; 220 MB) of the corpus.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3">Submitted Retrieval Experiments with EuroGOV</head><p>As mentioned in the introduction, one multilingual index was created. In order to generate a slim index we assembled a multilingual The bases for this list were the stopwordlists supplied by the University of Neuchatel<ref type="foot" target="#foot_0">1</ref> and a list developed specifically for the Czech language (Hofman Miquel 2005). All lists were combined and revised into one file. This multilingual stopwordlist covers twelve languages and was used for the indexing process of the corpus.</p><p>For our retrieval experiments, we created three different multilingual indexes. Two were created with the Lucene StandardAnalyzer 2 , which does not implement any linguistic processing apart from word segmentation. The first index covered the whole corpus whereas the second index cut off the indexing process after a maximum of 200 characters for each individual document. Due to this approach, the sizes of the indexes varies from 5 GB to 700 MB.</p><p>The third index was created with a NGram Analyzer also applied to multilingual ad-hoc retrieval before <ref type="bibr">(Hackl et al. 2005</ref>). Because of performance and time restrictions the trigram approach was only applied to the title field of the individual documents in the corpus files. As a result the size of the index is down to 300 MB which led to a very quick and stable performance at retrieval time. These three indexes are the foundation for our experiments. As a main retrieval engine, we used Lucene 1.4 3 . Some of the basic code for retrieval and n-gram analysis was adopted from previous CLEF ad-hoc experiments <ref type="bibr">(Hackl et al. 2005)</ref>. Six different baseline runs were submitted. We did not use any of the metadata that was supplied by the topics due to time and resource constraints. Our monolingual queries were created with the title field of the topic whereas the multilingual queries were based on the monolingual title field and the translation language English field. Both types of queries were sent to one multilingual index. Results are shown in table <ref type="table" target="#tab_0">1</ref>. Looking at the results of the submitted runs it becomes clear that the trigram index did not confirm the expectations. On average, the monolingual runs differ from the multilingual runs by about 0.0162 MRR points.</p><p>Having those results in mind the method of indexing the corpus with the Lucene StandardAnalyzer turned out to be more effective than the trigram strategy. The post experiments will illustrate this effect more clearly.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3">Post Submission Experiments with EuroGOV</head><p>For our post experiments we decided to generate another trigram index covering the whole corpus. The purpose of this experiment was to confirm the results from the official runs or to improve them by providing a better or more a complete index respectively. We also wanted to see if through boosting of the individual query fields (title &amp; translation language English) the difference between the mono-and multilingual runs could be compensated or even improved. As table 2 shows quite obviously, even a more complete index was not able to improve the MRR of the trigram runs. The results declined by approx. 50 %. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">Conclusion and Outlook</head><p>For the first web track at CLEF we intended to tune our system to be able to cope with a large amount of data. We suceeded in returning valid results for several runs.</p><p>In future experiments, we intend to step beyond the baseline runs and try to involve the metadata that is being provided by the WebCLEF topics. We also want to include advanced quality measures into consideration. Link based quality measures seem to be integral part of commercial search engines. They have been evaluated at the web track at TREC <ref type="bibr" target="#b6">(Hawking 2000)</ref>. Advanced quality measures take more features into account, especially information and design aspects <ref type="bibr" target="#b10">(Mandl 2005)</ref>.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_0"><head>Table 1 .</head><label>1</label><figDesc>WebCLEF 2005 results University of Hildesheim</figDesc><table><row><cell></cell><cell cols="5">UHi3TiMo UHi3TiMu UHiScoMo UHiScoMu UHiSMo</cell><cell>UHiSMu</cell></row><row><cell>Mean reciprocal rank</cell><cell>0.0373</cell><cell>0.0274</cell><cell>0.1301</cell><cell>0.1147</cell><cell>0.1603</cell><cell>0.137</cell></row><row><cell>Average success at 1</cell><cell>0.0219</cell><cell>0.0146</cell><cell>0.1024</cell><cell>0.0932</cell><cell>0.1261</cell><cell>0.1097</cell></row><row><cell>Average success at 5</cell><cell>0.0512</cell><cell>0.0402</cell><cell>0.1627</cell><cell>0.1353</cell><cell>0.2011</cell><cell>0.1627</cell></row><row><cell>Average success at 10</cell><cell>0.064</cell><cell>0.0494</cell><cell>0.1883</cell><cell>0.1609</cell><cell>0.2194</cell><cell>0.1927</cell></row><row><cell>Average success at 20</cell><cell>0.075</cell><cell>0.064</cell><cell>0.2322</cell><cell>0.192</cell><cell>0.2523</cell><cell>0.2249</cell></row><row><cell>Average success at 50</cell><cell>0.1024</cell><cell>0.0878</cell><cell>0.2505</cell><cell>0.2157</cell><cell>0.287</cell><cell>0.2578</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_1"><head>Table 2 .</head><label>2</label><figDesc>Results of the trigram index runIn the second part of our post experiments we took the four indexes we had generated, and modified the weights of the query fields. The ratio for the two query fields were 10 to 1 and vice versa. The results that are shown in table 3 and 4 show that by boosting the title field of the query the results improve by 0.0144 MRR points on average. Applying this procedure, the performance of the multilingual run based on the StandardAnalyzer Index results in higher MRR values. The boosted multilingual run has a better result than any monolingual run and is the best run of all our experiments.</figDesc><table><row><cell></cell><cell cols="2">UHi3Mo UHi3Mu</cell></row><row><cell>Mean reciprocal rank</cell><cell>0.0169</cell><cell>0.0099</cell></row><row><cell>Average success at 1</cell><cell>0.0091</cell><cell>0.0037</cell></row><row><cell>Average success at 5</cell><cell>0.0238</cell><cell>0.0183</cell></row><row><cell>Average success at 10</cell><cell>0.0366</cell><cell>0.0238</cell></row><row><cell>Average success at 20</cell><cell>0.042</cell><cell>0.0311</cell></row><row><cell>Average success at 50</cell><cell>0.0548</cell><cell>0.0402</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_2"><head>Table 3 .</head><label>3</label><figDesc>Translation language English field Boost 10 to 1</figDesc><table><row><cell></cell><cell cols="4">UHi3MuBo110 UHi3TiMuBo110 UHiScoMuBo110 UHiSMuBo110</cell></row><row><cell>Mean reciprocal rank</cell><cell>0.0063</cell><cell>0.0139</cell><cell>0.0677</cell><cell>0.0811</cell></row><row><cell>Average success at 1</cell><cell>0.0018</cell><cell>0.0091</cell><cell>0.053</cell><cell>0.0658</cell></row><row><cell>Average success at 5</cell><cell>0.0073</cell><cell>0.0165</cell><cell>0.0786</cell><cell>0.0987</cell></row><row><cell>Average success at 10</cell><cell>0.0146</cell><cell>0.0238</cell><cell>0.1079</cell><cell>0.1133</cell></row><row><cell>Average success at 20</cell><cell>0.0201</cell><cell>0.0293</cell><cell>0.1207</cell><cell>0.1316</cell></row><row><cell>Average success at 50</cell><cell>0.0402</cell><cell>0.0512</cell><cell>0.128</cell><cell>0.1444</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_3"><head>Table 4 .</head><label>4</label><figDesc>Title</figDesc><table><row><cell></cell><cell></cell><cell>field Boost 10 to 1</cell><cell></cell><cell></cell></row><row><cell></cell><cell cols="4">UHi3MuBo101 UHi3TiMuBo101 UHiScoMuBo101 UHiSMuBo101</cell></row><row><cell>Mean reciprocal rank</cell><cell>0.0172</cell><cell>0.0379</cell><cell>0.1307</cell><cell>0.1608</cell></row><row><cell>Average success at 1</cell><cell>0.0091</cell><cell>0.0219</cell><cell>0.1042</cell><cell>0.1298</cell></row><row><cell>Average success at 5</cell><cell>0.0256</cell><cell>0.053</cell><cell>0.1609</cell><cell>0.1974</cell></row><row><cell>Average success at 10</cell><cell>0.0329</cell><cell>0.0622</cell><cell>0.1883</cell><cell>0.2176</cell></row><row><cell>Average success at 20</cell><cell>0.0439</cell><cell>0.075</cell><cell>0.2285</cell><cell>0.245</cell></row><row><cell>Average success at 50</cell><cell>0.053</cell><cell>0.1042</cell><cell>0.2486</cell><cell>0.2724</cell></row></table></figure>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_0">Stopwordlists: http://www.unine.ch/Info/clef/ verified August 11th</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2005" xml:id="foot_1">2 Lucene StandardAnalyzer: http://lucene.apache.org verified on August 11th 2005 3 Lucene: http://lucene.apache.org verified August 11th 2005</note>
		</body>
		<back>
			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<monogr>
		<title/>
		<author>
			<persName><forename type="first">Arvind</forename><surname>Arasu</surname></persName>
		</author>
		<imprint/>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<monogr>
		<title/>
		<author>
			<persName><forename type="first">Junghoo</forename><surname>Cho</surname></persName>
		</author>
		<imprint/>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<monogr>
		<title/>
		<author>
			<persName><forename type="first">Hector</forename><forename type="middle">;</forename><surname>Garcia-Molina</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Andreas</forename><surname>Paepcke</surname></persName>
		</author>
		<imprint/>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">Searching the Web</title>
		<author>
			<persName><forename type="first">Sriram</forename><surname>Raghavan</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">ACM Transactions on Internet Technology</title>
		<imprint>
			<biblScope unit="volume">1</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page" from="2" to="43" />
			<date type="published" when="2001">2001</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<monogr>
		<title/>
		<author>
			<persName><forename type="first">René</forename><surname>Hackl</surname></persName>
		</author>
		<imprint/>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">Mono-and Cross-lingual Retrieval Experiments at the University of Hildesheim</title>
		<author>
			<persName><forename type="first">Thomas</forename><forename type="middle">;</forename><surname>Mandl</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Christa</forename><surname>Womser-Hacker</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Multilingual Information Access for Text, Speech and Images: Results of the Fifth CLEF Evaluation Campaign</title>
		<title level="s">Lecture Notes in Computer Science</title>
		<editor>et al</editor>
		<meeting><address><addrLine>Berlin</addrLine></address></meeting>
		<imprint>
			<publisher>Springer</publisher>
			<date type="published" when="2005">2005</date>
			<biblScope unit="volume">3491</biblScope>
			<biblScope unit="page" from="165" to="169" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b6">
	<analytic>
		<title level="a" type="main">Overview of the TREC-9 Web Track</title>
		<author>
			<persName><forename type="first">David</forename><surname>Hawking</surname></persName>
		</author>
		<ptr target="http://trec.nist.gov/pubs/trec9/t9_proceedings.html" />
	</analytic>
	<monogr>
		<title level="m">The Ninth Text Retrieval Conference (TREC-9)</title>
				<meeting><address><addrLine>Gaithersburg, Maryland</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2000-11">2000. November 2000</date>
			<biblScope unit="page" from="500" to="249" />
		</imprint>
		<respStmt>
			<orgName>National Institute of Standards and Technology</orgName>
		</respStmt>
	</monogr>
	<note>NIST Special Publication</note>
</biblStruct>

<biblStruct xml:id="b7">
	<monogr>
		<title level="m" type="main">Informationslinguistische Ressourcen für das Information Retrieval in der tschechischen Sprache im Rahmen des Cross Language Evaluation Forums</title>
		<author>
			<persName><forename type="first">Hofman</forename><surname>Miquel</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Laura</forename></persName>
		</author>
		<imprint>
			<date type="published" when="2005">2005</date>
		</imprint>
		<respStmt>
			<orgName>CLEF ; Information Science, University of Hildesheim</orgName>
		</respStmt>
	</monogr>
	<note type="report_type">Master Thesis</note>
</biblStruct>

<biblStruct xml:id="b8">
	<monogr>
		<title level="m" type="main">Web Information Retrieval am Beispiel des WEB-GOV Korpus</title>
		<author>
			<persName><forename type="first">Niels</forename><surname>Jensen</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2005">2005a</date>
		</imprint>
		<respStmt>
			<orgName>Information Science, University of Hildesheim</orgName>
		</respStmt>
	</monogr>
	<note type="report_type">Master Thesis</note>
</biblStruct>

<biblStruct xml:id="b9">
	<analytic>
		<title level="a" type="main">Mehrsprachiges Information Retrieval mit einem WEB-Korpus</title>
		<author>
			<persName><forename type="first">Niels</forename><surname>Jensen</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings Vierter Hildesheimer Information Retrieval und Evaluierungsworkshop</title>
				<editor>
			<persName><forename type="first">Thomas</forename><forename type="middle">;</forename><surname>Mandl</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">Christa</forename><surname>Womser-Hacker</surname></persName>
		</editor>
		<meeting>Vierter Hildesheimer Information Retrieval und Evaluierungsworkshop<address><addrLine>HIER</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2005">2005b. 2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b10">
	<analytic>
		<title level="a" type="main">The quest for the best pages on the web</title>
		<author>
			<persName><forename type="first">Thomas</forename><surname>Mandl</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Information Service &amp; Use. To appear McNamee</title>
				<meeting><address><addrLine>Paul</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2005">2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b11">
	<analytic>
		<title level="a" type="main">Character N-Gram Tokenization for European Language Text Retrieval</title>
		<author>
			<persName><forename type="first">James</forename><surname>Mayfield</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Information Retrieval</title>
		<imprint>
			<biblScope unit="volume">7</biblScope>
			<biblScope unit="issue">1/2</biblScope>
			<biblScope unit="page" from="73" to="98" />
			<date type="published" when="2004">2004</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b12">
	<monogr>
		<title/>
		<author>
			<persName><forename type="first">Börkur</forename><surname>Sigurbjörnsson</surname></persName>
		</author>
		<imprint/>
	</monogr>
</biblStruct>

<biblStruct xml:id="b13">
	<analytic>
		<title level="a" type="main">Blueprint of a Cross-Lingual Web Retrieval Collection</title>
		<author>
			<persName><forename type="first">Jaap</forename><forename type="middle">;</forename><surname>Kamps</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Maarten</forename><surname>De Rijke</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Journal of Digital Information Management</title>
		<imprint>
			<biblScope unit="volume">3</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page" from="9" to="13" />
			<date type="published" when="2005">2005a</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b14">
	<monogr>
		<title/>
		<author>
			<persName><forename type="first">Börkur</forename><surname>Sigurbjörnsson</surname></persName>
		</author>
		<imprint/>
	</monogr>
</biblStruct>

<biblStruct xml:id="b15">
	<monogr>
		<title level="m" type="main">Overview of WebCLEF</title>
		<author>
			<persName><forename type="first">Jaap</forename><forename type="middle">;</forename><surname>Kamps</surname></persName>
		</author>
		<author>
			<persName><surname>De Rijke ; Maarten</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2005">2005b. 2005</date>
		</imprint>
	</monogr>
	<note>In this volume</note>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
