<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Geographic Information Retrieval Track Overview</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fredric Gey</string-name>
          <email>gey@berkeley.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ray Larson</string-name>
          <email>ray@sims.berkeley.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark Sanderson</string-name>
          <email>m.sanderson@sheffield.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kerstin Bischoff</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Mandl</string-name>
          <email>mandl@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christa Womser-Hacker</string-name>
          <email>womser@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Diana Santos</string-name>
          <email>Diana.Santos@sintef.no</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paulo Rocha</string-name>
          <email>Paulo.Rocha@di.uminho.pt</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Information Studies, University of Sheffield</institution>
          ,
          <addr-line>Sheffield</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Giorgio M. Di Nunzio</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Information Science, University of Hildesheim</institution>
          ,
          <country country="DE">GERMANY</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Linguateca</institution>
          ,
          <addr-line>SINTEF ICT</addr-line>
          ,
          <country country="NO">NORWAY</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of California</institution>
          ,
          <addr-line>Berkeley, CA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2006</year>
      </pub-date>
      <abstract>
        <p>After being a pilot track in 2005, GeoCLEF advanced to be a regular track within CLEF 2006. The purpose of GeoCLEF is to test and evaluate cross-language geographic information retrieval (GIR): retrieval for topics with a geographic specification. For GeoCLEF 2006, twenty-five search topics were defined by the organizing groups for searching English, German, Portuguese and Spanish document collections. Topics were translated into English, German, Portuguese, Spanish and Japanese. Several topics in 2006 were significantly more geographically challenging than in 2005. Seventeen groups submitted 149 runs (up from eleven groups and 117 runs in GeoCLEF 2005). The groups used a variety of approaches, including geographic bounding boxes, named entity extraction and external knowledge bases (geographic thesauri and ontologies and gazetteers).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Existing evaluation campaigns such as TREC and CLEF have not, prior to 2005, explicitly evaluated
geographical relevance. The aim of GeoCLEF is to provide the necessary framework in which to evaluate GIR
systems for search tasks involving both spatial and multilingual aspects. Participants are offered a TREC style ad
hoc retrieval task based on existing CLEF collections. GeoCLEF 2005 was run as a pilot track to evaluate
retrieval of multilingual documents with an emphasis on geographic search on English and German document
collections. Results were promising, but it was felt that more work needed to be done to identify the research and
evaluation issues surrounding geographic information retrieval from text. Thus 2006 was the second year in
which GeoCLEF was run as a track within CLEF. For 2006, two additional document languages were added to
GeoCLEF, Portuguese and Spanish. GeoCLEF was a collaborative effort by research groups at the University of
California, Berkeley (USA) , the University of Sheffield (UK), University of Hildesheim (Germany), Linguateca
(Norway and Portugal), and University of Alicante (Spain). Seventeen research groups (increased from eleven in
2005) from a variety of backgrounds and nationalities submitted 149 runs (up from 117 in 2005) to GeoCLEF.
Geographical Information Retrieval (GIR) concerns the retrieval of information involving some kind of spatial
awareness. Given that many documents contain some kind of spatial reference, there are examples where
geographical references (geo-references) may be important for IR. For example, to retrieve, re-rank and visualize
search results based on a spatial dimension (e.g. “find me news stories about riots near Dublin City”). In addition
to this, many documents contain geo-references expressed in multiple languages which may or may not be the
same as the query language. For example, the city of Cologne (English) is also Köln (German), Colónia in
Portuguese from Portugal, Colônia in Brazilian Portuguese, and Colonia (Spanish). Queries with names such as
this may require an additional translation step to enable successful retrieval.</p>
      <p>For 2006, Spanish and Portuguese, in addition to German and English, were added as document languages, while
topics were developed in all four languages with topic translations provided for the other languages. In addition
the National Institute of Informatics of Tokyo, Japan translated the English version of the topics to Japanese.
There were two Geographic Information Retrieval tasks: monolingual (English to English, German to German,
Portuguese to Portuguese and Spanish to Spanish) and bilingual (language X to language Y, where X or Y was
one of English, German, Portuguese or Spanish and additionally X could be Japanese).</p>
    </sec>
    <sec id="sec-2">
      <title>Document collections used in GeoCLEF</title>
      <p>The document collections for this year's GeoCLEF experiments are all newswire stories from the years 1994 and
1995 used in previous CLEF competitions. Both the English and German collections contain stories covering
international and national news events, therefore representing a wide variety of geographical regions and places.
The English document collection consists of 169,477 documents and was composed of stories from the British
newspaper The Glasgow Herald (1995) and the American newspaper The Los Angeles Times (1994). The
German document collection consists of 294,809 documents from the German news magazine Der Spiegel
(1994/95), the German newspaper Frankfurter Rundschau (1994) and the Swiss news agency SDA (1994/95).
Although there are more documents in the German collection, the average document length (in terms of words in
the actual text) is much larger for the English collection. In both collections, the documents have a common
structure: newspaper-specific information like date, page, issue, special filing numbers and usually one or more
titles, a byline and the actual text. The document collections were not geographically tagged or contained any
other location-specific information. For Portuguese, GeoCLEF 2006 utilized two newspaper collections,
spanning over 1994-1995, for respectively the Portuguese and Brazilian newspapers Público (106,821
documents) and Folha de São Paulo (103,913 documents). Both are major daily newspapers in their countries.
Not all material published by the two newspapers is included in the collections (mainly for copyright reasons),
but every day is represented. The collections are also distributed for IR and NLP research by Linguateca as the
CHAVE collection (www.linguateca.pt/CHAVE/, see URL for DTD and document examples).</p>
    </sec>
    <sec id="sec-3">
      <title>Generating Search Topics</title>
      <p>A total of 25 topics were generated for this year’s GeoCLEF. Topic creation was shared among the four
organizing groups, each group creating initial versions of their proposed topics in their language, with
subsequent translation into English. In order to support topic development, Ray Larson indexed all collections
with his Cheshire II document management system and this was made available to all organizing groups for
interactive exploration of potential topics. While the aim had been to prepare an equal number of topics in each
language, ultimately only two topics (GC026 and GC027) were developed in English. Other original language
numbers were German, 8 topics (GC028 to GC035), Spanish, 5 topics (GC036 to GC040) and Portuguese, 10
topics (GC041 to GC050). This section will discuss the processes taken to create the spatially-aware topics for
the track.</p>
      <sec id="sec-3-1">
        <title>Topic generation</title>
        <p>In GeoCLEF 2005 some criticism arose about the lack of geographical challenges of the topics (favouring
keyword-based approaches) and the German task was inherently more difficult because several topics had no
relevant documents in the German collections. Therefore geographical and cross-lingual challenge and equal
distribution across language collections was considered central during topic generation. Topics should vary
according to the granularity and kind of geographic entity and should require adequate handling of named
entities within the process of translation (e.g. regarding decompounding, transliteration or translation).
For English topic generation, Fred Gey simply took two topics he had considered in the past (Wine regions
around rivers in Europe and Cities within 100 kilometers of Frankfurt, Germany) and developed them. The
latter topic (GC027) evolved into an exact specification of the latitude and longitude of Frankfurt am Main (to
distinguish it from Frankfurt an der Oder) in the narrative section. Interactive exploration verified that
documents could be found which satisfied these criteria on the basis of geographic knowledge by the proposer
(i.e. the Rhine and Moselle valleys of Germany and cities Heidelberg, Koblenz, Mainz, and Mannheim near
Frankfurt).</p>
        <p>
          The German group at Hildesheim started with brain storming on interesting geographical notions and looking for
potential events via the Cheshire II Interface, we unfortunately had to abandon all smaller geographic regions
soon. Even if a suitable number of relevant documents could be found in one collection, most times there were
few or no respective documents in the other language collections. This may not be surprising, because within the
domain of news criteria like (inter)national relevance, prominence, elite nation or elite person besides proximity,
conflict/negativism and continuity etc. (for an overview see Eidlers[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]) are assumed to affect what will become a
news article. Thus, the snow conditions or danger of avalanches in Grisons (canton in Switzerland) may be
reported frequently by the Swiss news agency SDA or even German newspapers, whereas the British or
American newspapers may not see the relevance for their audience. In addition, the geographically interesting
issue of tourism in general is not well represented in the German collection. As a result well known places and
larger regions as well as international relevant or dramatic concepts had to be focused on, although this may not
reflect all user needs for GIR systems (see also Kluck &amp; Womser-Hacker[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]). In order not to favor systems
relying purely on keywords we concentrated on more difficult geographic entities like historical or political
names used to refer geographically to a certain region and imprecise regions like the Ruhr or the Middle East.
Moreover some topics should require the use of external geographic knowledge e.g. to identify cities onshore of
the Sea of Japan or the Tropics. The former examples introduce ambiguity or translation challenges as well.
Ruhr could be the river or the area in Germany and the Middle East may be translated to German Mittlerer
Osten, which is nowadays often used, but would denote a slightly different region. The naming of the Sea of
Japan is difficult as it depends on the Japanese and Western perspective, whereas in Korea it would be named
East Sea (of Korea). After checking such topic candidates for relevant documents in other collections we
proposed eight topics, which we thought would contribute to a topic set varying in thematic content and system
requirements.
        </p>
        <p>
          The GeoCLEF topics proposed by the Portuguese group (a total of 10) were discussed between Paulo Rocha and
Diana Santos, according to an initial typology of possible geographical topics (see below for a refined one) and
after having scrutinized the frequency list of proper names in both collections, manually identifying possible
places of interest. Candidate topics were then checked in the collections, using the Web interface to the AC/DC
project[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], to investigate whether they were well represented. We included some interesting topics from a
Portuguese (language) standpoint, including "ill-defined" or at least little known regions in an international
context, such as norte de Portugal (North of Portugal) or Nordeste brasileiro (Brazilian Northeast). Basically,
they are very familiar and frequently used concepts in Portuguese, but have not a purely geographical
explanation. Rather, they have a strongly cultural and historical motivation. We also inserted a temporally
dependent topic (outdated Champion's Cup, now Champion's League – and already in 1994-1995 as well, but
names continue their independent life in newspapers and in folk's stock of words). This topic is particularly
interesting, since it in addition concerns "European" football, where one of the partners is
(non-geographicallyEuropean) Israel.
        </p>
        <p>We also strove to find topics which included more geographical relations than mere "in" (homogeneous region),
as well as different location types as far as grain and topology are concerned. As to the first concern, note that
although "shipwrecks in the Atlantic Ocean" seem to display an ordinary "in"-relation, shipwrecks are often near
the coasts, and the same is still more applicable about topics such as "fishing in Newfoundland", where it is
presupposed that you fish on the sea near the place (or that you are concerned with the impact of fishing to
Newfoundland). Likewise, anyone who knows what ETA stands for would at once expect that "ETA's activities
in France" would be mostly located in the French Basque country (and not anywhere in France).
For the second concern, that of providing different granularity and/or topology, note that the geographical span
of forest fires is clearly different from that of lunar of solar eclipses (a topic suggested by the German team). As
to form of the region, the "New England universities" topic circumscribes the "geographical region" to a set of
smaller conceptual "regions", each represented by a university. Incidentally, this topic displays another
complication, because it involves a multiword named entity: not only "New England" is made up of two different
words but both are very common and have a specific meaning on its own (in English and Portuguese alike). This
case is further interesting because it would be as natural to say New England in Portuguese as Nova Inglaterra,
given that the name is not originally Portuguese.</p>
        <p>We should also report interesting problems caused by translation into Portuguese from topics originally stated in
other languages (they are not necessarily translation problems, but were spotted because we had to look into the
particular cases of those places or expressions). For example, Middle East can be equally translated by Próximo
Oriente and Médio Oriente, and it is not politically neutral how precisely in that area some places are described.
For example, we chose to use the word Palestina (together with Israel) and leave out Gaza Strip (which,
depending on political views, might be considered a part of both). What is interesting here is that the political
details are absolutely irrelevant for the topic in question (which deals with archaeological findings), but the
GeoCLEF organizers (in common) decided to specify a lower level, or a higher precision description, of every
location/area mentioned, in the narrative, so that a list of Middle East countries and regions had to be supplied,
and agreed upon.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Format of topic description</title>
        <p>The format of GeoCLEF 2006 differed from that of 2005. No explicit geographic structure was used this time,
although such a structure was discussed by the organizing groups. Two example topics are shown in Figure 1.
&lt;top&gt;
&lt;num&gt;GC027&lt;/num&gt;
&lt;EN-title&gt;Cities within 100km of Frankfurt&lt;/EN-title&gt;
&lt;EN-desc&gt;Documents about cities within 100 kilometers of the city of Frankfurt in
Western Germany&lt;/EN-desc&gt;</p>
        <p>&lt;EN-narr&gt;Relevant documents discuss cities within 100 kilometers of Frankfurt am Main
Germany, latitude 50.11222, longitude 8.68194. To be relevant the document must describe
the city or an event in that city. Stories about Frankfurt itself are not
relevant&lt;/ENnarr&gt;
&lt;/top&gt;
&lt;top&gt;
&lt;num&gt; GC034 &lt;/num&gt;
&lt;EN-title&gt; Malaria in the tropics &lt;/EN-title&gt;
&lt;EN-desc&gt; Malaria outbreaks in tropical regions and preventive vaccination &lt;/EN-desc&gt;
&lt;EN-narr&gt; Relevant documents state cases of malaria in tropical regions and possible
preventive measures like chances to vaccinate against the disease. Outbreaks must be of
epidemic scope. Tropics are defined as the region between the Tropic of Capricorn,
latitude 23.5 degrees South and the Tropic of Cancer, latitude 23.5 degrees North. Not
relevant are documents about a single person's infection.&lt;/EN-narr&gt;
&lt;/top&gt;
As can be seen, after the brief descriptions within the title and description tags, the narrative tag contains detailed
description of the geographic detail sought and the relevance criteria.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Several kinds of geographical topics</title>
        <p>
          We came up with a tentative classification of topics according to the way they depend on place (in other words,
according to the way they can be considered "geographic"), which we believe to be one of the most interesting
results of our participation in the choice and topic formulation for GeoCLEF. Basically, this classification was
done as an answer to the overall too simplistic assumption of first GeoCLEF[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], namely the separation between
subject and location as if the two were independent and therefore separable pieces of information. (Other
comments to the unsuitability of the format used in GeoCLEF can be found in Santos and Cardoso[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], and will
not be repeated here.)
While it is obvious that in some (simple) cases geographical topics can be modeled that way, there's much more
to place and to the place of place in the meaning of a topic than just that, as we hope this categorization can help
making clear:
1 non-geographic subject restricted to a place (music festivals in Germany) [only kind of topic in
        </p>
        <p>GeoCLEF 2005]
2 geographic subject with non-geographic restriction (rivers with vineyards) [new kind of topic added in</p>
        <p>
          GeoCLEF 2006]
3 geographic subject restricted to a place (cities in Germany)
4 non-geographic subject associated to a place (independence, concern, economic handlings to
favour/harm that region, etc.) Examples: independence of Quebec, love for Peru (as often remarked, this
is frequently, but not necessarily, associated to the metonymical use of place names)
5 non-geographic subject that is a complex function of place (for example, place is a function of topic)
(European football cup matches, winners of Eurovision Song Contest)
geographical relations among places (how are the Himalayas related to Nepal? Are they inside? Do the
Himalaya mountains cross Nepal's borders? etc.)
geographical relations among (places associated to) events (Did Waterloo occur more north than the
battle of X? Were the findings of Lucy more to the south than those of the Cromagnon in Spain?)
relations between events which require their precise localization (was it the same river that flooded last
year and in which killings occurred in the XVth century?)
Note that we here are not even dealing with the obviously equally relevant interdependence of the temporal
dimension, already mentioned above, and which was actually extremely conspicuous in the preliminary
discussions among this year's organizing teams, concerning the denotation of "former Eastern bloc countries"
and "former Yugoslavia" now (that is, in 1994-1995). In a way, as argued in Santos and Chaves[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], which
countries or regions to accept as relevant depends ultimately on the user intention (and need). Therefore, pinning
down the meaning of a topic depends on geographical, temporal, cultural, and even personal constraints, that are
intertwined in a complex way, and more often than not do not allow a clear separation. To be able to make sense
of these complicated interactions and arrive at something relevant for a user by employing geographical
reasoning seems one of the challenges that lies ahead in future GeoCLEF tracks.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Approaches to Geographic Information Retrieval</title>
      <p>The participants used a wide variety of approaches to the GeoCLEF tasks, ranging from basic IR approaches
(with no attempts at spatial or geographic reasoning or indexing) to deep NLP processing to extract place and
topological clues from the texts and queries. Specific techniques used included:
•
•
•
•
•
•
•
•
•</p>
      <p>Ad-hoc techniques (blind feedback, German word decompounding, manual query expansion)
Gazetteer construction (GNIS, World Gazetteer)
Gazetteer-based query expansion
Question-answering modules utilizing passage retrieval
Geographic Named Entity Extraction
Term expansion using Wordnet
Use of geographic thesauri (both manually and automatically constructed)
Resolution of geographic ambiguity</p>
      <p>NLP – part-of-speech tagging</p>
    </sec>
    <sec id="sec-5">
      <title>Relevance assessment</title>
      <p>English assessment was shared by Berkeley and Sheffield Universities. German assessment was done by the
University of Hildesheim, Portuguese assessment by Linguateca, and Spanish assessment by University of
Alicante. All organizing groups utilized the DIRECT System provided by the University of Padua. The Padua
system allowed for automatic submission of runs by participating groups and for automatic assembling of the
GeoCLEF assessment pools by language.</p>
      <sec id="sec-5-1">
        <title>English relevance assessment</title>
        <p>The English document pool extracted from 73 monolingual and 12 bilingual (language X to) English runs
consisted of 17,9xx documents to be reviewed and judged by our 5 assessors or about 3,600 documents per
assessor. In order to judge topic GC027 (Cities within 100km of Frankfurt), Ray Larson used data from the
GeoNames Information System along with the Cheshire II geographic distance calculation function, to extract
and prepare a spreadsheet of populated places whose latitude and longitude was within a distance of 100 km of
the latitude and longitude of Frankfurt. This spreadsheet contained 5342 names and was made available to all
groups doing assessment. If a document in the pool contained the name of a German city or town, it was checked
against the spreadsheet to see if it was within 100km of Frankfurt. Thus documents with well-known names
(Mannheim, Heidelberg) were easily recognized, but Mecklenberg (where the German Grand Prix auto race is
held) was not so easily recognized. In reading the documents in the pool, we were surprised to find many Los
Angeles Times documents about secondary school sports events and scores in the pool. A closer examination
revealed that these documents contained the references to American students who had the same family name as
German cities and towns. It is clear that geographic named entity disambiguation from text still needs some
improvement.</p>
      </sec>
      <sec id="sec-5-2">
        <title>German relevance assessment</title>
        <p>For the pool of German monolingual and bilingual runs X2German 14.094 documents from the newspaper
Frankfurter Rundschau, the Swiss news agency SDA and the news magazine Spiegel had to be assessed. Every
assessor had to judge a number of assigned topics. Decisions on dubious cases were left open and then discussed
within the group and/or the other language co-ordinators. Since many topics had clear, predefined criteria as
specified in title, description and narrative, searching first the key concepts and their synonyms within the
documents and then identifying their geographical reference led to rejecting the bulk of documents as irrelevant.
Depending on the geographic entity asked for, manual expansion, e.g., the country names of the Middle East and
their capitals, was done to query the DIRECT System provided by the University of Padua. Of course, such a list
could never be complete and available resources would not be comprehensive enough to capture all possible
expansions (e.g. we could not verify the river Code on the island of Java). Thus skimming over the text was often
necessary to capture the documents main topic and geographical scope.</p>
        <p>While judging relevance was generally easier for the short news agency articles of SDA with their headlines,
keywords and restriction to one issue, Spiegel articles took rather long to judge, because of their length and
essay-like stories often covering multiple events etc. without a specific narrow focus. Many borderline cases for
relevance resulted from uncertainties about how broad/narrow a concept term should be interpreted and how
explicit the concept must be stated in the document (e.g. do parked cars destroyed by a bomb correspond to a car
bombing? Are attacks on foreign journalists and the Turkish invasion air attacks to be considered relevant as
fulfilling the concept of combat?). Often it seems that for a recurring news issue it is assumed that facts are
already known, so they are not explicitly cited. To keep the influence of order effects minimal is critical here.
Similarly, assessing relevance regarding the geographical criterion brought up a discussion on specificity wrt
implicit inclusion. In all cases, reference to the required geographic entity had to be explicitly made, i.e., a
document reporting about Fishing in the Norwegian Sea or the Greenland Sea without mentioning e.g. a certain
coastal city in Greenland or Newfoundland was not considered relevant. Moreover, the borders of oceans and its
minor seas are often hard to define (e.g. does Havana, Cuba border the Atlantic Ocean?). Figuring out the
location referred to was frequently difficult, when the city mentioned first in an article could have been the
domicile of the news agency or/and the city some event occurred in. This was especially true for GC040 active
volcanoes and for GC027 cities within 100km from Frankfurt, with Frankfurt being the domicile of the
Frankfurter Rundschau, which formed part of the collection. Problems with fuzzy spatial relations or imprecise
regions on the other hand did not figure very prominently as they were defined in the extended narratives (e.g.
“near” Madrid includes only Madrid and its outskirts) and the documents to be judged did not contain critical
cases. However, one my have argued against the decision to exclude all districts of Frankfurt as they do not
form own cities, but have a common administration.</p>
        <p>The topic on cities around Frankfurt (GC027) was together with GC050 about cities along the Danube and the
Rhine the most difficult one to judge. Although a list of relevant cities containing more than 4000 names was
provided by Ray Larson, this could not be used efficiently for relevance assessment to query the DIRECT
system. Moreover, the notion of an event or a description made assessment even more time-consuming. We
queried about 40 or 50 prominent relevant cities and actually read every document except tabular listings of
sports results or public announcements in tabular form. Since the Frankfurter Rundschau is also a regional
newspaper, articles on nearby cities, towns and villages are frequent. Would one consider the selling of parking
meters to another town an event? Or a public invitation to fruit picking or the announcement of a new vocational
training as nurse? As the other assessors did not face such a problem, we decided to be rather strict, i.e. an event
must be something popular, public and have a certain scope or importance (not only for a single person or a
certain group) like concerts, strikes, flooding or sports. In a similar manner, we agreed on a narrower
interpretation of the concept of description for GC026 and for GC050 as something unique or characteristic to a
city like statistical figures, historical reviews or landmarks. What would be usually considered a description was
not often found due to the kind of collection, likewise relevant documents for GC045 tourism in Northeast Brazil
were also few. While the SDA news agency articles will not treat travelling or tourism, such articles may
sometimes be found in Frankfurter Rundschau or Spiegel, but there is no special section on that issue.
Finally, for topic GC027 errors within the documents from the Frankfurter Rundschau will have influenced
retrieval results: some articles have duplicates (sometimes even up to four versions), different articles thrown
together in one document (e.g. one about Frankfurt and one about Wiesbaden), sentences or passages of articles
are missing. Thus a keyword approach may have found many relevant documents, because Frankfurt was
mentioned somewhere in the document.</p>
      </sec>
      <sec id="sec-5-3">
        <title>Portuguese Relevance Assessment</title>
        <p>Details of Portuguese group’s assessment are as follows: The assessor tried to find the best collection of
keywords – based on the detailed information in the narrative and his/her knowledge of the geographical
concepts and subjects involved – and queried the DIRECT system. Often there was manual refinement of the
query after finding new spellings in previous hits (note that our collections are written in two different varieties
of Portuguese). For example, for topic GC050, "cities along the Danube and the Rhine", the following (final)
query was used: Danúbio Reno Ulm Ingolstadt Regensburg Passau Linz Krems Viena Bratislava Budapeste
Vukovar Novi Sad Belgrado Drobeta-Turnu Severin Vidin Ruse Brăila Galaţi Tulcea Braila Galati Basel
Basiléia Basileia Estrasburgo Strasbourg Karlsruhe Carlsruhe Mannheim Ludwigshafen Wiesbaden Mainz
Koblenz Coblença Bona Bonn Colónia Colônia Cologne Düsseldorf Dusseldorf Dusseldórfia Neuss Krefeld
Duisburg Duisburgo Arnhem Nederrijn Arnhemia Nijmegen Waal Noviomago Utrecht Kromme Rijn Utreque
Rotterdam Roterdão. A similar strategy was used for cities within 100 km from Frankfurt am Main (GC027),
where both particular cities were mentioned, as well as words like cidade (city), Frankfurt, distância (distance),
and so on. Obviously, the significant passages for all hits were read, to assess whether the document actually
mentioned cities near Frankfurt.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>GeoCLEF Performance</title>
      <p>1</p>
      <sec id="sec-6-1">
        <title>Participants and Experiments</title>
        <p>As shown in Table 1, a total of 17 groups from 8 different countries submitted results for one or more of the
GeoCLEF tasks - an increase on the 13 participants of last year. A total of 149 experiments were submitted,
which is an increase on the 117 experiments of 2005. There is almost no variation in the average number of
submitted runs per participant: from 9 runs/participant of 2005 to 8.7 runs/participant of this year.</p>
        <p>Four different topic languages were used for GeoCLEF bilingual experiments. As always, the most popular
language for queries was English; German and Spanish tied for the second place. Note that Spanish is a new
collection added this year. The number of bilingual runs by topic language is shown in Table 4.
Monolingual retrieval was offered for the following target collections: English, German, Portuguese, and
Spanish. As can be seen from Table 3, the number of participants and runs for each language was quite similar,
with the exception of English, which has the greatest participation. Table 5 shows the top five groups for each
target collection, ordered by mean average precision. Note that only the best run is selected for each group, even
if the group may have more than one top run. The table reports: the short name of the participating group; the run
identifier, specifying whether the run has participated in the pool or not; the mean average precision achieved by
the run; and the performance difference between the first and the last participant. Table 5 regards runs using title
+ description fields only.</p>
        <p>Note that the top five participants contain both “newcomer” groups (i.e. groups that had not previously
participated in GeoCLEF) and “veteran” groups (i.e. groups that had participated in previous editions of
GeoCLEF), with the exception of monolingual Portuguese where only “veteran” groups were subscribed. Both
pooled and not pooled runs are in the best entries for each track.
Figures 2 to 5 show the interpolated recall vs. average precision for top participants of the monolingual tasks.
The bilingual task was structured in four subtasks (X → DE, EN, ES or PT target collection). Table 6 shows the
best results for this task with the same logic of Table 5. Note that the top five participants contain both
“newcomer” groups and “veteran” groups, with the exception of monolingual Portuguese and Spanish where
only “veteran” groups were subscribed.</p>
        <p>For bilingual retrieval evaluation, a common method is to compare results against monolingual baselines:
• X Æ DE: 70% of best monolingual German IR system</p>
      </sec>
      <sec id="sec-6-2">
        <title>Statistical Testing</title>
        <p>
          We used the MATLAB Statistics Toolbox, which provides the necessary functionality plus some additional
functions and utilities. We use the ANalysis Of VAriance (ANOVA) test. ANOVA makes some assumptions
concerning the data be checked. Hull [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] provides details of these; in particular, the scores in question should be
approximately normally distributed and their variance has to be approximately the same for all runs. Two tests
for goodness of fit to a normal distribution were chosen using the MATLAB statistical toolbox: the Lilliefors test
[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] and the Jarque-Bera test [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. In the case of the GeoCLEF tasks under analysis, both tests indicate that the
assumption of normality is violated for most of the data samples (in this case the runs for each participant).
In such cases, a transformation of data should be performed. The transformation for measures that range from 0
to 1 is the arcsin-root transformation:
        </p>
        <p>
          arcsin( x )
which Tague-Sutcliffe [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] recommends for use with precision/recall measures.
        </p>
        <p>The following tables, from Table 8 to Table 13, summarize the results of this test. All experiments,
regardless the topic language or topic fields, are included. Results are therefore only valid for comparison of
individual pairs of runs, and not in terms of absolute performance. Each table shows the overall results where all
the runs that are included in the same group do not have a significantly different performance. All runs scoring
below a certain group performs significantly worse than at least the top entry of the group. Likewise all the runs
scoring above a certain group perform significantly better than at least the bottom entry in that group. Each table
contains also a graph which shows participants' runs (y axis) and performance obtained (x axis). The circle
indicates the average performance while the segment shows the interval in which the difference in performance
is not statistically significant; for each graph the best group is highlighted.</p>
        <p>Note that there are no tables for Bilingual German and Bilingual Portuguese since, according to the Tukey
T, all the experiments of these tasks belong to the same group.</p>
      </sec>
      <sec id="sec-6-3">
        <title>Conclusions and Future Work</title>
        <p>The test collection developed for GeoCLEF is the first GIR test collection available to the GIR research
community. GIR is receiving increased notice both through the GeoCLEF effort as well as due to the GIR
workshops held annually since 2004 in conjunction with SIGIR or CIKM. At the GIR06 workshop recently held
in conjunction with SIGIR 2006 in Seattle, 14 groups participated and 16 full papers were presented, as well as a
keynote address by John Frank of MetaCarta, and a summary of GeoCLEF 2005 presented by Ray Larson. Six of
the groups also were participants in GeoCLEF 2005 or 2006. Six of the full papers presented at the GIR
workshop used GeoCLEF collections. Most of these papers used the 2005 collection and queries, although one
group used queries from this year’s collection with their own take on relevance judgments. Of particular interest
to the organizers of GeoCLEF are the 8 groups working in the area of GIR who are not yet participants in
GeoCLEF. All attendees at the GIR06 workshop were invited to participate in GeoCLEF for 2007.</p>
      </sec>
      <sec id="sec-6-4">
        <title>Acknowledgments:</title>
        <p>The English assessment was by the GeoCLEF organizers was volunteer labor – none of us has funding for
GeoCLEF work. Assessment was performed by Hans Barnum, Nils Bailey, Fredric Gey, Ray Larson, and Mark
Sanderson. German assessment was done by Claudia Bachmann, Kerstin Bischoff, Thomas Mandl, Jens
Plattfaut, Inga Rill and Christa Womser-Hacker of University of Hildesheim. Portuguese assessment was done
by Paulo Rocha, Luís Costa, Luís Cabral, Susana Inácio, Ana Sofia Pinto, António Silva and Rui Vilela, all of
Linguateca (thanks to grant POSI/PLP/43931/2001 from Portuguese FCT, co-financed by POSI), and Spanish by
Andrés Montoyo Guijarro, Oscar Fernandez, Zornitsa Kozareva, Antonio Toral of University of Alicante.
Japanese translation of the English topics was provided by Noriko Kando. The future direction and scope of
GeoCLEF will be heavily influenced by funding and the amount of volunteer effort available.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>W. J.</given-names>
            <surname>Conover</surname>
          </string-name>
          . Practical Nonparametric Statistics. John Wiley and Sons, New York, USA,
          <year>1971</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Eilders</surname>
          </string-name>
          .
          <article-title>The role of news factors in media use</article-title>
          .
          <source>Technical report, Wissenschaftszentrum Berlin für Sozialforschung gGmbH (WZB)</source>
          .
          <source>Forschungsschwerpunkt Sozialer Wandel, Institutionen und Vermittlungsprozesse des Wissenschaftszentrums Berlin für Sozialforschung</source>
          . Berlin,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Gey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Joho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Clough</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Petras</surname>
          </string-name>
          .
          <article-title>GeoCLEF: the CLEF 2005 crosslanguage geographic information retrieval track overview</article-title>
          .
          <source>In Cross-Language Evaluation Forum: CLEF 2005. Springer (Lecture Notes in Computer Science LNCS 4022)</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>David</given-names>
            <surname>Hull</surname>
          </string-name>
          .
          <article-title>Using statistical testing in the evaluation of retrieval experiments</article-title>
          . In Robert Korfhage, Edie Rasmussen, and Peter Willett, editors,
          <source>Proceedings of the 16th Annual International ACM SIGIR Conference on Research and Development on Information Retrieval (SIGIR</source>
          <year>1993</year>
          ), pages
          <fpage>329</fpage>
          -
          <lpage>338</lpage>
          , New York, USA,
          <year>1993</year>
          . ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Judge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. E.</given-names>
            <surname>Griffiths</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lutkepohl</surname>
          </string-name>
          and
          <string-name>
            <given-names>T. C.</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>Introduction to the Theory and Practice of Econometrics</article-title>
          . John Wiley and Sons, New York, USA, 2nd edition,
          <year>1988</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kluck</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Womser-Hacker</surname>
          </string-name>
          .
          <article-title>Inside the evaluation process of the cross-language evaluation forum (CLEF): Issues of multilingual topic creation and multilingual relevance assessment</article-title>
          .
          <source>In Proceedings of the third International Conference on Language Resources and Evaluation</source>
          ,
          <string-name>
            <surname>LREC</surname>
          </string-name>
          ,
          <year>2002</year>
          ,
          <string-name>
            <given-names>Las</given-names>
            <surname>Palmas</surname>
          </string-name>
          , Spain, pages
          <fpage>573</fpage>
          -
          <lpage>576</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Diana</given-names>
            <surname>Santos</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Bick</surname>
          </string-name>
          .
          <article-title>Providing internet access to portuguese corpora: the ac/dc project</article-title>
          .
          <source>In Maria Gavrilidou</source>
          , George Carayannis, Stella Markantonatou, Stelios Piperidis, and Gregory Stainhauer, editors,
          <source>Proceedings of the Second International Conference on Language Resources and Evaluation</source>
          ,
          <string-name>
            <surname>LREC</surname>
          </string-name>
          <year>2000</year>
          , (Athens, 31 May-2
          <source>June</source>
          <year>2000</year>
          ), pages
          <fpage>205</fpage>
          -
          <lpage>210</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Diana</given-names>
            <surname>Santos</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Cardoso</surname>
          </string-name>
          . Portuguese at CLEF 2005:
          <article-title>Reflections and challenges</article-title>
          .
          <source>In Cross Language Evaluation Forum: Working Notes for the CLEF 2005 Workshop (CLEF</source>
          <year>2005</year>
          )
          <article-title>(Vienna</article-title>
          , Austria,
          <fpage>21</fpage>
          -
          <issue>23</issue>
          <year>September 2005</year>
          ),
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Diana</given-names>
            <surname>Santos</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Chaves</surname>
          </string-name>
          .
          <article-title>The place of place in geographical information retrieval</article-title>
          . In Chris Jones and Ross Purves, editors,
          <source>Workshop on Geographic Information Retrieval (GIR06)</source>
          ,
          <fpage>SIGIR06</fpage>
          , Seattle, 10
          <source>August</source>
          <year>2006</year>
          , pages
          <fpage>5</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Tague-Sutcliffe</surname>
          </string-name>
          .
          <article-title>The Pragmatics of Information Retrieval Experimentation, Revisited</article-title>
          . In K. Sparck Jones and P. Willett, editors,
          <source>Readings in Information Retrieval</source>
          , pages
          <fpage>205</fpage>
          -
          <lpage>216</lpage>
          . Morgan Kaufmann Publishers, Inc., San Francisco, California, USA,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>