<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>University of Groningen at GeoCLEF 2007</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Geoffrey Andogah</string-name>
          <email>g.andogah@rug.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gosse Bouma</string-name>
          <email>g.bouma@rug.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computational Linguistics Group, Centre for Language and Cognition Groningen (CLCG), University of Groningen</institution>
          ,
          <addr-line>Groningen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the approach of the University of Groningen to GeoCLEF task for CLEF 2007. We used geographic scope based approach to rank documents. This paper describes non-geographic similarity, geographic similarity and combined similarity measures employed to approach GeoCLEF task for CLEF 2007. The motivation for our participation was to test geographic scope (geo-scope) based relevance ranking for geographic information retrieval (GIR). We participated in monolingual English task and our evaluation result shows no significant improvement for geo-scope based approach. Geographic Knowledge Base. We used the World Gazetteer1, GEOnet Names Server2 (GNS), Wikipedia3 and WordNet4 as the bases for our Geographic Knowledge Base (GKB) for several reasons: free availability, multilingual, coverage of most popular and major places, etc. Geographic Tagger. Alias-I LingPipe5 was used to detect named entities (location, person and organisation), geographic concepts (continent, region, country, city, town, village, etc.), spatial relations (near, in, south of, north west, etc.) and locative adjectives (e.g. Ugandan).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Approach</title>
      <sec id="sec-2-1">
        <title>Resources</title>
        <p>Lucene Search Engine. Apache Lucene6 is a high-performance, full-featured
text search engine library written entirely in Java. Lucene’s default similarity
measure is derived from the vector space model (VSM). Lucene was used to
index and search both indexes.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Document Indexing</title>
        <p>The reference document collection provided for experimentation was indexed
using Lucene. Document HEADLINE and TEXT contents were combined to
create document content for indexing (see Table 1 for details). Before indexing,
the documents were processed with the Porter stemmer and the default Lucene
English stopword list.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Non-Geographic Similarity Measure. We used the Apache Lucene IR li</title>
        <p>brary to perform non-geographic search. Lucene’s default similarity measure is
derived from the vector space model (VSM). The Lucene similarity score formula
combines several factors to determine the document score for a query [4]:
N onSim(q, d) =</p>
        <p>
          X tf (t in d) . idf (t) . bst . lN (t.f ield in d)
t in q
(
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
where, tf (t in d) is the term frequency factor for term t in document d, idf (t) is
the inverse document frequency of term t, bst is the field boost set during
indexing and lN (t.f ield in d) is the normalization value of a field given the number
of terms in the field.
        </p>
        <p>Geographic Similarity Measure. As in [2, 5], we chose to use geographic
scopes assigned to queries and documents to perform geographic relevance
ranking of documents. Our geographic scope resolver [1] assign multiple scopes to
documents ranking the scopes from the most relevant to the least relevant, thereby
associating a document with multiple scopes or associating a scope with
several documents. We limit geographic scopes (geo-scopes) of population centers
6 http://jakarta.apache.org/lucene
to continent, continent-directional (e.g. western Europe, eastern Africa, eastern
Europe, etc.), country, country-directional (e.g. north-of Italy), province7, and
province-directional (e.g. northern California) level.</p>
        <p>Equation 2 depicts our geographic similarity measure formula between query
q and document d:
where;</p>
        <p>GeoSim(q, d) =
(SF × W T S
0
if SF &gt; 0
otherwise
otherwise
W T S =</p>
        <p>X pwt(q,s) × log(1 + wt(d,s))
and, where; Nq is the number of scopes in the query scope set, Nd is the number
of scopes in the document scope set, N(d,q) is the number of document scopes
present in query scope set, wt(q,s) is the weight assigned to scope s in query q
by the scope resolver and wt(d,s) is the weight assigned to scope s in document
d by the scope resolver. For a given query, Nq is invariable whilst Nd and N(d,q)
vary per document retrieved. The motivation for designing Equation 3 is to
mitigate effects of arbitrarily large variations of Nd and Nq to a reasonable level.
The values of Equation 3 are within 0.0 &lt; SF ≤ 0.5. Equation 4 as arranged
provides the best overall performance for Equation 2 (that is, applying square
root weighting to wt(q,s) and logarithmic weighting to wt(d,s)).</p>
        <p>Geo-IR Similarity Measure. The final similarity score formula is directly
derived from Equation 1 and Equation 2 similarity score formulae.</p>
        <p>Sim(q, d) = λT N onSim(q, d) + λG GeoSim(q, d)
λT + λG = 1
where; λT is the non-geographic interpolation factor and λG is the geographic
interpolation factor. Before the ranked list for non-geographic and geographic
relevance ranking are linearly combined, their respective scores are normalized
to [0, 1].
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Query Formulation for Official GeoCLEF Runs</title>
        <p>
          This section describes the University of Groningen official runs for GeoCLEF
2007. In particular we describe which topic components are used for query
formulation and which similarity measures were used to perform relevance ranking.
7 Here a province represents first order administrative division of a country.
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          )
(
          <xref ref-type="bibr" rid="ref4">4</xref>
          )
(
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
(6)
Topic Categorization. In [3], geographic topics are categorized into eight
according to the way they depend on a place (e.g. UK, St. Andrew, etc.), geographic
subject (e.g. city, river, etc.) or geographic relation (e.g. north of, western, etc.).
GeoCLEF 2007 topics generation followed similar classification, and in our
experiment we grouped the topics into two: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) topics whose geographic scopes can
easily be resolved to a place (GROUP1), and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) topics whose geographic scopes
cannot be resolved to a place (GROUP2).
        </p>
        <p>We performed geographic expansion on the following members of GROUP1
– 51, 59, 60, 61, 63, 65, 66, 70. The motivation for geographic expansion on
these topics is that they lack sufficient geographic information or the geographic
information provided are too ambiguous. For example, topic 59 is expanded by
adding the names of major cities in Bolivia, Columbia, Ecuador and Peru. The
Lucene boost factor of 0.45F is assigned to placenames for geographic query
expansion while the boost factor for placenames in the original query is left at
the default value of 1.0F.</p>
        <p>Members of GROUP2 are topics – 56, 67, 68, 72. These topics fall under
geographic subject with non-geographic restriction with exception of topic 72
which is more complex. Resolving geographic scope of these topics to a specific
place is a non trivial undertaking. The most reasonable scope for these topics is
geographic subject scope such as lake, river, beach, city, etc. For example, topic
56 concern documents with scope lake.</p>
      </sec>
      <sec id="sec-2-5">
        <title>CLCGGeoEET00, CLCGGeoEETD00 and CLCGGeoEETDN00. Queries</title>
        <p>for these runs are formulated by the content of topic TITLE (T), TITLE-DESC
(TD) and TITLE-DESC-NARR (TDN) tags respectively. GROUP1 topics are
ranked according to Equation 5 with λT = 1.0 and λG = 0.0. GROUP2 topics
are ranked according to Equation 1. However, the query for CLCGGeoEETDN00
was mistakenly formulated by the content of topic TITLE instead of TDN.</p>
        <p>CLCGGeoEETDN00P is CLCGGeoEETDN00 with query formulated by topic
TDN tag content.</p>
        <p>CLCGGeoEETDN01. The query for this run should have been formulated by
the content of topic TDN tags, however, the official result submitted erroneously
used TITLE tag content. GROUP1 topics are ranked according to Equation 5
with λT = 0.85 and λG = 0.15. GROUP2 topics are ranked according to
Equation 1.</p>
        <p>CLCGGeoEETDN01P is CLCGGeoEETDN01 with query formulated by topic
TDN tag content.</p>
        <p>CLCGGeoEETDN01B. The query for this run should have been formulated
by the content of topic TDN tags, however, the official result submitted
erroneously used TITLE tag content. GROUP1 topics are ranked according to
Equation 5 with λT = 0.85 and λG = 0.15.</p>
        <p>For GROUP2 topics, we scan each document retrieved and ranked according
to Equation 1 for geographic types (geo-types) as well as determine the
geotypes of geographic names found in the documents. Each geo-type found in the
document is assigned a weight. For documents containing query geo-type, we
add the geo-type weight to Lucene score and then re-rank documents based on
the new score.</p>
        <p>CLCGGeoEETDN01BP is CLCGGeoEETDN01B with query formulated by
topic TDN tag content.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation and Future Work</title>
      <p>
        Our future work will focus on investigating: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) better geographic similarly
measure formulae, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) the use of geographic scopes selected by the searcher from
the returned documents for relevance feedback and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) term-relevance feedback
based on geographic terms (e.g.placenamess) extracted by the searcher
afterexaminingg retrieved documents. We are already testing some of these ideas and
results are promising.
We described non-geographic similarity, geographic similarity and combined
similarity measures employed to approach GeoCLEF task for CLEF 2007. We tested
geographic scope (geo-scope) based relevance ranking for geographic information
retrieval (GIR). Our evaluation result shows no significant improvement for
geoscope based approach in monolingual English task.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>This work is supported by NUFFIC within the framework of Netherlands
Programme for the Institutional Strengthening of Post-secondary Training
Education and Capacity (NPT) under project titled ”Building a sustainable ICT
training capacity in the public universities in Uganda”.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>G.</given-names>
            <surname>Andogah</surname>
          </string-name>
          , G. Bouma,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nerbonne</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Koster</surname>
          </string-name>
          .
          <article-title>Resolving geographical scope of documents with Lucene</article-title>
          . To appear in the near future,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>L.</given-names>
            <surname>Andrade</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Silva</surname>
          </string-name>
          .
          <article-title>Relevance ranking for geographic IR</article-title>
          . In Workshop on Geographical Information Retrieval, SIGIR'
          <volume>06</volume>
          ,
          <year>August 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>F.</given-names>
            <surname>Gey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bischoff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Womser-Hacker</surname>
          </string-name>
          .
          <article-title>GeoCLEF 2006: the CLEF 2006 Cross-Language Geographic Information Retrieval Track Overview</article-title>
          .
          <source>In Working Notes for CLEF 2006 Workshop (CLEF</source>
          <year>2006</year>
          ), Alcante, Spain,
          <year>September 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>O.</given-names>
            <surname>Gospodnetic</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Hatcher</surname>
          </string-name>
          . Lucene in Action. Manning Publications Co., 206 Bruce Park Avenue, Greenwich, CT
          <volume>06830</volume>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>B.</given-names>
            <surname>Martins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Silva</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Andrade</surname>
          </string-name>
          .
          <article-title>Indexing and ranking in Geo-IR systems</article-title>
          . In Workshop on Geographical Information Retrieval, SIGIR'05,
          <string-name>
            <surname>November</surname>
          </string-name>
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>