<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>TALP-UPC at MediaEval 2014 Placing Task: Combining Geographical Knowledge Bases and Language Models for Large-Scale Textual Georeferencing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniel Ferrés</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Horacio Rodríguez</string-name>
          <email>horaciog@cs.upc.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>TALP Research Center Computer Science Department Universitat Politècnica de Catalunya</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>16</fpage>
      <lpage>17</lpage>
      <abstract>
        <p>This paper describes our Georeferencing approaches, experiments, and results at the MediaEval 2014 Placing Task evaluation. The task consists of predicting the most probable geographical coordinates of Flickr images and videos using its visual, audio and metadata associated features. Our approaches used only Flickr users textual metadata annotations and tagsets. We used four approaches for this task: 1) an approach based on Geographical Knowledge Bases (GeoKB), 2) the Hiemstra Language Model (HLM) approach with Re-Ranking, 3) a combination of the GeoKB and the HLM (GeoFusion). 4) a combination of the GeoFusion with a HLM model derived from the English Wikipedia georeferenced pages. The HLM approach with Re-Ranking showed the best performance within 10m to 1km distances. The GeoFusion approaches achieved the best results within the margin of errors from 10km to 5000km.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The MediaEval 2014 Placing task requires that
participants use systems that automatically assign geographical
coordinates (latitude and longitude) to Flickr photos and
videos using one or more of the following data: Flickr
metadata, visual content, audio content, and social information
(see [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for more details about this evaluation). The Placing
Task training data consists of 5,000,000 geotagged photos
and 25,000 geotagged videos, and the test data consists of
500,000 photos and 10,000 videos. Evaluation of results is
done by calculating the distance from the actual point
(assigned by a Flickr user) to the predicted point (assigned by
a participant). Runs are evaluated nding how many videos
were placed at least within some threshold distances.
      </p>
    </sec>
    <sec id="sec-2">
      <title>SYSTEM DESCRIPTION</title>
      <p>
        place names, stopwords lists, and an English Dictionary.
The system uses the following rules from Toponym
Disambiguation techniques [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to get the geographical focus of the
photo/video: 1) select the most populated place that is not
a state, country or continent and has its state apearing in
the text, 2) otherwise select the most populated place that is
not a state, country or continent and has its country
apearing in the text, 3) otherwise select the most populated state
that has its country apearing in the text 4) otherwise apply
population heuristics.
      </p>
      <p>
        2) Hiemstra Language Model (HLM) with Re-Ranking.
This approach uses the Terrier2 Information Retrieval (IR)
software (version 3.0) with the HLM weighting model [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
The HLM default lambda ( ) parameter value in Terrier
(0.15) was used. See in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] more details about the Terrier
implementation of the HLM weighting model (version 1 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]).
The indexing of the metadata subsets were done with the
coordinates as a document number and some metadata elds
(Title, Description and User Tags) as the document text.
For each unique coordinate a document was created with all
the textual metadata elds content of all the photos/videos
that pertain to this coordinate.
      </p>
      <p>The indexing process uses a multilingual stopwords list to
lter the tokens that are indexed. The following metadata
elds (lowercased) from the photos/videos were used for the
query: User tags, Title and Description. A Re-Ranking
process is applied after the IR process. For each topic their rst
1000 retrieved coordinates pairs from the IR software are
used. From them we selected the subset of coordinates pairs
with a score equal or greater than the two-thirds (66.66%)
of the score of the coordinates pair ranked in rst position.
Then for each geographical coordinates pair of the subset
we sum its associated score (provided by the IR software)
and the score of their neighbours at a threshold distance
(e.g. 100km). Then we select the one with the maximum
weighted sum.</p>
      <p>3) GeoFusion: Hiemstra Language Model with
ReRanking and GeoKB. This approach is applied by
combining the results of the GeoKB approach and the IR
approach with Re-Ranking. From the GeoKB system are
selected the predicted coordinates that come only from the
heuristics 1, 2 and 3 (avoiding predictions from the
population heuristics rules). When the GeoKB rules (applied in
priority order: 1, 2, and 3) do not match then the predictions
are selected from the HLM approach with Re-Ranking
2Terrier. http://terrier.org
4) GeoFusion+GeoWiki: GeoFusion combined with
a HLM model of Georeferenced Wikipedia pages.
This is the only improvement with respect to the system
used at MediaEval 2011. This approach uses a set of 857,574
Wikipedia georeferenced pages3 that were indexed with
Terrier. The coordinates of the top ranked georeferenced Wikipedia
page are used as a prediction. The predictions from the
georeferenced Wikipedia based HLM model are used only in
case that the HLM model with Re-Ranking based on the
Training data gives an score lower than 7.0. This
threshold was found empirically training with the MediaEval 2011
test set. The system uses the coordinates of one of the most
photographied places in the world as a prediction when the
approaches cannot give a prediction.</p>
    </sec>
    <sec id="sec-3">
      <title>EXPERIMENTS AND RESULTS</title>
      <p>We designed a set of four experiments for the MediaEval
2014 Placing Task (Main Task) test set of 510,000 Flickr
photos and videos (see results in Figure 1 and Table 1):
1. The experiment run1 used the HLM approach with
Re-Ranking up to 100 km and the MediaEval 2014
training set metadata as a training data. From a set
of 5,050,000 photos and videos of the MediaEval 2014
training set, a set of 3,057,718 coordinates pairs with
related metadata info were created as textual
documents and then indexed with Terrier.
2. The experiment run3 used the GeoKB approach.
3. The experiment run4 used the GeoFusion approach
with the MediaEval training corpora.
4. The experiment run5 used the GeoFusion approach
with the MediaEval training corpora in combination
with the English Wikipedia georeferenced pages HLM
model.</p>
    </sec>
    <sec id="sec-4">
      <title>CONCLUSIONS</title>
      <p>We used four approaches at MediaEval 2014 Placing Task.
The GeoFusion approaches achieved the best results in the
experiments clearly outperforming the other approaches. These
approaches achieve the best results because combine high
precision rules based on Toponym Disambiguation
heuristics and predictions that come from an HLM models. The
GeoKB rules used in the GeoFusion approach achieved 81.17%
of accuracy (131,207 of 161,628 photos/videos) predicting
up to 100km. The most di cult cases for prediction with
our textual based approach are the ones with few textual
3http://de.wikipedia.org/wiki/Wikipedia:WikiProjekt\_Georeferenzierung/Hauptseite/
Wikipedia-World/en
30 %
20 %
10 %
information and tags. In this evaluation we tried an
approach that uses sometimes the English Wikipedia
Georefenced pages to handle these cases. The GeoFusion+GeoWiki
approach (that uses an HLM model of English Wikipedia
georeferenced pages) does not generally o ers better
performance than the original GeoFusion approach. This approach
only improved very slightly the results for estimations at
10km. The HLM approach with Re-Ranking obtained the
best results in the 10m to 1km range because the model
takes some bene ts of relating non-geographical descriptive
keywords and place names appearing in the geographical
coordinates' associated metadata.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work has been supported by the Spanish Research
Department (SKATER Project: TIN2012-38584-C06-01). TALP
Research Center is recognized as a Quality Research Group
(2014 SGR 1338) by AGAUR, the Research Department of
the Catalan Government.
5.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thomee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Friedland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Borth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Elizalde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gottlieb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Carrano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pearce</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Poland</surname>
          </string-name>
          .
          <article-title>The Placing Task: A Large-Scale Geo-Estimation Challenge for Social-Media Videos and Images</article-title>
          .
          <source>In Proceedings of the 3rd ACM International Workshop on Geotagging and Its Applications in Multimedia</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Ferres</surname>
          </string-name>
          and
          <article-title>Horacio Rodr guez</article-title>
          . TALP at MediaEval 2010 Placing Task:
          <article-title>Geographical Focus Detection of Flickr Textual Annotations</article-title>
          .
          <source>In Working Notes of the MediaEval 2010 Workshop</source>
          , Pisa, Italy,
          <year>October 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Ferres</surname>
          </string-name>
          and
          <article-title>Horacio Rodr guez. Georeferencing Textual Annotations and Tagsets with Geographical Knowledge and Language Models</article-title>
          . In Actas de la
          <source>SEPLN</source>
          <year>2011</year>
          , Huelva, Spain,
          <year>September 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ferres</surname>
          </string-name>
          and
          <string-name>
            <surname>H.</surname>
          </string-name>
          <article-title>Rodr guez</article-title>
          . TALP at MediaEval 2011 Placing Task:
          <article-title>Georeferencing Flickr Videos with Geographical Knowledge and Information Retrieval</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2011 Workshop</source>
          , Santa Croce in Fossabanda, Pisa, Italy, September 1-
          <issue>2</issue>
          ,
          <year>2011</year>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          .
          <article-title>Using Language Models for Information Retrieval</article-title>
          .
          <source>PhD thesis</source>
          , Enschede, Netherlands,
          <year>January 2001</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>