<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Working Notes for the Placing Task at MediaEval 2013</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudia Hauff</string-name>
          <email>c.hau @tudelft.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bart Thomeey</string-name>
          <email>bthomee@yahoo-inc.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michele Trevisiol</string-name>
          <email>trevi@yahoo-inc.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Delft University of Technology</institution>
          ,
          <addr-line>Delft</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Yahoo! Research</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>18</fpage>
      <lpage>19</lpage>
      <abstract>
        <p>This paper provides a description of the MediaEval 2013 Placing Task. The primary task of location estimation asks participants to place images on the world map, that is, to automatically estimate the latitude/longitude coordinates at which a photograph was taken. The newly introduced secondary task of placeability prediction asks participants to estimate the error of their predicted location. Annotating images with this kind of geographical location tag, or geotags, has a number of applications in personalization, recommendation, crisis management and archiving. Currently, the vast majority of images online are not labelled with this kind of data. This task encourages participants to nd innovative ways of automatically geo-labelling images while at the same time providing a measure of their algorithm' accuracy. This year's data were drawn from Flickr. In comparison to previous editions of this task, the test set has not only increased drastically in size but has also been derived according to di erent assumptions in order to model a more realistic use-case scenario.</p>
      </abstract>
      <kwd-group>
        <kwd>location prediction</kwd>
        <kwd>geotags</kwd>
        <kwd>image labelling</kwd>
        <kwd>benchmark</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        This task challenges participants to develop techniques
to automatically annotate images using their visual content
and selected, associated textual metadata. In particular, we
wish to see those taking part to extend and improve upon the
work of previous placing tasks at MediaEval and elsewhere
in the community [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7">1, 7, 3, 2, 6, 5, 4</xref>
        ].
      </p>
      <p>Images (and videos) that do contain latitude/longitude
coordinates have usually been annotated in one of two ways:
automatically by the device or manually by the user in a
post-processing step. An increasing number of devices (e.g.,
camera or camera-equipped mobile phone) are available that
can automatically encode geotags, using satellite-based
positioning systems, mobile cell towers or look-up of the
coordinates of local Wi-Fi networks. Users are also becoming more
All authors contributed equally to this work.
yThe European Commission FP7/2007-2013 supported
this research through the LiMoSINe project under Grant
#288024.
aware of the value of adding such data manually, as shown
by the increase in photo management software and portals
that allow users to annotate, browse and search according to
location (e.g., Flickr, Apple's iPhoto and Aperture, Google
Picasa WebAlbums).</p>
      <p>However, newly uploaded digital media and images in
particular, with any form of geographical data, are still
relatively rare compared to the total quantity uploaded. There
is also a signi cant amount of data that has already been
uploaded that does not currently have geotags. These
observations provide the motivation for this task.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>DATA</title>
      <p>The dataset for the Placing Task 2013 is di erent from the
previous editions. It contains nearly nine million Flickr
images released by the owners with Creative Commons License.</p>
      <p>All the images are geo-tagged with the Flickr accuracy level
16. On Flickr, di erent accuracy levels exist, the highest
being 16 which means that the location is accurate at the
street-level. Note though, that 16 is also Flickr's default
accuracy level when no accuracy level is provided; a fact that
can introduce some noise.</p>
      <p>The test data is available in ve di erent sizes, in order to
allow participation in the task even without very powerful
computational resources.</p>
      <p>When splitting the data into development and test sets,
we ensured that none of the users appear in both sets. This
means, that a user either contributes all his crawled images
to the development set or to the test set. We assume this to
be a more realistic setup, compared to previous years, where
users' contributions to both sets were usually mixed. In our
scenario, we thus focus on users who have not yet geotagged
a single one of their images.</p>
      <p>For each image, a set of selected metadata elements and
visual features are provided to the participants. The
images' metadata was crawled through the Flickr API. The
visual features (derived for the image version pointed to by
photoURL, 500 pixels on the longest side) were extracted
with the open-source LIRE library version 0.9.31, a content
based image retrieval library. The default parameter
settings were used. Note that, whilst the images themselves
are not distributed in this task, they are publicly accessible
on Flickr and the provided metadata contains links to the
source images.</p>
      <p>A detailed overview of the provided features for both the
development and test data is given in Table 1. Note that the
1LIRE, https://code.google.com/p/lire/
Visual features
photoID, userID, photoURL, associated tags, date taken, date uploaded, number of views, geotag
accuracy, licenseID
names of the visual features correspond to the LIRE classes
with which they were extracted.</p>
      <sec id="sec-2-1">
        <title>Development Data.</title>
        <p>The development data consists of approximately 8:5
million images. In addition to the features listed in Table 1,
the latitude/longitude coordinates at which each image was
taken is also provided to the participants.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Test Data.</title>
        <p>The test data consists photoIDs for which the primary
(location estimation) and secondary (placeability prediction)
tasks should be executed. We have developed ve test sets
of di erent sizes (between 5300 and 262; 000 images)
following the Russian dolls approach, i.e. the larger test sets
contain all images of the smaller test sets. This means that
the participants should only use one of the ve possible test
sets, the largest one that they are able to process.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>TASK DETAILS</title>
      <sec id="sec-3-1">
        <title>Location Estimation Task.</title>
        <p>The location estimation task is the primary task and the
same as in previous years: given an image, estimate its
location on the world map and provide a point estimate in terms
of a pair of latitude/longitude coordinates.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Placeability Task.</title>
        <p>This (optional) secondary task asks participants to
estimate the error of the locations predicted. Is it possible to
automatically estimate the accuracy of the predicted
location (derived for the location estimation task)?</p>
        <p>To provide an intuition, consider the following basic
procedure that could be employed to estimate the error: lets
assume that the location estimation approach computes a
ranked list of locations (and the top location is returned as
estimated location). If the top n locations are distributed
all over the globe, the method may have low con dence and
thus the estimated error would be high. On the other hand,
if the top n locations are spatially very close (e.g., having a
standard deviation of a few kilometres), then the method's
con dence in the location estimate would be high and the
error thus low. In this task, for each test image, the
participants are asked to specify an estimate of the error in
kilometres.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Runs.</title>
        <p>Participants may submit between two and ve runs. They
can make use of the provided metadata and visual features,
as well as external resources (e.g. gazetteers, dictionaries,
Web corpora), depending on the run type. The rst required
run allows for the free use of the provided data but no
additional resources. For the second required run only visual
features may be used. Participants can also submit three
optional runs which only have one restriction: it is not
allowed to crawl/use the geotags of the items contained in the
test set. To summarize:
1.
2.</p>
        <p>run (required): Only the provided data (metadata
and/or visual features) may be used.</p>
        <p>run (required): Only visual features may be used.
3.-5. runs (optional): Anything is accepted, except for
crawling the exact items contained in the test set.
4.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>EVALUATION</title>
      <p>The geo-coordinates associated with the images of the test
set (not provided to the participants) will be used as the
ground truth. For the primary task, the evaluation will be
carried out in a series of widening circles: f1; 10; 100; 1000g
kilometres. If a reported location is found within a given
circle radius, it is counted as correctly localised. The accuracy
over each circle is reported. Additionally, the median error
(in kilometres) over the test set is computed, that is, the
maximum error distance that 50% of the test data achieves.
To take into account the geographic nature of the task, the
Haversine distance is used.</p>
      <p>For the placeability task, we rely on the linear and rank
correlation (Kendall's Tau) coe cients to compare the
ability of the algorithms to estimate the error correctly.
Specifically, we correlate the true error distance in kilometres (as
determined for the primary task) with the predicted error
distance. A high correlation coe cient indicates that the
algorithm is able to infer the accuracy of the estimation.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Adam</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Kelm</surname>
          </string-name>
          .
          <source>Working Notes for the Placing Task at MediaEval</source>
          <year>2012</year>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Crandall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Backstrom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Huttenlocher</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Kleinberg</surname>
          </string-name>
          .
          <article-title>Mapping the World's Photos</article-title>
          .
          <source>In WWW</source>
          , pages
          <volume>761</volume>
          {
          <fpage>770</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Hau</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.-J.</given-names>
            <surname>Houben</surname>
          </string-name>
          .
          <article-title>Placing images on the world map: a microblog-based enrichment approach</article-title>
          .
          <source>In SIGIR</source>
          , pages
          <volume>691</volume>
          {
          <fpage>700</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hays</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Efros</surname>
          </string-name>
          .
          <article-title>IM2GPS: estimating geographic information from a single image</article-title>
          .
          <source>In CVPR</source>
          , volume
          <volume>05</volume>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N. O</given-names>
            <surname>'Hare</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Murdock</surname>
          </string-name>
          .
          <article-title>Modeling locations with social media</article-title>
          .
          <source>Inf</source>
          . Retr.,
          <volume>16</volume>
          (
          <issue>1</issue>
          ):
          <volume>30</volume>
          {
          <fpage>62</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Serdyukov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Murdock</surname>
          </string-name>
          , and R. van Zwol.
          <source>Placing Flickr Photos on a Map</source>
          .
          <source>In SIGIR</source>
          , pages
          <volume>484</volume>
          {
          <fpage>491</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Trevisiol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jegou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Delhumeau</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Gravier. Retrieving</surname>
          </string-name>
          geo
          <article-title>-location of videos with a divide &amp; conquer hierarchical multimodal approach</article-title>
          .
          <source>In ICMR, pages 1{8</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>