<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>RECOD Working Notes for Placing Task MediaEval 2011</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lin Tzy Li</string-name>
          <email>lintzyli@ic.unicamp.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jurandy Almeida</string-name>
          <email>jurandy.almeida@ic.unicamp.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ricardo da S. Torres</string-name>
          <email>rtorres@ic.unicamp.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Computing, University of Campinas - UNICAMP 13083-852</institution>
          ,
          <addr-line>Campinas, SP -</addr-line>
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <fpage>1</fpage>
      <lpage>2</lpage>
      <abstract>
        <p>This work is developed in the context of placing task at MediaEval 2011. It consists in automatically assigning geographical coordinates to a set of videos. Our group proposed an architecture design for the multimodal geocoding. In this paper, we focused on implementing a simple content-based approach, which is part of the proposed framework. The reported results show our strategy compared to those from previous year participant using only visual content to accomplish this task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>The geographic information is present in people's daily
life, thus it is not surprising that there is a huge amount
of data on the Web about geographical entities and a great
interest in localizing them on maps. That information is
often enclosed in digital objects (e.g., documents, image,
and videos). Once they are geocoded (i.e., associated to a
latitude or longitude), one can perform geographical queries.</p>
      <p>
        Current solutions for geocoding multimedia material are
usually based on textual information [
        <xref ref-type="bibr" rid="ref2 ref6">2, 6</xref>
        ]. Such a strategy
depends on the human intervention to tag textual
descriptions of the data. However, there is a lack of objectivity and
completeness of those descriptions, since the understanding
of the visual content of multimedia data may change
according to the experience, and perception of each subject,
not to mention lexical and geographical problems in
recognizing place names [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This opens new venues for the
investigation of methods that use image/video content in the
geocoding process. Furthermore, data fusion/rank
aggregation approaches could be also used for combining evidences
found in both textual and visual content.
      </p>
      <p>In this paper, we present an approach for visual
contentbased geocoding, although we aim to explore the
combination of textual and visual content of digital objects in order
to improve their geocoding. The idea here is to test how
well video similarity in term of its motion sequence would
t our purposes of predicting their location.</p>
      <p>
        This work is developed in the context of Placing Task at
MediaEval 2011. The goal of such a task is to automatically
assign geographical coordinates (latitude and longitude) to
a set of annotated videos. More details regarding data, task,
and evaluation are described in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>We thank FAPESP, CNPq, and CAPES for nancial support.
20-3'4%5"'$3/"-%</p>
      <p>&lt;"+%
% ="&gt;'*,.%
:; ?9%
&lt;
&lt;"+.+("(%
20("+%
&amp;'$.6%
70("+%
8'*(0('$"-%4'$B4+*@)%
"70("*."-)%-.+/"%
20("+%A0$6%4'$B4+*@)%%
$'@-)%('$")%("-./01,+*%
8'*(0('$"-%4'$B4+*@)%%
$'@-)%('$")%("-./01,+*%
C%&gt;'$.6%-.+/"%
C%.+*D("*."%-.+/"%
8'*(0('$"-%4'$B4+*@)%
"70("*."-)%-.+/"%
Data fusion
8+&gt;E0*"%
@"+.+(0*@%
/"-34$-%
8'*(0('$"-%4'$B4+*@)%
"70("*."-)%-.+/"%
2.</p>
    </sec>
    <sec id="sec-2">
      <title>THE PROPOSED FRAMEWORK</title>
      <p>The proposed architecture for dealing with multimodal
geocoding is composed by three modules (Figure 1): (1)
text-based geocoding; (2) content-based geocoding; and (3)
data fusion/rank aggregation-based geocoding. The rst
module is in charge of geocoding based solely on textual
part of the digital object. Content-based geocoding module
is responsible for dealing with and geocoding based on its
visual content. Finally, the rank aggregation-based module
combines the results generated by the previous modules and
gives the nal result of the geocoding. The idea is to rely
on text and image whenever possible.</p>
      <p>In this paper, we focused on the second module, exploring
a method to identify similar videos whose visual content
indicates where those videos were lmed. Although it is
allowed to use all the metadata associated to the given video,
such as descriptions and tags provided by users, we focused
on geocoding based on visual features of the videos.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Extracting &amp; Comparing Visual Features</title>
      <p>
        Instead of using any keyframe visual features provided by
the organizers, we adopted a simple and fast algorithm to
compare video sequences described in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It consists of three
main steps: (1) partial decoding; (2) feature extraction; and
(3) signature generation.
      </p>
      <p>
        For each frame of an input video, motion features are
extracted from the video stream. For that, 2 2 ordinal
matrices are obtained by ranking the intensity values of the four
luminance (Y) blocks of each macroblock. This strategy
is employed for computing both the spatial feature of the
4-blocks of a macroblock and the temporal feature of
corresponding blocks in three frames (previous, current, and
next). Each possible combination of the ordinal measures
is treated as an individual pattern of 16-bits (i.e., 2-bits for
each element of the ordinal matrices). Finally, the
spatiotemporal pattern of all the macroblocks of the video
sequence are accumulated to form a normalized histogram.
For a detailed discussion of this procedure, refer to [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>The comparison of histograms can be performed by any
vectorial distance function like Manhattan (L1) or Euclidean
(L2) distances. In this work, we compare video sequences
by using the histogram intersection, which is de ned as
d(HV1 ; HV2 ) =</p>
      <p>Pi min(HVi1 ; HVi2 ) ;</p>
      <p>P i
i HV1
where HV1 and HV2 are the histograms extracted from the
videos V1 and V2, respectively. This function returns a real
value ranging from 0 for situations in which those histograms
are not similar at all, to 1 when they are identical.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Geocoding the Visual Content</title>
      <p>We used 10,216 videos from the development set released
by Placing Task organizer as geo-pro les against which each
test video was compared to.</p>
      <p>In order to assess how well we did, only relying on visual
content during the development phase, we extracted the
visual content of each provided video, then we compared all
videos of the set against each other, and nally, for each
video, we produced a list of videos ordered by similarity in
descending order. Considering that a query video always is
the best match to itself, thus it will be the rst in this list,
we took the second video from the top list as the one that
will transfer its known lat/long to the query video.</p>
      <p>For the test result, applying the visual feature extraction
and the similarity computation explained previously, each
video in test set (5,347) was compared with those in the
development set. Then, for each test video, an ordered list of
similar videos from the development set was produced along
with its similarity score to that given test video. Finally,
we picked the most similar video of this list as the one that
will transfer its known lat/long to the query test video, and
reported that lat/long as the one to be given to test video.</p>
    </sec>
    <sec id="sec-5">
      <title>EXPERIMENTAL RESULTS</title>
      <p>
        For this task, we performed one submission for the run
that considered just visual content. The evaluation results
are shown in Table 1. Note that, by relying just on video
similarity based on its visual content, our algorithm will hit
79.45% only when accepting an error of 10,000 km between
the ground truth and the assigned point. However, when
considering 100 km of error, it predicts lat/long correctly
for only 2.71%. These results underperform those from the
reference algorithm for this task (winner of the last year),
which just analyzes user-contributed tags for predicting the
geotag of a video (73.6% of videos are within 100 km) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        However, we are interested in comparing to other results
using only video content to accomplish the placing task. For
instance, Kelm et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], who also reported their results
when only visual content of test videos were used to predict
their location on Earth, have used visual features of the
development set for training a multi-class SVM classi er with
RBF kernel. Their best results were achieved by a
hierarchical clustering with a diameter threshold of 100 km, which
determined 317 classes for the SVM with the descriptors
CED, FCTH, and Gabor. They presented their results for
video's location correctly predicted within radius of 50 km,
100 km, 200 km, 750 km, and 2,500 km.
      </p>
      <p>
        In order to compare our results to those presented by Kelm
et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], we aggregated the evaluation results presented
in Table 1 according to their experimental protocol. This
regrouping was possible due to the placing task organizers,
who made available to all participants of that task: their
tool to calculate the distance (Haversine distance formula)
from ground truth to the estimated location for each result;
and the test videos ground truth.
      </p>
      <p>
        Table 2 compares our approach with the results reported
by Kelm et al. (adopted from their Table 6) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Notice that
our method, although simpler, shows high precision
relative to their clustering-and-classi cation method. The key
advantage of our technique is its computational e ciency.
Unlike them, we did not use any data to train any classi er.
4.
      </p>
    </sec>
    <sec id="sec-6">
      <title>CONCLUSIONS</title>
      <p>Relying just on video content to estimate its location still
poses a challenge. It seems that this task requires using
textual information found in video metadata such as
descriptions, user tags, external knowledges bases as shown by
some related works.</p>
      <p>Our method used the video similarity between videos in
development set and those in test set to estimate location of
those. The similarity in this work is given by motion
patterns extracted from the video streams. This algorithm is
simple and achieved comparable results to those more
complex presented by previous work that also was based just on
video visual clues.</p>
      <p>We believe that we can improve the results by developing
new video similarities approaches as well as new combining
methods for image and textual evidences in the context of
geocoding digital objects.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. J.</given-names>
            <surname>Leite</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Torres</surname>
          </string-name>
          .
          <article-title>Comparison of video sequences with histograms of motion patterns</article-title>
          .
          <source>In ICIP</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C. B.</given-names>
            <surname>Jones</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Purves</surname>
          </string-name>
          .
          <article-title>Geographical information retrieval</article-title>
          .
          <source>Int. J. Geogr. Inf. Sci.</source>
          ,
          <volume>22</volume>
          (
          <issue>3</issue>
          ):
          <volume>219</volume>
          {
          <fpage>228</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kelm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schmiedeke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Sikora</surname>
          </string-name>
          <article-title>. Multi-modal, Multi-resource Methods for Placing Flickr Videos on the Map</article-title>
          .
          <source>In ICMR, pages 52:1{52:8</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Soleymani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Serdyukov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rudinac</surname>
          </string-name>
          , et al.
          <article-title>Automatic tagging and geotagging in video collections and communities</article-title>
          .
          <source>In ICMR, pages 51:1{51:8</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Larson</surname>
          </string-name>
          .
          <article-title>Geographic information retrieval and digital libraries</article-title>
          .
          <source>In ECDL</source>
          , volume
          <volume>5714</volume>
          /
          <year>2009</year>
          , pages
          <fpage>461</fpage>
          {
          <fpage>464</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Gallagher</surname>
          </string-name>
          .
          <article-title>Geotagging in multimedia and computer vision{a survey</article-title>
          .
          <source>Multimedia Tools Appl.</source>
          ,
          <volume>51</volume>
          (
          <issue>1</issue>
          ):
          <volume>187</volume>
          {
          <fpage>211</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Murdock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Serdyukov</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Kelm</surname>
          </string-name>
          .
          <article-title>Working Notes for the Placing Task at MediaEval 2011</article-title>
          . In MediaEval,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>