=Paper=
{{Paper
|id=None
|storemode=property
|title=RECOD Working Notes for Placing Task MediaEval 2011
|pdfUrl=https://ceur-ws.org/Vol-807/Li_UNICAMP_Placing_me11wn.pdf
|volume=Vol-807
|dblpUrl=https://dblp.org/rec/conf/mediaeval/LiAT11
}}
==RECOD Working Notes for Placing Task MediaEval 2011==
RECOD Working Notes for Placing Task MediaEval 2011 ∗
Lin Tzy Li, Jurandy Almeida, and Ricardo da S. Torres
Institute of Computing, University of Campinas – UNICAMP
13083-852, Campinas, SP – Brazil
{lintzyli, jurandy.almeida, rtorres}@ic.unicamp.br
ABSTRACT <"+1'/-0*@B 8'*(0('$"-%4'$B4+*@)%
!"#$% "70("*."-)%-.+/"%
This work is developed in the context of placing task at <"+.+(0*@%
MediaEval 2011. It consists in automatically assigning geo- &"$'('$')%'**+$',+*-)%
graphical coordinates to a set of videos. Our group proposed ("-./01,+*-%
<"+%+*$+4+@F%
an architecture design for the multimodal geocoding. In this <"+% !6"-'3/0%
paper, we focused on implementing a simple content-based 20-3'4%5"'$3/"-% =">'*,.% <'G"H""/-%
<:;%
?9%
approach, which is part of the proposed framework. The Data fusion
reported results show our strategy compared to those from 89:;%
8+>E0*"%
previous year participant using only visual content to ac- 20("+% &'$.6% 8'*(0('$"-%4'$B4+*@)%
@"+.+(0*@%
"70("*."-)%-.+/"%
complish this task. 1/+."--0*@% 70("+% /"-34$-%
Categories and Subject Descriptors 8'*(0('$"-%4'$B4+*@)%
H.3.3 [Information Search and Retrieval] <"+.+("(% "70("*."-)%-.+/"%
8'*(0('$"-%4'$B4+*@)%%
20("+% 20("+%A0$6%4'$B4+*@)%% $'@-)%('$")%("-./01,+*%
$'@-)%('$")%("-./01,+*% C%>'$.6%-.+/"%
C%.+*D("*."%-.+/"%
1. INTRODUCTION
The geographic information is present in people’s daily
life, thus it is not surprising that there is a huge amount Figure 1: Multimodal geocoding proposal
of data on the Web about geographical entities and a great
interest in localizing them on maps. That information is 2. THE PROPOSED FRAMEWORK
often enclosed in digital objects (e.g., documents, image, The proposed architecture for dealing with multimodal
and videos). Once they are geocoded (i.e., associated to a geocoding is composed by three modules (Figure 1): (1)
latitude or longitude), one can perform geographical queries. text-based geocoding; (2) content-based geocoding; and (3)
Current solutions for geocoding multimedia material are data fusion/rank aggregation-based geocoding. The first
usually based on textual information [2, 6]. Such a strategy module is in charge of geocoding based solely on textual
depends on the human intervention to tag textual descrip- part of the digital object. Content-based geocoding module
tions of the data. However, there is a lack of objectivity and is responsible for dealing with and geocoding based on its
completeness of those descriptions, since the understanding visual content. Finally, the rank aggregation-based module
of the visual content of multimedia data may change ac- combines the results generated by the previous modules and
cording to the experience, and perception of each subject, gives the final result of the geocoding. The idea is to rely
not to mention lexical and geographical problems in recog- on text and image whenever possible.
nizing place names [5]. This opens new venues for the in- In this paper, we focused on the second module, exploring
vestigation of methods that use image/video content in the a method to identify similar videos whose visual content
geocoding process. Furthermore, data fusion/rank aggrega- indicates where those videos were filmed. Although it is
tion approaches could be also used for combining evidences allowed to use all the metadata associated to the given video,
found in both textual and visual content. such as descriptions and tags provided by users, we focused
In this paper, we present an approach for visual content- on geocoding based on visual features of the videos.
based geocoding, although we aim to explore the combina-
tion of textual and visual content of digital objects in order
to improve their geocoding. The idea here is to test how
2.1 Extracting & Comparing Visual Features
well video similarity in term of its motion sequence would Instead of using any keyframe visual features provided by
fit our purposes of predicting their location. the organizers, we adopted a simple and fast algorithm to
This work is developed in the context of Placing Task at compare video sequences described in [1]. It consists of three
MediaEval 2011. The goal of such a task is to automatically main steps: (1) partial decoding; (2) feature extraction; and
assign geographical coordinates (latitude and longitude) to (3) signature generation.
a set of annotated videos. More details regarding data, task, For each frame of an input video, motion features are ex-
and evaluation are described in [7]. tracted from the video stream. For that, 2×2 ordinal matri-
ces are obtained by ranking the intensity values of the four
∗
We thank FAPESP, CNPq, and CAPES for financial support. luminance (Y) blocks of each macroblock. This strategy
is employed for computing both the spatial feature of the
4-blocks of a macroblock and the temporal feature of cor-
Copyright is held by the author/owner(s). responding blocks in three frames (previous, current, and
MediaEval 2011 Workshop, September 1-2, 2011, Pisa, Italy next). Each possible combination of the ordinal measures
Table 1: Results using only videos visual content (distance between ground truth and estimated)
Radius (km) 1 10 20 50 100 200 500 1000 2000 5000 10000
Dev set: % in range 14.42 16.02 16.44 16.93 17.51 18.36 21.20 26.04 34.76 46.76 84.29
Test set: % in range 0.21 1.12 1.59 1.93 2.71 3.33 6.08 12.16 22.11 37.78 79.45
is treated as an individual pattern of 16-bits (i.e., 2-bits for determined 317 classes for the SVM with the descriptors
each element of the ordinal matrices). Finally, the spatio- CED, FCTH, and Gabor. They presented their results for
temporal pattern of all the macroblocks of the video se- video’s location correctly predicted within radius of 50 km,
quence are accumulated to form a normalized histogram. 100 km, 200 km, 750 km, and 2,500 km.
For a detailed discussion of this procedure, refer to [1]. In order to compare our results to those presented by Kelm
The comparison of histograms can be performed by any et al. [3], we aggregated the evaluation results presented
vectorial distance function like Manhattan (L1 ) or Euclidean in Table 1 according to their experimental protocol. This
(L2 ) distances. In this work, we compare video sequences regrouping was possible due to the placing task organizers,
by using the histogram intersection, which is defined as who made available to all participants of that task: their
P i i tool to calculate the distance (Haversine distance formula)
i min(HV1 , HV2 ) from ground truth to the estimated location for each result;
d(HV1 , HV2 ) = P i ,
i HV 1 and the test videos ground truth.
Table 2 compares our approach with the results reported
where HV1 and HV2 are the histograms extracted from the
by Kelm et al. (adopted from their Table 6) [3]. Notice that
videos V1 and V2 , respectively. This function returns a real
our method, although simpler, shows high precision rela-
value ranging from 0 for situations in which those histograms
tive to their clustering-and-classification method. The key
are not similar at all, to 1 when they are identical.
advantage of our technique is its computational efficiency.
2.2 Geocoding the Visual Content Unlike them, we did not use any data to train any classifier.
We used 10,216 videos from the development set released Table 2: Our regrouped test results vs. Kelm et al.
by Placing Task organizer as geo-profiles against which each Radius (km) 50 100 200 750 2500
test video was compared to. Our approach % 1.93 2.71 3.33 9.18 24.48
In order to assess how well we did, only relying on visual Kelm et al. % 3.38 5.26 6.23 10.65 19.92
content during the development phase, we extracted the vi-
sual content of each provided video, then we compared all 4. CONCLUSIONS
videos of the set against each other, and finally, for each
Relying just on video content to estimate its location still
video, we produced a list of videos ordered by similarity in
poses a challenge. It seems that this task requires using
descending order. Considering that a query video always is
textual information found in video metadata such as de-
the best match to itself, thus it will be the first in this list,
scriptions, user tags, external knowledges bases as shown by
we took the second video from the top list as the one that
some related works.
will transfer its known lat/long to the query video.
Our method used the video similarity between videos in
For the test result, applying the visual feature extraction
development set and those in test set to estimate location of
and the similarity computation explained previously, each
those. The similarity in this work is given by motion pat-
video in test set (5,347) was compared with those in the
terns extracted from the video streams. This algorithm is
development set. Then, for each test video, an ordered list of
simple and achieved comparable results to those more com-
similar videos from the development set was produced along
plex presented by previous work that also was based just on
with its similarity score to that given test video. Finally,
video visual clues.
we picked the most similar video of this list as the one that
We believe that we can improve the results by developing
will transfer its known lat/long to the query test video, and
new video similarities approaches as well as new combining
reported that lat/long as the one to be given to test video.
methods for image and textual evidences in the context of
geocoding digital objects.
3. EXPERIMENTAL RESULTS
For this task, we performed one submission for the run
that considered just visual content. The evaluation results
5. REFERENCES
[1] J. Almeida, N. J. Leite, and R. S. Torres. Comparison of
are shown in Table 1. Note that, by relying just on video video sequences with histograms of motion patterns. In
similarity based on its visual content, our algorithm will hit ICIP, 2011.
79.45% only when accepting an error of 10,000 km between [2] C. B. Jones and R. S. Purves. Geographical information
the ground truth and the assigned point. However, when retrieval. Int. J. Geogr. Inf. Sci., 22(3):219–228, 2008.
considering 100 km of error, it predicts lat/long correctly [3] P. Kelm, S. Schmiedeke, and T. Sikora. Multi-modal,
for only 2.71%. These results underperform those from the Multi-resource Methods for Placing Flickr Videos on the
reference algorithm for this task (winner of the last year), Map. In ICMR, pages 52:1–52:8, 2011.
which just analyzes user-contributed tags for predicting the [4] M. Larson, M. Soleymani, P. Serdyukov, S. Rudinac, et al.
Automatic tagging and geotagging in video collections and
geotag of a video (73.6% of videos are within 100 km) [4]. communities. In ICMR, pages 51:1–51:8, 2011.
However, we are interested in comparing to other results [5] R. R. Larson. Geographic information retrieval and digital
using only video content to accomplish the placing task. For libraries. In ECDL, volume 5714/2009, pages 461–464, 2009.
instance, Kelm et al. [3], who also reported their results [6] J. Luo, D. Joshi, J. Yu, and A. Gallagher. Geotagging in
when only visual content of test videos were used to predict multimedia and computer vision–a survey. Multimedia Tools
their location on Earth, have used visual features of the de- Appl., 51(1):187–211, 2011.
velopment set for training a multi-class SVM classifier with [7] A. Rae, V. Murdock, P. Serdyukov, and P. Kelm. Working
RBF kernel. Their best results were achieved by a hierarchi- Notes for the Placing Task at MediaEval 2011. In
MediaEval, 2011.
cal clustering with a diameter threshold of 100 km, which