<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">RECOD Working Notes for Placing Task MediaEval 2011 *</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Lin</forename><forename type="middle">Tzy</forename><surname>Li</surname></persName>
							<email>lintzyli@ic.unicamp.br</email>
							<affiliation key="aff0">
								<orgName type="department">Institute of Computing</orgName>
								<orgName type="institution">University of Campinas -UNICAMP</orgName>
								<address>
									<addrLine>-852</addrLine>
									<postCode>13083</postCode>
									<settlement>Campinas</settlement>
									<region>SP</region>
									<country key="BR">Brazil</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Jurandy</forename><surname>Almeida</surname></persName>
							<email>jurandy.almeida@ic.unicamp.br</email>
							<affiliation key="aff0">
								<orgName type="department">Institute of Computing</orgName>
								<orgName type="institution">University of Campinas -UNICAMP</orgName>
								<address>
									<addrLine>-852</addrLine>
									<postCode>13083</postCode>
									<settlement>Campinas</settlement>
									<region>SP</region>
									<country key="BR">Brazil</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Ricardo</forename><surname>Da</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">Institute of Computing</orgName>
								<orgName type="institution">University of Campinas -UNICAMP</orgName>
								<address>
									<addrLine>-852</addrLine>
									<postCode>13083</postCode>
									<settlement>Campinas</settlement>
									<region>SP</region>
									<country key="BR">Brazil</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">S</forename><surname>Torres</surname></persName>
							<email>rtorres@ic.unicamp.br</email>
							<affiliation key="aff0">
								<orgName type="department">Institute of Computing</orgName>
								<orgName type="institution">University of Campinas -UNICAMP</orgName>
								<address>
									<addrLine>-852</addrLine>
									<postCode>13083</postCode>
									<settlement>Campinas</settlement>
									<region>SP</region>
									<country key="BR">Brazil</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">RECOD Working Notes for Placing Task MediaEval 2011 *</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">8F788C7893755C5AFF13CB44FA6FF73E</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-24T01:28+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>H</term>
					<term>3</term>
					<term>3 [Information Search and Retrieval]</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>This work is developed in the context of placing task at MediaEval 2011. It consists in automatically assigning geographical coordinates to a set of videos. Our group proposed an architecture design for the multimodal geocoding. In this paper, we focused on implementing a simple content-based approach, which is part of the proposed framework. The reported results show our strategy compared to those from previous year participant using only visual content to accomplish this task.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">INTRODUCTION</head><p>The geographic information is present in people's daily life, thus it is not surprising that there is a huge amount of data on the Web about geographical entities and a great interest in localizing them on maps. That information is often enclosed in digital objects (e.g., documents, image, and videos). Once they are geocoded (i.e., associated to a latitude or longitude), one can perform geographical queries.</p><p>Current solutions for geocoding multimedia material are usually based on textual information <ref type="bibr" target="#b6">[2,</ref><ref type="bibr" target="#b10">6]</ref>. Such a strategy depends on the human intervention to tag textual descriptions of the data. However, there is a lack of objectivity and completeness of those descriptions, since the understanding of the visual content of multimedia data may change according to the experience, and perception of each subject, not to mention lexical and geographical problems in recognizing place names <ref type="bibr">[5]</ref>. This opens new venues for the investigation of methods that use image/video content in the geocoding process. Furthermore, data fusion/rank aggregation approaches could be also used for combining evidences found in both textual and visual content.</p><p>In this paper, we present an approach for visual contentbased geocoding, although we aim to explore the combination of textual and visual content of digital objects in order to improve their geocoding. The idea here is to test how well video similarity in term of its motion sequence would fit our purposes of predicting their location.</p><p>This work is developed in the context of Placing Task at MediaEval 2011. The goal of such a task is to automatically assign geographical coordinates (latitude and longitude) to a set of annotated videos. More details regarding data, task, and evaluation are described in <ref type="bibr" target="#b11">[7]</ref>. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">THE PROPOSED FRAMEWORK</head><p>The proposed architecture for dealing with multimodal geocoding is composed by three modules (Figure <ref type="figure" target="#fig_0">1</ref> In this paper, we focused on the second module, exploring a method to identify similar videos whose visual content indicates where those videos were filmed. Although it is allowed to use all the metadata associated to the given video, such as descriptions and tags provided by users, we focused on geocoding based on visual features of the videos.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1">Extracting &amp; Comparing Visual Features</head><p>Instead of using any keyframe visual features provided by the organizers, we adopted a simple and fast algorithm to compare video sequences described in <ref type="bibr">[1]</ref>. It consists of three main steps: (1) partial decoding; (2) feature extraction; and (3) signature generation.</p><p>For each frame of an input video, motion features are extracted from the video stream. For that, 2×2 ordinal matrices are obtained by ranking the intensity values of the four luminance (Y) blocks of each macroblock. This strategy is employed for computing both the spatial feature of the 4-blocks of a macroblock and the temporal feature of corresponding blocks in three frames (previous, current, and next). Each possible combination of the ordinal measures is treated as an individual pattern of 16-bits (i.e., 2-bits for each element of the ordinal matrices). Finally, the spatiotemporal pattern of all the macroblocks of the video sequence are accumulated to form a normalized histogram.</p><p>For a detailed discussion of this procedure, refer to <ref type="bibr">[1]</ref>.</p><p>The comparison of histograms can be performed by any vectorial distance function like Manhattan (L1) or Euclidean (L2) distances. In this work, we compare video sequences by using the histogram intersection, which is defined as</p><formula xml:id="formula_0">d(HV 1 , HV 2 ) = i min(H i V 1 , H i V 2 ) i H i V 1 ,</formula><p>where HV 1 and HV 2 are the histograms extracted from the videos V1 and V2, respectively. This function returns a real value ranging from 0 for situations in which those histograms are not similar at all, to 1 when they are identical.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2">Geocoding the Visual Content</head><p>We used 10,216 videos from the development set released by Placing Task organizer as geo-profiles against which each test video was compared to.</p><p>In order to assess how well we did, only relying on visual content during the development phase, we extracted the visual content of each provided video, then we compared all videos of the set against each other, and finally, for each video, we produced a list of videos ordered by similarity in descending order. Considering that a query video always is the best match to itself, thus it will be the first in this list, we took the second video from the top list as the one that will transfer its known lat/long to the query video.</p><p>For the test result, applying the visual feature extraction and the similarity computation explained previously, each video in test set (5,347) was compared with those in the development set. Then, for each test video, an ordered list of similar videos from the development set was produced along with its similarity score to that given test video. Finally, we picked the most similar video of this list as the one that will transfer its known lat/long to the query test video, and reported that lat/long as the one to be given to test video.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">EXPERIMENTAL RESULTS</head><p>For this task, we performed one submission for the run that considered just visual content. The evaluation results are shown in Table <ref type="table" target="#tab_0">1</ref>. Note that, by relying just on video similarity based on its visual content, our algorithm will hit 79.45% only when accepting an error of 10,000 km between the ground truth and the assigned point. However, when considering 100 km of error, it predicts lat/long correctly for only 2.71%. These results underperform those from the reference algorithm for this task (winner of the last year), which just analyzes user-contributed tags for predicting the geotag of a video (73.6% of videos are within 100 km) <ref type="bibr" target="#b8">[4]</ref>.</p><p>However, we are interested in comparing to other results using only video content to accomplish the placing task. For instance, Kelm et al. <ref type="bibr" target="#b7">[3]</ref>, who also reported their results when only visual content of test videos were used to predict their location on Earth, have used visual features of the development set for training a multi-class SVM classifier with RBF kernel. Their best results were achieved by a hierarchical clustering with a diameter threshold of 100 km, which determined 317 classes for the SVM with the descriptors CED, FCTH, and Gabor. They presented their results for video's location correctly predicted within radius of 50 km, 100 km, 200 km, 750 km, and 2,500 km.</p><p>In order to compare our results to those presented by Kelm et al. <ref type="bibr" target="#b7">[3]</ref>, we aggregated the evaluation results presented in Table <ref type="table" target="#tab_0">1</ref> according to their experimental protocol. This regrouping was possible due to the placing task organizers, who made available to all participants of that task: their tool to calculate the distance (Haversine distance formula) from ground truth to the estimated location for each result; and the test videos ground truth.</p><p>Table <ref type="table" target="#tab_1">2</ref> compares our approach with the results reported by Kelm et al. (adopted from their Table <ref type="table">6</ref>) <ref type="bibr" target="#b7">[3]</ref>. Notice that our method, although simpler, shows high precision relative to their clustering-and-classification method. The key advantage of our technique is its computational efficiency. Unlike them, we did not use any data to train any classifier. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">CONCLUSIONS</head><p>Relying just on video content to estimate its location still poses a challenge. It seems that this task requires using textual information found in video metadata such as descriptions, user tags, external knowledges bases as shown by some related works.</p><p>Our method used the video similarity between videos in development set and those in test set to estimate location of those. The similarity in this work is given by motion patterns extracted from the video streams. This algorithm is simple and achieved comparable results to those more complex presented by previous work that also was based just on video visual clues.</p><p>We believe that we can improve the results by developing new video similarities approaches as well as new combining methods for image and textual evidences in the context of geocoding digital objects.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>Figure 1 :</head><label>1</label><figDesc>Figure 1: Multimodal geocoding proposal</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head></head><label></label><figDesc>): (1) text-based geocoding; (2) content-based geocoding; and (3) data fusion/rank aggregation-based geocoding. The first module is in charge of geocoding based solely on textual part of the digital object. Content-based geocoding module is responsible for dealing with and geocoding based on its visual content. Finally, the rank aggregation-based module combines the results generated by the previous modules and gives the final result of the geocoding. The idea is to rely on text and image whenever possible.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_0"><head>Table 1 :</head><label>1</label><figDesc>Results using only videos visual content (distance between ground truth and estimated)</figDesc><table><row><cell>Radius (km)</cell><cell>1</cell><cell>10</cell><cell>20</cell><cell>50</cell><cell>100</cell><cell>200</cell><cell>500</cell><cell>1000 2000 5000 10000</cell></row><row><cell cols="9">Dev set: % in range 14.42 16.02 16.44 16.93 17.51 18.36 21.20 26.04 34.76 46.76 84.29</cell></row><row><cell cols="2">Test set: % in range 0.21</cell><cell>1.12</cell><cell>1.59</cell><cell>1.93</cell><cell>2.71</cell><cell>3.33</cell><cell cols="2">6.08 12.16 22.11 37.78 79.45</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_1"><head>Table 2 :</head><label>2</label><figDesc>Our regrouped test results vs. Kelm et al.</figDesc><table><row><cell>Radius (km)</cell><cell>50</cell><cell>100 200</cell><cell>750</cell><cell>2500</cell></row><row><cell cols="5">Our approach % 1.93 2.71 3.33 9.18 24.48</cell></row><row><cell>Kelm et al. %</cell><cell cols="4">3.38 5.26 6.23 10.65 19.92</cell></row></table></figure>
		</body>
		<back>
			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<monogr>
		<idno>-./01,</idno>
		<ptr target="%-.+/&quot;%&lt;" />
		<title level="m">01,+*% 8&apos;*(0(&apos;$&quot;-%4&apos;$B4+*@)%% $&apos;@-)%(&apos;$&quot;)%</title>
				<imprint/>
	</monogr>
	<note>+%A0$6%4&apos;$B4+*@)%% $&apos;@-)%(&apos;$&quot;)%. :;% 8&apos;*(0(&apos;$&quot;-%4&apos;$B4+*@)% &quot;70</note>
</biblStruct>

<biblStruct xml:id="b1">
	<monogr>
		<title level="m" type="main">0(&apos;$&quot;-%4&apos;$B4+*@)%</title>
		<imprint/>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<analytic>
		<title level="a" type="main">0%</title>
		<idno>-&apos;3/</idno>
	</analytic>
	<monogr>
		<title level="m">&apos;*(0(&apos;$&quot;-%4&apos;$B4+*@)% &quot;70</title>
				<imprint/>
	</monogr>
	<note>+%+*$+4+@F% !6</note>
</biblStruct>

<biblStruct xml:id="b3">
	<monogr>
		<title level="m" type="main">/-% Data fusion</title>
		<author>
			<persName><forename type="first">G</forename></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename></persName>
		</author>
		<imprint/>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<monogr>
		<title/>
		<author>
			<persName><surname>References</surname></persName>
		</author>
		<imprint/>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">Comparison of video sequences with histograms of motion patterns</title>
		<author>
			<persName><forename type="first">J</forename><surname>Almeida</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><forename type="middle">J</forename><surname>Leite</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">S</forename><surname>Torres</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">ICIP</title>
				<imprint>
			<date type="published" when="2011">2011</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b6">
	<analytic>
		<title level="a" type="main">Geographical information retrieval</title>
		<author>
			<persName><forename type="first">C</forename><forename type="middle">B</forename><surname>Jones</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">S</forename><surname>Purves</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Int. J. Geogr. Inf. Sci</title>
		<imprint>
			<biblScope unit="volume">22</biblScope>
			<biblScope unit="issue">3</biblScope>
			<biblScope unit="page" from="219" to="228" />
			<date type="published" when="2008">2008</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b7">
	<analytic>
		<title level="a" type="main">Multi-modal, Multi-resource Methods for Placing Flickr Videos on the Map</title>
		<author>
			<persName><forename type="first">P</forename><surname>Kelm</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Schmiedeke</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Sikora</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">ICMR</title>
		<imprint>
			<biblScope unit="volume">52</biblScope>
			<biblScope unit="page">8</biblScope>
			<date type="published" when="2011">2011</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b8">
	<analytic>
		<title level="a" type="main">Automatic tagging and geotagging in video collections and communities</title>
		<author>
			<persName><forename type="first">M</forename><surname>Larson</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Soleymani</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Serdyukov</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Rudinac</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">ICMR</title>
		<imprint>
			<biblScope unit="volume">51</biblScope>
			<biblScope unit="page">8</biblScope>
			<date type="published" when="2011">2011</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b9">
	<analytic>
		<title level="a" type="main">Geographic information retrieval and digital libraries</title>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">R</forename><surname>Larson</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">ECDL</title>
		<imprint>
			<biblScope unit="volume">5714</biblScope>
			<biblScope unit="page" from="461" to="464" />
			<date type="published" when="2009">2009. 2009</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b10">
	<analytic>
		<title level="a" type="main">Geotagging in multimedia and computer vision-a survey</title>
		<author>
			<persName><forename type="first">J</forename><surname>Luo</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Joshi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Yu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Gallagher</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Multimedia Tools Appl</title>
		<imprint>
			<biblScope unit="volume">51</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page" from="187" to="211" />
			<date type="published" when="2011">2011</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b11">
	<monogr>
		<title level="m" type="main">Working Notes for the Placing Task at MediaEval 2011</title>
		<author>
			<persName><forename type="first">A</forename><surname>Rae</surname></persName>
		</author>
		<author>
			<persName><forename type="first">V</forename><surname>Murdock</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Serdyukov</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Kelm</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2011">2011</date>
			<publisher>MediaEval</publisher>
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
