<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Synchronization of Multi-User Event Media (SEM) at MediaEval 2014: Task Description, Datasets, and Evaluation</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Nicola</forename><surname>Conci</surname></persName>
							<email>nicola.conci@unitn.it</email>
							<affiliation key="aff0">
								<orgName type="institution">DISI -University of Trento Trento</orgName>
								<address>
									<country key="IT">Italy</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Francesco</forename><surname>De Natale</surname></persName>
							<email>francesco.denatale@unitn.it</email>
							<affiliation key="aff1">
								<orgName type="institution">DISI -University of Trento Trento</orgName>
								<address>
									<country key="IT">Italy</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Vasileios</forename><surname>Mezaris</surname></persName>
							<email>bmezaris@iti.gr</email>
							<affiliation key="aff2">
								<orgName type="institution">CERTH -ITI Thermi</orgName>
								<address>
									<country key="GR">Greece</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">Synchronization of Multi-User Event Media (SEM) at MediaEval 2014: Task Description, Datasets, and Evaluation</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">54678147FC61A3DC4C06948316796BC8</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-24T16:11+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>In this paper we provide an overview of the Synchronization of Multi-User Event Media (SEM) Task that is part of the 2014 MediaEval Benchmark for Multimedia Evaluation. The SEM task is presented this year for the first time in MediaEval, and poses a new challenge, namely the temporal alignment of a series of photo galleries that relate to the same event but have been collected by different users. Besides aligning the pictures on a common timeline, participants are also required to detect the sub-events attended by the users and to group the pictures accordingly. The task is validated on two different datasets related to the 2010 and 2012 Winter / Summer Olympic games, each dataset comprising a variable number of pictures, galleries, and subevents.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">INTRODUCTION</head><p>Content creation is more and more a collective experience. People attending large social events (a soccer match, a concert), but also personal-scale ones (a wedding, a birthday party) collect dozens of photos and video clips with their smartphones, tablets, cameras, and more recently social cameras. Such information is later exchanged in a number of different ways, including shared repositories, clouds, social networks, etc. In this way, different media galleries are made available to each other, making it possible for any user who attended, or is simply interested to the event, to create his own view of it through summaries, stories, personalized albums <ref type="bibr" target="#b1">[1]</ref> <ref type="bibr" target="#b2">[2]</ref>. However, such a large amount of data turns out to be unstructured and heterogeneous and, even if it would be possible to collect it on the same hard drive, the variability in terms of content, naming, archiving strategies makes it impossible to organize all the event-related material in a simple yet effective manner.</p><p>In this respect, a major issue is the need of aligning and presenting the various media galleries captured during an event in a consistent way <ref type="bibr" target="#b3">[3]</ref>. As a matter of fact, the time and location information attached to the captured media (timestamp, GPS) can be wrong, inaccurate or even missing (for instance, due to wrong setting of the clock/calendar, different time-zone, modification or removal of tags). Similarly, this is also a common situation in historical events and photo archives, where timestamps and especially loca-tion information is rarely available. In some other cases, images might be processed offline for post-production, thus losing the correct temporal information. In such cases, creating a single timeline could turn out to be complicated and challenging, with a concrete risk of representing the event in a misleading way.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">TASK DESCRIPTION</head><p>In our scenario we imagine a number of users (10+) attending the same event and taking photos and videos with different non-synchronized devices (smartphones, handheld cameras, DSLRs, tablets). Each user contributes to the task with one gallery, which includes an arbitrary number of photos, possibly covering just a part of the event, with variable density of acquisitions (single photos are also possible). Assuming that these users would like to merge their photo galleries in a single event-related collection, the best temporal alignment among the galleries should be found, so as to correctly report and preserve the temporal evolution of the event. Furthermore, considering the high variability in terms of acquisition devices, we cannot expect the clocks of each device of the same user to be synchronized, neither in terms of precision, nor in terms of the time zone set by the users. Furthermore, in some cases, also the location data could be unavailable (not all devices have a GPS onboard), further reducing the available information about the captured event. In view of creating a single timeline, these factors may considerably hinder the quality of the alignment, thus different solutions should be envisaged, encompassing the joint analysis of temporal data, position information, and visual similarity.</p><p>Therefore, the SEM task expects teams to provide the estimated time offset between different galleries of pictures collected by different users and cameras. The goal can be summarised as follows: given a set of image collections (galleries) taken by different users/devices at the same event, find the best (relative) time alignment among them at gallery level, and detect the significant sub-events over the whole event collection.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">DATASETS</head><p>For this challenge we make available two different datasets, consisting of a collection of images gathered from Flickr and made available under Creative Commons license. Both datasets refer to well known and structured sport events, namely the Olympic Games held in London in 2012 and the Vancouver Winter Olympic Games of 2010. We have cho- sen to work with these two events because on the one hand they exhibit a clear and organized schedule with precise timing. On the other hand they still exhibit a high variability in terms of visual content, due to the common features across different competitions in the same discipline, as well as strong similarities in the environments, in which the pictures are collected, making the synchronization a non-trivial task. As far as this task is concerned, the images within a gallery are consistent in terms of timestamp, and might include the GPS information. Therefore the temporal offsets are at gallery level thus assuming that every user uses one single device for acquisition.</p><p>The dataset collected from the London Olympics includes 2124 images, divided into 37 galleries. The first gallery comprises a subset of the data provided in the development set and is defined as the reference gallery. The dataset collected from the Vancouver Winter Olympic Games includes 1351 pictures representing most of the competitions, divided into 35 galleries with a variable number of pictures in each gallery. Also in this case, the first gallery is set as the reference. Fig. <ref type="figure" target="#fig_0">1</ref> shows a few samples of the two datasets.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">METRICS AND EVALUATION</head><p>Two objective metrics will be used to evaluate the results:</p><p>• time synchronization error</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>• sub-event detection error</head><p>As far as the first metric is concerned, the goal of the participants is to maximize the number of galleries for which the synchronization error is below a predefined threshold, and to minimize the time shift of those galleries. The synchronization error for a gallery Gi with respect to the reference Gr is defined as ∆Eir = ∆Tir − ∆T * ir , where ∆T * ir is the delay between Gi and Gr calculated on the ground truth. The threshold ∆Emax depends on the duration of the subevents in the dataset, and represents the maximum accepted time lapse within which we consider a gallery as reasonably well-synchronized.</p><p>As far the metrics for evaluation are concerned, we have considered for the temporal alignment the precision (Eq. 1) and accuracy (Eq. 2). For the quality of the clustering, we use the Rand Index (RI), as from Eq. 3, the Jaccard index (JI) Eq. 4 , and the F1 score Eq. 5, where P and R represent the Precision and Recall, respectively.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>P recision</head><formula xml:id="formula_0">= M N − 1 = Card (∆Eir &lt; ∆Emax) N − 1 (1) Accuracy = 1 − N −1 i=1 ∆Eir (N − 1)∆Emax<label>(2)</label></formula></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>RI =</head><p>T P + T N T P + T N + F P + F N</p><p>(3) JI = T P T P + F P + F N (4)</p><formula xml:id="formula_1">F 1 = 2P R P + R (5)</formula><p>Precision measures the number of galleries (M ) over the total number of galleries (N − 1, excluding the reference), that have been correctly synchronized, namely those galleries, for which the alignment error with respect to the reference gallery, is below a threshold. With the accuracy we instead evaluate the capabilities of the teams in minimizing the average time lapse calculated over the M synchronized galleries, normalized with respect to the maximum accepted time lapse.</p><p>The synchronization task provides a basis for the clustering task. Once the galleries are synchronized, it is possible to cluster the whole event collection to detect sub-events occurring within the entire event, for instance, the single competitions, or the ceremonies of the different disciplines. Sub-events are defined in a neutral and unbiased way (e.g., making reference to the calendar/schedule of the event) and coded into the ground truth. We measure the performance of the sub-event clustering over the whole synchronized collection of media. In this case, we use the three performance indicators reported above, namely RI, JI, and F1. In the formulation we define a true positives (TP), in case two images related to the same sub-event are associated the same cluster, and the true negative (TN), when two images associated to different sub-events are assigned to two different clusters). False positives (FP) occur instead when two images are assigned to the same cluster although belonging to different sub-events.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>Figure 1 :</head><label>1</label><figDesc>Figure 1: Sample images taken from the two datasets.</figDesc></figure>
		</body>
		<back>

			<div type="acknowledgement">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.">ACKNOWLEDGMENTS</head><p>This work was supported in part by the EC under contract FP7-600826 ForgetIT. We would like to thank Anastasia Ioannidou, Alain Malacarne, and Alessio Xompero for their precious help in collecting the images for the dataset.</p></div>
			</div>

			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<monogr>
		<title/>
		<author>
			<persName><surname>References</surname></persName>
		</author>
		<imprint/>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<analytic>
		<title level="a" type="main">Content-based synchronization for multiple photos galleries</title>
		<author>
			<persName><forename type="first">M</forename><surname>Broilo</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Boato</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><forename type="middle">De</forename><surname>Natale</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings -International Conference on Image Processing, ICIP</title>
				<meeting>-International Conference on Image Processing, ICIP</meeting>
		<imprint>
			<date type="published" when="2012">2012</date>
			<biblScope unit="page" from="1945" to="1948" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<analytic>
		<title level="a" type="main">Jointly aligning and segmenting multiple web photo streams for the inference of collective photo storylines</title>
		<author>
			<persName><forename type="first">G</forename><surname>Kim</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><forename type="middle">P</forename><surname>Xing</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition, CVPR &apos;13</title>
				<meeting>the 2013 IEEE Conference on Computer Vision and Pattern Recognition, CVPR &apos;13</meeting>
		<imprint>
			<date type="published" when="2013">2013</date>
			<biblScope unit="page" from="620" to="627" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">Photo stream alignment and summarization for collaborative photo collection and sharing</title>
		<author>
			<persName><forename type="first">J</forename><surname>Yang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Luo</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Yu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Huang</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">IEEE Transactions on</title>
		<imprint>
			<biblScope unit="volume">14</biblScope>
			<biblScope unit="issue">6</biblScope>
			<biblScope unit="page" from="1642" to="1651" />
			<date type="published" when="2012-12">Dec 2012</date>
		</imprint>
	</monogr>
	<note>Multimedia</note>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
