<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Synchronization of Multi-User Event Media (SEM) at MediaEval 2014: Task Description, Datasets, and Evaluation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicola Conci</string-name>
          <email>nicola.conci@unitn.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco De Natale</string-name>
          <email>francesco.denatale@unitn.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vasileios Mezaris</string-name>
          <email>bmezaris@iti.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CERTH - ITI</institution>
          ,
          <addr-line>Thermi</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>DISI - University of Trento</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>16</fpage>
      <lpage>17</lpage>
      <abstract>
        <p>In this paper we provide an overview of the Synchronization of Multi-User Event Media (SEM) Task that is part of the 2014 MediaEval Benchmark for Multimedia Evaluation. The SEM task is presented this year for the rst time in MediaEval, and poses a new challenge, namely the temporal alignment of a series of photo galleries that relate to the same event but have been collected by di erent users. Besides aligning the pictures on a common timeline, participants are also required to detect the sub-events attended by the users and to group the pictures accordingly. The task is validated on two di erent datasets related to the 2010 and 2012 Winter / Summer Olympic games, each dataset comprising a variable number of pictures, galleries, and subevents.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Content creation is more and more a collective
experience. People attending large social events (a soccer match,
a concert), but also personal-scale ones (a wedding, a
birthday party) collect dozens of photos and video clips with
their smartphones, tablets, cameras, and more recently
social cameras. Such information is later exchanged in a
number of di erent ways, including shared repositories, clouds,
social networks, etc. In this way, di erent media galleries
are made available to each other, making it possible for any
user who attended, or is simply interested to the event, to
create his own view of it through summaries, stories,
personalized albums [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ][
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. However, such a large amount of data
turns out to be unstructured and heterogeneous and, even if
it would be possible to collect it on the same hard drive, the
variability in terms of content, naming, archiving strategies
makes it impossible to organize all the event-related material
in a simple yet e ective manner.
      </p>
      <p>
        In this respect, a major issue is the need of aligning and
presenting the various media galleries captured during an
event in a consistent way [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. As a matter of fact, the time
and location information attached to the captured media
(timestamp, GPS) can be wrong, inaccurate or even
missing (for instance, due to wrong setting of the clock/calendar,
di erent time-zone, modi cation or removal of tags).
Similarly, this is also a common situation in historical events
and photo archives, where timestamps and especially
location information is rarely available. In some other cases,
images might be processed o ine for post-production, thus
losing the correct temporal information. In such cases,
creating a single timeline could turn out to be complicated and
challenging, with a concrete risk of representing the event in
a misleading way.
2.
      </p>
    </sec>
    <sec id="sec-2">
      <title>TASK DESCRIPTION</title>
      <p>In our scenario we imagine a number of users (10+)
attending the same event and taking photos and videos with
di erent non-synchronized devices (smartphones, handheld
cameras, DSLRs, tablets). Each user contributes to the task
with one gallery, which includes an arbitrary number of
photos, possibly covering just a part of the event, with variable
density of acquisitions (single photos are also possible).
Assuming that these users would like to merge their photo
galleries in a single event-related collection, the best
temporal alignment among the galleries should be found, so as
to correctly report and preserve the temporal evolution of
the event. Furthermore, considering the high variability in
terms of acquisition devices, we cannot expect the clocks of
each device of the same user to be synchronized, neither in
terms of precision, nor in terms of the time zone set by the
users. Furthermore, in some cases, also the location data
could be unavailable (not all devices have a GPS onboard),
further reducing the available information about the
captured event. In view of creating a single timeline, these
factors may considerably hinder the quality of the alignment,
thus di erent solutions should be envisaged, encompassing
the joint analysis of temporal data, position information,
and visual similarity.</p>
      <p>Therefore, the SEM task expects teams to provide the
estimated time o set between di erent galleries of pictures
collected by di erent users and cameras. The goal can be
summarised as follows: given a set of image collections
(galleries) taken by di erent users/devices at the same event,
nd the best (relative) time alignment among them at gallery
level, and detect the signi cant sub-events over the whole
event collection.
3.</p>
    </sec>
    <sec id="sec-3">
      <title>DATASETS</title>
      <p>For this challenge we make available two di erent datasets,
consisting of a collection of images gathered from Flickr
and made available under Creative Commons license. Both
datasets refer to well known and structured sport events,
namely the Olympic Games held in London in 2012 and the
Vancouver Winter Olympic Games of 2010. We have
chosen to work with these two events because on the one hand
they exhibit a clear and organized schedule with precise
timing. On the other hand they still exhibit a high
variability in terms of visual content, due to the common features
across di erent competitions in the same discipline, as well
as strong similarities in the environments, in which the
pictures are collected, making the synchronization a non-trivial
task. As far as this task is concerned, the images within a
gallery are consistent in terms of timestamp, and might
include the GPS information. Therefore the temporal o sets
are at gallery level thus assuming that every user uses one
single device for acquisition.</p>
      <p>The dataset collected from the London Olympics includes
2124 images, divided into 37 galleries. The rst gallery
comprises a subset of the data provided in the development set
and is de ned as the reference gallery. The dataset
collected from the Vancouver Winter Olympic Games includes
1351 pictures representing most of the competitions, divided
into 35 galleries with a variable number of pictures in each
gallery. Also in this case, the rst gallery is set as the
reference. Fig. 1 shows a few samples of the two datasets.</p>
    </sec>
    <sec id="sec-4">
      <title>METRICS AND EVALUATION</title>
      <p>Two objective metrics will be used to evaluate the results:
time synchronization error
sub-event detection error</p>
      <p>As far as the rst metric is concerned, the goal of the
participants is to maximize the number of galleries for which the
synchronization error is below a prede ned threshold, and
to minimize the time shift of those galleries. The
synchronization error for a gallery Gi with respect to the reference
Gr is de ned as Eir = Tir Tir , where Tir is the
delay between Gi and Gr calculated on the ground truth.
The threshold Emax depends on the duration of the
subevents in the dataset, and represents the maximum accepted
time lapse within which we consider a gallery as reasonably
well-synchronized.</p>
      <p>As far the metrics for evaluation are concerned, we have
considered for the temporal alignment the precision (Eq. 1)
and accuracy (Eq. 2). For the quality of the clustering, we
use the Rand Index (RI), as from Eq. 3, the Jaccard index
(JI) Eq. 4 , and the F1 score Eq. 5, where P and R represent
the Precision and Recall, respectively.</p>
      <p>P recision =</p>
      <p>M
N
1
=</p>
      <p>Card ( Eir &lt;</p>
      <p>N 1</p>
      <p>Emax)
Accuracy = 1
(N</p>
      <p>PiN=11 Eir</p>
      <p>1) Emax
RI =</p>
      <p>T P + T N
T P + T N + F P + F N
(1)
(2)
(3)</p>
      <p>JI =</p>
      <p>T P
T P + F P + F N</p>
      <p>Precision measures the number of galleries (M ) over the
total number of galleries (N 1, excluding the reference),
that have been correctly synchronized, namely those
galleries, for which the alignment error with respect to the
reference gallery, is below a threshold. With the accuracy we
instead evaluate the capabilities of the teams in minimizing
the average time lapse calculated over the M synchronized
galleries, normalized with respect to the maximum accepted
time lapse.</p>
      <p>The synchronization task provides a basis for the
clustering task. Once the galleries are synchronized, it is possible
to cluster the whole event collection to detect sub-events
occurring within the entire event, for instance, the single
competitions, or the ceremonies of the di erent disciplines.
Sub-events are de ned in a neutral and unbiased way (e.g.,
making reference to the calendar/schedule of the event) and
coded into the ground truth. We measure the performance
of the sub-event clustering over the whole synchronized
collection of media. In this case, we use the three performance
indicators reported above, namely RI, JI, and F1. In the
formulation we de ne a true positives (TP), in case two
images related to the same sub-event are associated the same
cluster, and the true negative (TN), when two images
associated to di erent sub-events are assigned to two di erent
clusters). False positives (FP) occur instead when two
images are assigned to the same cluster although belonging to
di erent sub-events.</p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was supported in part by the EC under
contract FP7-600826 ForgetIT. We would like to thank
Anastasia Ioannidou, Alain Malacarne, and Alessio Xompero for
their precious help in collecting the images for the dataset.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Broilo</surname>
          </string-name>
          , G. Boato, and F. De Natale.
          <article-title>Content-based synchronization for multiple photos galleries</article-title>
          .
          <source>In Proceedings - International Conference on Image Processing, ICIP</source>
          , pages
          <year>1945</year>
          {
          <year>1948</year>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kim</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. P.</given-names>
            <surname>Xing</surname>
          </string-name>
          .
          <article-title>Jointly aligning and segmenting multiple web photo streams for the inference of collective photo storylines</article-title>
          .
          <source>In Proceedings of the 2013 IEEE Conference on Computer Vision</source>
          and Pattern Recognition,
          <source>CVPR '13</source>
          , pages
          <fpage>620</fpage>
          {
          <fpage>627</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Huang</surname>
          </string-name>
          .
          <article-title>Photo stream alignment and summarization for collaborative photo collection and sharing</article-title>
          . Multimedia, IEEE Transactions on,
          <volume>14</volume>
          (
          <issue>6</issue>
          ):
          <volume>1642</volume>
          {
          <fpage>1651</fpage>
          ,
          <string-name>
            <surname>Dec</surname>
          </string-name>
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>