<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Synchronization of Multi-User Event Media at MediaEval 2015: Task Description, Datasets, and Evaluation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vasileios Mezaris CERTH - ITI Thermi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Greece bmezaris@iti.gr</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Francesco De Natale DISI - University of Trento Trento</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Mike Matton VRT</institution>
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Nicola Conci DISI - University of Trento Trento</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>14</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>The objective of this paper is to provide an overview of the Synchronization of Multi-User Event Media (SEM) Task, which is part of the MediaEval Benchmark for Multimedia Evaluation. The SEM task was initially presented at MediaEval in 2014, with the goal of proposing a challenge in aligning multiple users' photo galleries related to the same event but with unreliable timestamps. Besides aligning the pictures on a common timeline, participants were also required to detect the sub-events and cluster the pictures accordingly. For 2015 we have decided to extend the task also to other types of media, thus including audio and video information for a more complete and diversi ed representation of the analyzed event.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The ever increasing number of devices for the collection of
personal data (smartphones, portable cameras, audio recorders)
has lead to the generation of huge amount of data, which
can be either stored for personal records or shared among
friends, relatives, or social networks. In all cases, being able
to arrange such a vast amount of media is of critical
importance both for indexing, categorization, and retrieval. This
makes it possible for any user who attended, or is simply
interested in the event, to recreate the event according to
his personal experience, namely through summaries, stories,
personalized albums [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ][
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        However, it turns out that such a large amount of data
is often unstructured and heterogeneous. The strong
variability (and sometimes similarity) in terms of content and
archiving strategies makes it di cult to manually organize
all the event-related material in a simple yet e ective
manner. In this respect, it would be desirable to nd a
consistent way of presenting the media galleries captured during
an event [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This task is not trivial, since timing and
location information attached to the captured media (mostly
timestamps and GPS) could be inaccurate or missing [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>This lack of information is even more accentuated in case
people use devices that do not have a direct connection to
the Internet, thus requiring manual setting of the clock,
especially after a battery discharge or replacement. In fact,
in the case the temporal information is not represented
correctly, there is a concrete risk of a misleading interpretation
of the media collection, with high probability of losing part
of the semantics of the event, due to a bad alignment along
the temporal axis. Under such conditions, videos could be
of great help, since they contain both audio and visual
information that could be extremely relevant in providing
additional details about the ongoing event compared to the sole
presence of audio and still pictures.</p>
      <p>
        The SEM task presented in 2014 was dealing only with
still pictures and the results provided by the di erent teams
are de nitely encouraging. Participating teams competed
tackling the problem in di erent ways. The authors in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
proposed an approach based on the extraction of visual
features (SIFT) to nd the image pairs across the galleries that
exhibit strong similarity. Then a non-homogeneous linear
equation system is constructed to constrain the time o
sets between the galleries based on these matching pairs to
determine an approximate solution. Sansone et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
relied their implementation on the use of a Markov Random
Field to nd the best correspondences between the images
belonging to two di erent photo galleries. Zaharieva et al.
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] proposed two multimodal approaches that employ both
visual and time information for the synchronization of
different images galleries. The rst approach relies on the
pairwise comparison of images in order to link di erent galleries,
while in the second approach Xmeans clustering is applied,
and the time o sets are estimated by calculating the
average time di erences within the clusters. Apostolidis et al.
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] also proposed a method relying on the combination of
different visual features, and using the images exhibiting the
strongest similarity to compute the galleries o sets.
2.
      </p>
    </sec>
    <sec id="sec-2">
      <title>TASK DESCRIPTION</title>
      <p>In our scenario we imagine a number of users attending
the same event and taking photos and videos with di erent
non-synchronized devices (smartphones, compact cameras,
DSLRs, tablets). Each user contributes to the task with
one gallery, which includes an arbitrary number of photos,
audio les and videos. Assuming that users would like to
merge their photo galleries in a single event-related
collection, the best temporal alignment among the galleries should
be found, so as to correctly report and preserve the
temporal evolution of the event. Furthermore, considering the high
variability in terms of acquisition devices, we cannot expect
the clocks of each device to be synchronized, neither in terms
of precision, nor in terms of the time zone set by the users.
In addition, in some cases, also the location data could be
unavailable (not all devices have a GPS onboard), further
reducing the chances of a correct event reconstruction. In
fact, these factors may considerably hinder the quality of
the alignment, thus di erent solutions should be envisaged,
encompassing the joint analysis of temporal data, position
information, and audio-visual similarity.</p>
      <p>The SEM task expects teams to provide the estimated
time o set between di erent galleries of pictures collected by
di erent users and cameras. The goal can be summarised as
follows: given a set of media collections (galleries) taken by
di erent users/devices at the same event, nd the best
(relative) time alignment among them at gallery level, and detect
the signi cant sub-events over the whole event collection.</p>
    </sec>
    <sec id="sec-3">
      <title>DATASETS</title>
      <p>For this challenge we make available four di erent datasets,
exhibiting di erent challenges. The rst dataset is related
to the Tour de France 2014. It consists of images taken
during the event and collected from Flickr. The dataset is
split into 33 galleries. The dataset covers the entire
competition. Some images are also provided with GPS information
together with the timestamp. A second dataset concerns
the famous exhibition held every year in California, namely
NAMM 2015. The data-set consists of 420 images and 32
videos, split into 19 galleries. Each user gallery contains
a variable number of media (ranging from 12 to 49). All
images are downloaded from Flickr, while videos are
downloaded from YouTube. The Spring Party Salesiani 2015 is
a dataset collected by the organizers, and recorded during
a students' party held in Trento, Italy. It is composed of
videos and pictures captured by the attendees during the
event. Also in this case a gallery corresponds to the user's
device, and media are complemented with the corresponding
time-stamps. The last dataset, Salford Test Shoot includes
403 audio and 58 video les. Time-codes are available for
most of the media. All datasets are provided with the
corresponding ground truth, extracted by considering the
acquisition time of the media and manually veri ed to check
the consistency with respect to the captured event. The
datasets related to the Tour de France 2014, NAMM 2015,
and Spring Party Salesiani 2015 include material subject
to Creative Commons license and are freely available for
download 1. The Salford dataset is instead accessible via
the ICoSOLE project website2.</p>
    </sec>
    <sec id="sec-4">
      <title>METRICS AND EVALUATION</title>
      <p>Each submission will be evaluated in terms of: i) time
synchronization error, and ii) sub-event detection error.</p>
      <p>Concerning the rst one, the goal of the participants is to
maximize the number of galleries for which the
synchronization error is below a prede ned threshold Emax, and to
minimize the time shift of those galleries. The
synchronization error for a gallery Gi with respect to the reference Gr
is de ned as Eir = Tir Tir , where Tir and Tir
1mmlab.disi.unitn.it/MediaEvalSEM2015
2https://icosole.lab.vrt.be/viewer/home
are the delay between Gi and Gr calculated on the
participants' submission and ground truth, respectively. The
threshold Emax depends on the duration of the sub-events
in the dataset, and represents the maximum accepted time
lapse within which we consider a gallery as reasonably
wellsynchronized. We use the above quantities in order to
estimate the synchronization precision (Eq. (1)) and accuracy
(Eq. (2)):
(1)
(2)
P recision =</p>
      <p>M
N
1
=</p>
      <p>Card ( Eir &lt;</p>
      <p>N 1</p>
      <p>Emax)
Accuracy = 1
(N</p>
      <p>Precision measures the number of galleries (M ) over the
total number of galleries (N 1, excluding the reference),
that have been correctly synchronized. With the accuracy
we instead evaluate the capabilities of the teams in
minimizing the average time lapse calculated over the M
synchronized galleries, normalized with respect to the maximum
accepted time lapse.</p>
      <p>The synchronization task provides a basis for the
clustering task. Once the galleries are synchronized, it is possible
to cluster the whole event collection to detect sub-events
occurring within the entire event. Sub-events are de ned in
a neutral and unbiased way (e.g., making reference to the
calendar/schedule of the event) and coded into the ground
truth. We measure the performance of the sub-event
clustering over the whole synchronized collection of media. For
this, we use the Jaccard index J I and the clustering F1 score
(Eq. (3)), where for computing the latter we use P and R,
which represent the Precision and Recall, respectively.</p>
      <p>J I =</p>
      <p>T P
T P + F P + F N
;</p>
      <p>F 1 =
2P R
P + R
(3)</p>
      <p>In the formulation above we declare a true positive (TP)
when two images related to the same sub-event are put in
the same cluster, and a true negative (TN) when two images
belonging to di erent sub-events are assigned to two di
erent clusters). False positives (FP) occur instead when two
images are assigned to the same cluster although belonging
to di erent sub-events.
5.</p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSIONS</title>
      <p>In this paper we have presented the Synchronization of
Multi-User Event Media task held at MediaEval 2015. The
competing teams will be evaluated considering four datasets
collected by the organizers, and made available online
together with the corresponding ground truth. For the
evaluation both the synchronization and the clustering
performances will be evaluating, by measuring the galleries o set
and computing the F1 score, respectively.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was supported in part by the EC under contract
FP7-600826 ForgetIT. We would like to thank Alessio
Xompero and Kostantinos Apostolidis for their precious help in
collecting and annotating the images for the datasets used
in the task.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Apostolidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Papagiannopoulou</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Mezaris</surname>
          </string-name>
          . CERTH at
          <article-title>MediaEval 2014 Synchronization of Multi-User Event Media Task</article-title>
          .
          <source>In Proc. MediaEval 2014 Workshop</source>
          , CEUR vol.
          <volume>1263</volume>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Broilo</surname>
          </string-name>
          , G. Boato, and F. De Natale.
          <article-title>Content-based Synchronization for Multiple Photos Galleries</article-title>
          .
          <source>In Proc. IEEE Int. Conf. on Image Processing (ICIP)</source>
          , pages
          <year>1945</year>
          {
          <year>1948</year>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Conci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. D.</given-names>
            <surname>Natale</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Mezaris</surname>
          </string-name>
          .
          <article-title>Synchronization of multi-user event media (SEM) at MediaEval 2014: Task description, datasets, and evaluation</article-title>
          .
          <source>In Proc. MediaEval 2014 Workshop</source>
          , CEUR vol.
          <volume>1263</volume>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kim</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. P.</given-names>
            <surname>Xing</surname>
          </string-name>
          .
          <article-title>Jointly Aligning and Segmenting Multiple Web Photo Streams for the Inference of Collective Photo Storylines</article-title>
          .
          <source>In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)</source>
          , pages
          <fpage>620</fpage>
          {
          <fpage>627</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Nowak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Thaler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Stiegler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Bailer</surname>
          </string-name>
          .
          <article-title>JRS at Event Synchronization Task</article-title>
          .
          <source>In Proc. MediaEval 2014 Workshop</source>
          , CEUR vol.
          <volume>1263</volume>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Sansone</surname>
          </string-name>
          , G. Boato, and
          <string-name>
            <given-names>M.-S.</given-names>
            <surname>Dao</surname>
          </string-name>
          .
          <article-title>Synchronizing Multi-User Photo Galleries with MRF</article-title>
          .
          <source>In Proc. MediaEval 2014 Workshop</source>
          , CEUR vol.
          <volume>1263</volume>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Huang</surname>
          </string-name>
          .
          <article-title>Photo Stream Alignment and Summarization for Collaborative Photo Collection and Sharing</article-title>
          . Multimedia, IEEE Transactions on,
          <volume>14</volume>
          (
          <issue>6</issue>
          ):
          <volume>1642</volume>
          {
          <fpage>1651</fpage>
          ,
          <string-name>
            <surname>Dec</surname>
          </string-name>
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zaharieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Riegler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. Del</given-names>
            <surname>Fabro</surname>
          </string-name>
          .
          <article-title>Multimodal Synchronization of Image Galleries</article-title>
          .
          <source>In Proc. MediaEval 2014 Workshop</source>
          , CEUR vol.
          <volume>1263</volume>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>