<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Search and Hyperlinking Task at MediaEval 2013</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maria Eskevich</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robin Aly</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roeland Ordelman</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shu Chen</string-name>
          <email>shu.chen4@mail.dcu.ie</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gareth J.F. Jones</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CNGL Centre for Global Intelligent Content, Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>INSIGHT Centre for Data Analytics, Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Twente</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>18</fpage>
      <lpage>19</lpage>
      <abstract>
        <p>The Search and Hyperlinking Task formed part of the MediaEval 2013 evaluation campaign. The Task consisted of two sub-tasks: (1) answering known-item queries from a collection of roughly 1200 hours of broadcast TV material, and (2) linking anchors within the known-item to other parts of the video collection. We provide an overview of the task and the data sets used.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The increasing amount of digital multimedia content
available is inspiring new scenarios of user interaction. The
Search and Hyperlinking Task at MediaEval 2013 envisioned
the following scenario: a user is searching for a segment of
video that they know to be contained in a video collection
(henceforth the target \known-item"). If the user nds the
segment, he may wish to nd additional information about
some aspect of this segment. Computer systems should
support users in this use scenario by providing links to satisfy
the user's information needs. This use scenario is a re
nement of a similar task at MediaEval 2012, see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for an
overview of employed techniques. This paper describes the
experimental data set provided to task participants for
MediaEval 2013 and details of the two subtasks and their
evaluation.
      </p>
    </sec>
    <sec id="sec-2">
      <title>EXPERIMENTAL DATASET</title>
      <p>The dataset for both subtasks was a collection of 1,260
hours of video provided by the BBC. The average length
of a video was roughly 30 minutes and most videos were
in the English language. The collection was used both for
training and testing of systems. The BBC kindly provided
human generated textual metadata and manual transcripts
for each video. Participants were also provided with the
output of two automatic speech recognition (ASR) systems
and visual analysis. We describe these information sources
in the following subsections.</p>
    </sec>
    <sec id="sec-3">
      <title>Speech recognition transcripts</title>
      <p>
        The audio was extracted from the video stream using the
mpeg software toolbox (sample rate = 16,000Hz, number
of channels = 1). Based on this data, two sets of ASR
transcripts were created:
(i) All audio les were transcribed by LIMSI-CNRS/Vocapia1
using the VoxSigma vrbs trans system (version eng-usa 4.0)
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The models used by the system have been updated with
partial support from the Quaero program [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        (ii) The LIUM system2 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is based on the CMU Sphinx
project, and was developed to participate in the evaluation
campaign of the International Workshop on Spoken
Language Translation 2011. LIUM generated an English
transcript for each audio le successfully processed. These
results consist of: (i) one-best hypotheses in NIST CTM
format, (ii) word lattices in SLF (HTK) format, following a
4-gram topology, and (iii) confusion networks, in an ATT
FSM-like format.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Video cues</title>
      <p>In addition to spoken content, visual descriptions of video
content can potentially help for searching and hyperlinking.
We provided the participants with shot boundaries, one
extracted keyframe per shot, as well as the outputs of concept
detectors (see below) and face detectors (see below) for these
keyframes.</p>
      <p>
        For each video, shot boundaries were determined and a
single key frame per shot was extracted by a system kindly
provided by Technicolor [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The extracted frame was the
most stable I-frame within its shot. In total, the system
extracted approximately 1,200,000 shots/keyframes.
Concept detection scores for a list of concepts were provided.
These concepts were selected by extracting keywords from
metadata and spoken content. We used the on-the- y video
detector Visor, which was kindly provided by the Computer
Vision Group of University of Oxford [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. To make the
condence scores comparable over multiple detectors, we used
them as variables in a logistic regression framework, which
ensures the scores lie in the range [0 : 1]. We set the logistic
regression parameters to the expected value of the
parameters from over 374 detectors on the internet archive collection
used in TRECVid 2011.
      </p>
      <p>
        The appearance of faces in videos can be helpful
information for search and linking. INRIA [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] kindly provided
possible bounding boxes in keyframes with a con dence score
that the bounding box contains a face. Additionally, the
tool also contained for each bounding box, the n most
similar faces (bounding boxes) in the dataset.
1http://www.vocapia.com/
2http://www-lium.univ-lemans.fr/en/content/languageand-speech-technology-lst
      </p>
    </sec>
    <sec id="sec-5">
      <title>USER STUDY</title>
      <p>
        For the de nition of realistic queries and anchors, we
conducted a study with 30 users between the ages of 18 and
30. By browsing the collection, the users selected items,
a segment of a video with a start and an end time, that
were interesting to them. The users were then instructed
to consider these items as a known-item which they have
to re nd. We asked the users to formulate text and visual
queries that they would use in a search engine to carry our
their re nding. The study resulted in 50 known-items and
corresponding multimodal queries. Subsequently, we asked
the users to mark so-called anchors, or segments, related to
other items from within the collection within the
knownitem for which they would like to see links. A second session
of the study was conducted after the Task participants
submitted their results. A set of users partially overlapping
with the rst group (17 participants) were presented with
the selected anchors and with the hyperlinks proposed by
the participants. The users had to assess the suitability of
the proposed hyperlinks. Returning users assessed the
anchors that they de ned themselves. The reader can nd a
more elaborate description of this user study in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
    <sec id="sec-6">
      <title>4. SEARCH SUBTASK</title>
      <p>We are interested in cross-comparison of one method being
applied on all three types of transcripts. Thus we required
the participants to submit up to 5 di erent approaches or
their combinations, each being tested on all three transcripts.</p>
      <p>
        We used the following three metrics in order to
evaluate the submissions of the workshop participants: mean
reciprocal rank (MRR), mean generalized average precision
(mGAP) and mean average segment precision (MASP). MRR
assesses the ranking of the relevant units. mGAP [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] rewards
techniques that not only nd the relevant items earlier in the
ranked output list, but also are closer to the ideal point to
begin playback (the \jump-in" point) of the relevant
content. MASP [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] takes into account the ranking of the results
and the length of both relevant and irrelevant segments that
need to be listened to before reaching the relevant content.
      </p>
    </sec>
    <sec id="sec-7">
      <title>LINKING SUBTASK</title>
      <p>For the Hyperlinking subtask, the workshop participants
were provided with the so-called anchors created by the users
in the user study at the BBC and had to generate link
targets. To be more concrete, the participants had to return
a list of potential video segment link targets ranked by the
likelihood of being relevant to the anchor or to the anchor in
the context of corresponding known-item segment (though
always independently of the initial known-item query).</p>
      <p>To evaluate the linking subtask we used crowdsourcing via
Amazon's Mechanical Turk platform3, whereas the second
stage of user study at BBC allowed us to assess the reliability
of the crowdsourcing results.</p>
      <p>Due to time and resource constraints, we chose a random
subset of 30 anchors out of initial 98 for the formal task
assessment. For these anchors and potential links, we used a
pooling method to group the videos from the top 10 ranks
of no more than 5 submitted runs of each of the
participants. Submission were selected to maximize the diversity
of the linking methods used in the pools to be assessed. This
resulted in 9195 anchor-target pairs, that represented 7637
di erent pairs for crowdsoucing assessment. Users at BBC
studies evaluated only 1 run per each participant which
resulted in 2081 pairs, with 2078 being diverse. The manual
assessment of these links resulted in the ground truth used
to calculate precision at xed rank cuto s and MAP for all
the participants runs. Both mturk and BBC ground truths
were released to the participants for further performance
analysis.</p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGEMENTS</title>
      <p>This work was supported by Science Foundation Ireland
(Grant 08/RFP/CMS1677) Research Frontiers Programme
2008 and (Grant 07/CE/I1142) as part of the Centre for
Next Generation Localisation (CNGL) project at DCU, and
by the funding from the European Commission's 7th
Framework Programme (FP7) under AXES ICT-269980. The user
studies were executed in collaboration with Jana Eggink and
Andy O'Dwyer from BBC Research, to whom the authors
are greatful.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Linking inside a video collection: what and how to measure? In WWW (Companion Volume)</article-title>
          , pages
          <fpage>457</fpage>
          {
          <fpage>460</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Chat eld</surname>
          </string-name>
          , V.
          <string-name>
            <surname>Lempitsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Vedaldi</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <article-title>The devil is in the details: an evaluation of recent feature encoding method</article-title>
          .
          <source>In British Machine Vision Conference (BMVC</source>
          <year>2011</year>
          ), Dundee, United Kingdom,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R. G.</given-names>
            <surname>Cinbis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Verbeek</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <article-title>Unsupervised Metric Learning for Face Identi cation in TV Video</article-title>
          .
          <source>In Proceedings of ICCV</source>
          <year>2011</year>
          , Barcelona, Spain,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nadeem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Guinaudeau</surname>
          </string-name>
          , G. Gravier,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sebillot</surname>
          </string-name>
          , T. de Nies, P. Debevere, R. Van de Walle, P. Galuscakova,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pecina</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Larson</surname>
          </string-name>
          .
          <article-title>Multimedia information seeking through search and hyperlinking</article-title>
          .
          <source>In Proceedings of the 3rd ACM conference on International conference on multimedia retrieval</source>
          ,
          <source>ICMR'13</source>
          , pages
          <fpage>287</fpage>
          {
          <fpage>294</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Magdy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          .
          <article-title>New metrics for meaningful evaluation of informally structured speech retrieval</article-title>
          .
          <source>In Proceedings of ECIR 2012</source>
          , pages
          <fpage>170</fpage>
          {
          <fpage>181</fpage>
          ,
          <string-name>
            <surname>Barcelona</surname>
          </string-name>
          , Spain,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.-L.</given-names>
            <surname>Gauvain</surname>
          </string-name>
          .
          <source>The Quaero Program: Multilingual and Multimedia Technologies. IWSLT</source>
          <year>2010</year>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.-L.</given-names>
            <surname>Gauvain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lamel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Adda. The LIMSI Broadcast</surname>
          </string-name>
          <article-title>News transcription system</article-title>
          .
          <source>Speech Communication</source>
          ,
          <volume>37</volume>
          (
          <issue>1-2</issue>
          ):
          <volume>89</volume>
          {
          <fpage>108</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Massoudi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lefebvre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Demarty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Oisel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Chupeau</surname>
          </string-name>
          .
          <article-title>A video ngerprint based on visual digest and local ngerprints</article-title>
          .
          <source>In International Conference on Image Processing (ICIP</source>
          <year>2006</year>
          ), pages
          <fpage>2297</fpage>
          {
          <fpage>2300</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Pecina</surname>
          </string-name>
          , P. Ho mannova,
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Oard</surname>
          </string-name>
          .
          <article-title>Overview of the CLEF 2007 cross-language speech retrieval track</article-title>
          .
          <source>In Proceedings of CLEF 2007</source>
          , pages
          <fpage>674</fpage>
          {
          <fpage>686</fpage>
          . Springer,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rousseau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bougares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Deleglise</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schwenk</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Estev</surname>
          </string-name>
          .
          <article-title>LIUM's systems for the IWSLT 2011 Speech Translation Tasks</article-title>
          .
          <source>In Proceedings of IWSLT</source>
          <year>2011</year>
          , San Francisco, USA,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>