<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Search and Hyperlinking Task at MediaEval 2014</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maria Eskevich</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robin Aly</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David N. Racca</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roeland Ordelman</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shu Chen</string-name>
          <email>shu.chen4@mail.dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gareth J.F. Jones</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CNGL Centre for Global Intelligent Content, School of Computing, Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>INSIGHT Centre for Data Analytics, Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Twente</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>17</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>The Search and Hyperlinking Task at MediaEval 2014 is the third edition of this task. As in previous versions, it consisted of two sub-tasks: (i) answering search queries from a collection of roughly 2700 hours of BBC broadcast TV material, and (ii) linking anchor segments from within the videos to other target segments within the video collection. For MediaEval 2014, both sub-tasks were based on an ad-hoc retrieval scenario, and were evaluated using a pooling procedure across participants submissions with crowdsourcing relevance assessment using Amazon Mechanical Turk.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>The full value of the rapidly growing archives of newly
produced digital multimedia content and digitalisation of
previously created analog audio and video material will only
be realized with the development of technologies that allow
users to explore them through search and retrieval of
potentially interesting content.</p>
      <p>The Search and Hyperlinking Task at MediaEval 2014
envisioned the following scenario: a user is searching for
relevant segments within a video collection that address a
certain topic of interest expressed in a query. If the user nds
a segment which is relevant to their initial information need
expressed through the query, they may wish to nd
additional information about some aspect of this segment.</p>
      <p>
        The task framework asks participants to create systems
that support the search and linking aspects of the task. The
use scenario is the same as in the Search and Hyperlinking
task 2013 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] with the main di erence being that the search
sub-task has changed from known-item to ad-hoc. This
paper describes the experimental data set provided to task
participants for MediaEval 2014, details of the two sub-tasks,
and their evaluation.
      </p>
    </sec>
    <sec id="sec-2">
      <title>EXPERIMENTAL DATASET</title>
      <p>The dataset for both subtasks was a collection of 4021
hours of videos provided by the BBC, which we split into
a development set of 1335 hours, which coincided with the
test collection used in the 2013 edition of this task, and a
test set of 2686 hours. The average length of a video was
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Audio Analysis</title>
      <p>The audio was extracted from the video stream using the
mpeg software toolbox (sample rate = 16,000Hz, no. of
channels = 1). Based on this data, the transcripts were
created using the following ASR approaches and provided
to participants:</p>
      <p>
        (i) LIMSI-CNRS/Vocapia1, which uses the VoxSigma vrbs trans
system (version eng-usa 4.0) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Compared to the
transcripts created for the 2013 edition of this task, the
system's models had been updated with partial support from
the Quaero program [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        (ii) The LIUM system2 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], is based on the CMU Sphinx
project. The LIUM system provided three output formats:
(1) one-best transcripts in NIST CTM format, (2) word
lattices in SLF (HTK) format, following a 4-gram topology, and
(3) confusion networks in a format similar to ATT FSM.
      </p>
      <p>
        (iii) The NST/She eld system3 is trained on multi-genre
sets of BBC data that does not overlap with the collection
used for the task, and uses deep neural networks [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The
ASR transcript contains speaker diarization, similar to the
LIMSI-CNRS/Vocapia transcipts.
      </p>
      <p>
        Additionally, prosodic features were extracted using the
OpenSMILE tool version 2.0 rc1 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]4. The following list
of prosodic features were calculated over sliding windows of
10 milliseconds: root mean squared (RMS) energy, loudness,
probability of voicing, fundamental frequency (F0),
harmonics to noise ratio (HNR), voice quality, and pitch direction
(classes falling, at, raising, and direction score). Prosodic
information was provided for the rst time in 2014 to
encourage participants to explore its potential value for the
Search and Hyperlinking sub-tasks.
1http://www.vocapia.com/
2http://www-lium.univ-lemans.fr/en/content/languageand-speech-technology-lst
3http://www.natural-speech-technology.org
4http://opensmile.sourceforge.net/
      </p>
      <p>
        The computer vision groups at University of Leuven (KUL)
and University of Oxford (OXU) provided the output of
concept detectors for 1537 concepts from ImageNet5 using
different training approaches. The approach by KUL uses
examples from ImageNet as positive examples [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], while OXU
uses an on-the- y concept detection approach, which
downloads training examples through Google image search [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>USER STUDY</title>
      <p>
        In order to create realistic queries and anchors for our
test set, we conducted a study with 28 users between aged
between 18 and 30 from the general public around London,
U.K. The study was similar to our previous study carried out
for MediaEval 2013 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], with the main di erence being the
focus on information needs with multiple relevant segments.
The study focused on a home user scenario, and for this, to
re ect the current wide usage of computer tablets,
participants used a version of the AXES video search system [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] on
iPads to search and browse within the video collection. The
user study consisted of the following steps: i) a participant
de ned an information need using natural language, ii) they
searched the test set with a shorter query, one they might
use to search of the Youtube video repository, 3) after
selecting several possible relevant segments, they de ned anchor
points or regions within each segment and stated what kind
of links they would expect for this anchor.
      </p>
      <p>Users were then instructed to de ne queries that they
expected to have more than one relevant video segment in the
collection. These queries consisted of several terms, and
were used as input to a standard online search engine, e.g.
\sightseeing london". The study resulted in 36 ad-hoc search
queries for the test set. The development set for the task
consisted of 50 known-item queries from the MediaEval 2013
Search and Hyperlinking task.</p>
      <p>
        Subsequently, as in the 2013 studies, we asked the
participants to mark so-called anchors, or segments they would like
to see links to, within some the segments that are relevant
to the issued search queries. The reader can nd a more
elaborate description of this user study design in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>REQUIRED RUNS SUBMISSIONS AND</title>
    </sec>
    <sec id="sec-6">
      <title>EVALUATION PROCEDURE FOR THE</title>
    </sec>
    <sec id="sec-7">
      <title>SEARCH AND LINKING SUB-TASKS</title>
      <p>For the 2014 task, as well as ad hoc search, we were
interested in cross-comparison of methods being applied across all
four provided transcripts: one manual and 3 ASR. Thus, we
allowed participants to submit up to 5 di erent approaches
or their combinations, each being tested on all four
transcripts, for both sub-tasks. In case any of the groups based
their methods on video features only, they could submit this
type of run in addition as well.</p>
      <p>
        To evaluate the submissions of the search and linking
subtasks a pooling method was used to select submitted
segments and link targets for relevance assessment. The top-N
ranks of all submitted runs were evaluated using
crowdsourcing technologies. We report precision oriented metrics, such
as precision at various cuto s and mean average precision
(MAP), using di erent approaches to take into account
segment overlap, as described in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGEMENTS</title>
      <p>This work was supported by Science Foundation Ireland
(Grants 08/RFP/CMS1677 and 12/CE/I2267) and Research
Frontiers Programme 2008 (Grant 07/CE/I1142) as part
of the Centre for Next Generation Localisation (CNGL)
project at DCU, and by the funding from the European
Commission's 7th Framework Programme (FP7) under AXES
ICT-269980. The user studies were executed in
collaboration with Jana Eggink and Andy O'Dwyer from BBC
Research, to whom the authors are grateful.
6.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          .
          <article-title>Adapting binary information retrieval evaluation metrics for segment-based retrieval tasks</article-title>
          .
          <source>Technical Report 1312</source>
          .
          <year>1913</year>
          , ArXiv e-prints,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Linking inside a video collection: what and how to measure? In WWW (Companion Volume)</article-title>
          , pages
          <fpage>457</fpage>
          {
          <fpage>460</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>K.</surname>
          </string-name>
          <article-title>Chat eld and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          . Visor:
          <article-title>Towards on-the- y large-scale object category retrieval</article-title>
          .
          <source>In Computer Vision{ACCV</source>
          <year>2012</year>
          , pages
          <fpage>432</fpage>
          {
          <fpage>446</fpage>
          . Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          .
          <article-title>The Search and Hyperlinking Task at MediaEval 2013</article-title>
          .
          <source>In Proceedings of the MediaEval 2013 Workshop</source>
          , Barcelona, Spain,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Eyben</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Weninger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gross</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <article-title>Recent developments in opensmile, the munich open-source multimedia feature extractor</article-title>
          .
          <source>In Proceedings of ACM Multimedia</source>
          <year>2013</year>
          , pages
          <fpage>835</fpage>
          {
          <fpage>838</fpage>
          ,
          <string-name>
            <surname>Barcelona</surname>
          </string-name>
          , Spain.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.-L.</given-names>
            <surname>Gauvain</surname>
          </string-name>
          .
          <article-title>The Quaero Program: Multilingual and Multimedia Technologies</article-title>
          .
          <source>In Proceedings of IWSLT 2010</source>
          , Paris, France,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.-L.</given-names>
            <surname>Gauvain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lamel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Adda. The LIMSI Broadcast</surname>
          </string-name>
          <article-title>News transcription system</article-title>
          .
          <source>Speech Communication</source>
          ,
          <volume>37</volume>
          (
          <issue>1-2</issue>
          ):
          <volume>89</volume>
          {
          <fpage>108</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Lanchantin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J. F.</given-names>
            <surname>Gales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Long</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Quinnell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Renals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Saz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Seigel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Swietojanski</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. C.</given-names>
            <surname>Woodland</surname>
          </string-name>
          .
          <article-title>Automatic transcription of multi-genre media archives</article-title>
          .
          <source>In Proceedings of the First Workshop on Speech, Language and Audio in Multimedia (SLAM@INTERSPEECH)</source>
          , volume
          <volume>1012</volume>
          <source>of CEUR Workshop Proceedings</source>
          , pages
          <volume>26</volume>
          {
          <fpage>31</fpage>
          . CEUR-WS.org,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>McGuinness</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chat eld</surname>
          </string-name>
          , O. Parkhi,
          <string-name>
            <given-names>R.</given-names>
            <surname>Arandjelovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Douze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kemman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kleppe</surname>
          </string-name>
          , P. van der Kreeft,
          <string-name>
            <given-names>K.</given-names>
            <surname>Macquarrie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ozerov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. E. O</given-names>
            <surname>'Connor</surname>
          </string-name>
          , F. De Jong, A. Zisserman,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Perez</surname>
          </string-name>
          .
          <article-title>The AXES research video search system</article-title>
          .
          <source>In Proceedings of the IEEE ICASSP</source>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rousseau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Deleglise</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Esteve</surname>
          </string-name>
          .
          <article-title>Enhancing the ted-lium corpus with selected data for language modeling and more ted talks</article-title>
          .
          <source>In The 9th edition of the Language Resources and Evaluation Conference (LREC</source>
          <year>2014</year>
          ), Reykjavik, Iceland, May
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Tommasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tuytelaars</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Caputo</surname>
          </string-name>
          .
          <article-title>A testbed for cross-dataset analysis</article-title>
          .
          <source>CoRR, abs/1402.5923</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>