<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LinkedTV at MediaEval 2013 Search and Hyperlinking Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>M. Sahuguet</string-name>
          <email>sahuguet@eurecom.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>B. Huet</string-name>
          <email>huet@eurecom.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>B. Cˇ ervenková</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E. Apostolidis</string-name>
          <email>apostolid@iti.gr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>V. Mezaris</string-name>
          <email>bmezaris@iti.gr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Stein</string-name>
          <email>daniel.stein@iais.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>S. Eickeler</string-name>
          <email>stefan.eickeler@iais.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J.L. Redondo Garcia</string-name>
          <email>redondo@eurecom.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>R. Troncy</string-name>
          <email>troncy@eurecom.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>L. Pikora</string-name>
          <email>lukas.pikora@vse.cz</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Eurecom</institution>
          ,
          <addr-line>Sophia Antipolis</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fraunhofer IAIS</institution>
          ,
          <addr-line>Sankt Augustin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Information Technologies Institute</institution>
          ,
          <addr-line>Thessaloniki</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Economics</institution>
          ,
          <addr-line>Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>18</fpage>
      <lpage>19</lpage>
      <abstract>
        <p>This paper aims at presenting the results of LinkedTV's rst participation to the Search and Hyperlinking task at MediaEval challenge 2013. We used textual information, transcripts, subtitles and metadata, and we tested their combination with automatically detected visual concepts. Hence, we submitted various runs to compare diverse approaches and see the improvement when adding visual information.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        This paper describes the framework used by the LinkedTV
team to tackle the problem of Search and Hyperlinking
inside a video collection [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The applied techniques originate
from the LinkedTV project1, which aims at integrating TV
and internet experience, by enabling the user to access
additional information and media resources aggregated from
diverse sources, thanks to automatic media annotation.
      </p>
    </sec>
    <sec id="sec-2">
      <title>PRE-PROCESSING STEP</title>
      <p>
        Concept detection was performed on the key-frames of
the video, following the approach in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], while the algorithm
for Optical Character Recognition (OCR) described in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
was used for text localization. Moreover, for each video,
we extracted keywords from the provided subtitles, based
on the algorithm presented in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Finally, we grouped the
prede ned video shots into bigger segments (scenes), based
on the visual similarity and the temporal consistency among
them, using the method introduced in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
3.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>OUR FRAMEWORK</title>
    </sec>
    <sec id="sec-4">
      <title>Lucene indexing</title>
      <p>
        We indexed all available data in a Lucene index at
different granularities: video level, scene level, shot level and
segments created using sliding window algorithm [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Documents were represented by both textual elds (for a text
search) and oating point elds (for the visual concepts).
3.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>From visual cues to detected concepts</title>
      <p>
        Text search is straightforward with Lucene, by using the
default text search based on TF-IDF values. In order to
incorporate visual information to the search, we mapped
keywords extracted from the visual cues query (using Alchemy
API2) to visual concepts using a semantic word distance
based on Wordnet synsets [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. When visual concepts were
detected in the query, we enriched the textual query by range
queries on the values of the corresponding visual concepts.
3.3
      </p>
      <p>We concatenated textual and visual queries to perform
the text query. Two strategies were adopted: we either used
segments indexed in the Lucene engine, or performed queries
creating segments on the y, by merging video segments
based on their score.</p>
      <p>Performing a text query on the video index often returned
the relevant video in the top of the list. Hence, some runs
rst restrict the pool of videos that are going to be searched
to a small number, and then perform additional queries for
smaller segments inside this pool.</p>
      <p>
        We submitted 9 runs in total:
scenes-C : Scene search using textual and visual cues.
scenes-noC : Same as previous using textual cues only (no
visual cues) for comparison purposes.
part-sc-C : Partial scenes search from shot boundary using
textual and visual cues following three steps: ltering of
the list of videos; querying for shots inside each video;
ordering them by score. As a shot is a unit that is too small
to be returned to a viewer, we completed the segment with
the end of the scene that includes this shot.
part-sc-noC : Same as previous using textual cues only.
cl10-C : Temporal clustering of shots within a video using
text and visual cues in the following manner: ltering out
the set of videos to search; computing scores for every
shot in the video; clustering together shots closer than 10
seconds apart (scores were added to form the nal score).
cl10-noC :Same as previous using text search only.
scenes-S or scenes-U or scenes-I : Scene search using only
textual cue from transcript or subtitle, no metadata.
SW-60-I or SW-60-S : Search over segments created by
the sliding window algorithm for LIMSI/Vocapia [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
transcripts and subtitles, where the size of the sliding windows
is 60.
      </p>
      <p>SW-40-U : Same as above for LIUM transcript with sliding
window size of 40.</p>
      <p>A rst approach consisted in reusing the search
component with the scene approach and the shot clustering
approach. A query was crafted from the anchor: the text
query was made by extracting keywords from the subtitle
aligned at start time and end time of the anchor. Visual
concepts scores were extracted from the keyframes of shots
contained in the anchor. If the anchor was constituted by
more than one shot, we took for each concept the highest
score over all shots.</p>
      <p>
        A second approach made use of MorelikeThis Solr
component (MLT) combined with Entityclassi er.eu annotation
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We created a temporary document from the query as
the root for searching similar documents, and performed the
search over segments from the LIMSI transcripts created
using sliding windows and enriched with synonyms.
4.1
      </p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS</title>
    </sec>
    <sec id="sec-7">
      <title>Search task</title>
      <p>The results of the search task are listed in Table 1. We
rst notice than given the same conditions, subtitles perform
signi cantly better than any of the transcripts, which is an
expected outcome. It is also interesting to note that using
the visual concepts in the query slightly increases the results
for all measures (e.g., clustering10-C vs clustering10-noC).</p>
      <p>Overall, the best approaches are those using scenes and
sliding windows. Scene based approaches retrieve a higher
number of correct relevant segments within a time window
of 60 seconds (higher MRR), but they are not the most
precise in terms of start and end time, compared to the sliding
windows approach (as suggested by mGAP and MASP).
4.2</p>
    </sec>
    <sec id="sec-8">
      <title>Hyperlinking task</title>
      <p>The results are listed in Table 2. For both LA and LC
condictions, runs using scenes outperform other runs for all
metrics. The MoreLikeThis/Entityclassi er.eu approach comes
second. As expected, using the context increases the
precision when hyperlinking video segments. It is also notable
that the precision at rank n decreases when n increases.</p>
    </sec>
    <sec id="sec-9">
      <title>CONCLUSION</title>
      <p>This paper presented our framework and results at the
MediaEval Search and Hyperlinking task. From our runs,
it is clear that scene segmentation is the approach with the
best performances. Therefore, this approach should be
studied more in depth, a potential improvement being to re ne
the segmentation using semantics or speakers information.
Also, we see here that this task bene ts from the use of
visual information present in the video. Hence, those two axes
should be the next steps to study for a future challenge.
6.</p>
    </sec>
    <sec id="sec-10">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was supported by the European Commission under
contracts FP7-287911 LinkedTV and FP7-318101 MediaMixer.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dojchinovski</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Kliegr</surname>
          </string-name>
          .
          <article-title>Entityclassi er.eu: Real-Time Classi cation of Entities in Text with Wikipedia</article-title>
          . In H. Blockeel,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kersting</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nijssen</surname>
          </string-name>
          , and F. Zelezny, editors,
          <source>Machine Learning and Knowledge Discovery in Databases</source>
          , volume
          <volume>8190</volume>
          of Lecture Notes in Computer Science, pages
          <volume>654</volume>
          {
          <fpage>658</fpage>
          . Springer Berlin Heidelberg,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Gareth J.F.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          .
          <article-title>The Search and Hyperlinking Task at MediaEval 2013</article-title>
          . In MediaEval 2013 Workshop, Barcelona, Spain, October
          <volume>18</volume>
          -19
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          , G. Jones,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wartena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Verschoor</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          .
          <article-title>Comparing retrieval e ectiveness of alternative content segmentation methods for Internet video search</article-title>
          .
          <source>In Content-Based Multimedia Indexing (CBMI)</source>
          ,
          <year>2012</year>
          10th International Workshop on, pages
          <volume>1</volume>
          {
          <issue>6</issue>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Lamel</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.-L.</given-names>
            <surname>Gauvain</surname>
          </string-name>
          .
          <article-title>Speech processing for audio indexing</article-title>
          .
          <source>In Advances in Natural Language Processing (LNCS 5221)</source>
          , pages
          <fpage>4</fpage>
          <lpage>{</lpage>
          15. Springer,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lin</surname>
          </string-name>
          . An
          <string-name>
            <surname>Information-Theoretic De</surname>
          </string-name>
          nition of Similarity.
          <source>In Proceedings of the Fifteenth International Conference on Machine Learning, ICML '98</source>
          , pages
          <fpage>296</fpage>
          {
          <fpage>304</fpage>
          , San Francisco, CA, USA,
          <year>1998</year>
          . Morgan Kaufmann Publishers Inc.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sidiropoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mezaris</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Kompatsiaris.</surname>
          </string-name>
          <article-title>Enhancing Video concept detection with the use of tomographs</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Image Processing (ICIP)</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sidiropoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mezaris</surname>
          </string-name>
          , I. Kompatsiaris,
          <string-name>
            <given-names>H.</given-names>
            <surname>Meinedo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bugalho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I.</given-names>
            <surname>Trancoso</surname>
          </string-name>
          .
          <article-title>Temporal Video Segmentation to Scenes Using High-Level Audiovisual Features</article-title>
          .
          <source>IEEE Transactions on Circuits and Systems for Video Technology</source>
          ,
          <volume>21</volume>
          (
          <issue>8</issue>
          ):
          <volume>1163</volume>
          {
          <fpage>1177</fpage>
          ,
          <string-name>
            <surname>Aug</surname>
          </string-name>
          .
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Eickeler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bardeli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Apostolidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mezaris</surname>
          </string-name>
          , and
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Muller. Think Before You Link { Meeting Content Constraints when Linking Television to the Web</article-title>
          .
          <source>In Proc. NEM Summit</source>
          , Nantes, France, Oct.
          <year>2013</year>
          . to appear.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Tscho</surname>
          </string-name>
          <article-title>pel and</article-title>
          <string-name>
            <given-names>D.</given-names>
            <surname>Schneider</surname>
          </string-name>
          .
          <article-title>A lightweight keyword and tag-cloud retrieval algorithm for automatic speech recognition transcripts</article-title>
          .
          <source>In Proceedings of the 11th Annual Conference of the International Speech Communication Association ISCA (INTERSPEECH)</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>