<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multimedia Content Processing and Retrieval in the REVEAL THIS setting</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stelios Piperidis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harris Papageorgiou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Katerina Pastra</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Netousek</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eric Gaussier</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tinne Tuytelaars</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Crestani</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francis Bodson</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chris Mellor</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute for Language and Speech Processing</institution>
          ,
          <addr-line>Athens</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Katholieke Universiteit Leuven R&amp;D</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>SAIL LABS Technology AG</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>TVEyes UK Ltd</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>The REVEAL THIS project (</institution>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Strathclyde</institution>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>Xerox - The Document Company S.A.S</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>- The explosion of multimedia digital content and the development of technologies that go beyond traditional broadcast and TV have rendered access to such content important for all end-users of these technologies. REVEAL THIS develops content processing technology able to semantically index, categorise and cross-link multiplatform, multimedia and multilingual digital content, providing the system user with search, retrieval, summarisation and translation functionalities.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Index Terms—audio-image-text analysis, cross-media linking
and indexing, cross-media categorisation, cross-media
summarisation, cross-lingual translation</p>
      <p>I. INTRODUCTION
T</p>
      <p>
        HE development of methods and tools for content-based
organization and filtering of the large amount of
multimedia information that reaches the user is a key issue
for its effective consumption. Despite recent technological
progress in the new media and the Internet, the key issue
remains “how digital technology could add value to
information channels and systems” [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>REVEAL THIS aims at answering this question by
tackling the following scientific and technological challenges:
• enrichment of multilingual multimedia content with
semantic information like topics, speakers, actors, facts,
categories
• establishment of semantic links between pieces of
information presented in different media and languages
• development of cross-media categorization and
summarization engines
• deployment of cross-language information retrieval and
machine translation to allow users to search for and retrieve
information according to their language preferences.</p>
      <p>
        As depicted in Figure 1, the REVEAL THIS system
comprises a number of single-media and multimedia
technologies that can be grouped, for presentation purposes,
in the following subsystems: (i) Content Analysis &amp; Indexing
(CAIS), (ii) Cross-media Categorisation (CCS), (iii)
Crossmedia Summarisation (CSS), (iv) Cross-lingual Translation
(CLTS), and (v) Cross-media Content Access and Retrieval.
The CAIS subsystem consists of technologies and components
for medium-specific analysis:
• Speech processing (SPC) – involving speech recognition,
speaker identification and speaker turn detection
• Image analysis and categorization (IAC)- involving shot
and keyframe extraction, low-level visual feature extraction,
image categorisation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
• Face analysis (FDIC) – involving face recognition &amp;
identification [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
• Text processing (TPC)- involving named entity, term &amp; fact
extraction, topic detection
• Cross-media Indexing (CMIC) – catering for the
establishment of links between all above-mentioned
metadata for a multimedia file using a modified TF-IDF and
a Dempster-Shafer based approach [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
The metadata/indices produced by the above components are
aligned, synchronized, linked to the corresponding points of
the source material (text, audio and video) and encoded in
MPEG7. Information suggested by audio processing (speaker
turns) and topic detection is taken into account to segment the
audiovisual or audio files into segments, or what one could
call “stories” i.e. thematic sections of the document.
Categorisation, summarisation and translation of multimedia
documents themselves make use of part of these metadata.
      </p>
      <p>B. Cross-media categorisation</p>
      <p>
        The categorization subsystem considers documents
containing not only text or images but a combination of
different types of media (text, image, speech, video). A
multiple-view fusion method is adopted, which builds 'on top'
of two single-media categorizers, a textual and an image
categorizer, without the need to re-train them. Data annotated
manually for both textual and image categories is used for
training the cross-media categorizer. In that set, dependencies
between single-media category systems are exploited in order
to refine the categorization decisions made [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>C. Cross-media summarisation</title>
      <p>
        The cross-media summarisation subsystem (CSS) determines
and presents the most salient parts according to the users’
profiles and interests by fusing video, audio and textual
metadata. It comprises three major components: the
textualbased summarization (TS), the visual-based summarization
(VS), and the cross-media summarization components,
aiming at fusing the two analyses and creating a
selfcontained object. Building on the MEAD development
platform [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the TS component extracts the top-ranked
sentences of a story: for each sentence, a salience score is
computed as a weighted sum of summary-worthy features.
      </p>
      <p>
        The VS component comprises the scene segmentation,
scene clustering &amp; labelling modules [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The scene
segmenter segments the video sequence into scenes. Scene
boundaries are detected as local minima in the visual
coherence function with each scene corresponding, ideally, to
a story of the video. Scene clustering caters for simple
applications that need a few indicative images. Keyframes of
the scene are clustered into larger parts, from which a
prototypical image is chosen. Clustering is repeated
iteratively to acquire a hierarchical cluster tree. The
prototypes of these clusters can be seen as representative
images of the scene. Going a step further, scene labelling is
invoked, for creating structured views of a file; currently, it is
news programmes that can be browsed in such a way,
allowing the user to watch all “anchor”, “interview”, and
“reportage” segments. The module labels all shots of a file
accordingly, by exploiting a belief propagation network.
Finally, the CSS brings all these pieces of information
together, providing visualisation interfaces
(SMIL/HTML+TIME) that enable the user to preview
multimedia objects effectively, before downloading them.
      </p>
    </sec>
    <sec id="sec-3">
      <title>D. Cross-lingual Translation</title>
      <p>The CLTS subsystem allows users to query documents written
in different languages, to categorise content expressed in
different languages and to preview language specific
summaries. A bilingual lexicon extraction module is used to
generate lexical equivalences for query translation purposes,
but also to replace keywords in a target language, in case a
document is linguistically not well formed (e.g. output from a
speech recognizer) and thus, not effectively translated. Last, a
statistical machine translation module is responsible for
providing translations of the textual part of the summaries
produced by the Cross-Media Summarization Subsystem.</p>
    </sec>
    <sec id="sec-4">
      <title>E. Usability Evaluation</title>
      <p>
        Apart from technical evaluation of the system components (cf
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]), the integrated REVEAL THIS prototype is
currently being evaluated by prospective users following a
task-based approach. A pool of about 30 users for each
application (pull and push) has been created. Based on typical
search sessions of these users, appropriate search tasks have
been created for the users to undertake using the REVEAL
THIS prototype. Feedback from the evaluation will guide the
final system refinements. The technology developed is
envisaged to contribute to a content management platform
that can be used by content providers, to add value to their
content, and directly by end users, for accessing multimedia
information.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Pastra</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Piperidis</surname>
          </string-name>
          ,
          <article-title>"Video Search: New Challenges in the Pervasive Digital Video Era"</article-title>
          ,
          <source>Journal of Virtual Reality and Broadcasting</source>
          , in press
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Perronnin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dance</surname>
          </string-name>
          , G. Csurka, M. Bressan, “
          <article-title>Adapted Vocabularies for Generic Visual Categorization”</article-title>
          ,
          <source>European Conference on Computer Vision</source>
          (ECCV), Graz, Austria,
          <year>2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>M. De Smet</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Fransens</surname>
            ,
            <given-names>L. Van Gool</given-names>
          </string-name>
          ,
          <article-title>"A generalised EM approach for 3D model based face recognition under occlusions"</article-title>
          ,
          <source>in Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)</source>
          , New York, USA,
          <year>2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Yakici</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          ,
          <article-title>"Cross-media Indexing in the Reveal This prototype"</article-title>
          ,
          <source>in Proceedings of the LREC workshop on "Crossing media for improved information access"</source>
          , Genoa, Italy,
          <year>2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Renders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Gaussier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Goutte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pacull</surname>
          </string-name>
          , G.Csurka,
          <article-title>"Categorization in multiple category systems"</article-title>
          ,
          <source>Proceedings of the 23rd International Conference on Machine Learning (ICML)</source>
          , Pittsburgh, USA,
          <year>2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Georgantopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Goedeme</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lounis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Papageorgiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tuytelaars</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. Van Gool</surname>
          </string-name>
          ,
          <article-title>"Cross-media summarization in a retrieval setting"</article-title>
          ,
          <source>in Proceedings of the LREC 2006 workshop on "Crossing media for improved information access"</source>
          , Genoa, Italy,
          <year>2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Simard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Cancedda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Cavestro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dymetman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Gaussier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Goutte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yamada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Langlais</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Mauser</surname>
          </string-name>
          ,
          <article-title>"Translating with Noncontiguous Phrases"</article-title>
          ,
          <source>In Proceedings of HLT/EMNLP</source>
          , Vancouver, Canada,
          <year>2005</year>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>