<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SALERO - Semantic Audiovisual Entertainment Reusable Objects</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Werner Haas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Georg Thallinger</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pedro Cano</string-name>
          <email>pcano@iua.upf.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Charlie Cullen</string-name>
          <email>charlie.cullen@dit.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tobias Bürger</string-name>
          <email>tobias.buerger@deri.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IP SALERO is partially funded under FP 6 of the European Commission within the IST Workprogramme 2004 (IST FP6-2004-027122). W. Haas, G. Thallinger, are with JOANNEUM RESEARCH</institution>
          ,
          <addr-line>Graz, Austria (phone:</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>- The Integrated Project SALERO aims to advance the state of the art in digital media to the point where it becomes possible to create audiovisual content for cross-platform delivery using intelligent content tools, with greater quality at lower cost, to provide audiences with more engaging entertainment and information at home or on the move. SALERO will build on and extend research in media technologies, web semantics and context based image retrieval, to reverse the trend toward everincreasing cost of creating media.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Index Terms— audiovisual intelligent objects, content creation,
context aware behaviour</p>
    </sec>
    <sec id="sec-2">
      <title>I. VISION &amp; OBJECTIVES</title>
      <p>
        SALERO’s [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] overall vision is to define and develop
‘intelligent content’ for media production, consisting of
multimedia objects with context-aware behaviour for
selfadaptive use and delivery across different platforms.
‘Intelligent Content’ should enable the creation and re-use of
complex, compelling media by artists who need to know little
of the technical aspects of how the tools that they use actually
work.
      </p>
      <p>Complete realisation of SALERO’s vision is a long-term
goal. This gives rise to three overarching R&amp;D
objectives:
Address characters, objects, sounds, language sets and
behaviours,
Research into methodologies for creating and finding
intelligent content,
Develop toolsets to create, manage, edit, retrieve and
deliver content objects.</p>
    </sec>
    <sec id="sec-3">
      <title>II. INTELLIGENT CONTENT CREATION</title>
      <p>The first goal is to obtain a better understanding of the
relations between media types, genres, workflows and styles
as a pre-requisite to the adaptation and transfer of content
elements across productions and platforms. To this end,
metadata, media semantics and ontologies need to be
analysed, researched and developed that define the parameters
necessary for the creation and manipulation of semantically
aware media objects of various types. Practical methods of
context-based information retrieval will be researched that
simplify the location and retrieval of characters, sounds,
images, movements or behaviours from very large datasets
and media storage systems. Improved methods and tools for
language processing and speech synthesis, as a means of
supporting the generation of multilingual media content, need
to be developed.</p>
      <sec id="sec-3-1">
        <title>A. Media Semantics and Ontologies</title>
        <p>The objective of this research strand is twofold: the main
objective is to devise a machine process able description for
the semantic features of a multimedia object and the context it
should be used in. This will be tackled by building up a set of
ontologies taking into account a layered approach – using an
appropriate representation technique for every level of
metainformation – that relies as much as possible on current
description standards for multimedia. The second objective is
to design and implement necessary tools and applications to
build up, maintain and query ontologies for multimedia
objects.</p>
      </sec>
      <sec id="sec-3-2">
        <title>B. Media Forms, Programme Styles &amp; Structures</title>
        <p>
          Media objects have different specification needs on
different platforms from game consoles and online services to
DVD, television and cinema, related to the overall expressive
and stylistic objectives of the production. Audience
expectations are related to the production genre, be it a
western, soap opera, comedy, tragedy, thriller,
actionadventurer, a medieval sword &amp; magic MMPORG (Massively
Multi-Player Online Role-Playing Game). The genre
expectations (whether the engagers are passive watchers of
Jerry Seinfeld on television, or active console game players
represented on the screen by Lara Croft) are elegantly
expressed by the active questions of Philip Parker [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>C. Context Based Information Retrieval</title>
        <p>We expect that the intelligent content elements developed
by SALERO will adapt themselves to the context of the
production. We will therefore need to research ways of
defining, creating (or locating), managing and delivering
content objects of different kinds in a range of contexts. We
address the context-based retrieval of media objects within a
media production environment. The idea is to re-use objects,
motion data and other production related data for the creation
of new production materials. It is not only about finding and
re-using elements at production time, but also using retrieval
technologies for the creation of interactive media productions.
That means for example that a character may react or adapt to
a scene in a way that is based on the users input. A number of
factors affect the use of retrieval techniques within media
production environments. The most important one is that they
should be integrated into the production environment:
retrieval should happen as part of the programme development
and not as a cumbersome, extra activity. We also need
intelligent context sensitive retrieval mechanisms that identify
both user context and task context.</p>
      </sec>
      <sec id="sec-3-4">
        <title>D. Speech and Language</title>
        <p>
          The aim of this activity is to enable programmes created in
one language to be re-purposed and/or synthesised in another
language or dialect by researching and developing a ‘speech
corpus/concordancer’. A speech corpus[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], tagged for various
features such as rhythm, pitch contours, intensity contours and
emotional dimensions [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] will be used to inform lip-synching,
character animation and TTS synthesis stages by establishing
emotional rules - initially for English - with which to
potentially repurpose ‘neutral’ or ‘nearest match’ speech
segments in the database for the other language or dialect.
Once the rules have been established for English, they will be
adapted for Spanish or Catalan.
        </p>
        <p>This requires a framework for defining voice stereotypes
for age, genre, emotional dimension etc and a suitable tagging
system for corpus transcripts-initially for Catalan, Spanish and
English. In the case of English, tagging for speech rhythms
and other acoustic features within the recorded speech clips is
seen to play a crucial role in developing a more natural
corpus. The tagging of the resultant speech corpus will be
applied to rule-based analysis, synthesis, lip-synching and
character animation.</p>
      </sec>
      <sec id="sec-3-5">
        <title>E. Characters, Characteristics &amp; Effects</title>
        <p>The research in this activity deals with both visual objects
(such as characters) and audio objects (such as effects,
speech). It provides grounding work for the linking of visual,
audio, and behavioural objects, whose initial intelligence is
expected to be increased along the lifetime of the project. It
develops along different levels, from the low level provision
of basic affordable rendering engines for media; through
intermediate level, such as the modelling and animation of
characters; to high level aspects, e.g. programme generators.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>III. TOOLSETS, DEMONSTRATION &amp; TRAINING</title>
      <p>Software toolkits, software systems, plug-ins and interfaces
will be developed that allow the control of appearances,
sounds, semantic behaviour and properties of intelligent
content objects for media production and post-production, and
can be used in conjunction with existing industry programs.
They will be validated and evaluated through a series of
experimental productions, based on scenarios defined by
artists and creative media professionals.</p>
      <p>Results will be promoted by a broad initiative, developing
demonstration test beds and training structures for
professionals and researchers, as well as by addressing the
relevant standardisation bodies.</p>
    </sec>
    <sec id="sec-5">
      <title>IV. RELATED WORK</title>
      <p>
        A number of research groups are dealing with ontology
based description of multimedia items often by applying
reasoning to low level features extracted, e.g. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Use of
ontology languages for media annotation has been
investigated in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] for video and in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] for audio.
Deployment of semantic technologies in media production in
tools used every day by the media professional has been rarely
investigated.
      </p>
      <p>
        The SMaRT networking cluster [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] (which SALERO is
member of) combines research in the fields of: Semantic Web,
Multimedia and Signal Analysis to address emerging research
challenges in Semantic Multimedia.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] SALERO web page</article-title>
          , http://www.SALERO.eu
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Parker</surname>
          </string-name>
          , “
          <article-title>The art</article-title>
          and science of screenwriting”,
          <source>Intellect</source>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.F.</given-names>
            <surname>Campbell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Meinardi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Richardson</surname>
          </string-name>
          , “
          <article-title>Let the corpus speak!”, 40th IATEFL Annual Conference</article-title>
          and Exhibition,
          <fpage>9</fpage>
          -
          <lpage>12</lpage>
          April 2006, Harrogate, UK.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.R.</given-names>
            <surname>Scherer</surname>
          </string-name>
          ,
          <article-title>Emotion as a multi-component process: A model and some cross cultural data</article-title>
          .
          <source>Review of Personality and Social Psychology</source>
          ,
          <year>1984</year>
          . 5: p.
          <fpage>37</fpage>
          -
          <lpage>63</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dasiopoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mezaris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Kompatsiaris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.K.</given-names>
            <surname>Papastathis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.G.</given-names>
            <surname>Strintzis</surname>
          </string-name>
          <article-title>: "Knowledge-assisted semantic video object detection"</article-title>
          ,
          <source>IEEE Transactions on Circuits and Systems on Video Technology</source>
          , Vol.
          <volume>15</volume>
          , No.
          <volume>10</volume>
          , pp.
          <fpage>1210</fpage>
          -
          <lpage>1224</lpage>
          , (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Heflin</surname>
          </string-name>
          , “OWL Web Ontology Language:
          <article-title>Use cases</article-title>
          and requirements”,
          <source>W3C Recommendation</source>
          , http://www.w3.org/TR/webontreq/, (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V.</given-names>
            <surname>Mezaris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kompatsiaris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Boulgouris</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Strintzis</surname>
          </string-name>
          , “
          <article-title>Real-time compressed domain spatiotemporal segmentation and ontologies for video indexing and retrieval”</article-title>
          ,
          <source>IEEE Transactions on Circuits and Systems on Video Technology</source>
          , Vol.
          <volume>14</volume>
          , No.
          <issue>5</issue>
          , pp.
          <fpage>606</fpage>
          -
          <lpage>620</lpage>
          , (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>J.R. van Ossenbruggen; F.-M. Nack</surname>
            ; and
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Hardman</surname>
          </string-name>
          <article-title>; “That obscure object of desire: multimedia metadata on the web (Part II)”</article-title>
          , IEEE Multimedia, Vol.
          <volume>12</volume>
          , No.
          <issue>1</issue>
          , pp.
          <fpage>54</fpage>
          -
          <lpage>63</lpage>
          , (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Koppenberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Le</given-names>
            <surname>Groux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ricard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Wack</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Herrera</surname>
          </string-name>
          ,
          <year>2005</year>
          . “
          <article-title>Nearest-neighbour automatic sound classification with a wordnet taxonomy”</article-title>
          .
          <source>Journal of Intelligent Information Systems</source>
          Vol.
          <volume>24</volume>
          .2 pp.
          <fpage>99</fpage>
          -
          <lpage>111</lpage>
          (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>[10] http://kspace.qmul.net:8080/kspace/kspacesmartcluster.jsp</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>