<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Interlinking Media Archives with the Web of Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dietmar Glachs</string-name>
          <email>dietmar.glachs@salzburgresearch.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sebastian Schaffert</string-name>
          <email>sebastian.schaffert@salzburgresearch.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoph Bauer</string-name>
          <email>christoph.bauer@salzburgresearch.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>SalzburgResearch Forschungsgesellschaft m.b.H.</institution>
          ,
          <addr-line>Salzburg, Austria Österreichischer Rundfunk, Wien, Österreich</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <fpage>17</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>Today's enterprises heavily rely upon accurate, consistent, and timely access to data. However, company data is typically scattered across multiple databases and file shares in a multitude of forms and versions. Moreover, an increasing amount of valuable background information is available outside the companies' influence and control. This situation is typical for many enterprise information integration scenarios, also in Austria's largest broadcasting media archive. Our demonstration argues for an information integration approach that uses semantic web principles to interlink archival media content of the Austrian Broadcasting Corporation (ORF) with the web of data and with internal knowledge resources to facilitate semantic search and to increase the user experience of browsing and discovering media content in the daily production workflow.</p>
      </abstract>
      <kwd-group>
        <kwd>Linked Enterprise Data</kwd>
        <kwd>Linked Media</kwd>
        <kwd>Semantic Media Archive</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The Linked Open Data (LOD) community project was initiated in 2007 by the W3C
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and proposes the usage of standards like the Resource Description Framework
(RDF) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for publishing datasets on the web in order to make them available for
interlinking [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The number of datasets available, commonly referred to as the Linked
Data Cloud1 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], is still growing and provides enterprises with the opportunity to
interlink enterprise data with background information or to allow for disambiguation of
concepts. Enterprises however still hesitate to use Linked Data in their value chain.
Based on experiences with industrial partners, the main barriers in the adoption of
Linked Data are (i) a rather new technology since accessing data from the Linked
Data cloud is still cumbersome; (ii) the lack of complete solutions because Linked
Data is still considered read-only and metadata-only whilst enterprise data is highly
dynamic and increasingly includes multimedia content and (iii) the need of adapting
established enterprise processes when using linked data [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        With this article we propose the integration of large datasets available on the web
by following the Linked Data principles as outlined in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] to enhance closed
enterprise content with additional information from the Linked Open Data cloud [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This
demo uses the Linked Media Framework (LMF2) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], a platform for enterprise
information integration. Based on Linked Data as well as Apache Stanbol3 fo r content
analysis, the LMF shows how to eliminate the entry barriers when using Lin ked Data
in enterprises.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Semantic Media Archive</title>
      <p>The Austrian Public Broadcaster’s (ORF – Österreichischer Rundfunk) arch ive is the
central repository for all video and audio material created by the ORF in the last 60
years and contains a vast amount of media content in different formats. Th e primary
objective of the archive is to preserve audio/video content for potential future use and
make it accessible to editors. When archiving new content, several archiv ing tools
restricted to expert users are used; FESAD4 as an example is used to man age video
based content. However, for journalists, editors and program planners the archiving
division uses a web based tool for federated search and investigation. For now, the
work of describing the clips (e.g. annotating the content) is actually carried out solely
by members of the archiving division. The users of the search tool currently cannot
modify/annotate content in order to improve data quality or search confidence.</p>
      <p>The main objective for the ORF is therefore to (i) provide additional information to
the end users like editors and journalists, (ii) to allow simple annotation means which
are not restricted to the archiving division and (iii) provide/integrate seman tic search
facilities for improved search results. As an integrated solution we integrated the LMF
as Linked Media Server in the Archival Toolset of the ORF. In addition to the
existing tools, the LMF provides extended semantic search facilities and also allows for
interlinking of archival content with publicly available linked data sources. As shown
in Fig. 1, the LMF extends the search tool mARCo by adding itself as an additional
data source and by providing means for annotating mARCO search results. The
annotations are then subject of future searches in mARCo.
2.1</p>
      <sec id="sec-2-1">
        <title>Annotating Media Content</title>
        <p>
          When browsing search results, editors or journalists are enabled to annotate the
content. With the help of a special annotation plugin, the formerly “read-only” search
result page becomes editable by injecting the annotation features into the w eb page.
Parts of the page such as the content description are analyzed by Apache Stanbol5. As
a result, eligible resource annotations are provided to the user as shown in Fig. 2.
By selecting a suggestion, the journalist can review the proposal and finally annotate
the content. The Linked Media Framework stores the annotation by means of
SPARQL Update [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and also collects the available properties of the referenced
resource and thus makes the information immediately available for semantic search.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Semantic Media Search</title>
        <p>The search experience can be improved by facilitating the semantic relatio ns of the
archived data. By using the semantic concepts of the data which are either p roduction
related (e. g. moderator, editor, program etc.) or content related (e.g. persons named
or in video, content description, location of the clips content), it is possible to provide
a faceted search as shown in Fig. 3, for example to narrow down the search, the user
may select one or more facet properties shown in the search interface.
5
http://incubator.apache.org/stanbol
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>DEMO OUTLINE</title>
      <p>
        The Linked Media Framework (LMF) serves as the backend whereas the both clients
for search and annotation are lightweight JavaScript implementations using RESTful
webservices for the communication with the backend service. The LMF is a service
oriented framework which uses semi-structured data representation (RDF) and HTTP
URLs as uniform resource identifier to store and identify resources, as recommended
for Linked Data [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The demo we show at the conference will first show the
Semantic Search Component as it is a fundamental part of the LMF and demonstrates the
power and flexibility of using Semantic Web technologies for search and retrieval.
We will then use a VIE bookmarklet6 for the annotation of a typical ORF search result
page which relies on concepts from DBPedia7 and an internal SKOS8 based thesaurus.
Accepting proposed annotations with the LMF will immediately influence the search
results and optionally add new concepts to an internal company thesaurus. In the
production scenario, the LMF will also be tightly connected with the mARCo search
facility and therefore will be part of the federated search component.
      </p>
      <p>The LMF integrates/connects the linked data cloud as possible sources for
background information and finally enables annotation by storing selected concepts in the
(local) Linked Data server by means of SPARQL Update statements. In particular this
annotation functionality will be subject of the demonstration given at I-Semantics to
first show the where we will preload the LMF with a selection of news articles out of
the Austrian Broadcasters Archive. The demonstration will also cover how the news
articles are presented to journalists for annotation. Finally, the demonstration of the
search interface is also available online at the NewMediaLabs demonstration site9.
4</p>
    </sec>
    <sec id="sec-4">
      <title>CONCLUSION</title>
      <p>The potential of Linked Data in general and the Linked Media Framework as a
platform for supporting semantic search has been proven in several projects. With this
demonstration we aimed to outline its potential for the use in an Enterprise
Information Integration scenario where Linked Data technology is used to support users in
their daily work and to improve the amount and quality of content annotation. The
latter directly leads to an improved search result with respect to precision which is a
fundamental requirement in the news domain. Because of the smooth integration in
existing processes, the functionality is offered as an optional add-on to the users. The
improved search results as well as the provided background information are the
inducement for the users to use the offered functionality. In contrast to the increasing
number of semantic web case studies10, the demonstrated scenario Linked Media
Framework allows the publication of structured information as Linked Data and also
6 http://szabyg.github.com/vie-annotation-bookmarklet/
7 http://dbpedia.org
8 http://www.w3.org/2004/02/skos/
9 http://labs.newmedialab.at/ORF/orf/search/index.html
10 http://www.w3.org/2001/sw/sweo/public/UseCases/
enables the full read-write management of the published data and in particular enables
the full roundtrip of annotations for further usage during search and retrieval.
5</p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGMENTS</title>
      <p>The media content enhancement and the semantic search described in this paper
were planned and developed in the Austrian research centre "Salzburg NewMediaLab
- The Next Generation" (SNML-TNG). The centre is funded by the Austrian Federal
Ministry of Economy, Family and Youth (BMWFJ), the Austrian Federal Ministry for
Transport, Innovation and Technology (BMVIT) and the Province of Salzburg. The
demo content is taken from the ORF archive by courtesy of the Austrian Broacasting
Corporation. The development of the LMF has been inspired by the needs &amp; requests
of our industrial partners. As a result, the Linked Media Framework currently serves
several real-world scenarios.
6</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Linking</given-names>
            <surname>Open Data</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>W3C SWEO Community Project</article-title>
          . Retrieved from http://esw.w3.org/topic/SweoIG/TaskForces/CommunityProjects/LinkingOpenData
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. RDF: G. Klyne and
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Carroll</surname>
          </string-name>
          .
          <article-title>Resource description framework (RDF): Concepts and abstract syntax</article-title>
          .
          <source>Technical report, W3C</source>
          ,
          <fpage>2</fpage>
          <lpage>2004</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>How to Publish Linked Data on the Web</article-title>
          . Retrieved from http://www4.wiwiss.fuberlin.de/bizer/pub/LinkedDataTutorial
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Linked Data - The Story So Far</article-title>
          .
          <source>International Journal on Semantic Web and Information Systems</source>
          ,
          <volume>4</volume>
          (
          <issue>2</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          . Elsevier. Retrieved from http://www.citeulike.org/user/omunoz/article/5008761
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Wood</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>Linking Enterprise Data</article-title>
          .
          <source>ISBN 978-1-4419-7664-2. DOI 10</source>
          .007/978-1-
          <fpage>4419</fpage>
          -7665-9
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>Linked Data - Design Issues</article-title>
          . Retrieved from http://www.w3.org/DesignIssues/LinkedData.html
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ayers</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raimond</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Interlinking Open Data on the Web (Poster)</article-title>
          .
          <source>In 4th European Semantic Web Conference (ESWC2007)</source>
          , pages
          <fpage>802</fpage>
          -
          <lpage>815</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kurz</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schaffert</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bürger</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>LMF - A Framework for Linked Media</article-title>
          .
          <source>In: Workshop for Multimedia on the Web (MMWeb2011).</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Damjanovic</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kurz</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Westenthaler</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Behrendt</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gruber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Schaffert</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Semantic enhancement: The key to massive and heterogeneous data pools</article-title>
          .
          <source>In Proceeding of the 20th International IEEE ERK (Electrotechnical and Computer Science) Conference</source>
          <year>2011</year>
          , Portoroz, Slovenia.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. Prudތ hommeaux, E., &amp;
          <string-name>
            <surname>Seaborne</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>SPARQL Query Language for RDF. W3C working draft</article-title>
          . Retrieved from http://www.w3.org/TR/rdf-sparql-query
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>