<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Dealing E ciently with Ontology-Enhanced Linked Data for Multimedia</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oliver Gries</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ralf Moller</string-name>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anahita Na ssi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maurice Rosenfeld</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kamil Sokolski</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sebastian Wandelt</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Amirkabir University of Technology</institution>
          ,
          <addr-line>Tehran</addr-line>
          ,
          <country country="IR">Iran</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Hamburg University of Technology</institution>
          ,
          <addr-line>Hamburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Humboldt University of Technology</institution>
          ,
          <addr-line>Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Lufthansa Industry Solutions</institution>
          ,
          <addr-line>Hamburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Nordex Energy GmbH</institution>
          ,
          <addr-line>Hamburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Lubeck</institution>
          ,
          <addr-line>Lubeck</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In order to provide automatic ontology-based multimedia annotation for producing linked data, scalable high-level media interpretation processes on (video) streams are required. In this paper we shortly describe an abductive media interpretation agent, and based on a Multimedia Content Ontology we introduce partitioning techniques for huge sets of time-related annotation assertions such that interpretation as well as retrieval processes refer to manageable sets of metadata.</p>
      </abstract>
      <kwd-group>
        <kwd>linked data</kwd>
        <kwd>multimedia interpretation</kwd>
        <kwd>stream processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        A large amount of multimedia content is available on the Web, and these days
appropriate multimedia documents can hardly be systematically found using
keyword-based search. Therefore, the eld of Linked Data has emerged [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Linked data are also called rich semantic media in the literature [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. These
research elds investigate the derivation and management of symbolic descriptions
for multimedia content. Symbolic descriptions are anchored at various parts of
a multimedia object, and they can be used to link various (parts of) multimedia
objects. Hence the term linked data has emerged. Symbolic content descriptions
approximate human-level interpretations of media content, and, therefore, can be
used for systematic document retrieval based on high-level topic-based queries.
Retrieval based on linked data can be enhanced if retrieval processes are based
on ontologies, namely a domain ontology and a general ontology for describing
document structure and content (see, e.g., [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]).
      </p>
      <p>To some extent linked data can be automatically derived using existing
datadriven media analysis systems. However, there still exists a gap between, for
instance, low-level image/video analysis and high-level image/video interpretation,
not to mention human-level understanding. Thus, analysis-level results obtained
from state-of-the-art tools have to be augmented with more abstract symbolic
descriptions. This is accomplished in an automatic process which we call media
interpretation. Recent research in the area of ontology-based media
interpretation has shown enormous advances, and we assume that media interpretation
processes can safely generate linked data, to be used in ontology-oriented media
retrieval processes. For brevity, linked-data generation is also called (automatic)
annotation in this paper, and we focus on videos as multimedia objects in order
to be as concrete as possible.</p>
      <p>The sheer amount of assertions for appropriately describing the content of
large media objects makes media interpretation as well as annotation-based
retrieval increasingly di cult. In this paper we advance the state-of-the-art in
several areas by: 1. Proposing a description language for video annotations that
supports scalable high-level reasoning about video content (interpretation as
well as retrieval). 2. Explaining ontology-based reasoning techniques for an
annotation agent, which is used to compute high-level interpretations of videos.
3. Showing how to support decomposition-based scalability for reasoning in the
context of long streams of video content.</p>
      <p>
        There already exist various proposals for annotation languages (see, e.g.,
early approaches based on MPEG-7 [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ] or newer ones dedicated to knowledge
management [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]). However, none of the languages has been developed while
keeping in mind scalable stream-based reasoning w.r.t. an ontology (rather than
mere data retrieval). Reasoning is used for media interpretation, which is a
service being used for computer-aided semantics annotation of multimedia (see the
EU project CASAM http://www.casam-project.eu/). The CASAM
Multimedia Content Ontology introduced in this paper (called MCO for short) is an
extension and modi cation of a previous multimedia ontology described in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>Scalability is a signi cant issue in at least two respects. If we talk about
interpretation of a video document, then, on the one hand, there is the time
dimension to be considered. On the other hand, another dimension is the
interpretation depth. As we have argued above, we can assume that interpretation
is based on explicit (symbolic) low-level information for each perceptive unit. A
perceptive unit, such as a video shots, is called \segment" in MCO. The aim
of interpretation is to compute high-level information for a segment given the
knowledge acquired so far. Thus, in our annotation language we need to be able
to represent time information as well as to support the ability to draw
conclusions on higher levels of interpretation. Hence, the notion of a segment has to be
appropriately de ned using an ontology, and assertions representing
interpretation results at various levels of detail have to be attached to segments using an
appropriate annotation language.</p>
      <p>It can easily be seen that this kind of two-dimensional streaming scenario,
with multiple streams for multiple modalities, yields a signi cant growth of
assertions over time. Although our low-level annotation language is based on a
description logic for which e cient typical-case reasoning systems are known,
we need to exploit new partitioning techniques to break down the data
descriptions used for interpretation into smaller pieces to be handled over time. This is
even more important if low-level results become available for time frames in an
asynchronous way (maybe with substantial time delays according to the
intricacies of certain tools for di erent modalities). In order to improve scalability,
we identify and use locality in the video stream. Given an annotation of a video
stream, we use split-operations to compute so called \islands," which are
sufcient for reasoning with respect to the current state of annotation. Running
interpretation tasks on separate islands (instead of the whole set of assertions)
improves performance signi cantly.</p>
      <p>
        The remaining part of the paper is structured as follows. Section 2 presents
our motivation for choosing the description logic ALHf R+ (D) as an ontology
language for de nition of an annotation language MCO. In Section 3, we
introduce the Multimedia Content Ontology in detail and explain how ALHf R+ (D)
is used to represent the content of multimedia documents. We discuss scalability
issues and present solutions for partitioning large sets of assertions into
manageable \islands" or \chunks" in Section 4 such that interpretation processes
run e ciently. Several examples demonstrate the e ectiveness of the proposed
techniques. We conclude in Section 5. Due to space constraints, retrieval is not
covered in this paper (We refer to [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] for details).
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Representation of Multimedia Content</title>
      <p>
        In order to describe multimedia documents in terms of annotations (stored as
metadata), the Moving Pictures Experts Group (MPEG) has speci ed the ISO
standard Multimedia Content Description Interface, also denoted as MPEG-7
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In this framework, XML descriptions of multimedia data are associated with
content, with the objective to allow for e cient search and retrieval of
multimedia documents. The MPEG-7 schema language provides for restrictions on
valid media descriptions, for which XML query languages are de ned. If, in the
context of linked data, the inherent problem of XML query languages comes into
play, namely the lack of facilities for querying the name of a relation of which a
certain annotation tuple is an element, then RDF-based representations are
bene cial. Proposals for using RDF in the context of MPEG-7-alike representations
have been discussed in the literature as well (e.g., [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]), and retrieval languages
such as SPARQL can be used to nd media objects based on RDF content
descriptions. RDF query answering with respect to ontologies means that data
(tuples, or triples to be more precise) that can be inferred w.r.t. the ontology
are implicitly added to what is given explicitly in the RDF annotation. E cient
query engines might not materialize implicit tuples, though. In this case, given
the implicit tuples (deductive closure), more media objects are likely to be found
if ontologies come into play for query answering. Note that w.r.t. an ontology, a
set of RDF triples can also become inconsistent. Inconsistencies can be detected
with reasoning engines, but this is not relevant throughout this paper (although
inconsistencies also restrict the set of possible annotations in the same way as
an XML schema restricts the set of valid annotations).
      </p>
      <p>
        Ontology languages such as RDFS or OWL2 are languages which have a
formal semantics, a feature that is bene cial for formally de ning decision
problems and checking correctness of corresponding decision procedures (aka
inference algorithms). However, speci c ontologies (aka knowledge bases) speci ed
using an expressive ontology language reveal more of the \semantics" of media
documents as mere RDF triples do. Ontologies achieve this by adding (lots of)
implicit content description tuples to the annotations given explicitly as part of
the documents' annotations. Thus, we have \semantics" in the sense of formal
semantics of a representation language and \semantics" in the sense of implicit
tuples added to the explicitly given ones. Many papers in the Semantic Web
literature amalgamate these two kinds of \semantics", suggesting that semantics
in the sense of content descriptions come for free using formal representation
languages. For the latter notion of \semantics," we prefer the name content
description in order not to confuse the reader. Content descriptions do not come
for free but must be derived using media interpretation processes, which require
dedicated knowledge bases for interpretation knowledge [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>In MPEG-7, a multimedia document is related to its modality speci c
content, composed of, e.g., video, audio, or text, and each of these parts consists
of a set of segments specifying \regions" of modality speci c data. At rst, we
believe it is necessary to be able to specify more general resp. more speci c
concepts and roles (e.g. that video content also is of type multimedia content) for
building up a taxonomy. Further, it is important to be able to specify concept
disjointness. In order to represent relations (e.g. from content to segments) roles
can be speci ed, whose domain and range usually are constrained to a speci c
concept (e.g., the role hasMediaDecomposition is constrained to only relate
instances of multimedia content to multimedia segments) and which are possibly
functional or transitive. In addition, we propose that for modality speci c
concepts the range of roles is further restricted to modality speci c concepts (e.g.
audio content is only allowed to be related to audio segments). Finally, for
representing multimedia content it is usually necessary to be able to specify concrete
domains such as integers or strings.</p>
      <p>
        We argue that this expressivity is su cient for the representation and
interpretation of multimedia content for a large range of problems. For example,
we propose to abandon existential restrictions on the right side of inclusion
axioms, since we believe that it is not required to constrain multimedia content
descriptions to consist of \anonymous" individuals of a speci c type (which
cannot be retrieved explicitly [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]). The respective DL is denoted ALHf R+ (D)
(restricted attributive concept language with role hierarchies, functional roles,
transitive roles and concrete domains). We made several experiments with the
DL reasoner RacerPro [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] strongly indicating that reasoning with ALHf R+ (D)
is e cient.
      </p>
      <p>
        We now shortly introduce the descripton logic (DL) nomenclature. A DL
signature is a tuple S = (CN; RN; AN; IN), where CN = fA1; :::; Ang is the
set of concept names (we also use A for concept names in the sequel). RN =
fR1; :::; Rmg is the set of role names. Further, AN is a set of concrete domain
attributes (i.e., roles whose range is a concrete domain). The signature also
contains a component IN indicating a set of individuals. A DL knowledge base
OS = (T ; A), de ned with respect to a signature S, is comprised of a
terminological component T (called Tbox ) and an assertional component A (called
Abox ). In the following we just write O if the signature is clear from context.
An ALHf R+ (D) Tbox is a set of axioms A1 v A2 and R1 v R2 (atomic
subsumption), A1 v :A2 (disjointness), 9R:&gt; v A and &gt; v 8R:A (domain and
range restrictions on roles), &gt; v ( 1 R) (functional roles), T rans(R)
(transitive roles) and A1 v 8R:A2 (local range restrictions on roles). An Abox A i is a
set of concept assertions A(a) and role assertions R(a; b), where A is a concept
name, R is a role name, and a; b represent individuals. Aboxes can also contain
equality (a = b) and inequality assertions (a 6= b) as well as attribute assertions
of the form Attr(a; val) where Attr is an attribute and val is either a string
or an integer (with the obvious denotation). For a detailed introduction to the
incorporation of concrete domains into DLs, to the semantics of concepts and
roles, as well as an introduction to the satis ability conditions for axioms and
assertions we refer to [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], respectively. Standard DL decision problems
are also formally de ned in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] (e.g., computing the concept and role hierarchies,
as well as concept-based and conjunctive instance retrieval).
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>The Multimedia Content Ontology</title>
      <p>
        In this section, the CASAM Multimedia Content Ontology is presented to an
extent that the solution to scalabilitiy problems can be understood. The full MCO
can be found at http://www.sts.tu-harburg.de/casam/mco.owl. In contrast
to approaches transforming the complete MPEG-7 standard to RDFS [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] or
OWL [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], our approach is inspired by using only those parts of MPEG-7
describing a general structure for multimedia documents. The main objective is to
e ectively exploit quantitative and qualitative time information in order to relate
co-occurring observations. Co-occurrences are detected either within the same
or between di erent modalities regarding the video shots. In the following, we
focus on axioms relating concept and role names required for these capabilities.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Concept Hierarchy and Role Hierarchy</title>
        <p>In Fig. 3.1 the concept hierarchy (on the left) and the role hierarchy (on the
right) of the multimedia ontology is shown. A complete multimedia document
is represented by the concept MultimediaDocument . Only one instance of type
MultimediaDocument should be speci ed for a document to be annotated.</p>
        <p>Individuals which are instances of (subconcepts of) MultimediaContent
represent di erent modalities of the video. The concept VideoContent represents
the video modality and holds all video segments. In the same way AudioContent
holds all segments from the audio modality.</p>
        <p>TextContent represents text paragraphs associated with certain segments or the
whole video. Auxiliary text documents that are related with the whole
annotated video are represented by the subconcept AuxiliaryContent . During the
annotation process, a user can make free text annotations, which describe the
whole multimedia document or a single segment (shot) of the video. These
free text annotations are represented by GlobalUserAnnotationContent resp.
LocalUserAnnotationContent . As speech, recognized in the video, is transformed
into text, the concept SpeechRecognitionContent is also subsumed by TextContent .</p>
        <p>To represent parts of the content, MultimediaContent instances can be
decomposed into MultimediaSegment instances. TextSegment refers to words in the
text modality. The concept SegmentLocator is used to specify start and end of
segments. The concrete values of start and end represent temporal position
information for audio and video, or denote character positions for text. BoundingBox
is used to determine the position of a recognized object in a video frame. All
concepts within the same hierarchy level are disjoint.</p>
        <p>The role hasLogicalDecomposition decomposes the whole media document
into the di erent parts by relating instances of type MultimediaDocument with
modality speci c content description individuals (that is, instances of the concept
M ultimediaContent). An individual of the type MultimediaContent is
associated to its segments by the role hasMediaDecomposition. To relate an individual
of type MultimediaSegment with its locators, the role hasSegmentLocator is used.</p>
        <p>
          The roles nextTextContent and nextTextSegment are used to specify the
order in the text paragraphs resp. words. Both are transitive roles. Subroles of
correlatesWith can be used to represent associations between content
descriptors. A TextContent instance can be related to a VideoSegment using the role
belongsTo. We also use a small subset of the Allen relations [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] to relate video
segment, or, more precisely, the locators associated with video segments. Note
that we do not require reasoning on Allen relations since the corresponding
relations are generated based on quantitative data. While o (overlap) describes an
intersection between audio and video locators, m (meets) describes the
alignment of two video or two audio segments. Please note that we compute
(qualitative) relations such as o using (quantitative) information about locator objects.
Quantitative information is given in terms of restrictions on values for attributes
hasStart and hasEnd (see Section 3.3).
        </p>
        <p>The role depicts is used to establish a mapping from individuals of the
Multimedia Content Ontology to observations from the domain ontology that were
extracted by analysis modules. In a similar way as depicts, hasInterpretation
provides a map to individuals that were generated as a part of interpretations
of observations. To represent the aggregating characteristic of high-level
interpretations, the role associatedWith is used to related high-level interpretations
with other interpretations or directly with observations.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Range Restrictions</title>
        <p>Range restrictions on roles constrain the corresponding role llers to be of a
speci c type. For example,</p>
        <p>&gt; v 8hasM ediaDecomposition:M ultimediaSegment
de nes the range restrictions on the role hasM ediaDecomposition such that the
role ller is constrained to be of type M ultimediaSegment.</p>
        <p>Local range restrictions constrain the range of roles further when the role is
applied to a speci c concept. The local range restriction</p>
        <sec id="sec-3-2-1">
          <title>AudioContent v 8hasM ediaDecomposition:AudioSegment</title>
          <p>speci es that the range of the role hasM ediaDecomposition associated with the
concept AudioContent is further restricted to AudioSegment.
3.3</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Attribute Values</title>
        <p>The attributes hasStart and hasEnd are used to specify time information of
video or audio segments. For example:</p>
        <p>AudioLocator(as1); AudioLocator(al1); hasSegmentLocator(as1; al1)
hasStart(al1; \00:43"); hasEnd(al1; \00:52")
de nes the starting time of an AudioSegment, e.g., as1 by specifying concrete
values to its corresponding AudioLocator al1. Integer values are used to specify
character positions to identify words in larger text strings. Also regarding the
text modality, the property hasConcreteValue is used to associate strings to
instances of speci c types such as CityN ame.</p>
        <p>
          Given quantitative information about start and end time, qualitative
relations between locator instances are computed by the media interpretation agent.
From the potential 13 qualitative relations de ned by Allen [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] we explicitly
represent o (overlaps) , d (during), and m (meets) between segment locators. As
shown in the next section the main motivation for switching from a
quantitative temporal representation to a qualitative one is to achieve scalability. Given
the M I Agent introduced above (see 4.1 for more details), qualitative relations
allows to partition the interpretation Abox(es) that always grow(s) over time.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Scalable Video Interpretation</title>
      <p>
        As we have seen, for improving shot-based video annotation, interpretations are
computed for co-occurrences of locator individuals according to temporal
information. In the course of the video interpretations, Aboxes grow signi cantly.
If, e.g., a particular video segment is focused on because there is a new
assertion coming in, referring to this video segment, only very few other assertions
are relevant. Large parts of Aboxes (e.g., for temporally far away parts of the
video) need not be processed. In this section we formalize the subdivision of large
Aboxes into meaningful parts (partitions) such that reasoning problems are
handled in the same way as with the large Abox. Reasoning on the small partitions,
also called island reasoning, is known to improve reasoning performance
significantly [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. We start with the introduction of some important aspects of the
media interpretation agent.
4.1
      </p>
      <sec id="sec-4-1">
        <title>The Multimedia Interpretation Agent</title>
        <p>
          In [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], an agent for Multimedia Interpretation (MI Agent ) was introduced. It
uses a probabilistic interpretation engine which, among others, is based upon
abduction. The idea is to generate explanations for observations in the form of
hypothesized Abox assertions. Given the added assertions and a set of rules as
part of the agent's knowledge base, the observations are then entailed. The agent
computes assertions that \support" the observations. The MI Agent receives
percepts in the form of assertions that represent the ongoing video analysis
and annotation process. The assertions are received in a streaming way by the
MI Agent in small bunches, which we formalize as sets here.
        </p>
        <p>Each is added to the Abox that the agent maintains. Subsequently, a set
of forward-chaining rules is applied. The general form of these rules is</p>
        <p>Q1(Y1); : : : ; Qn(Yn) ! P (X)
vc1: VideoContent
hasMediaDecomposition</p>
        <p>
          …
vs1: VideoSegment
vs2: VideoSegment
…
vl2: VideoLocator
… …
… …
… as2: AudioSegment
as1: AudioSegment
depicts hasSegmentLocator o
where Q1; : : : ; Qn; P denote concept or role names and underlined letters denote
(possible) tuples of all-quanti ed variables with the condition that each variable
appearing in P (X) does also appear in at least one Qi(Yi). In order to be able
to apply a rule, appropriate individuals have to be substituted for the variables.
Conclusions P (i) are then added to the Abox. For the conclusions P (i) the
M I Agent seeks further explanation using an abduction process. The main idea
is to backward chain a set of rules of the form introduced above. Due to space
restrictions, this process cannot be explained in detail, and we refer to [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. In
any case, if the Aboxes get larger and large, performance will degrade if there
are no speci c techniques employed.
        </p>
        <p>Example 1 A car is shown in a video shot, represented by an assertion Car(c1),
and there is the sound of a door slam, represented by an assertion DoorSlam(ds1).
The car and the door slam are associated to video and audio segments, resp.
Those, in turn, are associated with locator objects. Now assume that the car and
the door slam co-occur, i.e., the locator objects for the audio segment is located
during the video segment. Figure 2 depicts a complete scenario for the example.</p>
        <p>Using relations between time points, one might use rules to de ne a during
relation as a view based on the quantitative temporal information for the locator
objects. However, using relations between time points, in principle, every locator
might be associated with every other locator, and thus the agent can hardly
partition the large Abox into smaller parts. Therefore, we have designed the
agent in such a way that it adds qualitative relations such as overlaps (o), during
(d), and meets (m) to make certain temporal information explicit that is hidden
in the quantitative locator time speci cations. The motivation for the agent to
switch to the more verbose qualitative representation is that the input Abox
becomes partitionable.</p>
        <p>Qualitative temporal relations are used in forward-chaining rules to compute
assertions that are then explained by the agents (see above). For instance, based
on the forward-chaining rule</p>
        <sec id="sec-4-1-1">
          <title>8x; xl; y; yl; w; z V ideoSegment(x); hasSegmentLocator(x; xl);</title>
          <p>V ideoLocator(xl); AudioSegment(y); hasSegmentLocator(y; yl);
AudioLocator(yl); d(yl; xl); depicts(x; w); depicts(y; z);</p>
          <p>
            Car(w); DoorSlam(z) ! CarDoorSlam(w; z)
the role assertion CarDoorSlam(c1; ds1) (marked with an ellipse in Figure 2) is
generated and added to the Abox. This new assertion is seen as a speci c
observation that requires an explanation [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]. Possible explanations, e.g., are car entry
or car exit events, which might be represented using assertions CarEntry(i1) or
CarExit(i2), where i1 and i2 are new individuals. Both individuals are
associated with the car end the door slam individuals (role associatedW ith, see above).
Inevitably, in the course of explanation generation, the Abox grows again
significantly. This leads to very large Aboxes (imagine the annotation of a two-hour
movie) and the application of forward-chaining rules (as well as the abduction
process) will be very ine cient, since complex joins for huge relations can hardly
be avoided in order to check whether rules are applicable (and to compute the
bindings for variables). Pretty soon, the video description Abox does not t into
main memory any longer. In the following, we present a proposal to overcome
the problem of Aboxes becoming too large.
4.2
          </p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Island Reasoning</title>
        <p>
          As stated before, the input can be considered as a stream. The information
content derived from a stream is collected over time and stored together with
the interpretations in an Abox or in multiple ones, respectively, if more than one
interpretation is possible. These Aboxes are put to the previously introduced
agenda A. The more knowledge is gathered, the larger those Aboxes become and
the longer it takes to complete all computations such as applying the
forwardchaining rules or arranging the interpretation process itself. Current
state-of-theart DL reasoning systems cannot deal with this amount of information any more,
because they rely on in-memory structures. To overcome this problem, in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]
island-based reasoning for ALCHI ontologies is proposed as a solution. In the
meantime the island approach is extended to SHIQ(D) by a more ne-grained
syntactical analysis. Since SHIQ(D) is a more expressive description logic than
ALHf R+ (D), the mechanism is also applicable for our annotation language.
        </p>
        <p>
          The underlying idea is that only a small subset of concept and role assertions
called island is necessary to perform instance checking for a particular given
individual i and a given (complex) concept C. The approach chosen here is to
identify role assertions which can be used during the application of a tableau
algorithm for instance checking [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] (note that (T ; A) ? C(i) can be reduced
to checking whether (T [ A [ f:C(i)g) is unsatis able via a tableau algorithm).
First, the ontology is transformed into a normal form, called shallow normal
form. For the details of the transformation please refer to [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Given the shallow
normal form, a so-called 8-info structure for an ontology O is used to determine
which concepts are (worst-case) propagated over role assertions in an Abox. This
helps to de ne a notion of separability. The following de nition of O-separability
is used to determine the importance of role assertions in a given Abox A.
De nition 1. Given an ontology O = (T ; A), a role assertion R(a; b) is called
O-separable, if we have IN C(O) i IN C(hT ; A2gi), where
        </p>
        <p>A2 = A n fR(a; b)g [ fR(a; b0); R(a0; b)g [ fb0 : Cjb : C 2 Ag [ fa0 : Cja : C 2 Ag;
s.t. a0 and b0 are fresh individual names and IN C(O) denotes an inconsistent
ontology O.</p>
        <p>Example 1 (cont.) Applying the de nition of O-separability to the Abox
depicted in Figure 2, islands are computed as shown in Figure 4. Instead of applying
all possible substitutions, the forward-chaining rule does only need to be applied
to the island with the locators vl1 and al1 in order to add CarDoorSlam(c1; ds1).
This enables parallel processing for abduction and retrieval scenarios. However,
given the local range restriction for AudioContent, if as1 is not explicitly
specied as AudioSegment but rather as M ultimediaSegment, the de nition of
Oseparability would be violated for hasM ediaDecomposition(ac1; as1)|so that
the respective island would be larger than before.</p>
        <p>
          This general Abox modularization approach has proven well regarding
scalability issues. For more details, in particular regarding a theoretical and practical
underpinning of the island approach, we refer the reader to [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
        </p>
        <p>hasLogicalDecomposition
vc1‘: VideoContent
md1‘‘: MultimediaDocument
hasLogicalDecomposition
vc1: VideoContent
hasMediaDecomposition
vs1‘: VideoSegment … …
vc1‘‘: VideoContent
hasMediaDecomposition
vs1: VideoSegment
vl2‘: VideoLocator</p>
        <p>
          m
Under the consideration of MPEG-7 and [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], a multimedia content ontology
has been introduced that is represented with the DL ALHf R+ (D). The
expressiveness of this logic has proven to be su cient for the MCO arising from the
scenarios considered in the CASAM project. As the MCO covers most of the
relevant concepts, roles and attributes to be expected in video streaming scenarios,
the ALHf R+ (D) can be safely assumed to be su cient for similar multimedia
interpretation scenarios. Based on time information in Aboxes corresponding
to this MCO, a multimedia agent performs stream-based forward-chaining and
abductive backward chaining in order to obtain interpretation possibilities.
Partitioning techniques ensure that interpretation Aboxes can be decomposed into
manageable parts such that even large videos can be handled (Aboxes can be
swapped to secondary memory).
        </p>
        <p>Some initial experiments were performed to see how the approach behaves in
the CASAM context. The results are very promising and almost all roles were
O-separable after qualitative assertions were added to Aboxes such that
quantitative information is no longer required. Thus, switching from a quantitative to
a qualitative representation provides practical bene ts for the agent.</p>
        <p>
          Our work complements other work on stream reasoning, i.e., for e ciently
maintaining materialized views as described in [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. We show that in some cases
the views based on quantitative information can be avoided.
Hyvonen, E., ten Teije, A., Stuckenschmidt, H., Cabral, L., Tudorache, T., eds.:
The Semantic Web: Research and Applications. Volume 6088 of Lecture Notes in
Computer Science. Springer Berlin / Heidelberg (2010) 1{15
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Castano</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Espinosa</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferrara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karkaletsis</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaya</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Moller, R.,
          <string-name>
            <surname>Montanelli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petasis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wessel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Multimedia interpretation for dynamic ontology evolution</article-title>
          .
          <source>Journal of Logic and Computation</source>
          <volume>19</volume>
          (
          <issue>5</issue>
          ) (
          <year>2008</year>
          )
          <volume>859</volume>
          {
          <fpage>897</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Linked data - the story so far</article-title>
          .
          <source>International Journal on Semantic Web and Information Systems (IJSWIS)</source>
          ,
          <volume>5</volume>
          (
          <issue>3</issue>
          ) (
          <year>2009</year>
          )
          <volume>1</volume>
          {
          <fpage>22</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Espinosa</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaya</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Moller, R.:
          <article-title>The BOEMIE Semantic Browser: A semantic application exploiting rich semantic metadata</article-title>
          .
          <source>In: Proceedings of the Applications of Semantic Technologies Workshop</source>
          (AST-
          <year>2009</year>
          ), Lubeck, Germany. (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. ISO/IEC15938-5FCD:
          <article-title>Multimedia content description interface (MPEG-7) (</article-title>
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>HunterJ</surname>
          </string-name>
          , J.:
          <article-title>Adding multimedia to the semantic web: Building an MPEG-7 ontology</article-title>
          .
          <source>In: Proc. of the 1st Semantic Web Working Symposium</source>
          , Stanford University, California, USA. (
          <year>2001</year>
          ) pp.
          <volume>261</volume>
          {
          <fpage>283</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Staab</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franz</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , Gorlitz,
          <string-name>
            <given-names>O.</given-names>
            ,
            <surname>Saatho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Schenk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Sizov</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <article-title>Lifecycle knowledge management: Getting the semantics across in X-Media</article-title>
          .
          <source>In: Foundations of Intelligent Systems, 15th International Symposium, ISMIS</source>
          <year>2006</year>
          , Bari, Italy,
          <year>September 2006</year>
          . LNCS, Springer 1{
          <fpage>10</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dasiopoulou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dalakleidi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tzouvaras</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kompatsiaris</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <source>D3</source>
          .
          <article-title>4 - Multimedia ontologies</article-title>
          . Boemie, project deliverable, National Technical University of Athens (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Wandelt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Moller, R.:
          <article-title>Updatable island reasoning over alchi ontologies</article-title>
          .
          <source>In: Conference on Knowledge Engineering and Ontology Development (KEOD)</source>
          .
          <source>(2009) CEUR Workshop Proceedings</source>
          (Vol.
          <volume>477</volume>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Gries</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , Moller, R.,
          <string-name>
            <surname>Na</surname>
            <given-names>ssi</given-names>
          </string-name>
          , A.,
          <string-name>
            <surname>Rosenfeld</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sokolski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wessel</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A probabilistic abduction engine for media interpretation</article-title>
          . In Alferes, J.,
          <string-name>
            <surname>Hitzler</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lukasiewicz</surname>
          </string-name>
          , T., eds.
          <source>: Proc. International Conference on Web Reasoning and Rule Systems (RR-2010)</source>
          .
          <article-title>(</article-title>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Haarslev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , Moller, R.:
          <article-title>On the scalability of description logic instance retrieval</article-title>
          .
          <source>Journal of Automated Reasoning</source>
          <volume>41</volume>
          (
          <issue>2</issue>
          ) (
          <year>2008</year>
          )
          <volume>99</volume>
          {
          <fpage>142</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Haarslev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moeller</surname>
          </string-name>
          , R.:
          <article-title>Racer: A core inference engine for the semantic web</article-title>
          .
          <source>In: Proc. of the 2nd International Workshop on Evaluation of Ontology-based Tools, located at the 2nd International Semantic Web Conference ISWC</source>
          . (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Baader</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanschke</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A scheme for integrating concrete domains into concept languages</article-title>
          .
          <source>International Conference on Arti cial Intelligence</source>
          (
          <year>1991</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Baader</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calvanese</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuinness</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nardi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel-Schneider</surname>
          </string-name>
          , P.F., eds.:
          <source>The Description Logic Handbook: Theory, Implementation and Application</source>
          . Cambridge UP: Cambridge, NY. (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Celma</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Semantic integration and retrieval of multimedia metadata</article-title>
          .
          <source>In: Proc. of the 4th International Semantic Web Conference (ISWC)</source>
          , Galway,
          <string-name>
            <surname>Ireland.</surname>
          </string-name>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Allen</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          :
          <article-title>Maintaining knowledge about temporal intervals</article-title>
          .
          <source>Commun. ACM</source>
          <volume>26</volume>
          (
          <issue>11</issue>
          ) (
          <year>1983</year>
          )
          <volume>832</volume>
          {
          <fpage>843</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Wandelt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Moller, R.:
          <article-title>Island Reasoning for ALCHI Ontologies</article-title>
          .
          <source>In: Proceedings of the 5th International Conference on Formal Ontology in Information Systems (FOIS-04)</source>
          , IOS Press (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Wandelt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Moller, R.:
          <article-title>Towards abox modularization of semi-expressive description logics</article-title>
          .
          <source>Journal of Applied Ontology</source>
          <volume>7</volume>
          (
          <issue>2</issue>
          ) (
          <year>2012</year>
          )
          <volume>133</volume>
          {
          <fpage>167</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Barbieri</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Braga</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Della Valle</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grossniklaus</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Incremental reasoning on streams and rich background knowledge</article-title>
          . In Aroyo, L.,
          <string-name>
            <surname>Antoniou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>