<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Reasoning with Large Data Sets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Darko Anicic</string-name>
          <email>darko.anicic@deri.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Digital Enterprise Research Institute (DERI), University of Innsbruck</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Efficient reasoning is a critical factor for successful Semantic Web applications. In this context, applications may require vast volumes of data to be processed in a short time. We develop novel reasoning techniques which will extend current reasoning methods as well as existing database technologies in order to enable large scale reasoning. We propose advances and key design principles primarily in: making an efficient query execution plan as well as in memory, storage and recovery management. Our study is being implemented in Integrated Rule Inference System (IRIS) - a reasoner for Web Service Modeling Language.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The Web Service Modeling Language WSML1 is a language framework for
describing various aspects related to Semantic Web (SW) services. We are
developing IRIS2 to serve as a WSML reasoner which handles large workload efficiently.</p>
      <p>
        Current inference systems exploit reasoner methods developed rather for
small knowledge bases [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. These systems3, although utilize mature and efficient
relational database management systems (RDBMSs) and exploit a number of
their evaluation strategies (e.g., query planning, caching, buffering etc.), cannot
meet requirements for reasoning in complex SW applications. Reason for this
is found in the fact that database techniques are rather developed for explicitly
represented data, and need to be extended for dealing with implicit knowledge.
      </p>
      <p>
        In this work we investigate a framework which generalizes relational databases
by adding deductive capabilities to them. RDBMSs suffer some limitations w.r.t
the expressivity of their language. Full support for recursive views is one of them
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Further on, negation as failure is recognized as a very important
nonmonotonic property for the Semantic Web. RDBMSs, although deal with negation
as failure, can not select a minimal fixpoint that reflects the intended meaning
in situations where the minimal fixpoint may not be unique. Our framework,
although exceeding capabilities of RDBMSs, does not compromise their
performance.
      </p>
      <p>Current reasoners cannot cope with large data sets (i.e., relations larger than
system main memory). Hence a reasoner needs to deal effectively with portions of
1 WSML: http://www.wsmo.org/TR/d16/d16.1/v0.2/.
2 IRIS: http://sourceforge.net/projects/iris-reasoner/.
3 Reasoners which utilize persistant storage: KAON2, Aditi, InstanceStore, DLDB.
relations (possible distributed over many machines), and sophisticated strategies
for partition-level relation management are required. Consequently, a relevant
topic for our present and future work is: The development of effective
optimization algorithms as well as distribution and memory management strategies for
reasoning with large data sets.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Efficient Large Scale Reasoning: an Approach</title>
      <p>We will now give a short overview of our approach to achieving effective reasoning
with large data sets.</p>
      <p>
        Unlike other inference systems4, which utilize SQL to access existential
relations, we tightly integrate IRIS with its storage layer (i.e., rules are translated
into relational algebra expressions and SQL is avoided as an unnecessary
overhead). We extend embedded RDBMS query optimizer (which is rather designed
to be used for extensional data) for derived relations. The estimation of the size
and evaluation cost of the intensional predicates will be based on the adaptive
sampling method [
        <xref ref-type="bibr" rid="ref1 ref4 ref6">4, 1</xref>
        ], while the extensional data will be estimated using a
graph-based synopses of data sets similarly as in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Further on, for reasoning
with large relations, run time memory overflow may occur. Therefore in IRIS
we are developing novel techniques for a selective pushing of currently processed
tuples to disk. This technique will be further extended for data distributed over
many disks (e.g., a cluster of machines). Such techniques aim to enable IRIS to
effectively handle large workload which cannot fit in main memory of the system.
      </p>
      <p>Our framework comprises a recovery manager and thus features fault-tolerant
architecture. Using logging and replications we ensure that, when a crash occurs,
the system may continue with an ongoing operation without loss of previously
computed results.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgment</title>
      <p>I am grateful to Michael Kifer and my supervisors: Stijn Heymans and Dieter
Fensel for their help in the work conceptualization and insightful discussions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>M. E. Vidal</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Ruckhaus</surname>
            and
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Ruiz</surname>
          </string-name>
          .
          <article-title>Query evaluation and optimization in the semantic web</article-title>
          .
          <source>In ALPSWS2006 Workshop</source>
          , Washington, USA.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Dieter</given-names>
            <surname>Fensel</surname>
          </string-name>
          and Frank van Harmelen.
          <article-title>Unifying reasoning and search to web scale</article-title>
          .
          <source>IEEE INTERNET COMPUTING, page 3</source>
          , 2
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Michael</given-names>
            <surname>Kifer</surname>
          </string-name>
          , Arthur Bernstein, and
          <string-name>
            <surname>Philip</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lewis</surname>
          </string-name>
          .
          <source>Database Systems: An Application Oriented Approach. Addison-Wesley</source>
          , Boston, MA, USA,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Lipton</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Naughton</surname>
          </string-name>
          .
          <article-title>Query size estimation by adaptive sampling</article-title>
          .
          <source>In PODS '90</source>
          , NY, USA.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J.</given-names>
            <surname>Spiegel</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Polyzotis</surname>
          </string-name>
          .
          <article-title>Graph-based synopses for relational selectivity estimation</article-title>
          .
          <source>In SIGMOD '06</source>
          , NY, USA.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>4 KAON2, QUONTO, InstanceStore and DLDB exploit SQL for querying</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>