<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Adaptive Framework for RDF Stream Reasoning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Qiong Li</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiaowang Zhang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhiyong Feng</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guohui Xiao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Computer Science, Free University of Bozen-Bolzano</institution>
          ,
          <addr-line>Bolzano I-39100</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computer Science and Technology, Tianjin University</institution>
          ,
          <addr-line>Tianjin 300350</addr-line>
          ,
          <country country="CN">P. R. China</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Computer Software,Tianjin University</institution>
          ,
          <addr-line>Tianjin 300350</addr-line>
          ,
          <country country="CN">P. R. China</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Tianjin Key Laboratory of Cognitive Computing and Application</institution>
          ,
          <addr-line>Tianjin 300350</addr-line>
          ,
          <country country="CN">P.R. China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we propose an adaptive framework for RDF stream reasoning (PRSPR) in order to obtain more meaningful and valuable information, which is an extension of our previous work. Moreover, our work is a kind of plug-in framework which makes it more adaptive and flexible. Within this framework, not only can we apply all kinds of SPARQL query engines to process RDF streams, but also simultaneously support various inference engines for RDFS/OWL for stream reasoning. Finally, we experimentally evaluate the performance of PRSPR on YABench. The experiments show that PRSPR can still maintain the high performance with SPARQL query engines in RDF stream reasoning although there are some slight differences among them.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        RDF stream, as a new type of dataset, can model real-time and continuous information
in a wide range of applications, e.g., environmental monitoring and smart city [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] we already presented an adaptive framework PRSP to process RDF streams by
exploiting various SPARQL query engines in a brief way. However, if we want to obtain
more detailed and abounding information, it is necessary to address the problem of
performing reasoning for very dynamic inputs. There are many approaches about RDFS
reasoning over static RDF graphs, but rare over RDF streams. Liu et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] proposed
a method over static data to enhance the performance of rule-based OWL reasoning
on Spark by exploiting a locally optimal executable strategy. Chang et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] present a
approach to perform stream reasoning on RDF data using the GPU computing
architecture, but it reasons streams after processing streams.
      </p>
      <p>In this paper, we provide an adaptive framework for RDF stream reasoning named
PRSPR which is an extension of our PRSP. In our work, we further optimize PRSP
by improving the performance. On the other hand, we apply the locally optimal
executable strategy to implement stream reasoning. Therefore, PRSPR not only makes it
possible to use the high-performance SPARQL query engines to process large-volume
RDF streams, but also makes it easier for stream reasoning by applying all kinds of
RDF/RDFS and OWL entailment rules in a convenient way.</p>
    </sec>
    <sec id="sec-2">
      <title>Preliminaries</title>
      <p>RDF stream An RDF stream S is defined as ordered sequences of pairs, made of an
RDF triple and a timestamp : (hs; p; oi; ).</p>
      <p>Continuous Query Formally, a continuous SPARQL query Q can be taken as a 5-tuple
of the form:</p>
      <p>Q = [Reg; S; w; s; (Q)]
where
– Reg: the registration;
– S: the RDF stream registered;
– w: RANGE, i.e., the window size;
– s: STEP, i.e., the updating time of windows;
– (Q): a SPARQL query.</p>
      <p>RDF Schema RDF Schema (RDFS) is a set of classes with certain properties using
the RDF extensible knowledge representation data model, providing basic elements for
the description of ontologies, otherwise called RDF vocabularies, intended to structure
RDF resources.
3</p>
    </sec>
    <sec id="sec-3">
      <title>A framework for RDF stream reasoning</title>
      <p>PRSPR is a framework for querying and reasoning both RDF graphs and RDF streams
shown in Figure 1. Both continuous query and RDF streams as the input of PRSPR,
they continuously reason and process by exploiting the following four modules: Query
Preparation, Data Transformer, RDFS Reasoner and Query Execution. We give a
detailed description about the workflow below.</p>
      <p>RDF</p>
      <sec id="sec-3-1">
        <title>Stream</title>
      </sec>
      <sec id="sec-3-2">
        <title>Continuous</title>
      </sec>
      <sec id="sec-3-3">
        <title>Query RDF</title>
        <p>Data Transformer Graph</p>
      </sec>
      <sec id="sec-3-4">
        <title>RDFS</title>
      </sec>
      <sec id="sec-3-5">
        <title>Reasoner</title>
        <p>RDF(S)</p>
      </sec>
      <sec id="sec-3-6">
        <title>Graph</title>
      </sec>
      <sec id="sec-3-7">
        <title>Window</title>
      </sec>
      <sec id="sec-3-8">
        <title>Selector</title>
      </sec>
      <sec id="sec-3-9">
        <title>Query</title>
      </sec>
      <sec id="sec-3-10">
        <title>Preparation</title>
      </sec>
      <sec id="sec-3-11">
        <title>SPARQL</title>
      </sec>
      <sec id="sec-3-12">
        <title>Query RDF</title>
      </sec>
      <sec id="sec-3-13">
        <title>Graph</title>
      </sec>
      <sec id="sec-3-14">
        <title>Query</title>
      </sec>
      <sec id="sec-3-15">
        <title>Execution</title>
      </sec>
      <sec id="sec-3-16">
        <title>Result</title>
        <p>Query Preparation Continuous queries, as the input of Query Preparation mode,
generate two types of queries, namely, SPARQL query and Window Selector, which can be
addressed in the Query Execution and Data Transformer module respectively.
Data Transformer Data Transformer module processes RDF streams. It transforms
RDF streams into RDF graphs based on the window size and step size set at Window
Selector. And the output, RDF graphs are as the input of RDFS Reasoning module.
RDFS Reasoner RDFS Reasoner is responsible for reasoning RDF graphs based on the
rule-based knowledge ontology on Spark. The reasoning computes the deductive
closure of an ontology by applying RDF/RDFS entailment rules in order to make implicit
knowledge explicit, i.e., obtaining RDF(S) graphs.</p>
        <p>Query Execution PRSPR defines a unified interface for SPARQL query engines, which
makes it possible and easy for SPARQL query engines to process RDF streams.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments and Evaluations</title>
      <p>
        All centralized experiments were carried out on a machine running Linux, which has
4 CPUs with 6 cores and 64GB memory, and 4 machines with the same performance
for distributed experiments. The version of Spark is 1.5.2. We utilized YABench RSP
benchmark[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], and registered LUBM as the RDF streams. The complexity of the
scenarios was in the ascending order, from the least complex configuration (LUBM200)
that loaded roughly 28 million triples to the most complex configuration (LUBM1000)
that injected more than 130 million triples. In our experiments, we perform sliding
windows with a 2400-seconds-window which slides every 2300 seconds, and chose the
two queries over LUBM, i.e., Q6 and Q10. The reason is that the two queries can not
return results over LUBM unless the data has been inferred by reasoners. We use the
OWL-Horst rules[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] as the reasoner in our experiments.
300 400 500 1000
      </p>
      <p>LUBM/the number of university
gStore RDF-3X gStoreD TriAD S2RDF
(a) Triples loading time
300 400 500 1000</p>
      <p>LUBM/the number of university
gStore RDF-3X gStoreD TriAD S2RDF
(b) Query response time
300 400 500 1000</p>
      <p>LUBM/the number of university
gStore RDF-3X gStoreD TriAD S2RDF
(c) Engine execution time</p>
      <p>We compare the performance of the five different SPARQL query engines, including
two centralized engines (RDF-3X and gStore) and three distributed systems (gStoreD,
TriAD and S2RDF) in a unified way. The RDFS reasoning time (shown in table 1), is
increased in varying degrees with the growth of dataset. The processing time of the two
queries within PRSPR is shown in Fig 2 and Fig 3, respectively. gStoreD and S2RDF
can not process the whole triples in the window at the specific time over LUBM1000,
which results in some window triples lose and the results are incomplete. When the
load ranges from LUBM200 to LUBM1000, the triples loading time, query response
time and engine execution time are increasing. But we can also get that the distributed
engines (TriAD and S2RDF) have the better performance than centralized engines.
1 104 50 1 104
300 400 500 1000</p>
      <p>LUBM/the number of university
gStore RDF-3X gStoreD TriAD S2RDF
(a) Triples loading time
300 400 500 1000</p>
      <p>LUBM/the number of university
gStore RDF-3X gStoreD TriAD S2RDF
(b) Query response time
300 400 500 1000</p>
      <p>LUBM/the number of university
gStore RDF-3X gStoreD TriAD S2RDF
(c) Engine execution time</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In this paper, we present the PRSPR, as a plugin adaptable for various SPARQL query
engines and reasoning machines, which makes the system more adaptive. Moreover,
PRSPR is for both RDF streaming and stream reasoning, so that it can obtain more
valuable information. Therefore, it can also process large-volume RDF streams by
applying distributed SPARQL query engines.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We thank Bo Zhao for his assistance in the experiments of this paper. This work is
supported by the programs of the National Natural Science Foundation of China (61672377),
the National Key R&amp;D Program of China (2016YFB1000603), and the Key Technology
Research and Development Program of Tianjin (16YFZCGX00210).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Barbieri</surname>
            ,
            <given-names>D. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Braga</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Della Valle</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Grossniklaus</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Querying RDF streams with C-SPARQL</article-title>
          .
          <source>SIGMOD REC</source>
          .
          <volume>39</volume>
          (
          <issue>1</issue>
          ),
          <fpage>20</fpage>
          -
          <lpage>26</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kolchin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wetz</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesling</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Tjoa</surname>
            ,
            <given-names>A. M.:</given-names>
          </string-name>
          <article-title>YABench: A comprehensive framework for RDF stream processor correctness and performance assessment</article-title>
          .
          <source>In: Proc. of ICWE '16</source>
          , pp.
          <fpage>280</fpage>
          -
          <lpage>298</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>PRSP: A plugin-based framework for RDF stream processing</article-title>
          .
          <source>In: Proc. of WWW'17</source>
          , poster, pp.
          <fpage>815</fpage>
          -
          <lpage>816</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Liu</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Urbani</surname>
            <given-names>J</given-names>
          </string-name>
          , and Qi G.:
          <article-title>Efficient RDF stream reasoning with graphics processingunits (GPUs)</article-title>
          .
          <source>In: Proc. of WWW '14</source>
          , pp.
          <fpage>343</fpage>
          -
          <lpage>344</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Rao</surname>
          </string-name>
          , G.:
          <article-title>RORS: Enhanced Rule-Based OWL Reasoning on Spark</article-title>
          .
          <source>In: Proc. of APWeb '16</source>
          , pp.
          <fpage>444</fpage>
          -
          <lpage>448</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>