<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards E cient Processing of RDF Data Streams - Short Paper</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alejandro Llaves</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Javier D. Fernandez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oscar Corcho</string-name>
          <email>ocorchog@fi.upm.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Introduction to RDF Stream Processing</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ontology Engineering Group, Universidad Politecnica de Madrid</institution>
          ,
          <addr-line>Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the last years, there has been an increase in the amount of real-time data generated. Sensors attached to things are transforming how we interact with our environment. Extracting meaningful information from these streams of data is essential for some application areas and requires processing systems that scale to varying conditions in data sources, complex queries, and system failures. This paper describes ongoing research on the development of a scalable RDF streaming engine.</p>
      </abstract>
      <kwd-group>
        <kwd>RDF Stream Processing</kwd>
        <kwd>Adaptive Query Processing</kwd>
        <kwd>RDF Compression</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Extracting information from data streams is complex because of the
heterogeneity of the data, the rate of data generation, the high volumes, and the often
unclear data provenance. One use case would be real-time monitoring of
public transportation in a city. Here, decisions on unexpected events, such as a car
crash, should be taken on short time slots based on a set of spatio-temporal data
streams coming from di erent providers. For instance, by diverting a bus line
route. This means reasoning over data in the temporal order it is ingested by
the processing system. Linked Stream Data may help to integrate datasets from
di erent providers and to solve interoperability problems. Yet, remaining
challenges require solutions from a stream processing approach. In the remainder of
the paper we discuss some of these challenges (section 2), a general description
of our approach towards e cient processing of RDF streams (section 3), and
open questions to debate during the workshop (section 4).
2</p>
    </sec>
    <sec id="sec-2">
      <title>Challenges on RDF Stream Processing</title>
      <p>The development of scalable services for real-time stream processing involves
support for high throughput, management of complex queries, low latency
response, fault-tolerance, and statistics extraction, among others. Nowadays, cloud
services o er solutions to some of these problems, e.g. by applying elastic load
balancing in presence of input data bursts. We mainly focus on the e cient
processing of user queries over RDF streams (C.1), which requires
parallelization at the query operator level.</p>
      <p>Distributed computing refers to the processing of data in distributed systems.
The continuous transmission of data (C.2) between sources and processing
nodes, and among nodes may cause response delays that should be minimized
when possible.</p>
      <p>The integration of historical and real-time data with background
knowledge (C.3) is challenging in Web-scale environments. Many RSP systems
combine background knowledge with real-time processing, but historical data
management is often overlooked [15]. Traditional systems tend to store data
aggregates in order to save space, but with the capacity of current systems it
is possible (and recommended) to store all data in raw format and de ne views
on data batches [16]. The e cient management of historical data is essential to
detect trends in data, extract statistics, or compare old data with current data
to identify anomalies [15]. For instance, by aggregating metro users on the last
hours and compare the numbers to the average of users during the last days.
3</p>
      <p>E</p>
      <p>cient Processing of Queries over RDF Streams
This section gives an high level overview of our approach, which is motivated by
the challenges described above. We address three aspects of stream processing:
adaptivity in query processing, data compression, and the architecture choice
for integrating real-time and historical data.
3.1</p>
      <sec id="sec-2-1">
        <title>Adaptive query processing for data streams</title>
        <p>Heterogeneous data streams are generated from di erent sources, at di erent
rates, and include multiple domains. Our purpose is to build a distributed stream
processing engine capable of adapting to changing conditions while serving
complex continuous queries. Some sources already generate Linked Data streams [3].
Otherwise, we provide a layer serving an ontology-based access to non RDF data
stream sources. Adapters for various input formats, such as CSV or REST APIs,
are used to convert heterogeneous streams to RDF.</p>
        <p>To reach an e cient query processing over data streams (section 2, C.1) we
will focus on query execution planning. Traditional databases include a query
optimizer that designs an execution plan based on the registered query and data
statistics. In a distributed stream processing environment, there are several
aspects to contemplate: changing rates of the input data, failure of processing
nodes, and distribution of workload, among others. Adaptive Query Processing
(AQP) techniques [7] allow adjusting the query execution plan to varying
conditions of the data input, the incoming queries, and the system. Additionally, it is
used to correct query optimizer mistakes and cope with unknown statistics [2].</p>
        <p>First, we will analyze di erent strategies to process query operators, such as
JOIN or FILTER. There are various examples in the literature that use ordering
approaches for a more scalable stream processing, e.g. in the implementation of
JOIN operators [12] or in eviction strategies [11]. We will design Storm8
topologies to e ciently process a set of common operators based on parallelizable tasks.
Storm topologies are formed by spouts and bolts. Spouts are stream sources,
whereas bolts are stream processors. Trident is a high-level abstraction on top of
Storm that allows for stateful stream processing.9 For instance, a windowed join
can be implemented using stream snapshots, which in Trident are called states.
Depending on criteria such as the selectivity of the join or the size of snapshots,
the processing engine can decide on the more appropriate join operator in terms
of e ciency. Then, we will de ne a list of queries for a speci c use case and will
extend the topologies to t the queries. While new data is entering the system, a
dedicated bolt will manage stream data statistics to reassign a di erent topology
if data stream conditions vary. In Storm, the coordination between the master
node, which controls the assignation of tasks to spouts and bolts, and the worker
nodes is managed by Zookeeper.10 However, Zookeeper does not provide elastic
load balancing o -the-shelf. The use of cloud services for this purpose, such as
the one o ered by Amazon EC2, will be addressed in the near future.
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Compressing RDF streams</title>
        <p>To date, universal compressors (e.g. gzip) and speci c RDF compressors (e.g.
HDT [9]) are commonly used to reduce RDF exchange costs and delays on
the network. These approaches, though, consider a static view of RDF datasets,
8 http://storm.incubator.apache.org/
9 https://storm.incubator.apache.org/documentation/Trident-tutorial.html
10 http://zookeeper.apache.org/
disregarding the dynamic nature of RDF stream management. A recent work [10]
points out the importance of e cient RDF stream compression and proposes
an initial solution leveraging the inherent data redundancy of the RDF data
streams. Based on this proposal, we are currently working on a compressed data
structure speci cally designed for the particularities of dynamic RDF streams
[8]. In particular, we aim at providing a lightweight serialization of RDF streams
which i) minimizes the data exchange among processing nodes (section 2, C.2)
while ii) serving a small set of operators on the compressed data.
Our engine will address real-time processing on the Web of Things context
following the Lambda principles [16]. Lambda is a 3-layer architecture designed to
t the requirements of Big Data: a batch layer (using Hadoop11) stores all the
incoming data in an immutable master dataset and pre-computes batch views on
11 http://hadoop.apache.org/
historic data (section 2, C.3); a serving layer (NoSQL database) indexes views
on the master dataset; and a speed layer manages the real-time processing issues
and requests data views depending on incoming queries. The cloud platform to
deploy our engine should provide computing services that scale on demand, as
well as elastic load balancing.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Open Questions</title>
      <p>Next steps on the short term will address the implementation of an scalable RDF
stream processing system. Open questions that we would like to discuss at the
workshop are:
{ How does the order of tuple arrival a ect the parallelization of join processing
tasks?
{ Are the spatial (or spatio-temporal) properties of a tuple a dimension to
have into account for ordering? In this case, what in uence does it have on
reasoning tasks? And on parallelization tasks?
{ How does the out-of-order tuples a ect the processing of streams? In case of
discarding them, how to communicate this decision in the results?</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>
        The research leading to these results has received funding from the European
Union's Seventh Framework Programme (FP7/2007-2013) under grant
agreement no. 257641, PlanetData network of excellence, and from Ministerio de
Economa y Competitividad (Spain) under the project "4V: Volumen, Velocidad,
Variedad y Validez en la Gestin Innovadora de Datos" (TIN2013-46238-C4-2-R).
5. Calbimonte, J.P., Jeung, H., Corcho, O., Aberer, K.: Enabling Query
Technologies for the Semantic Sensor Web. International Journal on Semantic Web and
Information Systems 8(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), 43{63 (2012)
6. Compton, M., Barnaghi, P., Bermudez, L., Garcia-Castro, R., Corcho, O., Cox,
S., Graybeal, J., Hauswirth, M., Henson, C., Herzog, A., Huang, V., Janowicz,
K., Kelsey, W.D., Phuoc, D.L., Lefort, L., Leggieri, M., Neuhaus, H., Nikolov, A.,
Page, K., Passant, A., Sheth, A., Taylor, K.: The ssn ontology of the w3c semantic
sensor network incubator group. Web Semantics: Science, Services and Agents on
the World Wide Web 17(0) (2012)
7. Deshpande, A., Ives, Z., Raman, V.: Adaptive Query Processing. Foundations and
      </p>
      <p>
        Trends in Databases 1(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), 1{140 (Jan 2007)
8. Fernandez, J.D., Llaves, A., Corcho, O.: E cient RDF Interchange (ERI) Format
for RDF Data Streams (accepted). In: ISWC 2014 (2014)
9. Fernandez, J.D., Mart nez-Prieto, M.A., Gutierrez, C., Polleres, A., Arias, M.:
Binary RDF representation for publication and exchange (HDT). Web Semantics:
Science, Services and Agents on the World Wide Web 19(0), 22{41 (2013)
10. Fernandez Garc a, N., Arias Fisteus, J., Sanchez Fernandez, L., Fuentes-Lorenzo,
D., Corcho, O.: RDSZ : An approach for lossless RDF stream compression. In: In
11th European Semantic Web Conference (ESWC 2014). pp. 1{16. Crete, Greece
(2014)
11. Gao, S., Scharrenbach, T., Bernstein, A.: The CLOCK Data-Aware Eviction
Approach: Towards Processing Linked Data Streams with Limited Resources. In: The
11th Extended Semantic Web Conference. Lecture Notes in Computer Science,
Springer (May 2014)
12. Golab, L., Ozsu, M.T.: Processing Sliding Window Multi-joins in Continuous
Queries over Data Streams. In: Proceedings of the 29th International Conference
on Very Large Data Bases - Volume 29. pp. 500{511. VLDB '03, VLDB Endowment
(2003)
13. Le-phuoc, D., Nguyen, H., Quoc, M., Van, C.L., Hauswirth, M.: Elastic and
Scalable Processing of Linked Stream Data in the Cloud. In: Alani, H., Kagal, L.,
Fokoue, A., Groth, P., Biemann, C., Parreira, J.X., Aroyo, L., Noy, N., Welty, C.,
Janowicz, K. (eds.) The Semantic Web ISWC 2013, vol. 287305, pp. 280{297.
      </p>
      <p>
        Springer Berlin Heidelberg (2013)
14. Le-phuoc, D., Parreira, J.X., Hauswirth, M.: Linked Stream Data Processing. In:
Reasoning Web. Semantic Technologies for Advanced Query Answering, pp. 245{
289. Springer Berlin Heidelberg (2012)
15. Margara, A., Urbani, J., van Harmelen, F., Bal, H.: Streaming the web: Reasoning
over dynamic data. Web Semantics: Science, Services and Agents on the World
Wide Web 0(0) (0)
16. Marz, N., Warren, J.: Big Data: principles and best practices of scalable realtime
data systems. Manning Publications (2013)
17. Rinne, M., Nuutila, E., Seppo, T.: INSTANS: High-Performance Event Processing
with Standard RDF and SPARQL. In: Henson, C., Taylor, K., Corcho, O. (eds.)
Proceedings of the 5th International Workshop on Semantic Sensor Networks. pp.
81{96. CEUR-WS, Boston, Massachusetts, USA (2012)
18. Sequeda, J.F., Corcho, O.: Linked Stream Data: A Position Paper. In: Proceedings
of the 2nd International Workshop on Semantic Sensor Networks, SSN 09.
CEURWS, Washington, USA (2009)
19. Sheth, A., Henson, C., Sahoo, S.S.: Semantic Sensor Web. IEEE Internet
Computing 12(
        <xref ref-type="bibr" rid="ref4">4</xref>
        ), 78{83 (2008)
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Anicic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fodor</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rudolph</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stojanovic</surname>
          </string-name>
          , N.:
          <article-title>EP-SPARQL: A Uni ed Language for Event Processing and Stream Reasoning</article-title>
          .
          <source>In: Proceedings of the 20th International Conference on World Wide Web</source>
          . pp.
          <volume>635</volume>
          {
          <fpage>644</fpage>
          . WWW '11,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Babu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizarro</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Adaptive Query Processing in the Looking Glass</article-title>
          .
          <source>In: Proceedings of the Second Biennial Conference on Innovative Data Systems Research (CIDR)</source>
          ,
          <year>Jan</year>
          .
          <year>2005</year>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Balduini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Della</given-names>
            <surname>Valle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>DellAglio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Tsytsarau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Palpanas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Confalonieri</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          :
          <article-title>Social Listening of City Scale Events Using the Streaming Linked Data Framework</article-title>
          . In: Alani,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Kagal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Fokoue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Groth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Biemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Parreira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.X.</given-names>
            ,
            <surname>Aroyo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Noy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Welty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Janowicz</surname>
          </string-name>
          ,
          <string-name>
            <surname>K</surname>
          </string-name>
          . (eds.)
          <source>The Semantic Web ISWC</source>
          <year>2013</year>
          , pp.
          <volume>1</volume>
          {
          <fpage>16</fpage>
          . Springer Berlin Heidelberg (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Barbieri</surname>
            ,
            <given-names>D.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Braga</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valle</surname>
            ,
            <given-names>E.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grossniklaus</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>C-sparql: a continuous query language for rdf data streams</article-title>
          .
          <source>Int. J. Semantic Computing</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ),
          <volume>3</volume>
          {
          <fpage>25</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>