<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Emrooz: A Scalable Database for SSN Observations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Markus Stocker</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Narasinha Shurpali</string-name>
          <email>narasinha.shurpali@uef.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kerry Taylor</string-name>
          <email>kerry.taylor@acm.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>George Burba</string-name>
          <email>george.burba@licor.com</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mauno Ronkko</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mikko Kolehmainen</string-name>
          <email>mikko.kolehmaineng@uef.fi</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Biogeochemistry Research Group, Department of Environmental Science, University of Eastern Finland</institution>
          ,
          <addr-line>P.O. Box 1627, 70211 Kuopio</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>College of Engineering and Computer Science, Australian National University</institution>
          ,
          <addr-line>Acton ACT 2601</addr-line>
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Environmental Informatics Research Group, Department of Environmental Science, University of Eastern Finland</institution>
          ,
          <addr-line>P.O. Box 1627, 70211 Kuopio</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>LI-COR</institution>
          ,
          <addr-line>4647 Superior Street, Lincoln, Nebraska</addr-line>
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The design of ontologies for sensor data and metadata has received considerable attention. The most prominent is arguably the Semantic Sensor Network (SSN) ontology. For persistence and retrieval of sensor observations, systems that adopt the SSN ontology most obviously build on an RDF database (triple store). However, large volumes of collected sensor data can be challenging for RDF databases, as the evaluation of SPARQL queries for SSN observations quickly becomes prohibitively expensive. This is arguably due to the fact that triple stores are optimized to e ciently evaluate graph pattern queries, not time series interval queries. As our main contribution, we present Emrooz, a scalable database capable of consuming SSN observations represented in RDF and evaluating queries for SSN observations formulated in SPARQL. We present the Emrooz implementation on Apache Cassandra and Sesame and its performance compared to two state-of-the-art RDF databases. The results show that Emrooz query performance outperforms the two RDF databases by orders of magnitude with increasingly large datasets. We motivate the need for scalable databases for SSN observations on a case study in micrometeorology.</p>
      </abstract>
      <kwd-group>
        <kwd>Sensor Data</kwd>
        <kwd>Data Management</kwd>
        <kwd>Query Performance</kwd>
        <kwd>Ontology</kwd>
        <kwd>RDF</kwd>
        <kwd>SSN ontology</kwd>
        <kwd>Semantic Web</kwd>
        <kwd>Linked Data</kwd>
        <kwd>Emrooz</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Emrooz means `today' in Farsi and is the name of an open source database for
SSN [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] observations represented in RDF [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] capable of evaluating queries for SSN
observations formulated in SPARQL [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Emrooz builds on Apache Cassandra
and Sesame [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which serve in the implementation of Emrooz data and
knowledge stores, respectively. The Emrooz code repository is available on GitHub.5
      </p>
      <p>
        Over the past decade, several authors have proposed ontologies with
formalized vocabulary for describing sensors|their metadata such as observed
properties, operating ranges, and location|and sensor observations|the data collected
from sensors such as the observation value and the time at which the observation
was made. Compton et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] reviewed some of the e orts in the `semantic
specication of sensors'. Today, the most prevalent ontology for the domain of sensing
is arguably the SSN ontology. It has been widely adopted in the literature [6{14].
      </p>
      <p>Ontologies can facilitate the querying, integration and reuse of sensor data
as well as ease the management of large networks with heterogeneous
sensors. Whereas the volume of metadata about sensors is generally comparatively
small|and can thus be easily managed by RDF databases, speci cally triple
stores|the volume of data collected from sensors is often large.</p>
      <p>As we demonstrate in this paper, the performance of state-of-the-art RDF
databases in evaluating SSN observation queries quickly degrades with an
increasing number of observations. This constitutes an obvious and practical
problem for the adoption of the SSN ontology. If unable to answer queries quickly, a
technology that promises semantic interoperability, reasoning, and data linked
to metadata in graph data structures arguably remains a mere theoretical
curiosity. Database systems for SSN observations with good load and excellent query
performance are needed for the technology to become viable in practice.</p>
      <p>Our focus is on environmental sensor networks [15] and their use in earth
and environmental science research. Primarily for this community, we aim at
developing a database capable of consuming SSN observations and fast evaluation
of corresponding queries formulated in SPARQL. The proposed case study is in
micrometeorology, speci cally in monitoring of surface-atmosphere energy and
trace gas uxes using a typical LI-COR Eddy Covariance System. As we discuss
in more details in Section 3, such systems generate large volumes of data,
currently stored as les. Researchers are thus unable to readily retrieve a time series
of arbitrary time interval using a declarative query language. Emrooz attempts
to address this particular concern for scientists who measure surface-atmosphere
uxes by proposing an approach that merely commits to the SSN ontology and
is thus generic with regard to speci c sensors, their data and metadata. In
addition to advantages such as declarative querying and semantic interoperability
of data and metadata, the use of the SSN ontology in Emrooz frees individual
researchers from having to design models and database schemata for their sensor
data and metadata.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Implementation</title>
      <p>Emrooz de nes two main abstractions: the data store and the knowledge store.
The data store supports the persistence of sensor observations. Sensor
observations are represented following the SSN ontology as sets of RDF statements</p>
      <sec id="sec-2-1">
        <title>5 https://github.com/markusstocker/emrooz</title>
        <p>(i.e. triples). Accordingly, a sensor observation relates to the measured value,
the time at which the observation became available, the sensor that made the
observation, and the observed property and feature. The retrieval of sensor
observations is enabled by data store query handlers. A data store query handler
evaluates a set, Qso, of sensor observation queries, qso.</p>
        <p>The knowledge store manages sensor speci cations (metadata). A speci
cation for a sensor de nes the observed property and the sampling frequency.
The observed property relates to a feature. Sensors are speci ed by creating
and relating relevant individuals of SSN classes, typically using an editor such
as Protege.6 The resulting le can be loaded by the knowledge store. A
knowledge store can create query handlers. A knowledge store query handler owns
a data store query handler and evaluates a SPARQL SELECT query for SSN
observations, qssn.</p>
        <p>Queries qssn are formulated by some agent, e.g. a user, and must de ne a
time interval [t1; t2[. The evaluation of such queries occurs in three stages. First,
qssn is translated into a sensor observation query, qso. A sensor observation query
qso (s_; p_; f_; t_1; t_2) consists of parameters for the sensor, s_, property, p_, feature,
f_, and time interval, [t_1; t_2[. Values for these parameters are extracted from qssn
during translation. The parameters may be bound or unbound. Second, qso is
rewritten into a set of sensor observation queries, Qso, that may be a singleton
set. This is the case when qso has de ned sensor, property, and feature. If any
of these parameters is unde ned, then sensor speci cations managed by the
i
knowledge store are utilized to rewrite qso into queries qso 2 Qso so that (1)
each qsio matches the de ned parameters in qso and (2) the de ned tuple (s_; p_; f_)
matches a sensor speci cation. Third, a knowledge store query handler for qssn
and a data store query handler for Qso are composed. The knowledge store query
handler evaluates qssn on the results returned by the data store query handler
in evaluating Qso.</p>
        <p>Emrooz builds its knowledge store implementation on the Sesame
framework for RDF.7 Of particular interest to Emrooz are Sesame repositories and
SPARQL query parsing and evaluation. Sensor speci cations are managed by a
Sesame repository. Depending on the application, the repository may be volatile
or persistent, resident on the local machine or on a remote server. To evaluate
qssn, the knowledge store query handler implementation for Sesame utilizes a
volatile repository initialized with RDF statements returned by the composed
data store query handler.</p>
        <p>The data store implementation builds on Apache Cassandra.8 Sensor
observations are persisted to rows of a data table with schema consisting of partition
key (row key) of type ascii; clustering key (column name) of type timeuuid;
and column value of type blob. The partition key and the clustering key form a
compound primary key.</p>
      </sec>
      <sec id="sec-2-2">
        <title>6 http://protege.stanford.edu</title>
      </sec>
      <sec id="sec-2-3">
        <title>7 http://rdf4j.org/</title>
      </sec>
      <sec id="sec-2-4">
        <title>8 http://cassandra.apache.org</title>
        <p>The partition key consists of two dash-concatenated parts: a SHA-256 hex
string and a date-time string. The hex string is a digest of a message (string)
consisting of dash-concatenated identi ers (URIs) for a sensor, a property, and
a feature. The date-time string follows the pattern `yyyyMMddHHmm'. Given a
sensor observation and the speci cation for the related sensor, the date-time
string is computed from the observation result time, truncated to the year,
month, day, hour, or minute depending on the speci ed sampling frequency. For
instance, for sensors with sampling frequency ]1, 100] Hz the computed date-time
string is truncated to the hour. The date-time string thus limits the number of
sensor observations per partition key for any given (s_; p_; f_) tuple. For a sensing
device with sampling frequency 10 Hz each row holds 36000 sensor observations.</p>
        <p>Sensor observation result times determine clustering keys. Speci cally, given
a sensor observation, the corresponding result time is translated into a time
UUID, which acts as column name. Columns are ordered in time and support
fast time interval scans.</p>
        <p>Column values are byte arrays for the binary-encoded sets of RDF statements
corresponding to sensor observations. The representation of sensor observations
as sets of RDF statements is handled by an RDF entity representer, implemented
in Emrooz. The representer translates entities (Java objects) into corresponding
sets of RDF statements. Sets of RDF statements are then converted to byte
arrays using the Sesame binary RDF writer.</p>
        <p>In addition to the SSN ontology, Emrooz also adopts OWL-Time [16] for
the representation of temporal entities, GeoSPARQL [17] for the representation
of spatial entities, and the Quantities, Units, Dimensions and Data Types
Ontologies (QUDT) [18] for the representation of quantities and units, such as the
sampling frequency in sensor speci cations.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Case Study</title>
      <p>We evaluate Emrooz comparative performance with data of a typical LI-COR
Eddy Covariance System for the direct measurement of CO2, CH4, and H2O
uxes.</p>
      <p>Eddy covariance is a method to directly measure surface-atmosphere uxes
of energy and trace gases. It has been employed to monitor uxes over various
ecosystems and for diverse applications, also in climate change research where
CO2 and CH4 ux measurements by eddy covariance method support
determining whether the observed ecosystem is a carbon sink or source. Large data
volumes for surface-atmosphere uxes of energy and trace gases are managed by
platforms such as ICOS Carbon Portal. 9</p>
      <p>The installation consists of a LI-7500A Open Path CO2/H2O Gas Analyzer,
a LI-7700 Open Path CH4 Analyzer, and a sonic anemometer. The devices
operate at 10 Hz sampling frequency. Collected data is typically stored on a USB
drive of a LI-7550 Analyzer Interface Unit. The components of a LI-COR Eddy</p>
      <sec id="sec-3-1">
        <title>9 https://www.icos-cp.eu/</title>
        <p>Covariance System are often installed on a tripod, which thus acts as a platform
for the devices. For this study, we consider two gas analyzers, the property of
mole fraction, and three features for the monitored gases.</p>
        <p>The data are available in ZIP archive les. Each archive contains text les
with metadata about the site, instruments, and the data les as well as the
data for 30 min of measurement. The time period considered in our experiments
begins on January 7, 2015 and ends on May 26, 2015. The total number of
archive les is 6045. E ectively there should be 6720 archive les for the period
but the dataset is incomplete between March 3 and April 12, during which it
misses 675 archive les.</p>
        <p>For each archive le, the data le of interest is the one containing observation
values for CO2, H2O, and CH4. Except for a header spanning the rst few lines,
this data le consists of a 18 000 40 matrix. The number of rows is equivalent
to the number of 10 Hz samples in 30 min (10 60 30 = 18 000). Of this
matrix, we concentrate on the three columns for measured CO2 [ mol mol 1],
H2O [mmol mol 1], and CH4 [ mol mol 1] plus the two columns for date and
time. Thus, we expect 54 000 sensor observations per matrix, i.e. per 30 min
of measurement. For the January-May period the expected number of sensor
observations is 326 430 000. Considering that each sensor observation maps to
a set of 15 RDF statements (triples) the expected total number of processed
triples is approximately 4.9 billion.</p>
        <p>We evaluate the load and query performance of Emrooz on 10 subsets, for 30
minutes, 1 hour, 3 hours, 6 hours, 12 hours, 1 day, 7 days, 1 month, 3 months,
and the complete dataset (J-M). All subsets begin on January 7, except those
for 1 month (February) and 3 months (February-April). Note that the 3 months
subset is incomplete. Query performance is evaluated using a de ned query
qso (s_; p_; f_; t_1; t_2), whereby f_ is CO2 and the time interval [t_1; t_2[ is 10 min.
The expected result set size of each query is 6000. We evaluate the query
performance as the mean value of three runs per subset. Emrooz performance is
compared with two RDF databases: Stardog10 2.2 and Blazegraph11 1.5.1. For
all three systems, we use the integrated Sesame API to load and query sensor
observations. For both Stardog and Blazegraph we use persistent disk databases
(local triple stores). The disk databases are created rst and data is loaded in
transactions of maximally approximately 2 million triples. We use Apache
Cassandra 2.1.3, Sesame 2.8.1, and Emrooz 0.2.0. The evaluation is performed on a
Fujitsu CELSIUS W420 with an i7-3770 3:40 GHz CPU, 4 8 GB DDR3 1600
MHz DIMM memory modules, and 2 1 TB 7200 RPM SATA hard drives.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results and Discussion</title>
      <p>We rst provide an overview of subset sizes in terms of number of sensor
observations, corresponding triples, and distinct triples. Table 1 summarizes the
numbers. Sensor observations are represented as sets of triples, including triples
10 http://stardog.com
11 http://www.blazegraph.com/bigdata
asserting class membership of sensors, properties, and features. The number of
triples is always 15 times the number of observations. The set of distinct triples
is smaller because it excludes duplicate triples. We also observe that with the 6 h
subset the expected and actual number of sensor observations (and thus triples)
di er. We investigated the reason and found that the data le for January 7 at 4
a.m. misses data for 04:00:53.100. Hence the three missing sensor observations in
the 6 h subset. We suspect that this also explains di erences between expected
and actual number of sensor observations in other subsets.</p>
      <p>Figure 1 summarizes the load performance for the 10 subsets and Emrooz
compared to Stardog and Blazegraph. The gure shows that Emrooz is
outperformed on small datasets. However, with larger datasets Emrooz outperforms
both Stardog and Blazegraph. We attribute this behavior to the apparent
gradually increasing cost of committing transactions in Stardog and Blazegraph.</p>
      <p>Figure 2 summarizes the query performance for the 10 subsets and Emrooz
compared to Stardog and Blazegraph. With constant time at roughly 2:3 s,
Emrooz outperforms both triple stores|by several orders of magnitude for large
datasets. The query performance di erence between Emrooz and the two triple
stores is, however, not surprising. Given a de ned query qso (s_; p_; f_; t_1; t_2),
Apache Cassandra can e ciently retrieve the relevant set of triples by directly
addressing the row key and perform a range scan on column names. The
resulting set of triples is subsequently processed by Sesame using a (volatile)
memory store. In contrast, both Stardog and Blazegraph evaluate the de ned tuple
(s_; p_; f_) as a SPARQL basic graph pattern with expensive joins and resulting in
an intermediate result set corresponding to the complete time series eventually
ltered to the desired interval [t_1; t_2[.</p>
      <p>The results suggest that Emrooz query performance is independent of data
store size. However, query performance is dependent on query time interval
duration. The e ect of varying intervals ranging from 1 s to 60 min is shown in
Figure 3. The query with time interval duration 10 min executes in 2:21 s, which
30 m 1 h 3 h 6 h 12 h 1 d 7 d 1 m 3 m J-M</p>
      <p>Subsets
is comparable to Figure 2 for variable subsets. Queries with shorter time interval
duration evaluate faster while queries with longer time interval duration evaluate
slower. Several factors are at play, in particular the time required to evaluate
qssn on larger memory stores in Sesame post-processing and the time required
to iterate over result sets of increasing size.</p>
      <p>Compared to Stardog and Blazegraph, and comparable triple stores,
Emrooz has some constraints. Most obviously, Emrooz cannot manage arbitrary
RDF data. Furthermore, Apache Cassandra has no means to perform standard
reasoning tasks on SSN observations. O -the-shelf reasoning can only be
performed by Sesame, on the knowledge store and in post-processing query result
sets returned by Apache Cassandra. Some level of reasoning pushed down to the
Apache Cassandra data store can be implemented by Emrooz using the Sesame
knowledge store and query rewriting [19].</p>
      <p>Emrooz is currently capable of evaluating SPARQL queries with a basic
graph pattern for SSN observations, whereby the observation result time must
be constrained by a time interval [t_1; t_2[ speci ed as FILTER. The related sensor,
property, and feature may be bound or unbound. SPARQL features such as
aggregates and solution modi ers such as ORDER BY can be speci ed over [t_1; t_2[.</p>
      <p>Sesame post-processing adds overhead which can be avoided if applications do
not require SPARQL. SPARQL adds exibility, e.g. it enables selecting variables,
ordering or ltering results. However, in some applications this exibility may
not be required and does thus not justify the overhead. For instance, a data
portal may simply want to return the set of RDF statements matching the user
query qso (s_; p_; f_; t_1; t_2) and leave further processing to the user.</p>
      <p>The data in our case study arguably fall into the category of \particularly
hard cases" for triple stores. Assuming equally sized datasets, SSN observation
query evaluation on data collected from sensor networks with more sensors,
properties, and features but sampling at lower frequency are less expensive for
triple stores. This is because the tuple (s_; p_; f_), as well as its elements, are more
30 m 1 h 3 h 6 h 12 h 1 d 7 d 1 m 3 m J-M</p>
      <p>Subsets
selective. The intermediate result sets are smaller and basic graph pattern joins
less expensive. Furthermore, more diverse selectivity estimates for triple patterns
could give query optimizers more room to nd better query plans.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Related and Future Work</title>
      <p>A number of authors have developed RDF data management systems that build
on NoSQL stores such as Apache Cassandra. Cudre-Mauroux et al. [20] provide
a comparative evaluation for the load and query performance of several systems
that implement an RDF data management layer on top of a NoSQL system.
Of particular interest here is CumulusRDF [21], as it also builds on Apache
Cassandra and Sesame. However, these systems aim at being RDF databases and
thus implement indexes specialized for answering arbitrary SPARQL queries on
RDF. In contrast, Emrooz is designed for the management of SSN observations
represented in RDF and for the evaluation of SSN observation queries formulated
in SPARQL. Emrooz is thus designed to be a scalable time series database for
sensor observations represented in RDF according to the vocabulary de ned by
the SSN ontology.</p>
      <p>
        Authors who developed systems for (historical or streamed) sensor data
management have recognized that persisting large volumes of sensor data in an RDF
database is hardly viable. Presenting a platform designed to connect (semantic)
sensor data with data in the `Linked Data Cloud', Le-Phuoc et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] resort to a
relational database management system for historical sensor data management.
Describing a data warehouse for water resource management, Abecker et al. [22]
also propose a hybrid approach in which time series sensor data is managed by
a relational database system (PostGIS) whereas information objects with more
complex relationships are managed by an RDF database. The authors note that
\a complete `semanti cation' [...] of all data [...] seemed not feasible and
promising to us, especially regarding the measurement data."
] 6
s
[em 5
iT 4
3
2
1
01 s 30 s 1 m 5 m 10 m 20 m 30 m 40 m 50 m 60 m
      </p>
      <p>Query time interval</p>
      <p>NoSQL systems have been utilized to manage SSN observations, speci cally.
Wang et al. [23] present a Hadoop-based system designed to manage SSN
observations. The authors describe how their system stores SSN observations to
HBase which, however, features an index structure that is typical for RDF
statements, akin to the systems surveyed by Cudre-Mauroux et al. [20]. Wang et al.
evaluate the performance of various queries. However, to our understanding the
queries are not for SSN observations but rather for single triple patterns, e.g. a
pattern with bound subject and unbound predicate and object. As such, both
the indexing approach and the query performance evaluation are di erent from
those presented in this paper for Emrooz.</p>
      <p>
        There are several potentially interesting directions for future work. First,
we plan to extend the implementation so that it supports the management of
dataset observations represented following the RDF Data Cube (QB) Vocabulary
[24]. With this extension, Emrooz could thus manage not only raw sensor data
but also processed sensor data. For instance, CO2 ux sensor data are used to
compute Net Ecosystem Exchange (NEE). NEE data are a result of sensor data
processing and form a dataset; hence the di erent vocabulary. Combining the
SSN ontology and the QB vocabulary in systems has been demonstrated in the
literature, e.g. [
        <xref ref-type="bibr" rid="ref6">6, 25</xref>
        ].
      </p>
      <p>Second, Emrooz can be equipped with further features. Command line tools
could simplify user interaction with Emrooz. A RESTful service could expose
(load) and query functionality for client-server interaction over HTTP. A
browserbased client could support the visualization of time series. R12 and Matlab13
libraries could enable users to query and persist data from statistical
computing environments. Data managed by Emrooz could then be loaded into R data
12 http://www.r-project.org/
13 www.mathworks.com/products/matlab/
frames. Results of computations in R could be persisted in Emrooz. Finally,
Emrooz could be enhanced with more exibility in SPARQL query formulation, e.g.</p>
      <p>lter for a set of properties, as well as reasoning at query time. These
enhancements could be supported by extending the existing query rewriting mechanism.
A detailed analysis of the SPARQL expressivity covered in Emrooz may also be
of interest.</p>
      <p>Third, it is interesting to compare Emrooz performance with triple stores
that build on SQL and NoSQL databases, such as SDB14 and CumulusRDF [21],
respectively, as well as with non-triple stores, including SQL or OGC standards
compliant databases, such as PostgreSQL15 and 52North,16 respectively.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>We have presented Emrooz, a scalable database for SSN observations
represented in RDF capable of evaluating queries for SSN observations formulated in
SPARQL. We brie y discussed how Emrooz builds on Apache Cassandra and
Sesame for its implementations of a data store and a knowledge store,
respectively, and how the two stores interact. Emrooz is motivated by the following two
contrasting aspects. On one hand, the attractiveness of the RDF data model and
the SSN ontology for representing metadata about sensors and what is sensed,
as well as for representing data resulting in sensor measurement, is an argument
for adopting these technologies in systems. On the other hand, the most obvious
approach to SSN observations management using triple stores seems to fail on a
fundamental and important requirement, i.e. scalable fast evaluation of SSN
observation queries. As we demonstrated for two state-of-the-art triple stores, SSN
observation query evaluation becomes quickly prohibitive as store size grows
to tens of millions sensor observations. To serve client applications with SSN
observations in RDF is attractive for several reasons, including data linked to
metadata and formal descriptions of vocabulary semantics. However, for
practical viability the underlying SSN observations management system needs to be
designed for time series query evaluation.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>This research is funded by the Academy of Finland project \FResCo:
Highquality Measurement Infrastructure for Future Resilient Control Systems" (Grant
number 264060).
14 https://jena.apache.org/documentation/sdb/
15 http://www.postgresql.org/
16 http://52north.org/
14. Rinne, M., amd Robin Keskisarkka, E.B., Nuutila, E.: Event Processing in RDF.</p>
      <sec id="sec-7-1">
        <title>In Gangemi, A., Gruninger, M., Hammar, K., Lefort, L., Presutti, V., Scherp, A.,</title>
        <p>eds.: Proceedings of the 4th Workshop on Ontology and Semantic Web Patterns.</p>
      </sec>
      <sec id="sec-7-2">
        <title>Volume 1188., Sydney, Australia, CEUR-WS (October 2013)</title>
        <p>15. Martinez, K., Hart, J.K., Ong, R.: Environmental Sensor Networks. Computer
37(8) (2004) 50{56
16. Hobbs, J.R., Pan, F.: Time Ontology in OWL. Working draft, W3C (September
2006)
17. Perry, M., Herring, J.: OGC GeoSPARQL - A Geographic Query Language for</p>
      </sec>
      <sec id="sec-7-3">
        <title>RDF Data. Technical Report OGC 11-052r4, Open Geospatial Consortium Inc.</title>
        <p>(September 2012)
18. Hodgson, R., Keller, P.J., Hodges, J., Spivak, J.: Qudt { quantities, units,
dimensions and data types ontologies. Technical report, TopQuadrant, Inc (March
2014)
19. Kontchakov, R., Zakharyaschev, M.: An Introduction to Description Logics and</p>
      </sec>
      <sec id="sec-7-4">
        <title>Query Rewriting. In Koubarakis, M., Stamou, G., Stoilos, G., Horrocks, I., Ko</title>
        <p>laitis, P., Lausen, G., Weikum, G., eds.: Reasoning Web. Reasoning on the Web in
the Big Data Era. Volume 8714 of Lecture Notes in Computer Science. Springer</p>
      </sec>
      <sec id="sec-7-5">
        <title>International Publishing (2014) 195{244</title>
        <p>20. Cudre-Mauroux, P., Enchev, I., Fundatureanu, S., Groth, P., Haque, A., Harth,</p>
      </sec>
      <sec id="sec-7-6">
        <title>A., Keppmann, F.L., Miranker, D., Sequeda, J.F., Wylot, M.: NoSQL Databases</title>
        <p>for RDF: An Empirical Evaluation. In Alani, H., Kagal, L., Fokoue, A., Groth, P.,</p>
      </sec>
      <sec id="sec-7-7">
        <title>Biemann, C., Parreira, J.X., Aroyo, L., Noy, N., Welty, C., Janowicz, K., eds.: The</title>
      </sec>
      <sec id="sec-7-8">
        <title>Semantic Web { ISWC 2013. Volume 8219 of Lecture Notes in Computer Science.</title>
      </sec>
      <sec id="sec-7-9">
        <title>Springer Berlin Heidelberg (2013) 310{325</title>
        <p>21. Ladwig, G., Harth, A.: CumulusRDF: Linked Data Management on Nested
Key</p>
      </sec>
      <sec id="sec-7-10">
        <title>Value Stores. In Fokuoe, A., Liebig, T., Guo, Y., eds.: The 7th International Work</title>
        <p>shop on Scalable Semantic Web Knowledge Base Systems (SSWS 2011), Bonn,
Germany (October 2011) 30{42
22. Abecker, A., Brauer, T., Magoutas, B., Mentzas, G., Papageorgiou, N.,
Quenzer, M.: A Sensor and Semantic Data Warehouse for Integrated Water Resource</p>
      </sec>
      <sec id="sec-7-11">
        <title>Management. In Gomez, J.M., Sonnenschein, M., Vogel, U., Winter, A., Rapp,</title>
      </sec>
      <sec id="sec-7-12">
        <title>B., Giesen, N., eds.: Proceedings of the 28th Conference on Environmental Infor</title>
        <p>matics - Informatics for Environmental Protection, Sustainable Development and</p>
      </sec>
      <sec id="sec-7-13">
        <title>Risk Management (EnviroInfo 2014), Oldenburg, Germany, BIS-Verlag, Oldenburg</title>
        <p>(September 2014) 517{524
23. Wang, D., Zhang, X., Gao, H.: HDSW: Semantic Sensor Network System Based on</p>
      </sec>
      <sec id="sec-7-14">
        <title>Hadoop. International Journal of Multimedia and Ubiquitous Engineering 9(12)</title>
        <p>(2014) 61{72
24. Cyganiak, R., Reynolds, D., Tennison, J.: The RDF Data Cube Vocabulary.
Recommendation, W3C (January 2014)
25. Stocker, M., Baranizadeh, E., Portin, H., Komppula, M., Ronkko, M., Hamed,</p>
      </sec>
      <sec id="sec-7-15">
        <title>A., Virtanen, A., Lehtinen, K., Laaksonen, A., Kolehmainen, M.: Representing</title>
        <p>situational knowledge acquired from sensor data for atmospheric phenomena.
Environmental Modelling &amp; Software 58 (2014) 27{47</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Compton</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barnaghi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bermudez</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garca-Castro</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corcho</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cox</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Graybeal</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hauswirth</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henson</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herzog</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Janowicz</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelsey</surname>
            ,
            <given-names>W.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Phuoc</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lefort</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leggieri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neuhaus</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nikolov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Page</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Passant</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , K.:
          <article-title>The SSN ontology of the W3C semantic sensor network incubator group</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>17</volume>
          (
          <issue>0</issue>
          ) (
          <year>2012</year>
          )
          <volume>25</volume>
          {
          <fpage>32</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wood</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lanthaler</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>RDF 1.1 Concepts and Abstract Syntax</article-title>
          . Recommendation,
          <source>W3C (February</source>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seaborne</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>SPARQL 1.1 Query Language</article-title>
          . Recommendation,
          <source>W3C (March</source>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Broekstra</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kampman</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>van Harmelen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Sesame: A Generic Architecture for Storing and Querying RDF and RDF Schema</article-title>
          . In Horrocks, I.,
          <string-name>
            <surname>Hendler</surname>
          </string-name>
          , J., eds.:
          <source>The Semantic Web { ISWC 2002. Volume 2342 of Lecture Notes in Computer Science</source>
          . Springer Berlin Heidelberg (
          <year>2002</year>
          )
          <volume>54</volume>
          {
          <fpage>68</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Compton</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henson</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lefort</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neuhaus</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A Survey of the Semantic Speci cation of Sensors</article-title>
          . In Taylor, K.,
          <string-name>
            <surname>Ayyagari</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roure</surname>
          </string-name>
          , D.D., eds.
          <source>: Proceedings of the 2nd International Workshop on Semantic Sensor Networks</source>
          . Volume
          <volume>522</volume>
          .,
          <string-name>
            <surname>Washington</surname>
            <given-names>DC</given-names>
          </string-name>
          , USA, CEUR-WS (
          <year>October 2009</year>
          )
          <volume>17</volume>
          {
          <fpage>32</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Lefort</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bobruk</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haller</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , K.,
          <string-name>
            <surname>Woolf</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A Linked Sensor Data Cube for a 100 Year Homogenised Daily Temperature Dataset</article-title>
          . In Henson,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Taylor</surname>
          </string-name>
          , K.,
          <string-name>
            <surname>Corcho</surname>
          </string-name>
          , O., eds.
          <source>: Proceedings of the 5th International Workshop on Semantic Sensor Networks</source>
          . Volume
          <volume>904</volume>
          ., Boston, Massachusetts, CEUR-WS (
          <year>November 2012</year>
          )
          <volume>1</volume>
          {
          <fpage>16</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Le-Phuoc</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quoc</surname>
            ,
            <given-names>H.N.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parreira</surname>
            ,
            <given-names>J.X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hauswirth</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The Linked Sensor Middleware|Connecting the real world and the Semantic Web</article-title>
          .
          <source>In: Proceedings of the Semantic Web Challenge at the 10th International Semantic Web Conference</source>
          , Bonn, Germany (
          <year>October 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Muller, H.,
          <string-name>
            <surname>Cabral</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morshed</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>From RESTful to SPARQL: A Case Study on Generating Semantic Sensor Data</article-title>
          . In
          <string-name>
            <surname>Corcho</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henson</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barnaghi</surname>
          </string-name>
          , P., eds.
          <source>: Proceedings of the 6th International Workshop on Semantic Sensor Networks</source>
          . Volume
          <volume>1063</volume>
          ., Sydney, Australia, CEUR-WS (
          <year>October 2013</year>
          )
          <volume>51</volume>
          {
          <fpage>66</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Calbimonte</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jeung</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corcho</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aberer</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Deriving Semantic Sensor Metadata from Raw Measurements</article-title>
          . In Henson,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Taylor</surname>
          </string-name>
          , K.,
          <string-name>
            <surname>Corcho</surname>
          </string-name>
          , O., eds.
          <source>: Proceedings of the 5th International Workshop on Semantic Sensor Networks</source>
          . Volume
          <volume>904</volume>
          ., Boston, Massachusetts, USA, CEUR-WS (
          <year>November 2012</year>
          )
          <volume>33</volume>
          {
          <fpage>48</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gould</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , K.:
          <article-title>Linked Data Approach for Automated Failure Detection in Sewere Rising Mains Using Real-Time Sensor Data</article-title>
          .
          <source>In: Proceedings of the 11th International Conference on Hydroinformatics</source>
          , New York City, USA (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Representing and Inferring Events from Deforestation Observations</article-title>
          . In Gensel, J.,
          <string-name>
            <surname>Josselin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vandenbroucke</surname>
          </string-name>
          , D., eds.
          <source>: Proceedings of the AGILE'2012 International Conference on Geographic Information Science</source>
          , Avignon, France (
          <year>April 2012</year>
          )
          <volume>80</volume>
          /392{85/392
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Llaves</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuhn</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>An event abstraction layer for the integration of geosensor data</article-title>
          .
          <source>International Journal of Geographical Information Science</source>
          <volume>28</volume>
          (
          <issue>5</issue>
          ) (
          <year>2014</year>
          )
          <volume>1085</volume>
          {
          <fpage>1106</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , K.,
          <string-name>
            <surname>Leidinger</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Ontology-Driven Complex Event Processing in Heterogeneous Sensor Networks</article-title>
          . In Antoniou, G.,
          <string-name>
            <surname>Grobelnik</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simperl</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parsia</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plexousakis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leenheer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pan</surname>
          </string-name>
          , J., eds.:
          <source>The Semanic Web: Research and Applications. Volume 6644 of Lecture Notes in Computer Science</source>
          . Springer Berlin Heidelberg (
          <year>2011</year>
          )
          <volume>285</volume>
          {
          <fpage>299</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>