<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Short Paper: Assessing the Quality of Semantic Sensor Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chris Baillie</string-name>
          <email>c.baillie@abdn.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Edwards</string-name>
          <email>p.edwards@abdn.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edoardo Pignotti</string-name>
          <email>e.pignotti@abdn.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Corsar</string-name>
          <email>dcorsar@abdn.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computing Science and dot.rural Digital Economy Hub, University of Aberdeen</institution>
          ,
          <addr-line>Aberdeen</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Sensors are increasingly publishing observations to the Web of Linked Data. However, assessing the quality of such data remains a major challenge for agents (human and machine). This paper describes how Qual-O, a vocabulary for describing quality assessment, can be used to perform quality assessment on semantic sensor data.</p>
      </abstract>
      <kwd-group>
        <kwd>Semantic Web</kwd>
        <kwd>Linked Data</kwd>
        <kwd>ontology</kwd>
        <kwd>quality</kwd>
        <kwd>provenance</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The number of sensors publishing data to the Web has increased dramatically, a
trend that is set to accelerate further with the growth of the Internet of Things.
However, in order to identify reliable datasets agents must rst perform data
quality assessment [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Data quality is de ned in terms of ` tness for use' and
is assessed against one or more dimensions of quality, such as timeliness and
accuracy using a set of quality metrics [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Such metrics typically produce values
between 0 (indicating low quality) and 1 (high quality) by examining the context
around data [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Some sensors now publish their data to the Web of Linked Sensor
Data using vocabularies such as the Semantic Sensor Network (SSN) ontology [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
to describe the context around observations, such as the time observations were
generated and the feature the measured phenomenon (e.g. temperature) a ects
(e.g. a geographical region). In this paper we will demonstrate how a generic
quality assessment framework can use this context to facilitate data quality
assessment. Moreover, we will argue that this context should be further enriched
to include a description of sensor data provenance: a record of the agents,
entities, and activities involved in data derivation. Capturing such metadata using
a vocabulary such as the W3C PROV1 recommendation enables users to
better understand, trust, reproduce, and validate the data available on the Web
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Any quality assessment should thus examine data provenance as there are a
number of provenance dimensions that can a ect the quality of data. For
example, derivation: Was the observation derived from any other observations? and
attribution: Who was associated with the generation of this observation?.
      </p>
      <sec id="sec-1-1">
        <title>1 http://www.w3.org/TR/prov-overview/</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Qual-O: An Ontology for Quality Assessment</title>
      <p>In this section, we present the Qual-O ontology2. This model enables the
specication of quality assessments that can examine the semantic representation of
data (e.g. using SSN) and their provenance (e.g. using PROV). Moreover, this
model can be used to describe the provenance of such assessments. For the
remainder of this paper, we refer to the provenance used in assessment as subject
provenance and the provenance of past assessments as QA provenance.</p>
      <p>
        The concepts in Figure 1 preceded by the qual namespace characterise a
minimal set of concepts for describing quality assessment, in uenced by
existing quality vocabularies such as the DQM ontology [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. We adopt a di erent
approach to these vocabularies insofar as we base our model on the PROV
vocabulary to ensure it is capable of describing QA provenance. PROV is de ned
in terms of three main concepts: Entity (a physical, digital, conceptual, or other
kind of thing with some xed aspects), Activity (something that occurs over a
period of time and acts upon or with entities), and Agent (something that bears
some form of responsibility for an activity). However, Entity alone is insu cient
to characterise metrics and subjects as they are simply used by an Activity to
generate a further entity. Therefore, we use subclasses of prov:Role to describe
the function of an entity with respect to an activity, e.g. subject, metric, and
result. A qual:Subject is thus an Entity with a subject role, a qual:Metric is an
Entity with a metric role, and a qual:Result is an Entity with result role. It then
follows that a qual:Assessment is a kind of Activity that used one qual:Metric,
one qual:Subject, and generated one qual:Result. The relations between each qual
concept can also be de ned in terms of PROV: targets and guidedBy are
subproperties of prov:used and describe the relationship between an Assessment and
a Subject and Metric, respectively; resultOf is a sub-property of wasGeneratedBy
and attributes a Result to its Assessment.
      </p>
      <p>
        De ning a quality assessment model in this way has a number of advantages.
Firstly, using OWL2 RL3 enables the provenance of quality assessment to be
inferred based on the concepts used to de ne the assessment. Secondly, the
assessment activity is described including the time the assessment started and
ended, and the kind of platform that performed QA (e.g. some computational
service). Thirdly, we can attribute assessment results to the agent associated
with QA to form descriptions of why the assessment was performed using agent
intent [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] characterised as a set of goals and constraints (int namespace in Figure
1). For example, an agent can only use sensor data with a score of 0.75 to
achieve a goal, decide whether to take a jacket with them on a walk. Finally, we
can associate QA results with speci c subgraphs in the subject; for example, to
describe the precise property and value in either the subject description or its
provenance record. This could facilitate more complex QA re-use queries, e.g.
`select completed assessments with an accuracy score greater than 0.75 a ecting
location data generated by a GPS device and not mobile mast location '.
      </p>
      <sec id="sec-2-1">
        <title>2 http://sensornet.abdn.ac.uk/onts/Qual-O.ttl</title>
      </sec>
      <sec id="sec-2-2">
        <title>3 http://www.w3.org/TR/owl2-pro les/</title>
        <p>int:wasBasedOn
int:wasBasedOn
int:Goal
int:Constraint
prov:hadRole
prov:Role prov:wasDerivedFrom prov:Entity
qual:result qual:metric qual:subject</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Assessing the Quality of Semantic Sensor Data</title>
      <p>To investigate how our quality model performs it was rst necessary to obtain
some sensor data, described using a suitable ontology. To this end, we have
developed a number of sensing devices based on the Arduino electronics
prototyping platform. Each is equipped with sensors capable of describing, for example,
temperature, humidity, location and acceleration (via GPS). Figure 2 provides
an example of how we describe sensing devices and their observations using
the SSN ontology. The Arduino is an instance of Platform with a number of
attached SensingDevices. These devices produce ssn:Observations describing a
speci c Property (e.g. Temperature). The observed real-world phenomenon (e.g.
Edinburgh, UK as described by DBPedia4) is captured using FeatureOfInterest
and ObservationValue is used to describe the measured value.</p>
      <p>Using this framework, we have produced a number of datasets containing
sensor observations describing the environmental conditions in a number of use cases
in April 2013. CarCommuter (D1) describes a car journey between Ellon and
Aberdeen, UK; CityWalk (D2) describes a pedestrian journey in Edinburgh, UK;
TrainJourney (D3) describes a train journey between Aberdeen and Edinburgh;
CoastalWalk (D4) describes a recreational walk on a beach near Aberdeen; and
WeatherStation (D5) describes a 24 hour period in a xed location near
Aberdeen. We have developed a web service5 that enables the visualisation of each
dataset and provides the URL of each dataset's SPARQL endpoint. Clicking on
individual observations within a visualisation triggers assessment of the selected
observation as a Subject. The example in Figure 2 extends the logic of a Metric
using SPIN6 to calculate the average temperatures for the feature of interest</p>
      <sec id="sec-3-1">
        <title>4 http://www.dbpedia.org</title>
      </sec>
      <sec id="sec-3-2">
        <title>5 http://sensornet.abdn.ac.uk:8080/SensorBox</title>
      </sec>
      <sec id="sec-3-3">
        <title>6 http://www.spinrdf.org</title>
        <p>at the time the observation was produced. In this example, the observation
describes the temperature of Edinburgh, UK where the average of the high and low
temperature in April, according to DBPedia, is 7 C. The metric then states that
observation quality, in terms of Consistency, decreases the further its value is
from 7 C. In this example, the observation has a value of 22.4 C and therefore
is annotated with a quality score of 0.3125 (indicating a low quality
observation). It should be noted that this represents only one method of computing a
consistency score for this observation; other agents may have others depending
on their intended use for the data.</p>
        <p>int:Decision
"This data is not appropriate int:wasBasedOn
for my intended use"</p>
        <p>int:wasBasedOn
prov:Role
qual:result
prov:hadRole
prov:Entity
dtp:Rslt234</p>
        <p>int:Goal
"Decide whether to wear a jacket
while walking in Edinburgh"</p>
        <p>int:Constraint
"Onlyscuosreedoaft0a.7w5ithoragcroenasteisr"tency</p>
        <p>As noted earlier in the paper, documenting QA provenance can enable agents
to better understand the outcomes of quality assessment. Reasoning about
quality using Qual-O allows a reasoner to infer QA provenance as an assessment is
performed. Such a provenance record can link a Subject Entity with the Metric
Entity used to assess it via the Assessment Activity. Furthermore, an Agent 's
intent can also be captured (boxes with int namespace in Figure 2). In this
example, the agent has a constraint that they can only use data with a minimum
consistency score of 0.75. Quality assessment has produced a consistency score
of 0.3125 for this example observation and so it is reasonable to conclude that
this agent will not use this observation to achieve its goal, decide whether to
wear a jacket while walking in Edinburgh.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>To investigate the performance of our quality assessment framework, we
conducted a series of experiments using a Sun Fire X4100 M2 with two dual-core
AMD Opteron 2218 CPUs and 32GB of RAM running CentOS 5.8, Java SE1.6,
JENA 2.10 and SPIN 1.3. Each experiment used a set of quality metrics (Table 1)
to assess 600 observations across datasets 2, 4 and 5. For each metric we created
two SPIN rules: one that examines the context around observations described
using only SSN; the other uses subject provenance to assess quality based on the
observations the subject was derived from. Experiment 1 measured the
reasoning time required to apply quality metrics, expressed using Qual-O, to individual
observations. Figure 3 shows that using metric set 2 resulted in a signi cant
increase in reasoning time. This is only to be expected as the amount of metadata
to examine has increased. However, set 2 was able to identify quality issues that
set 1 could not as these issues were only present in the provenance record, e.g.
that an observation was derived from another, low quality, observation.
Experiment 2 compared the overhead of performing quality assessment (M1) with the
overhead of performing quality assessment and capturing its provenance (M2).
These results (Figure 4) demonstrate an increase in the reasoning time required
to document quality assessment provenance due to the reasoner having to infer
more triples during the assessment.</p>
      <p>Dimension Description</p>
      <sec id="sec-4-1">
        <title>Consistency Average April temperature7 for Aberdeen and Edinburgh is around 7 C.</title>
      </sec>
      <sec id="sec-4-2">
        <title>Consistency Average humidity8 for Aberdeen and Edinburgh is around 79%.</title>
        <sec id="sec-4-2-1">
          <title>Believability GPS observations should have been produced using at least 3 satellites.</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>Accuracy GPS observations should have an error margin less than 100 metres.</title>
        </sec>
        <sec id="sec-4-2-3">
          <title>Believability There are no areas of Scotland below sea level and dry.</title>
        </sec>
        <sec id="sec-4-2-4">
          <title>Accuracy At least 4 GPS satellites are required to measure altitude.</title>
        </sec>
        <sec id="sec-4-2-5">
          <title>Reasoning Time by Metadata Type</title>
        </sec>
        <sec id="sec-4-2-6">
          <title>Reasoning Time by Model</title>
          <p>)s 120
(m 100
e 80
im 60
T
D2</p>
          <p>D4</p>
        </sec>
        <sec id="sec-4-2-7">
          <title>Dataset D5 Set 1 Set 2</title>
          <p>s 100
)
m
(
e 50
m
i
T
M1
M2</p>
        </sec>
        <sec id="sec-4-2-8">
          <title>Metric Set</title>
          <p>Set 1
Set 2</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>7 http://dbpedia.org/page/Aberdeen and http://dbpedia.org/page/Edinburgh</title>
      </sec>
      <sec id="sec-4-4">
        <title>8 http://www.currentresults.com/Weather/UnitedKingdom/humidity-annual.php</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions &amp; Future Work</title>
      <p>
        Through the experimental results presented here and the experience gained
deploying our framework in a real-world application [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], we have been able to
demonstrate that quality assessment can be performed on sensor observations by
examining observations (described using SSN) and their provenance (described
using PROV). This should enable agents to make decisions about which datasets
(sensor or otherwise) are reliable based on the quality results produced by our
framework. Moreover, we have shown that inclusion of observation provenance
during quality assessment causes an increase in required reasoning time. While
this is a considerable increase, we argue that this is o set by the richer
descriptions of assessment provenance that can be achieved using this method. Agents
should be able to examine this provenance to better understand how
assessments were performed and potentially make decisions about re-use of existing
assessment results. Our future work will therefore investigate how existing
quality results can be re-used to potentially reduce the overhead associated with
QA. This work will include investigating how agents can use speci c elements of
QA provenance to make re-use decisions, e.g. if QA was performed by an agent
that they trust. Further, we aim to strengthen our evaluation by repeating our
experiments with di erent sets of metrics, di erent hardware speci cations, and
larger observation models.
      </p>
      <p>Acknowledgements The research described here is supported by the award made by the RCUK
Digital Economy programme to the dot.rural Digital Economy Hub; award reference: EP/G066051/1</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Baillie</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pignotti</surname>
          </string-name>
          , E.:
          <article-title>Quality reasoning in the semantic web</article-title>
          .
          <source>In: The Semantic Web - ISWC 2012. Lecture Notes in Computer Science</source>
          , vol.
          <volume>7650</volume>
          , pp.
          <volume>383</volume>
          {
          <issue>390</issue>
          (November
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cygniak</surname>
          </string-name>
          , R.:
          <article-title>Quality-driven information ltering using the wiqa policy framework</article-title>
          .
          <source>Journal of Web Semantics</source>
          <volume>7</volume>
          ,
          <issue>1</issue>
          {
          <fpage>10</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Compton</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>The SSN ontology of the W3C Semantic Sensor Network Incubator Group</article-title>
          , vol.
          <volume>17</volume>
          , pp.
          <volume>25</volume>
          {
          <fpage>32</fpage>
          . Web Semantics: Science,
          <source>Services and Agents on the World Wide Web</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Corsar</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baillie</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markovic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papangelis</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nelson</surname>
          </string-name>
          , J.:
          <article-title>Short paper: Citizen sensing within a real time passenger information system</article-title>
          .
          <source>In: 6th International Workshop on Semantic Sensor Networks</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Furber</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hepp</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Using semantic web resources for data quality management</article-title>
          .
          <source>In: 17th International Conference on Knowledge Engineering and Knowledge Management</source>
          . pp.
          <volume>211</volume>
          {
          <issue>225</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Miles</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Munroe</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moreau</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Prime: A methodology for developing provenance-aware applications</article-title>
          .
          <source>ACM Transactions on Software Engineering and Methodology</source>
          <volume>20</volume>
          (
          <issue>3</issue>
          ),
          <volume>39</volume>
          {46 (
          <year>June 2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Pignotti</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gotts</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polhill</surname>
          </string-name>
          , G.:
          <article-title>Enhancing work ow with a semantic description of scienti c intent</article-title>
          .
          <source>Journal of Web Semantics</source>
          <volume>9</volume>
          ,
          <issue>222</issue>
          {
          <fpage>244</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strong</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Beyond accuracy: what data quality means to data consumers</article-title>
          .
          <source>In: Journal of Management Information Systems</source>
          . vol.
          <volume>12</volume>
          , pp.
          <volume>5</volume>
          {
          <issue>33</issue>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>