<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>10th International Workshop on Uncertainty Reasoning for the Semantic Web (URSW 2014)</article-title>
      </title-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>37</fpage>
      <lpage>78</lpage>
      <kwd-group>
        <kwd>Proceedings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Foreword</title>
      <p>This volume contains the papers presented at the 10th International Workshop
on Uncertainty Reasoning for the Semantic Web (URSW 2014), held as a part of the
13th International Semantic Web Conference (ISWC 2014) at Riva del Garda, Italy,
October 19, 2014. 4 technical papers and 1 position paper were accepted at URSW
2014. Furthermore, there was a special session on Methods for Establishing Trust of
(Open) Data (METHOD 2014) including 1 technical paper and 2 position papers.
All the papers were selected in a rigorous reviewing process, where each paper was
reviewed by three program committee members.</p>
      <p>The International Semantic Web Conference is a major international forum for
presenting visionary research on all aspects of the Semantic Web. The International
Workshop on Uncertainty Reasoning for the Semantic Web provides an opportunity
for collaboration and cross-fertilization between the uncertainty reasoning
community and the Semantic Web community.</p>
      <p>We wish to thank all authors who submitted papers and all workshops
participants for fruitful discussions. We would like to thank the program committee
members for their timely expertise in carefully reviewing the submissions.</p>
      <p>September 2014</p>
      <p>URSW 2014</p>
    </sec>
    <sec id="sec-2">
      <title>Workshop Organization</title>
      <sec id="sec-2-1">
        <title>Organizing Committee</title>
        <p>Fernando Bobillo (University of Zaragoza, Spain)
Rommel Carvalho (Universidade de Brasilia, Brazil)
Paulo C. G. da Costa (George Mason University, USA)
Davide Ceolin (VU University Amsterdam, The Netherlands)
Claudia d'Amato (University of Bari, Italy)
Nicola Fanizzi (University of Bari, Italy)
Kathryn B. Laskey (George Mason University, USA)
Kenneth J. Laskey (MITRE Corporation, USA)
Thomas Lukasiewicz (University of Oxford, UK)
Trevor Martin (University of Bristol, UK)
Matthias Nickles (National University of Ireland, Ireland)
Michael Pool (Goldman Sachs, USA)</p>
      </sec>
      <sec id="sec-2-2">
        <title>Program Committee</title>
        <p>Fernando Bobillo (University of Zaragoza, Spain)
Rommel Carvalho (Universidade de Brasilia, Brazil)
Davide Ceolin (VU University Amsterdam, The Netherlands)
Paulo C. G. da Costa (George Mason University, USA)
Fabio Gagliardi Cozman (Universidade de Sa~o Paulo, Brazil)
Claudia d'Amato (University of Bari, Italy)
Nicola Fanizzi (University of Bari, Italy)
Marcelo Ladeira (Universidade de Bras lia, Brazil)
Kathryn B. Laskey (George Mason University, USA)
Kenneth J. Laskey (MITRE Corporation, USA)
Thomas Lukasiewicz (University of Oxford, UK)
Trevor Martin (University of Bristol, UK)
Alessandra Mileo (DERI Galway, Ireland)
Matthias Nickles (National University of Ireland, Ireland)
Je Z. Pan (University of Aberdeen, UK)
Rafael Pen~aloza (TU Dresden, Germany)
Michael Pool (Goldman Sachs, USA)
Livia Predoiu (University of Mannheim, Germany)
Guilin Qi (Southeast University, China)
David Robertson (University of Edinburgh, UK)
Thomas Scharrenbach (University of Zurich, Switzerland)
V
Giorgos Stoilos (National Technical University of Athens, Greece)
Umberto Straccia (ISTI-CNR, Italy)
Matthias Thimm (Universitat Koblenz-Landau, Germany)
Peter Vojtas (Charles University Prague, Czech Republic)</p>
      </sec>
      <sec id="sec-2-3">
        <title>Organizing Committee</title>
        <p>Davide Ceolin (VU University Amsterdam, The Netherlands)
Tom De Nies (Ghent University, Belgium)
Paul Groth (VU University Amsterdam, The Netherlands)
Olaf Hartig (University of Waterloo, Canada)
Stephen Marsh (University of Ontario Institute of Technology, Canada)</p>
      </sec>
      <sec id="sec-2-4">
        <title>Program Committee</title>
        <p>Edzard H g (Beuth University of Applied Sciences Berlin, Germany)
Rino Falcone (Institute of Cognitive Sciences and Technologies, Italy)
Jean-Francois Lalande (INSA Centre Val de Loire, France)
Zhendong Ma (Austrian Institute of Technology, Austria)
Uwe Nestmann (Technische Universitat Berlin, Germany)
Florian Skopik (Austrian Institute of Technology, Austria)
Gabriele Lenzini (University of Luxembourg, Luxembourg)
Erik Mannens (Ghent University, Belgium)
Matthias Flgge (Fraunhofer FOKUS, Germany)
Trung Dong Huynh (University of Southampton, UK)
Paolo Missier (Newcastle University, UK)
Khalid Belhajjame (Paris-Dauphine University, France)
James Cheney (University of Edinburgh, UK)
Christian Bizer (Mannheim University Library, Germany)
Yolanda Gil (University of Southern California, USA)
Daniel Garijo (Universidad Politecnica de Madrid, Spain)
Tim Lebo (Rensselaer Polytechnic Institute, USA)
Simon Miles (Kings College of London, UK)
Andreas Schreiber (German Aerospace Center, Cologne, Germany)
Kieron O'Hara (University of Southampton, UK)
Tim Storer (University of Glasgow, Scotland)
Babak Esfandiari (Carleton University, Canada)
Christian Damsgaard Jensen (Technical University of Denmark, Denmark)
Khalil El-Khatib (University of Ontario Institute of Technology, Canada)
Kai Eckert (Mannheim University Library, Germany)</p>
        <p>VII</p>
        <p>Table of Contents</p>
        <sec id="sec-2-4-1">
          <title>URSW 2014 Technical Papers</title>
          <p>{ A Probabilistic OWL Reasoner for Intelligent Environments</p>
          <p>David Aus n, Diego Lopez-De-Ipin~a and Federico Castanedo
{ Learning to Propagate Knowledge in Web Ontologies</p>
          <p>Pasquale Minervini, Claudia d'Amato, Nicola Fanizzi, Volker Tresp
{ Automated Evaluation of Crowdsourced Annotations in the Cultural
Heritage Domain
Archana Nottamkandath, Jasper Oosterman, Davide Ceolin, Wan
Fokkink
{ Probabilistic Relational Reasoning in Semantic Robot Navigation
Walter Toro, Fabio Cozman, Kate Revoredo, Anna Helena Reali
Costa</p>
        </sec>
        <sec id="sec-2-4-2">
          <title>URSW 2014 Position Papers</title>
          <p>{ Towards a Distributional Semantic Web Stack
Andre Freitas, Edward Curry, Siegfried Handschuh
1-12
13-24
25-36
37-48
49-52
53-54
55-66
67-72
73-78
Research Papers
Short Papers</p>
        </sec>
        <sec id="sec-2-4-3">
          <title>Special Session: Methods for Establishing Trust of (Open) Data</title>
          <p>Overview
{ Overview of METHOD 2014: the 3rd International Workshop on
Methods for Establishing Trust of (Open) Data</p>
          <p>Tom De Nies, Davide Ceolin, Paul Groth, Olaf Hartig, Stephen Marsh
{ Hashing of RDF Graphs and a Solution to the Blank Node Problem</p>
          <p>Edzard Hoe g, Ina Schieferdecker
{ Rating, Recognizing and Rewarding Metadata Integration and Sharing
on the Semantic Web</p>
          <p>Francisco Couto
{ Towards the De nition of an Ontology for Trust in (Web) Data
Davide Ceolin, Archana Nottamkandath, Wan Fokkink, Valentina
Maccatrozzo
IX</p>
        </sec>
      </sec>
      <sec id="sec-2-5">
        <title>A Probabilistic OWL Reasoner for Intelligent</title>
      </sec>
      <sec id="sec-2-6">
        <title>Environments</title>
        <p>David Aus´ın1, Diego L´opez-de-Ipin˜a1, and Federico Castanedo2
Abstract. OWL ontologies have gained great popularity as a context
modelling tool for intelligent environments due to their expressivity.
However, they present some disadvantages when it is necessary to deal
with uncertainty, which is common in our daily life and affects the
decisions that we take. To overcome this drawback, we have developed a
novel framework to compute fact probabilities from the axioms in an
OWL ontology. This proposal comprises the definition and description
of our probabilistic ontology. Our probabilistic ontology extends OWL
2 DL with a new layer to model uncertainty. With this work we aim
to overcome OWL limitations to reason with uncertainty, developing a
novel framework that can be used in intelligent environments.
1</p>
        <p>
          Introduction
In Ambient Intelligence applications, context can be defined as any data which
can be employed to describe the state of an entity (a user, a relevant object, the
location, etc.) [
          <xref ref-type="bibr" rid="ref16 ref6">6</xref>
          ]. How this information is modelled and reasoned over time is
a key component of an intelligent environment in order to assist users in their
daily activities or execute the corresponding actions. An intelligent environment
is any space in which daily activities are enhanced by computation [
          <xref ref-type="bibr" rid="ref14 ref4">4</xref>
          ].
        </p>
        <p>
          One of the most popular techniques for context modelling is OWL ontologies
[18]. They have been employed in several Ambient Intelligence projects such as
SOUPA[
          <xref ref-type="bibr" rid="ref13 ref3">3</xref>
          ], CONON[20] or CoDAMoS [15], to name a few.
        </p>
        <p>
          OWL is the common way to encode description logics in real world. However,
when the domain information contains uncertainty, the employment of OWL
ontologies is less suitable [
          <xref ref-type="bibr" rid="ref21">11</xref>
          ]. The need to handle uncertainty has created a
growing interest in the development of solutions to deal with it.
        </p>
        <p>As in other domains, uncertainty is also present in Ambient Intelligence [16]
and affects to the decision making process. This task requires context information
? This work is supported by the Spanish MICINN project FRASEWARE
(TIN201347152-C3-3-R)
in order to respond to the users’ needs. Data in Ambient Intelligence applications
are provided by several sensors and services in real time. Unfortunately, these
sensors can fail, run out of battery or be forgotten by the user, in the case of
wearable devices. On the other hand, the services can also be inaccessible due
to network connectivity problems or technical difficulties on the remote server.
Nonetheless, that unavailable information may be essential to answer correctly
user’s requirements.</p>
        <p>For this reason, we present a novel approach to deal with uncertainty in
intelligent environments. This work proposes a method to model uncertainty,
that combines OWL ontologies with Bayesian networks. The rest of this article
is organized as follows. The next section describes the problem that we address.
Section 3 explains the semantics and syntax of our proposal. Section 4 gives an
exemplary use case where our proposal is applied and describes how to model
it. Finally, section 5 summarizes this work and addresses the future work.
2</p>
        <p>Description of the Problem
In intelligent environments, the lack of information causes incomplete context
information and it may be produced by several causes:
– Sensors that have run out of batteries. Several sensors, such as wearable
devices, depend on batteries to work.
– Network problems. Sensors, actuators and computers involved in the
environment sensing and monitoring are connected to local networks that can
suffer network failures. In these cases, the context information may be lost,
although the sensors, actuators and computers are working properly.
– Remote services’ failures. Some systems rely on remote services to provide
a functionality or to gather context information.
– A system device stops working. Computer, sensors and actuators can suffer
software and hardware failures that hamper their proper operation.</p>
        <p>When one of these issues occurs, the OWL reasoner will infer conclusions
that are insufficient to attend the user’s needs. Besides, taking into account that
factors can improve several tasks carried in intelligent environments, such as
ontology-based activity recognition. For instance, WatchingTVActivity can be
defined as an activity performed by a Person who is watching the television in
a room:</p>
        <p>W atchingT V Activity ≡ ∃isDoneBy.(P erson u ∃isIn(Room u
∃containsAppliance.(T V u ∃isSwitched.(true))))
(1)</p>
        <p>If the user is watching the television and the system receives the values of all
the sensors, then it is able to conclude that the user’s current activity is of the
type WatchingTVActivity. In contrast, if the value of the presence sensor is not
available, then it is not possible to infer that the user is watching the television.</p>
        <p>In addition, sometimes there is not a rule of thumb to classify an individual
as a member of a class. For instance, we can classify the action that the system
has to perform regarding the current activity of the user. Thus, we can define
that the system should turn off the television, when the user is not watching it:
T urnOf f T V ≡ ∃requiredBy.(P erson u ∀isDoing.¬W atchingT V Activity
u∃hasAppliance.(T V u ∃isSwitched.(true)))(2)</p>
        <p>However, this concept definition does not accurately model the reality. The
action can fulfil every condition expressed in the TurnOffTV definition, but the
television should not be turned off. This situation may occur when the user goes
to the toilet or answers a call phone in another room, among others.</p>
        <p>
          In these cases in which the information of the domain comes with quantitative
uncertainty or vagueness, ontology languages are less suitable [
          <xref ref-type="bibr" rid="ref21">11</xref>
          ]. Uncertainty
is usually considered as the different aspects of the imperfect knowledge, such as
vagueness or incompleteness. In addition, the uncertainty reasoning is defined as
the collection of methods to model and reason with knowledge in which boolean
truth values are unknown, unknowable or inapplicable [19]. Other authors [
          <xref ref-type="bibr" rid="ref1 ref11">1</xref>
          ]
[
          <xref ref-type="bibr" rid="ref21">11</xref>
          ] consider that there are enough differences to distinguish between uncertainty
and vague knowledge. According to them, uncertainty knowledge is comprised
by statements that are either true or false, but we are not certain about them
due to our lack of knowledge. In contrast, vagueness knowledge is composed of
statements that are true to certain degree due to vague notions.
        </p>
        <p>In our work, we are more interested in the uncertainty caused by the lack
of information rather than the vague knowledge. For this reason, probabilistic
approaches are more suitable to solve our problem.
3</p>
        <p>
          Turambar Solution
Our proposal, called Turambar, combines a Bayesian network model with an
OWL 2 DL ontology in order to handle uncertainty. A Bayesian network [
          <xref ref-type="bibr" rid="ref23">13</xref>
          ]
is a graphical model that is defined as a directed acyclic graph. The nodes in
the model represent the random variables and the edges define the dependencies
between the random variables. Each variable is conditionally independent of its
non descendants given the value of its parents.
        </p>
        <p>Turambar is able to calculate the probability associated to a class, object
property or data property assertions. These probabilistic assertions have only
two constraints:
– The class expression employed in the class assertion should be a class.
– For positive and negative object property assertions, the object property
expression should be an object property.</p>
        <p>However, these limitations can be solved declaring a class equivalent to a class
expression or an object property as the equivalent of an inverse object property.
Examples of probabilistic assertions that can be calculated with Turambar are:
W atchingT V Activity(Action1) 0.8
isSwitched(T V 1, true) 1
(3)
(4)
(5)
The probabilistic object property assertion expressed in (3) states that John is
in Bedroom1 with a probability of 0.7. On the other hand the probabilistic class
assertion (4) describes that the Action1 is member of the class
WatchingTVActivity with a probability of 0.2. Finally, the probabilistic data property assertion
(5) defines that the television, TV1, is on with a probability of 1.0. The
probability associated to these assertions is calculated through Bayesian networks
that describe how other property and class assertions influence each other. In
Turambar, the probabilistic relationships should be defined by an expert. In
other words, the Bayesian networks must be generated by hand, since learning
a Bayesian network is out of the scope of this paper and it is not the goal of this
work.
3.1</p>
        <p>Turambar Functionality Definition
The classes, object properties and data properties of the OWL 2 DL ontology
involved in the probabilistic knowledge are connected to the random variables
defined in the Bayesian network model. For example, the OWL class
WatchingTVActivity is connected to at least one random variable, in order to be able
to calculate probabilistic assertions about that class. The set of data properties,
object properties and classes that are linked to random variables is called Vprob
and a member of Vprob, vprobi; such that vprobi ∈ Vprob.</p>
        <p>In Turambar, every random variable (RV) is associated to a Vprob and every
RV’s domain is composed of a set of functions that determine the values that a
random variable can take, such as V al(RV ) = {f1, f2...fn} and fi ∈ V al(RV ).
These functions require a property or class and individual to calculate the
probabilistic assertion, such as fi : a1, ex → result where a1 is an OWL individual,
ex, a class, data property or object property; result, a class assertion, object
property assertion, data property assertion or void (no assertion). In the case,
that every function in the domain of a random variable returns void, it means
that the random variable is auxiliary. In contrast, if any fi in the domain of
a random variable returns a probability associated to an assertion, then the
random variable is called final.</p>
        <p>For instance, the data property lieOnBedTime is linked to a random variable
named SleepTime whose domain is composed of two functions f1 that check if
the user has been sleeping for less than 8 hours and f2 function that checks if
the user has been sleeping for more than 8 hours. Both functions are not able to
generate assertions, so the random variable SleepTime is auxiliary. By contrast,
WatchingTVActivity class is linked to a random variable called WatchingTV
whose domain comprises f3 function that checks if an individual is member of
the class WatchingTVActivity (e.g. W atchingT V Activity(Activity1) 0.8) and
the f4 function which checks if an individual is a member of the complement of
WatchingTVActivity.</p>
        <p>It is also important to remark that a vprobi can be referenced from several
random variables. For example, the TurnOffTV depends on the user’s
impairments, so if the blind user is deaf, it is more likely that the television needs to be
turned off. Additionally, having an impairment also affects to the probability of
having another impairment: deaf people have a higher probability of also being
mute. In this case, we can link hasImpairment object property with two random
variables in order to model it.</p>
        <p>Apart from the conditional probability distribution, nodes connected
between them may have an associated node context. The context defines how
different random variables are related between them and the condition that
must fulfil. This context establishes an unequivocal relationship in which
every individual involved in that relationship should be gatherer before
calculating the probability of an assertion. If the relationship is not fulfilled then
the probabilistic query cannot be answered. For example, to estimate the
probability for the TurnOffTV, the reasoner needs to know who is the user and
in which room he is. For this case the relationship may be the following one
isIn(?user, ?room) ∧ requiredBy(?action, ?user) ∧ hasAppliance(?user, ?tv) ,
being ?user, ?action, ?tv and ?room variables. So, if we ask for the probability
that Action1 is member of TurnOffTV, such as Pr(TurnOffTV(Action1)), then
the first step to calculate it is to check its context. If everything is right the
evaluation of this relationship should return that the Action1 is required only
by one user who is only in one room and has only one television. Otherwise, the
probability cannot be calculated.</p>
        <p>Our proposal can be viewed as a SROIQ(D) extension that includes a
probabilistic function P r which maps role assertions and concept assertions to a value
between 0 and 1. The sum of the probabilities obtained for a random variable is
equal to 1. In contrast, the sum of probabilities for the set of assertions obtained
for a vprobi may be different from 1. For instance, the object property
hasImpairment is related to two random variables one to calculate the probability of
being deaf and another one to calculate the probability of being mute. If both
random variables have a domain with two functions, we can get four
probabilistic assertions that sums 2 instead of 1, but the sum of probabilities obtained in
one random variable is 1:
– Random variable deaf’s assertions: hasImpairment(J ohn, Deaf )0.8 and
¬hasImpairment(J ohn, Deaf )0.2.
– Random variable mute’s assertions: hasImpairment(J ohn, M ute)0.7 and
¬hasImpairment(J ohn, M ute)0.3.</p>
        <p>The probability of an assertion that exists in the OWL 2 DL ontology is
always 1 although the data property, object properties or class is not member of
Vprob. For example, if an assertion states that John is a Person (P erson(J ohn))
and we ask for the probability of this assertion, then its probability is 1, such
as P erson(J ohn)1. However, if the data property, object properties or class is
not member of Vprob and there is not an assertion in the OWL 2 DL ontology
that states it, then the probability for that assertion is unknown. We consider
that the probabilistic ontology is satisfied if the OWL 2 DL ontology is satisfied
and the Bayesian network model is not in contradiction with the OWL ontology
knowledge.
3.2</p>
        <p>Turambar Ontology Specification
In Turambar, a probabilistic ontology comprises an ordinary OWL 2 DL
ontology and the dependency description ontology that defines the Bayesian network
model.</p>
        <p>The ordinary OWL ontology imports the Turambar annotations ontology,
which defines the following annotations:
– turambarOntology annotation defines the URI of the dependency description
ontology.
– turambarClass annotation links OWL classes in the ordinary ontology to
random variables in the dependency description ontology.
– turambarProperty annotation connects OWL data properties and object
properties in the ordinary ontology to random variables in the dependency
description ontology.</p>
        <p>We choose to separate the Bayesian network definition from the ordinary
ontology in order to isolate the probabilistic knowledge definition from the OWL
knowledge. We define isolation as the ability of exposing an ontology with an
unique URI that locates the traditional ontology and the probabilistic one. So,
given the URI of a probabilistic ontology, a compatible reasoner loads the
ordinary ontology and the dependency description ontology it. In contrast, a
traditional reasoner only loads the ordinary ontology. So, if the Turambar probabilistic
ontology is loaded by a traditional reasoner, the traditional reasoner does not
have access to the knowledge encoded in the dependency description ontology.
In this way, we also want to promote the re-utilization of probabilistic ontologies
as simple OWL 2 DL ontologies by traditional systems and the interoperability
between our proposal and them.</p>
        <p>On the other hand, the dependency description ontology defines the
probabilistic model employed to estimate the probabilistic assertions. To model that
knowledge, it imports the Turambar ontology, which defines the vocabulary to
describe the probabilistic model. As the figure 1 shows, the main classes and
properties in the Turambar ontology are the following ones:
– Node class represents the nodes in Bayesian networks. Node instances are
defined as auxiliary random variables through the property isAuxiliar. The
hasProbabilityDistibution object property links Node instances with their
corresponding probability distributions and hasState object property
associates Node instances with their domains. Furthermore, hasChildren object
property and its inverse hasParent set the dependencies between Node
instances. Finally, hasContext object property defines the context for a node
and hasVariable object property, the value of the variable that the node
requires.
– MetaNode is a special type of Node that is employed with non functional
object properties and data properties. Its main functionality is to group
several nodes that share a context and are related to the same property. For
instance, in the case of the hasImpairment object property we can model a
MetaNode with two nodes: Deaf and Mute. Both nodes share the same
context but have different parents and states. The object property compriseNode
identifies the nodes that share a context.
– State class defines the values of random variables’ domain. In other words, it
describes the functions which generate probabilistic assertions. These
functions are expressed as a string through the data property stateCondition.
– ProbabilityDistribution class represents a probability distribution.
Probability distributions are given in form of conditional probability tables. Cells of
the conditional probability table are associated to the instances of
ProbabilityDistribution through hasProbability object property.
– Probability class represents a cell of a conditional probability table, such
P (x1 | x2, x3) = value, where x1, x2 and x3 are State individuals and value
is the probability value. x1 State is assigned to an instance of Probability
class through the hasValue object property and x2 and x3 conditions through
hasCondition object property. Finally, the data property hasProbabilityValue
sets the probability value for that cell.
– Context class establishes the relationships between the nodes of a Bayesian
network. Relationships between nodes are expressed as a SPARQL-DL query
through the data property relationship.
– Variable class represents the variables of the context. Their instances
identify the SPARQL-DL variables defined in the context SPARQL-DL query.
The variableName data property establishes the name of the variable. For
example, if the context has been defined with the following SPARQL-DL
expression: select ?a ?b where { PropertyValue(p:livesIn, ?a, ?b)} , then
we should create two instances of Variable with the variableName property
value of a and b, respectively.
– Plugin class defines a library that provides some functions that are referenced
by State class instances and are not included as member of the Turambar
core. The core functions are the following ones: (i) numbers and strings
comparison, (ii) ranges of number and string comparison, (iii) individual
instances comparison, (iv) boolean comparison, (v) class memberships
checking and (vi) the void assertion to define the probability that no assertion
involves an individual. Only i, iii and iv are able to generate probabilistic
assertions. Every function, except the void function, has their inverse function
to check if that value has been asserted as false.
4</p>
        <p>
          Related Works
We can classify probabilistic approaches to deal with uncertainty in two groups:
probabilistic description logics approaches and probabilistic web ontology
languages [
          <xref ref-type="bibr" rid="ref21">11</xref>
          ].
        </p>
        <p>
          In the first group, P-CLASSIC [
          <xref ref-type="bibr" rid="ref19 ref9">9</xref>
          ] extends description logic CLASSIC to add
probability. In contrast, Pronto [
          <xref ref-type="bibr" rid="ref18 ref8">8</xref>
          ] is a probabilistic reasoner for P-SROIQ, a
probabilistic extension of SROIQ. Pronto models probability intervals with its
custom OWL annotation pronto#certainty. Apart from the previously described
works, there are several other approaches that have been explained in different
surveys such as [
          <xref ref-type="bibr" rid="ref24">14</xref>
          ].
        </p>
        <p>
          In contrast, probabilistic web ontology languages combine OWL with
probabilistic formalisms based on Bayesian networks. Since our proposal falls under
this group, we will review in depth the most important works in this category:
BayesOWL, OntoBayes and PR-OWL. The BayesOWL [
          <xref ref-type="bibr" rid="ref17 ref7">7</xref>
          ] framework extends
OWL capacities for modelling and reasoning with uncertainty. It applies a set
of rules to transform the class hierarchy defined in an OWL ontology into a
Bayesian network. In the generated network there are two types of nodes:
concept nodes and L-Nodes. The former one represents OWL classes and the latter
one is a special kind of node that is employed to model the relationships defined
by owl:intersectionOf, owl:unionOf, owl:complementOf, owl:equivalentClass and
owl:disjointWith constructors. Concept nodes are connected between them by
directed arcs that link superclasses with their classes. On the other hand,
LNodes and concept nodes involved in a relationship are linked following the
rules established for each constructor. The probabilities are defined with the
classes PriorProb, for prior probabilities, and CondProb, for conditional
probabilities. For instance, BayesOWL [22] recognizes some limitations: (i) variables
should be binaries, (ii) probabilities should contain only one prior variable, (iii)
probabilities should be complete and (iv) in case of inconsistency the result may
not satisfy the constraints offered. BayesOWL approach is not valid for our
purpose, because it only supports uncertainty to determine the class membership
of an individual and this may not be enough for context modelling. For
example, sensors’ values may be represented as data and object properties values and
knowing the probability that a sensor has certain value may be very useful for
answering user’s needs.
        </p>
        <p>In contrast to BayesOWL, OntoBayes [21] focuses on properties. In
OntoBayes, every random variable is a data or object property. Dependencies between
them are described via the rdfs:dependsOn property. It supports to describe
prior and conditional probabilities, besides it contains a property to specify the
full disjoint probability distribution. Another improvement of OntoBayes over
BayesOWL is that it supports multi-valued random variables. However, it is not
possible to model relationships between classes in order to prevent errors when
extracting Bayesian network structure from ontologies. OntoBayes offers us a
solution for the limitation presented in BayesOWL regarding OWL properties,
but its lack of OWL class support makes it unsuitable for our goal.</p>
        <p>
          PR-OWL [
          <xref ref-type="bibr" rid="ref15 ref5">5</xref>
          ] is an OWL extension to describe complex bayesian models. It
is based on the Multi-Entity Bayesian newtworks (MEBN) logic. MEBN [
          <xref ref-type="bibr" rid="ref10 ref20">10</xref>
          ]
defines the probabilistic knowledge as a collection of MEBN fragments, named
MFrags. A set of MFrags configures a MTheory and every PR-OWL ontology
must contain at least one MTheory. To consider a MFrag set as a MTheory, it
must satisfy consistency constraints ensuring that it only exists a joint
probability distribution over MFrags’ random variables. In PR-OWL, probabilistic
concepts can coexist with non probabilistic concepts, but these are only
benefited by the advantages of the probabilistic ontology. Each MFrag is composed
of a set of nodes which are classified in three groups: resident, input and
context node. Resident nodes are random variables whose probability distribution
is defined in the MFrag. Input nodes are random variables whose probability
distribution is defined in a distinct MFrag than the one where is mentioned.
In contrast, context nodes specify the constraints that must be satisfied by an
entity to substitute an ordinary variable. Finally, node states are modelled with
the object property named hasPossibleValues.
        </p>
        <p>
          The last version of PR-OWL [
          <xref ref-type="bibr" rid="ref12 ref2">2</xref>
          ], PR-OWL 2, addresses the PR-OWL 1
limitations regarding to its compatibility with OWL: no mapping to properties of
OWL and the lack of compatibility with existing types in OWL. Although,
PROWL offers a good solution to deal with uncertainty, it does not provide some
characteristics that we covet for our systems, such as isolation.
        </p>
        <p>Our proposal is focused on computing the probability of data properties
assertions, object properties assertions and class assertions. This issue is only covered
by PR-OWL, because BayesOWL only takes into account class membership and
OntoBayes, object and data properties.</p>
        <p>In addition, we pretend to offer a way to keep the uncertainty information
isolated as much as possible from the traditional ontology. With this policy, we want
to ease the reutilization of our probabilistic ontologies by traditional systems that
do not offer support for uncertainty and the interoperability between them.
Furthermore, we aim to avoid that traditional reasoners load unnecessary
information about the probabilistic knowledge that they do not need. Thus, if we load the
Turambar probabilistic ontology located in http://www.example.org/ont.owl,
traditional OWL reasoners load only the knowledge defined in the ordinary OWL
ontology and do not have access to the probabilistic knowledge. In contrast,
Turambar reasoner is able to load the ordinary OWL ontology and the dependency
description ontology. The Turambar reasoner needs to access to the ordinary
OWL ontology to answer traditional OWL queries and to find the evidences of
the Bayesian networks defined in the dependency description ontology. It is also
important to clarify that a class or property can have deterministic assertions
and probabilistic assertions without duplicating them due to the links between
Bayesian networks’ nodes and OWL classes and properties through
turambarClass and turambarProperty, respectively. Thanks to this feature, a Turambar
ontology has a unique URI that allows it to be used as an ordinary OWL 2 DL
ontology without loading the probabilistic knowledge. This characteristic is not
offered by other approaches as far as we know.</p>
        <p>Another difference with other approaches is that we have taken into account
the extensibility of our approach through plug-ins to increase the basis
functionalities. We believe that it is necessary to offer a straightforward, transparent and
standard mechanism to extend reasoner functionality in order to cover
heterogeneous domains’ needs.</p>
        <p>During studies of the subject, we found that three issues were of paramount
importance, if we wanted to solve the problem. We will explain these first, before
investigating the existing work.</p>
        <p>
          Blank Node Identifiers: RDF graphs might contain blank nodes. Such nodes
do not have an Internationalized Resource Identifier (IRI) [
          <xref ref-type="bibr" rid="ref17 ref7">7</xref>
          ] assigned. Once
loaded by an RDF implementation, they are assigned local identifiers, which
are not transferable to other implementations. When trying to calculate a
hash value, this is a major issue as blank nodes cannot be deterministically
addressed. One can think of blank nodes as anonymous and having an
identity is a necessary pre-condition for calculating a hash value. Solving the
blank node issue is an algorithmic challenge.
        </p>
        <p>Order of Statements Calculation of hash values effectively serialises a RDF
graph structure into a string. RDF does not imply a certain order of
statements, e.g. the order of predicates attached to a single subject. For
serialisation purposes we need a deterministic order, or otherwise we might end
up with hash values differing between implementations. This issue can be
solved by adhering to a sort order when serialising the graph.</p>
        <p>
          Encoding RDF uses literals to store values in the graph. These literals need to
follow a common encoding, otherwise different hash values might be
calculated. The same goes for namespace prefixes or relative IRIs (both features of
the XML syntax for RDF [
          <xref ref-type="bibr" rid="ref18 ref8">8</xref>
          ]) –they need to be encoded with fully qualified
names when stored in memory, or, at the latest, when serialised.
        </p>
        <p>
          The most influential paper on the subject of RDF graph hashing was written
by Carroll in 2003 [
          <xref ref-type="bibr" rid="ref19 ref9">9</xref>
          ]. Building on earlier work [
          <xref ref-type="bibr" rid="ref10 ref20">10</xref>
          ], Carroll explains that the
runtime of any algorithm for generic hashing of RDF data is equivalent to the
graph isomorphism problem, which is not known to be N P-complete nor known
to be in P. Carroll then refrains from finding a generic solution to the problem
and details his algorithm, which runs in O(n log(n)), but re-writes RDF graphs
to a canonical format. The proposed algorithm works on the N-Triples format (a
concrete document syntax). As far as solving the blank node identity problem,
the article states: “Since the level of determinism is crucial to the workings of
the canonicalization algorithm, we start by defining a deterministic blank node
labelling algorithm. This suffers from the defect of not necessarily labeling all
the blank nodes.” [9, Section 4]. Carroll continued to work on the subject, for
example by publishing applications based on digitally signing graphs, together
with Bizer, Hayes, and Stickler [
          <xref ref-type="bibr" rid="ref21">11</xref>
          ], but did not seem to have designed a general
algorithm for hashing RDF graphs.
        </p>
        <p>
          Sayers and Karp, colleagues of Carroll, published two technical reports at
Hewlett-Packard that explains RDF hashing and applications thereof [
          <xref ref-type="bibr" rid="ref22 ref23">12,13</xref>
          ].
They identify four different ways to tackle the blank node problem [13, Section
3 ff.], namely:
Limit operations on the graph The idea is to maintain blank node
identifiers across implementations, which is not possible in an open world scenario.
Limit the graph itself Avoid the use of blank nodes. This is clearly not the
way for us, as we strive for a general solution to the problem.
Modify the graph Work around the problem by adding information about
blank node identity within the graph. Not possible for us, as we don’t want
to change the RDF graph.
        </p>
        <p>Change the RDF specification Generally assign globally unique identifiers
to blank nodes. This will most likely never happen.</p>
        <p>Apart from changing the RDF specification, none of these methods actually
solves the problem and they can be seen as workarounds.</p>
        <p>
          Other authors also investigated RDF graph hashing, e.g. Giereth [
          <xref ref-type="bibr" rid="ref24">14</xref>
          ] for the
purpose of encrypting fragments of a graph. Giereth works around the blank node
issue by modifying the graph. A more current approach by Kuhn and Dumontier
[15] proposes an encoding of hash values in URIs. The approach replaces blank
nodes with arbitrarily generated identifiers and thus needs to modify the RDF
data to work.
        </p>
        <p>Although Carroll did already provide pointers in the right direction, the final
idea for solving the issue can be attributed to Tummarello et al., who introduce
the concept of a “Minimum Self-Contained Graph” (MSG) [16]. A MSG is a
partitioning of a graph, so that each MSG contains at most one transitively
connected sub-graph of blank nodes. We apply this idea to construct a blank
node identity, as explained in the next section.
3</p>
        <p>The Algorithm
β1
β2</p>
        <p>A</p>
        <p>B
C
β4
β3</p>
        <p>Our approach to the blank node labeling problem relies on constructing the
identity of a blank node though its context. This is similar to the MSG principle
of Tummarello et al. [16, Section 2] and a logical continuation of the thoughts
of Carroll [9, Section 7.1]. For an example, see Fig. 1. The diagram shows three
nodes with IRIs (A, B, C) and four blank nodes (β1 ... β4). If we go on to define
the identity of a node as determined by its direct subjects, we can distinguish β2
and β4, but not yet the two other ones: there are both blank nodes pointing to
another blank node and the C node. We can only discern every node by taking
more of the context into account: not only direct neighbours, but neighbouring
nodes one hop away. Consequentially, the identity of a blank node can only be
constructed when following all of the transitive blank nodes, reachable from the
original one. Using this scheme, we are able to establish an identity for blank
nodes. Having identity for both: blank nodes, and IRI nodes, we can generate a
characteristic, implementation-independent string for both node types. As RDF
only allows these two types of nodes as subjects, we are now able to create a list
of so-called “subject strings”. To solve the issue of implementation-dependent
statement order, it is necessary to establish an overall ordering criterium over
this list. We are using a simple lexicographical ordering based on the unicode
value of each character in a string. As we are solely relying on subject nodes to
calculate the hash value, we need to also encode the predicates and objects of the
RDF graph into the subject strings, as well. This way of encoding graph structure
into the overall data used to calculate the hash value is a major difference to
other algorithms striving for the same goal. Usually, only flat triple lists are
processed and the hash function calculates a value for each triple. To encode the
graph structure we are using special delimiter symbols. Without these symbols,
differing graph topology might lead to the same string representation and thus,
the same hash value. For the final cryptographic calculation of the hash value,
we are using SHA-256 on UTF-8 [17] encoded data. SHA-256 was chosen as
the amount of characters in the overall data string seems to be moderate. It is
recommended to use SHA-256 for less than 264 bits of input [4, Section 1].
3.1</p>
        <p>Preconditions and Remarks
To calculate the digest of a single RDF graph g, the graph has to reside in
memory first, as we are not concerned with any network or file-based
representations of the graph. We require that the graph’s content is accessible as a set of
&lt; S, P, O &gt; triples 3. It is beneficial to have fast access to all subject nodes of
the graph, and to all properties of a node and we are using matching patterns to
express this type of access, e.g. &lt;?ns, ?, ? &gt; to denote any node ns that appears
in the role of a subject in the RDF triple data4. Literals and IRI identifiers need
to be stored as unicode characters. There are no restrictions on the blank node
identifiers, as we do not use them for calculation of the digest. For sake of clarity,
we present our algorithm broken down in four separate sub-aglorithms: A
function that calculates the hash value for a given graph (Algorithm 1), a procedure
that calculates the string representation for a given subject node (Algorithm 2),
another one for the string representation of the properties of a given subject
node (Algorithm 3), and a last one for calculation of the string representation
of an object node (Algorithm 4). These sub-algorithms call each other and thus,
could be combined in a single operation. It should be noted, that the algorithm
uses reentrance to establish the transitive relationship needed to assign identities
to the blank nodes.</p>
        <p>
          Our algorithm uses a number of different delimiter symbols with strictly
defined, constant values, which we assigned greek letters to. Table 1 gives an
overview of these symbols.
3 Triples with a Subject, P redicate, and Object
4 The question mark notation is inspired by the SPARQL query syntax [
          <xref ref-type="bibr" rid="ref13 ref3">3</xref>
          ]
Symbol Symbol Name Value (Unicode) Value Name
αs Subject start symbol { (U+007B) Left curly bracket
ωs Subject end symbol } (U+007D) Right curly bracket
αp Property start symbol ( (U+0028) Left parentheses
ωp Property end symbol ) (U+0029) Right parentheses
αo Object start symbol [ (U+005B) Left square bracket
ωo Object end symbol ] (U+005D) Right square bracket
β Blank node symbol ∗ (U+002A) Asterisk
        </p>
        <p>Table 1. Guide to delimiters and symbols used in the algorithms</p>
        <p>Use of these symbols is unproblematic in regard to their appearance as part
of the RDF content. As we use two symbols to delimit a scope in the string, we
can clearly distinguish between use as delimiter and use as content. As our goal
is the creation of a string representation for each subject node, the algorithm
makes heavy use of string concatenation and we are using the ⊕ symbol to
denote this operation. In several places, strings are nested between start and
stop delimiters, like this: α ⊕ string ⊕ ω.</p>
        <p>Some final remarks regarding the implementation before delving into the
specifics of the algorithm: We are using some variables (visitedN odes, g) as
parameters for functions and procedures. Of course, these should be better put
away as shared variables (e.g. attributes in an object). The variable result is
always local and needs to be empty at the start of each function. Furthermore,
there are some functions that dependent on a concrete implementation and are
quite trivial to use. We skip an in-depth discussion of those, e.g., predicates(...).
3.2</p>
        <p>Calculating the Hash Value
Algorithm 1 Calculating a hash value for a RDF graph g</p>
        <p>To calculate the hash value for a given RDF graph, Algorithm 1 is used. The
function takes a single parameter: the RDF graph g to use for calculation of
its hash value. At first the algorithm iterates over all of the subject nodes ns
that exist in g (lines 2–5). A subject node is any node that appears in the role
of a subject in a RDF triple contained in g. For each of the subject nodes we
create a data structure, called visitedN odes, which is used to record if we already
processed some blank node. This is necessary for termination of the construction
of the blank node identities. visitedN odes needs to be empty before calculating
the string representation for ns by calling the procedure encodeSubject(...) in
line 4, which is explained in the next section. After all subject nodes are encoded,
the resulting list needs to be sorted. Any sorting order could be used and as we
do not require specific semantics for this step, we are establishing an ordering
simply by comparing the unicode numbering of letters. The sorting operation
in line 6 is key to deterministically create a hash value, as the result would
otherwise build upon the (implementation-dependent) order of nodes in g. All of
the sorted subject-strings are then concatenated to form an overall result string,
while each single subject string is enclosed with the subject delimiter symbols
(line 8). Finally, the result string is subjected to a cryptographic hash function
and returned (line 10).
3.3</p>
        <p>Encoding the Subject Nodes
In RDF, subject nodes can be of two types: they can either be blank nodes or
IRIs [2, Section 3.1]. The procedure encodeSubject(...), shown in Algorithm 2
and used by Algorithm 1, needs to take care of this. The procedure has three
arguments: ns – the subject node to encode as a string, visitedN odes – our data
structure for tracking already visited nodes, and g – the RDF graph.</p>
        <p>The discrimination of types comes first: lines 2–8 process blank nodes and
lines 9–11 take care of IRIs. For the blank nodes, we have to distinguish between
the case where we already met a blank subject node (line 3–4), and the case where
we didn’t (line 5–7). In the case that the subject node was already encountered,
the graph traversal ends and we return an empty string. If the blank subject
node is hitherto unknown, then result is set to the blank node symbol β (see
Table 1) and the node is recorded in visitedN odes as being processed. If the
subject is not a blank node, but an IRI, we simply set result to be the IRI itself.
Only encoding the subject node itself is not sufficient for our purposes, as we
need to establish an identity based on the context of the current subject node. In
line 12 this process is triggered by calling the encodeP roperties(...) procedure
(see Algorithm 3) and concatenating the returned string with the existing result.
The result itself is returned in line 13.
3.4</p>
        <p>Encoding the Properties of a Subject Node
Algorithm 3 is responsible for encoding all of the properties of a given subject
node ns into a single string representation. We understand properties as the
predicates (p) and objects (o) that fullfil &lt; ns, ?p, ?o &gt;, where ns is a given
subject. Apart from ns, the procedure encodeP roperties(...) needs visitedN odes
as a second, and g as a third argument.</p>
        <p>Algorithm 3 Encode properties for a subject node</p>
        <p>The algorithmic structure reflects the complexity of graph composition using
RDF predicates: a subject node can be associated with multiple predicates, and
the predicates are allowed to be similar, if associated with different objects. We
use a two stage process to encode properties: First, all unique predicate IRIs of
the given subject node ns are retrieved and ordered (lines 2–3). We are
postulating a function predicates(...) that returns all predicate IRIs for a given subject
node ns by searching all triples for matches to &lt; ns, ?px, ? &gt;, extracting the IRI
of the identified predicate px, and eliminating double entries. In a second step,
we encode each predicate IRI and the set of objects associated with it (line 4–14).
The property encoding starts in line 5, where the result is concatenated with
the property start symbol αp and the predicate IRI. Subsequently, we retrieve
all object nodes (nodes that appear in triples with ns as subject and iri as
predicate), encode each object node using the procedure encodeObject(...), and store
their respective string representations in a objectStrings list (line 6–8). The
encodeObject(...) procedure is detailed in Algorithm 4. After collecting all the
encoded object strings, the resulting list has to be sorted (line 9). In line 10–12
each object string is appended to result, while taking care to enclose the string
in delimiter symbols. Encoding of a single property (one pass of the loop started
in line 4) ends with appending the property stop symbol ωp to result. Once the
procedure has encoded all properties it returns with the complete result string
in line 15.
3.5</p>
        <p>Encoding the Object Nodes
Processing the object nodes itself is trivial when compared to the property
encoding. Objects in RDF triples can be three things: an IRI, a literal, or a blank node
[2, Section 3.1]. The encodeObject(...) procedure needs to return an appropriate
string representation for each of these three cases. It takes three arguments: an
object node no, visitedN odes, and the RDF graph g.</p>
        <p>Algorithm 4 Encode an object node
1: procedure encodeObject(no, visitedN odes, g)
2: if no is a blank node then
3: return encodeSubject(no, visitedN odes, g)
4: else if no is a literal then
5: return literal representation of no
6: else
7: return IRI of no
8: end if
9: end procedure
. Re-enter Algorithm 2
. Consider language and type
. no has to be a IRI</p>
        <p>The three aforementioned cases are treated as follows. If no is a blank node,
we continue with re-entering Algorithm 2 (see line 3). The re-entrance allows
us to construct a path through neighbouring blank nodes. Together with having
potentially many object nodes associated with a single subject, this yields a
connected graph, similar to the MSGs of Tummarello et al. If no is a literal,
it is returned in a format according to [2, Section 3.3] in line 5, including any
language and type information. If no is neither a blank node, nor a literal, it
has to be an IRI and we return it verbatim. After all objects, properties, and
subjects have been encoded, all sub-algorithms have returned and Algorithm 1
terminates.
3.6</p>
        <p>An Example
Consider the RDF graph shown in Figure 1 at the beginning of Section 3. When
applying the algorithm to the given graph, we end up with the following string
before calculating the SHA-256 hash (see Algorithm 1, line 10):</p>
        <p>{∗(−[∗(−[A][C])][C])} {∗(−[∗(−[B][C])][C])} {∗(−[A][C])} {∗(−[B][C])}
Please note that this string is not valid RDF, as scheme and path information
are missing from the employed IRIs5 used by A, B, and C. Also the symbol “−”
is used to indicate an arbitrary IRI employed for all predicates.</p>
        <p>Thanks to the delimiters, it is quite easy to understand the string’s structure. The
four subject node strings for β1, β3, β2, and β4 (in this order) are encapsulated
between curly brackets each. A, B, and C only appear as object nodes and thus do
not trigger the creation of additional subject strings. Instead, they are encoded
as part of the blank node subject strings. Let’s take a look at the first subject
string for node β1: {∗(−[∗(−[A][C])][C])}. Apart from the curly brackets, the
string starts with the blank node symbol “∗”, followed by the properties of that
node, delimited in parentheses. The node has only a single predicate, used with
two different objects: [∗(−[A][C])] and [C]. If we would have additional predicate
types, there would also be further parentheses blocks. Objects are delimited by
square brackets and due to the re-entrant nature of the algorithm object strings
follow the same syntax as just discussed.
4</p>
        <p>
          Conclusion and Future Work
The presented algorithm calculates a hash value for RDF graphs including blank
nodes. It is not necessary to alter the RDF data or to record additional
information. It is not dependent on any concrete syntax. It solves the blank node labeling
problem by encoding the complete context of a blank node, including the RDF
graph structure, into the data used to calculate the hash value. The algorithm
has a runtime complexity of O(nn), which is consistent with current research on
algorithms for solving the graph isomorphism problem [
          <xref ref-type="bibr" rid="ref19 ref9">9</xref>
          ]. The worst-case
scenario is a fully meshed RDF graph of blank nodes, which does not seem to make
any sense whatsoever — usually, we would expect the amount of blank nodes in
a graph to be far smaller, thus the real execution speed to be less catastrophic.
We concentrated on solving the primary problem, not on runtime optimisations
and we are certain, that there is room for improvement in the given algorithm.
There are some obvious starting points for doing this. For example, there are
redundancies in the string representations for transitive blank node paths (the
string representations for β2 and β4 appear twice in the example given in Section
3.6). One could cache already computed subject-strings, pulling them from the
cache when needed. Also, the interplay between the SHA-256 digest computation
and the subject string calculations has not been researched in sufficient detail.
5 For example A instead of http://a/
It might be possible to reduce processing and storage overhead by combining
these two operations, calling the digest operation on each subject string and
combining the resulting values in order of the final sorted list. The sensibility
of such optimizations should largely depend on the susceptibility of the digest
algorithm implementation to the length of the given input strings. This, in turn,
depends on the usage of blank nodes in the input RDF data. More blank nodes
in the input data and more references between blank nodes means longer
subject strings. Consequentially, to come to a more substantial assessment of the
presented algorithm, we will need to study its performance on a number of real
(and larger) data sets.
        </p>
        <p>While we trust the general approach for solving the blank node labeling
problem through an assignment of identity based on the surrounding context
of the node, we did not proof that the algorithm works correctly. To assure
that it works properly, we did test it: on the one hand in regard to its ability
to process all possible RDF constructs using tests from W3C’s RDF test cases
recommendation [18], on the other hand in regard to the correctness of the blank
node labeling approach using manually constructed graphs. These graphs range
from simple, non cyclic ones with a single blank node to all possible permutations
of a fully meshed graph of blank nodes with variations on the attachment of IRI
nodes and predicate types.</p>
        <p>Acknowledgments. We would like to thank Thomas Pilger and Abdul Saboor
for their collection of material documenting the current state of the art and for
tireless implementation work.
11. Carroll, J. J., Bizer, C., Hayes, P., Stickler, P.: Named Graphs, Provenance and</p>
        <p>Trust. Proc. International World Wide Web Conference, pp. 613–622 (2005)
12. Sayers, C., Karp. A. H.: Computing the Digest of an RDF Graph. Hewlett-Packard</p>
        <p>Labs Technical Report HPL-2003-235 (2003)
13. Sayers, C., Karp. A. H.: RDF Graph Digest Techniques and Potential Applications.</p>
        <p>Hewlett-Packard Labs Technical Report HPL-2004-95 (2004)
14. Giereth, M.: On Partial Encryption of RDF-Graphs. Proc. of the Fourth
International Semantic Web Conference, LNCS 3729, pp. 308–322 (2005)
15. Kuhn, T., Dumontier, M.: Trusty URIs: Verifiable, Immutable, and Permanent
Digital Artifacts for Linked Data. Proc. Eleventh European Semantic Web
Conference, LNCS 8465, pp. 395–410 (2014)
16. Tummarello, G., Morbidoni, C., Puliti, P., Piazza, F.: Signing Individual Fragments
of an RDF Graph. International World Wide Web Conference, pp. 1020–1021 (2005)
17. Yergeau, F.: UTF-8, a Transformation Format of ISO 10646. Internet Engineering</p>
        <p>Task Force – Request for Comments 3629 (2003)
18. World Wide Web Consortium.: RDF Test Cases, W3C Recommendation, http:
//www.w3.org/TR/2004/REC-rdf-testcases-20040210/ (2004)
Rating, recognizing and rewarding metadata integration
and sharing on the semantic web</p>
        <p>Francisco M. Couto
LASIGE, Dept. de Informa´tica, Faculdade de Cieˆncias, Universidade de Lisboa, Portugal
fcouto@di.fc.ul.pt
Abstract. Research is increasingly becoming a data-intensive science, however
proper data integration and sharing is more than storing the datasets in a public
repository, it requires the data to be organized, characterized and updated
continuously. This article assumes that by rewarding and recognizing metadata sharing
and integration on the semantic web using ontologies, we are promoting and
intensifying the trust and quality in data sharing and integration. So, the proposed
approach aims at measuring the knowledge rating of a dataset according to the
specificity and distinctiveness of its mappings to ontology concepts.</p>
        <p>The knowledge ratings will then be used as the basis of a novel reward and
recognition mechanism that will rely on a virtual currency, dubbed KnowledgeCoin
(KC). Its implementation could explore some of the solutions provided by
current cryptocurrencies, but KC will not be a cryptocurrency since it will not rely
on a cryptographic proof but on a central authority whose trust depends on the
knowledge rating measures proposed by this article. The idea is that every time
a scientific article is published, KCs are distributed according to the knowledge
rating of the datasets supporting the article.</p>
        <p>
          Keywords: Data Integration, Data Sharing, Linked Data, Metadata, Ontologies
1
Research is increasingly becoming a data-intensive science in several areas, where
prodigious amounts of data can be collected from disparate resources at any time [
          <xref ref-type="bibr" rid="ref16 ref6">6</xref>
          ].
However, the real value of data can only be leveraged through its trust and quality, which
ultimately results in the acquisition of knowledge through its analysis. Since multiple
types of data are involved, often from different sources and in heterogeneous formats,
data integration and sharing are key requirements for an efficient data analysis. The
need for data integration and sharing has a long-standing history, and besides the big
technological advances it still remains an open issue. For example, in 1985 the
Committee on Models for Biomedical Research proposed a structured and integrated view
of biology to cope with the available data [
          <xref ref-type="bibr" rid="ref18 ref8">8</xref>
          ]. Nowadays, the BioMedBridges 1
initiative aims at constructing the data and service bridges needed to connect the emerging
biomedical sciences research infrastructures (BMSRI), which are on the roadmap of the
European Strategy Forum on Research Infrastructures (ESFRI). One common theme to
all BMSRIs is the definition of the principles of data management and sharing [
          <xref ref-type="bibr" rid="ref13 ref3">3</xref>
          ]. The
Linked Data initiative 2 already proposed a well-defined set of recommendations for
exposing, sharing and integrating data, information and knowledge using semantic web
technologies. In this paradigm data integration and sharing is achieved in the form of
links connecting the data elements themselves and adding semantics to them. Following
and understanding the links between data elements in publicly available Data Linked
stores (Linked Data Cloud) enables us to access the data and knowledge shared by
others. The Linked Data Cloud offers an effective solution to break down data silos;
however the systematic usage of these technologies requires a strong commitment from
the research community.
        </p>
        <p>
          Promoting the trust and quality of data through their proper integration and sharing
is essential to avoid the creation of silos that store raw data that cannot be reused by
others, or even by the owners themselves. For example, the current lack of incentive to
share and preserve data is sometimes so problematic, that there are even cases of authors
that cannot recover the data associated with their own published works [
          <xref ref-type="bibr" rid="ref15 ref5">5</xref>
          ]. However,
the problem is how to obtain a proactive involvement of the research community in data
integration and sharing. In 2009, Tim Berners-Lee gave a TED talk3, where he said:
“you have no idea the number of excuses people come up with to hang onto their data
and not give it to you, even though you’ve paid for it as a taxpayer.” Public funding
agencies and journals may enforce the data-sharing policies, but the adherence to them
is most of the times inconsistent and scarce [
          <xref ref-type="bibr" rid="ref1 ref11">1</xref>
          ]. Besides all the technological advances
that we may deliver to make data integration and sharing tasks easier, researchers need
to be motivated to do it correctly. For example, due to the Galileos strong commitment
to the advance of Science, he integrated the direct results of his observations of Jupiter
with careful and clear descriptions of how they were performed, which he shared in
Sidereus Nuncius [
          <xref ref-type="bibr" rid="ref14 ref4">4</xref>
          ]. These descriptions enabled other researchers not only to be aware
of Galileos findings but also to understand, analyze and replicate his methodology. This
is another situation that we could characterize with the famous phrase “That’s one small
step for a man, one giant leap for mankind.” Now let us imagine if we could extend
Galileos commitment to all the research community, the giant leap that it could bring to
the advance of science.
        </p>
        <p>Thus the commitment of the research community to data integration and sharing
is currently a major concern, and this explains why BMSRIs have recently included in
their definition of the principles of data management and sharing the following
challenge: “to encourage data sharing, systematic reward and recognition mechanisms are
necessary”. They suggest studying not only measurements of citation impact, but also
highlighting the importance to investigate other mechanisms as well. Systematic reward
and recognition mechanisms should motivate the researchers in a way that they become
strongly committed in sharing data, so others can easily understand and reuse it. By
doing so, we encourage the research community to improve previous results by
replicating the experiments and testing new solutions. However, before developing a reward
and recognition mechanism we must formally define: i) what needs to be rewarded and
recognized; ii) and measure its value in a quantitative and objective way.
2 http://linkeddata.org/
3 http://www.ted.com/talks/tim_berners_lee_on_the_next_web
Proper data integration and sharing is more than storing the datasets in a public
repository, it requires the data to be organized and characterized in a way that others can
find it and reuse it effectively. In an interview4 to Nature, Steven Wiley emphasized
that sharing data “is time-consuming to do properly, the reward systems aren’t there
and neither is the stick”. Not adding links to external resources hampers the efficient
retrieval and analysis of data, and therefore its expansion and update. Making a dataset
easier to find and access is also a way to improve its initial trust and quality, as more
studies analyze, expand and update it. Like the careful and clear descriptions provided
by Galileo, semantic characterizations in the form of metadata must also be present so
others can easily find the raw data and understand how it can be retrieved and explored.</p>
        <p>
          Metadata is a machine-readable description of the contents of a resource made
through linking the resource to the concepts that describe it. However, to fully
understand such diverse and large collections of raw data being produced, their metadata need
to be integrated in a non-ambiguous and computational amenable way [
          <xref ref-type="bibr" rid="ref19 ref23 ref9">9, 13</xref>
          ].
Ontologies can be loosely defined as “a vocabulary of terms and some specification of their
meaning” [
          <xref ref-type="bibr" rid="ref17 ref24 ref7">7, 14</xref>
          ]. If an ontology is accepted as a reference by the community (e.g.,
the Gene Ontology), then its representation of its domain becomes a standard, and data
integration and sharing facilitated. The complex process of enriching a resource with
metadata by means of semantically defined properties pointing to other resources often
requires human input and domain expertise. Thus, the proposed approach assumes that
by rewarding and recognizing metadata sharing and integration on the semantic web
using standard and controlled vocabularies, we are promoting and intensifying scientific
collaboration and progress.
        </p>
        <p>Figure 1 illustrates the
Semantic Web in action with two
datasets annotated with its
respective metadata using a
hypothetical Metal Ontology. A
dataset including Gold Market
Stats contains an ontology
mapping (e.g., an RDF triple) to
the concept Gold, and another
dataset Silver Market Stats
contains an ontology mapping to the
concept Silver. Given that Gold
and Silver are both coinage met- Fig. 1. An hypothetical metal ontology and dataset
mapals, a semantic search engine is pings.
capable of identifying as
relevant both datasets when asked
for market stats of coinage metals.</p>
        <p>Now, we need to define the value of metadata in terms of knowledge it provides
about a given dataset. Semantic interoperability is a key requirement in the realization
of the semantic web and it is mainly achieved through mappings between resources.
For example, all dataset mappings to ontology concepts are to some extent important to
enhance the retrieval of that dataset, but the level of importance varies across mappings.
The proposed approach assumes that metadata can be considered as a set of links where
all the links are equal, but some links are more equal than others (adaption of George
Orwells quote). Thus, the proposed approach aims at measuring the knowledge rating
of any given dataset through its mappings to concepts specified in an ontology.
3</p>
        <p>Knowledge rating
The proposed approach assumes that the metadata integration and sharing value of a
dataset, dubbed as knowledge rating, is proportional to the specificity and
distinctiveness of its mappings to ontology concepts in relation to all the others datasets in the
Linked Data Cloud.</p>
        <p>
          The specificity of a set of ontology concepts can be defined by the information
content (IC) of each concept, which was introduced by [
          <xref ref-type="bibr" rid="ref21">11</xref>
          ]. For example, intuitively the
concept dog is more specific than the concept animal. This can be explained because the
concept animal can refer to many distinct ideas, and, as such, carries a small amount of
information content when compared to the concept dog, which has a more informative
definition. The distinctiveness of a set of ontology concepts can be defined by its
conceptual similarity [
          <xref ref-type="bibr" rid="ref12 ref2 ref22">2,12</xref>
          ] to all the others sets of ontology concepts, i.e. a distinctiveness
of a dataset is high if there are no other semantically similar datasets available.
Conceptual similarity explores ontologies and the relationships they contain to compare their
concepts and, therefore, the entities they represent. Conceptual similarity enables us to
identify that arm and leg are more similar than arm and head, because an arm is a limb
and a leg is also a limb. Likewise, because an airplane contains wings, the two concepts
are more related to each other than wings is to boat.
        </p>
        <p>
          Most implementations of IC and conceptual similarity only span a single domain
specified by an ontology [
          <xref ref-type="bibr" rid="ref10 ref20">10</xref>
          ]. However, realistic datasets frequently use concepts from
distinct domains of knowledge, since reality is rarely unidisciplinary. So, the scientific
challenge is to propose innovative algorithms to calculate the IC and conceptual
similarity using multiple-domain ontologies to measure the specificity and distinctiveness
of a dataset. Similarity in a multiple ontology context will have to explore the links
between different ontologies. Such correspondences already exist for some ontologies
that provide cross-reference resources. When these resources are unavailable, ontology
matching techniques can be used to automatically create them.
4
        </p>
        <p>Reward and recognition mechanism
The reward and recognition mechanism can rely on the implementation of a new virtual
currency, dubbed KnowledgeCoin (KC), that will be specifically designed to promote
and intensify the usage of semantic web technologies for scientific data integration and
sharing. The idea is that every time a scientific article is published, KCs are distributed
according to the knowledge rating of the datasets supporting that article. Note that KCs
should by no means be a new kind of money and the design of KC transactions will
focus on the exchange of scientific data and knowledge.</p>
        <p>After developing the knowledge rating measures, they can be used to implement the
supply algorithm of a new virtual currency, KC. This will not only aim at validating the
usefulness of the proposed knowledge ratings but also deliver an efficient reward and
recognition mechanism to promote and intensify the usage of semantic web
technologies for scientific data integration and sharing. Unlike conventional cryptocurrencies,
the KCs will rely on a trusted central authority and not on a cryptographically proof.
But even without being a cryptocurrency, the KC will take advantage of the technical
solutions provided by existing cryptocurrencies, such as bitcoin5.</p>
        <p>The scientific challenge is to create a trusted central authority that issue new KCs
when new knowledge is created in the form of a scientific article, as long as it
references a supporting dataset properly integrated in the Linked Data Cloud. If there is no
reference to the dataset in the Linked Data Cloud no KCs will be issued. This way,
researchers will be incentivized to publically share the dataset, including the raw data
or at least a description of the raw data, in the Linked Data Cloud. If a dataset is shared
through the Linked Data Cloud then its level of integration will be measured by its
knowledge rating. This way, researchers will be encouraged to properly integrate their
data. The success of this mining process will rely on the trustworthiness of the
knowledge ratings, and therefore will further validate the developed measures.</p>
        <p>From recognition researchers may get reputation, and from reputation they may
get a reward. For example, researchers recognize the relevance of a research’s work
by citing it, and by having a high number of citations the researcher obtains a strong
reputation, which may in the end help him to be rewarded with a project grant. Thus,
KCs can be interpreted as a form of reputation that in the end can result in a reward.
However, we can also design and implement direct reward mechanisms through KCs
transactions as a way to establish a virtual marketplace of scientific data and knowledge
exchanges. The main scenario of a KCs transaction is to represent the exchange of
datasets identified by an URI from the data provider to the data consumer, which may
include recognition statements.
The design of the approach is ongoing work and its direction depends on a more detailed
analysis of many social and technical challenges that its implementation poses. For
example, some of the issues that need to be further studied and discussed: i) knowledge
ratings implementation, i.e. their validation, aggregation, performance, exceptions, and
extension to any mappings besides the ontological ones; ii) potential abuses, such as the
creation of spam mappings and other security threats; iii) central trusted authority for
the KC vs. the peer-to-peer mechanisms used by bitcoin; iv) use case scenarios for the
KC, e.g. exchange of datasets and their characterization based on KC transactions.</p>
        <p>In a nutshell, this paper presents the guidelines for delivering sound knowledge
rating measures to serve as the basis of a systematic reward and recognition mechanism
5 http://bitcoin.org/
based on KCs for improving the trust and quality of data through proper data integration
and sharing on the semantic web. The proposed idea aims to be the first step in providing
an effective solution towards data silos extinction.</p>
        <p>Acknowledgments
The anonymous reviewers for their valuable comments and suggestions. Work funded by the
Portuguese FCT through the LASIGE Strategic Project (PEst-OE/EEI/UI0408/2014) and SOMER
project (PTDC/EIA-EIA/119119/2010).</p>
        <p>References
Towards the Definition of an Ontology for Trust
in (Web) Data
Davide Ceolin, Archana Nottamkandath,
Wan Fokkink, and Valentina Maccatrozzo</p>
        <p>VU University, Amsterdam, The Netherlands
{d.ceolin,a.nottamkandath,w.j.fokkink,v.maccatrozzo}@vu.nl
Abstract. This paper introduces an ontology for representing trust that
extends existing ones by integrating them with recent trust theories.
Then, we propose an extension of such an ontology, tailored for
representing trust assessments of data, and we outline its specificity and its
relevance.</p>
        <p>
          Keywords: Trust, Ontology, Web Data, Resource Definition Framework (RDF)
1
1 Stephen Marsh, “Trust: Really, Really Trust”, IFIP Trust Management Conference
2014 Tutorial
Trust is a widely explored topic within a variety of computer science areas.
Here, we focus on those works directly touching upon the intersection of trust,
reputation and the Web. We refer the reader to the work of Sabater and Sierra
[19], Artz and Gil [
          <xref ref-type="bibr" rid="ref12 ref2">2</xref>
          ], and Golbeck [
          <xref ref-type="bibr" rid="ref22">12</xref>
          ] for comprehensive reviews about trust
in artificial intelligence, Semantic Web and Web respectively. Trust has also
been widely addressed in the agent systems community. Pinyol and Sabater-Mir
provide an up-to-date review of the literature in this area [18].
        </p>
        <p>
          We extend the ontology proposed by Alnemr et al. [
          <xref ref-type="bibr" rid="ref1 ref11">1</xref>
          ]. We choose it because:
(1) it focuses on the computational part of trust, rather than on social and agent
aspects that are marginal to our scope, and (2) it already presents elements that
are useful to represent computational trust elements. Nevertheless, we propose
to extend it to cover at least the main elements of the trust theory of O’Hara,
that are missing from their original ontology, and we highlight how these
extensions can be beneficial to model trust in (Web) data. Viljanen [21] envisions
the possibility to define an ontology for trust, but puts a particular emphasis
on trust between people or agents. Heath and Motta [
          <xref ref-type="bibr" rid="ref24">14</xref>
          ] propose an ontology
for representing expertise, thus allowing us to represent an important aspect of
trust, but again posing more focus on the agents rather than on the data. A
different point of view is taken by Sherchan et al. [20], who propose an ontology
for modeling trust in services.
        </p>
        <p>
          Goldbeck at al. [
          <xref ref-type="bibr" rid="ref23">13</xref>
          ], Cesare et al. [
          <xref ref-type="bibr" rid="ref15 ref5">5</xref>
          ] and Huang et al. [15] propose ontologies
for modeling trust in agents. Although these could be combined with the ontology
we propose (e.g., to model the trust in the author of a piece of data), for the
moment their focus falls outside of the scope of our work, that is trust in data.
        </p>
        <p>
          Trust has been modeled also in previous works of ours [
          <xref ref-type="bibr" rid="ref16 ref17 ref18 ref19 ref6 ref7 ref8 ref9">7,6,8,9</xref>
          ] using generic
models (e.g., the Open Annotation Model [
          <xref ref-type="bibr" rid="ref14 ref4">4</xref>
          ] or the RDF Data Cube
Vocabulary [
          <xref ref-type="bibr" rid="ref10 ref20">10</xref>
          ]). Here we aim at providing a specific model for representing trust.
3
        </p>
        <p>A Definition of Trust in Short
We recall here the main elements of “A Definition of Trust” by O’Hara [17], that
provide the elements of trust we use to extend the ontology of Alnemr et al.
Tw&lt;Y,Z,R(A),C&gt; (Trustworthiness) agent Y is willing, able and
motivated to behave in such a way as to conform to behaviour R, to the benefit
of members of audience A, in context C, as claimed by agent Z.</p>
        <p>Tr&lt;X,Y,Z,I(R(A),c),Deg,Warr&gt; (Trust attitude) X believes, with
confidence Deg on the basis of warrant Warr, that Y’s intentions, capacities and
motivations conform to I(R[A],c), which X also believes is entailed by R(A),
a claim about how Y will pursue the interests of members of A, made about
Y by a suitably authorised Z.</p>
        <p>
          X places trust in Y (Trust Action) X performs some action which
introduces a vulnerability for X, and which is inexplicable without the truth of
Trust attitude.
We are interested in enabling the sharing of the trust values regarding both
trust attitude and actions, along with their provenance. The ontology proposed
by Alnemr et al. [
          <xref ref-type="bibr" rid="ref1 ref11">1</xref>
          ] captures the basic computational aspects of these trust
values. However we believe that it lacks some peculiar trust elements that are
present in the theory of O’Hara, and thus we extend that ontology as shown
in Figure 12. Compared with the ontology of Alnemr et al., we provide some
important additions. We clearly identify the parts involved in the trust relation:
Trustor (source): every trust assessment is made by an agent (human or not),
that takes his decision based on his policy and on the evidence at his disposal;
Trustee (target): the agent or piece of information that is actually being
trusted or not trusted. This class replaces the generic “Entity” class, as it
emphasizes its role in the trust relation.
        </p>
        <p>We also distinguish between the attitude and the act of trusting.
Trust Attitude Object: it represents the graded belief held by the trustor in
the trustworthiness of the trustee and it is treated as a quality attribute
when deciding if to place trust in the trustee or not. It replaces the
reputation object defined by Alnemr et al. because it has a similar function to
it (quantifying the trust in something), but implements the trust attitude
relation defined by O’Hara that is more precise and complete (e.g. warranties
are not explicitly modeled by the reputation object);
Trust Action Object: the result of the action of placing trust. Placing trust
is an instantaneous action based on an “outright” belief. Therefore the trust
value is likely to be a Boolean value.</p>
        <p>Role, Context and Warranty: in the original ontology, the criterion is a
generic class that contextualizes the trust value. We specialize it, to be able to
model the role and the context indicated in the theory of O’Hara, as well as
the evidence on which the trust value is based, by means of the warranty.</p>
        <p>The trustworthiness relation is not explicitly modeled, since it falls outside
our current focus. We discuss this further in the following section. The remaining
elements of the model shown in Figure 1 are part of the original trust ontology.
These include a criterion for the trust value, and an algorithm that allows
combining observations (warranties) into a trust value (an algorithm is used also to
determine the value of the trust action). The trust attitude value corresponds
to the Deg element of the theory of O’Hara. We model the action both when it
is performed and when it is not. Both trust values are modeled uniformly.
5</p>
        <p>Modeling Trust in Data
In the previous section we provided an extended ontology that aims at capturing
the basic elements that are involved in the process of taking trust decisions. Here
we focus on the specificity of trusting (Web) data.
hasCriteria 1...*
hasTrustAttitudeValues</p>
        <p>Warranty</p>
        <p>*
hasWarranty
hasRole1
*</p>
        <p>Context</p>
        <p>*
hasContext rdfs:subClassof</p>
        <p>1 1
Criterion
1</p>
        <p>Description</p>
        <p>Name
* 1...*
CalculatedBy</p>
        <p>rdfs:subClassof
hasValue
part of
1...*</p>
        <p>1 *</p>
        <p>Trust Attitude Value
hasCriteria
hasTrustValues
hasCriteria</p>
        <p>Role</p>
        <p>*
collectedBy</p>
        <p>Assessment
rdf:type rdf:type</p>
        <p>Rating
Fig. 1. Extended Trust Ontology. We highlight in red the added elements and in yellow
the updated ones.</p>
        <p>
          Data are used as information carriers, so actually trusting data does not mean
to place trust in a sequence of symbols. It rather means to place trust in the
interpretation of such a sequence of symbols and on the basis of its trustworthiness.
For instance, consider a painting reproducing the city “Paris” and its annotation
“Paris”. To trust the annotation, we must have evidence that the painting
actually corresponds to the city Paris. But, to do so, we must: (1) give the right
interpretation to the word “Paris” (e.g., there are 26 US cities and towns named
“Paris”), and (2) check if one of the possible interpretations is correct. Both in
the case the picture represents another city or in the case the picture represents
a town named Paris which existence we ignored, we would not place trust in the
data, but for completely different reasons. One possible RDF representation of
the above example is: exMuseum:ParisPainting ex:depicts dbpedia:Paris,
where we take for granted the correctness of the subject and of the property
and, if we accept the triple, we do so because we believe in the correctness of the
object in that context (represented by the subject), and role (represented by the
property). We make use of the semantics of RDF 1.1 [22], from which we recall
the elements of a simple interpretation I of an RDF graph:
1. A non-empty set IR of resources, called the domain or universe of I.
2. A set IP, called the set of properties of I.
3. A mapping IEXT : IP → P (IR × IR).
4. A mapping IS: IRIs → (IR ∪ IP). An IRI (Internationalized Resource
Identifier [
          <xref ref-type="bibr" rid="ref21">11</xref>
          ]) is a generalization of a URI [
          <xref ref-type="bibr" rid="ref13 ref3">3</xref>
          ].
5. A partial mapping IL from literals into IR
Also, the following semantic conditions for ground graphs hold:
a. if E is a literal then I(E) = IL(E)
b. if E is an IRI then I(E) = IS(E)
c. if E is a ground triple s p o. then I(E) = true if I(p) ∈ IP and the pair
&lt;I(s),I(o)&gt; ∈ IEXT(I(p)) otherwise I(E) = false.
d. if E is a ground RDF graph then I(E) = false if I(E’) = false for some triple
E’ ∈ E, otherwise I(E) = true.
        </p>
        <p>Items 1, 2, 4 and 5 map the URIs, literals and the RDF triples to real-world
objects. We are particularly interested in Item 3, that maps the property of an
RDF triple to the corresponding real-world relation between subject and object.
Trusting a piece of data means to place trust in the information it carries, in a
given context. The trust context can be represented by means of the subject and
object of an RDF triple, so their semantic interpretation is assumed to be known
by the trustor. If the trustor trusts the triple, he believes that the interpretation
of the object o makes the interpretation of the triple s p o true:</p>
        <p>TrustAttitudetrustor(o|s, p) = Belief trustor(∃I(o) : &lt;I(s),I(o)&gt; ∈ IEXT (I(p))</p>
        <p>Belief is an operator that maps logical propositions to values that quantify
their believed truth, e.g., by means of subjective opinions [16] quantified in the
Deg value of the theory of O’Hara and based on evidence expressed by Warranty.</p>
        <p>By virtue of items c and d, we do not model explicitly the trustworthiness
relation defined by O’Hara: we consider an object o to be trustworthy by virtue
of the fact that it is part of an RDF triple that is asserted.</p>
        <p>Criterion</p>
        <p>Trustor (Source)
rdf:type
ex:me
rdf:type
hasContext
hasRole
rdf:object</p>
        <p>CO
Trustee (Target)
rdf:type
dbpedia:Paris</p>
        <p>hasSource rdf:type
hasContext TO1
hasTarget hasTrustAttitudeValues
0.8 calculatedBy
TrustAlgorithm
Fig. 2. Snapshot of the trust ontology, specialized for modeling data trustworthiness.</p>
        <p>Figure 2 presents a snapshot of the trust ontology modeling the example
above and adding a trust attitude value computed with a sample trust algorithm.
6</p>
        <p>Conclusion
In this paper we introduce an ontology for trust representation that extends
an existing model with recent trust theories. We specialize it in order to model
data-related trust aspects, and we motivate our design choices based on standard
RDF 1.1 semantics. This model is still at a very early stage, but it emerges from
previous research and from standard trust theories. In the future, it will be
extended, and evaluated in depth, also by means of concrete applications.
1. R. Alnemr, A. Paschke, and C. Meinel. Enabling reputation interoperability
through semantic technologies. In I-SEMANTICS, pages 1–9. ACM, 2010.
2. D. Artz and Y. Gil. A survey of trust in computer science and the semantic web.</p>
        <p>Journal of Semantic Web, 2007.
3. T. Berners-Lee, R. Fielding, and L. Masinter. Uniform Resource Identifier (URI):
Generic Syntax (RFC 3986). Technical report, IETF, 2005. http://www.ietf.
org/rfc/rfc3986.txt.
4. S. Bradshaw, D. Brickley, L. J. G. Castro, T. Clark, T. Cole, P. Desenne,
A. Gerber, A. Isaac, J. Jett, T. Habing, B. Haslhofer, S. Hellmann, J. Hunter,
R. Leeds, A. Magliozzi, B. Morris, P. Morris, J. van Ossenbruggen, S.
SoilandReyes, J. Smith, and D. Whaley. Open Annotation Core Data Model. http:
//www.openannotation.org/spec/core, 2012. W3C Community Draft.
5. S. J. Casare and J. S. Sichman. Towards a functional ontology of reputation. In</p>
        <p>AAMAS, pages 505–511. ACM, 2005.
6. D. Ceolin, A. Nottamkandath, and W. Fokkink. Automated Evaluation of
Annotators for Museum Collections using Subjective Logic. In IFIPTM 2012, pages
232–239. Springer, May 2012.
7. D. Ceolin, A. Nottamkandath, and W. Fokkink. Semi-automated Assessment of</p>
        <p>Annotation Trustworthiness. In PST 2013. IEEE, 2013.
8. D. Ceolin, A. Nottamkandath, and W. Fokkink. Efficient semi-automated
assessment of annotations trustworthiness. Journal of Trust Management, 1(1):3, 2014.
9. D. Ceolin, W. R. van Hage, and W. Fokkink. A Trust Model to Estimate the</p>
        <p>Quality of Annotations using the Web. In WebSci 2010. Web Science Trust, 2010.
10. R. Cyganiak, D. Reynolds, and J. Tennison. The rdf data cube vocabulary.
Technical report, W3C, 2014.
11. M. Dürst and M. Suignard. Internationalized Resource Identifiers (IRIs) (RFC
3987). Technical report, IETF, 2005. http://www.ietf.org/rfc/rfc3987.txt.
12. J. Golbeck. Trust on the World Wide Web: A Survey. Foundations and Trends in</p>
        <p>Web Science, 1(2):131–197, 2006.
13. J. Golbeck, B. Parsia, and J. A. Hendler. Trust networks on the semantic web. In</p>
        <p>CIA, pages 238–249. Springer, 2003.
14. T. Heath and E. Motta. The Hoonoh Ontology for describing Trust Relationships
in Information Seeking. In PICKME 2008. CEUR-WS.org, 2008.
15. J. Huang and M. S. Fox. An ontology of trust: Formal semantics and transitivity.</p>
        <p>In ICEC ’06, pages 259–270. ACM, 2006.
16. A. Jøsang. A logic for uncertain probabilities. Intl. Journal of Uncertainty,
Fuzziness and Knowledge-Based Systems, 9(3):279–212, 2001.
17. K. O’Hara. A General Definition of Trust. Technical report, University of</p>
        <p>Southampton, 2012.
18. I. Pinyol and J. Sabater-Mir. Computational trust and reputation models for open
multi-agent systems: A review. Artif. Intell. Rev., 40(1):1–25, June 2013.
19. J. Sabater and C. Sierra. Review on computational trust and reputation models.</p>
        <p>Artificial Intelligence Review , 24:33–60, 2005.
20. W. Sherchan, S. Nepal, J. Hunklinger, and A. Bouguettaya. A trust ontology for
semantic services. In IEEE SCC, pages 313–320. IEEE Computer Society, 2010.
21. L. Viljanen. Towards an ontology of trust. In Trust, Privacy, and Security in</p>
        <p>Digital Business, pages 175–184. Springer, 2005.
22. W3C. RDF 1.1 Semantics. http://www.w3.org/TR/2014/
REC-rdf11-mt-20140225/, 2014.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Ho¨fig, E. Supporting Trust via Open Information Spaces.
          <source>Proc IEEE 36th Annual Computer Software and Applications Conference</source>
          , pp.
          <fpage>87</fpage>
          -
          <lpage>88</lpage>
          , (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. World Wide Web Consortium.
          <source>: RDF 1.1 Concepts</source>
          and
          <string-name>
            <given-names>Abstract</given-names>
            <surname>Syntax</surname>
          </string-name>
          , W3C Recommendation, http://www.w3.org/TR/rdf11-concepts
          <string-name>
            <surname>/</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. World Wide Web Consortium.
          <source>: SPARQL 1</source>
          .
          <article-title>1 Query Language</article-title>
          , W3C Recommendation, http://www.w3.org/TR/sparql11-query/ (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. National Institute of Standards and Technology.:
          <article-title>Secure Hash Standard (SHS)</article-title>
          .
          <source>Federal Information Processing Standards Publication 180-4</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. World Wide Web Consortium.
          <article-title>: XML Signature Syntax</article-title>
          and
          <string-name>
            <surname>Processing (Second Edition</surname>
          </string-name>
          ),
          <source>W3C Recommendation</source>
          , http://www.w3.org/TR/xmldsig-core/ (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cloran</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Irwin</surname>
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>XML Digital Signature and RDF</article-title>
          .
          <source>Proc. Information Security South Africa</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Duerst</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suignard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Internationalized Resource Identifiers (IRIs). Internet Engineering Task</surname>
          </string-name>
          Force - Request
          <source>for Comments</source>
          <volume>3987</volume>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. World Wide Web Consortium.
          <source>: RDF 1</source>
          .
          <article-title>1 XML Syntax, W3C Recommendation</article-title>
          , http://www.w3.org/TR/rdf-syntax-grammar/ (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Carroll</surname>
            ,
            <given-names>J. J.</given-names>
          </string-name>
          :
          <source>Signing RDF Graphs. Proc. of the Second International Semantic Web Conference, LNCS 2870</source>
          , pp.
          <fpage>369</fpage>
          -
          <lpage>384</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Carroll</surname>
            ,
            <given-names>J. J.</given-names>
          </string-name>
          :
          <source>Matching RDF Graphs. Proc. of the First International Semantic Web Conference, LNCS 2342</source>
          , pp.
          <fpage>5</fpage>
          -
          <lpage>15</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alsheikh-Ali</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qureshi</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Mallah</surname>
            ,
            <given-names>M.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ioannidis</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          :
          <article-title>Public availability of published research data in high-impact journals</article-title>
          .
          <source>PloS one 6</source>
          (
          <issue>9</issue>
          ),
          <year>e24357</year>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          2.
          <string-name>
            <surname>Couto</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pinto</surname>
            ,
            <given-names>H.S.:</given-names>
          </string-name>
          <article-title>The next generation of similarity measures that fully explore the semantics in biomedical ontologies</article-title>
          .
          <source>Journal of bioinformatics and computational biology</source>
          <volume>11</volume>
          (
          <issue>05</issue>
          ) (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          3. ELIXIR,
          <article-title>EU-OPENSCREEN, BBMRI</article-title>
          , EATRIS, ECRIN, INFRAFRONTIER, INSTRUCT, ERINHA,
          <string-name>
            <surname>EMBRC</surname>
          </string-name>
          , Euro-BioImaging, LifeWatch, AnaEE, ISBE, MIRRI:
          <article-title>Principles of data management and sharing</article-title>
          at European Research Infrastructures (DOI:105281/zenodo8304, Feb
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          4.
          <string-name>
            <surname>Galilei</surname>
          </string-name>
          , G.:
          <article-title>Sidereus Nuncius, or The Sidereal Messenger</article-title>
          . University of Chicago Press (
          <year>1989</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          5.
          <string-name>
            <surname>Goodman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pepe</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blocker</surname>
            ,
            <given-names>A.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borgman</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cranmer</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crosas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Di</surname>
            <given-names>Stefano</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Gil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Groth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Hedstrom</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          , et al.:
          <article-title>Ten simple rules for the care and feeding of scientific data</article-title>
          .
          <source>PLoS computational biology</source>
          <volume>10</volume>
          (
          <issue>4</issue>
          ),
          <year>e1003542</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hey</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tansley</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tolle</surname>
            ,
            <given-names>K.M.:</given-names>
          </string-name>
          <article-title>The fourth paradigm: data-intensive scientific discovery (</article-title>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jasper</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uschold</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>A framework for understanding and classifying ontology applications</article-title>
          .
          <source>In: Proceedings 12th Int. Workshop on Knowledge Acquisition, Modelling, and Management KAW</source>
          . vol.
          <volume>99</volume>
          , pp.
          <fpage>16</fpage>
          -
          <lpage>21</lpage>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          8.
          <string-name>
            <given-names>National</given-names>
            <surname>Research</surname>
          </string-name>
          <article-title>Council (US). Committee on Models for Biomedical Research: Models for biomedical research: a new perspective</article-title>
          .
          <source>National Academies</source>
          (
          <year>1985</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          9.
          <string-name>
            <surname>Noy</surname>
            ,
            <given-names>N.F.</given-names>
          </string-name>
          :
          <article-title>Semantic integration: a survey of ontology-based approaches</article-title>
          .
          <source>ACM Sigmod Record</source>
          <volume>33</volume>
          (
          <issue>4</issue>
          ),
          <fpage>65</fpage>
          -
          <lpage>70</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          10.
          <string-name>
            <surname>Pedersen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pakhomov</surname>
            ,
            <given-names>S.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patwardhan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chute</surname>
            ,
            <given-names>C.G.</given-names>
          </string-name>
          :
          <article-title>Measures of semantic similarity and relatedness in the biomedical domain</article-title>
          .
          <source>Journal of biomedical informatics 40(3)</source>
          ,
          <fpage>288</fpage>
          -
          <lpage>299</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          11.
          <string-name>
            <surname>Resnik</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Using information content to evaluate semantic similarity in a taxonomy</article-title>
          .
          <source>In: Proceedings of the 14th international joint conference on Artificial intelligence-Volume</source>
          <volume>1</volume>
          . pp.
          <fpage>448</fpage>
          -
          <lpage>453</lpage>
          . Morgan Kaufmann Publishers Inc. (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          12. Rodr´ıguez,
          <string-name>
            <given-names>M.A.</given-names>
            ,
            <surname>Egenhofer</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.J.:</surname>
          </string-name>
          <article-title>Determining semantic similarity among entity classes from different ontologies. Knowledge and Data Engineering</article-title>
          , IEEE Transactions on
          <volume>15</volume>
          (
          <issue>2</issue>
          ),
          <fpage>442</fpage>
          -
          <lpage>456</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          13.
          <string-name>
            <surname>Uschold</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gruninger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Ontologies and semantics for seamless connectivity</article-title>
          .
          <source>ACM SIGMod Record</source>
          <volume>33</volume>
          (
          <issue>4</issue>
          ),
          <fpage>58</fpage>
          -
          <lpage>64</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          14.
          <string-name>
            <surname>Uschold</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gruninger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Ontologies: Principles, methods and applications</article-title>
          .
          <source>The knowledge engineering review</source>
          <volume>11</volume>
          (
          <issue>02</issue>
          ),
          <fpage>93</fpage>
          -
          <lpage>136</lpage>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>