<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Expressing No-Value Information in RDF</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Computer Science, Free University of Bozen-Bolzano</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>RDF is a data model to represent positive information. Consequently, it is not clear how to represent the non-existence of information in RDF. We present a technique to express such information in RDF and incorporate it into SPARQL query answering. Given an empty query answer, our technique can distinguish whether it is empty due to possibly incomplete information, or non-existent information.</p>
      </abstract>
      <kwd-group>
        <kwd>RDF</kwd>
        <kwd>negative knowledge</kwd>
        <kwd>nulls</kwd>
        <kwd>SPARQL Reason to visit</kwd>
        <kwd>To discover how no-value information can (1) be represented in RDF and (2) be used to infer SPARQL query emptiness</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The notion of no-value information was rst introduced in the relational
databases [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. There, the term `null value' was used, which may have di erent
meanings: there exists no value (i.e., non-existence); there exists a value but
it is unknown; or it is unknown whether a value exists. For the second case,
we can leverage RDF blank nodes, while for the third case, the Open World
Assumption (OWA) of RDF simply permits it. However, RDF cannot represent
the rst case, which is the no-value nulls, while in fact this no-value information
is useful to distinguish from incomplete information. Furthermore, by having
no-value information, an empty query answer can have two di erent meanings:
whether it is empty because of possibly incomplete information, or whether it is
truly empty from information that does not exist in the real-world.
      </p>
      <p>In this poster, we present a technique for representing no-value information
in RDF and incorporating such information into query answering. This
introduction is followed by a formalization of no-value information and query answering
in the presence of such information, a concrete RDF representation of no-value
information, and a discussion.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Formalization</title>
      <p>Preliminaries. Assume there are three pairwise disjoint in nite sets I (IRIs), L
(literals) and V (variables). A tuple (s; p; o) 2 I I (I [ L) is called a triple.
An RDF graph G consists of a nite set of triples.</p>
      <p>
        SPARQL is the standard query language for RDF [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The basic building
blocks of a SPARQL query are triple patterns, which look like RDF triples,
except that in each position, variables are also allowed. In this work, we focus
on the conjunctive fragment of SPARQL where queries are represented as basic
graph patterns (BGPs), that is, sets of triple patterns. The evaluation of a BGP
P over G is de ned as JP KG = f j P G and dom( ) = var (P ) g. Given a
query Q = (W; P ), where P is a BGP and W var (P ) is the set of distinguished
variables, the evaluation JQKG is the restriction of JP KG to W . Over Q, we de ne
the prototypical graph P~ as the graph resulting from mapping each variable
in P to a fresh IRI. The prototypical graph encodes any possible graph that
can satisfy the query. Furthermore, a CONSTRUCT query has the abstract form
(CONSTRUCT P1 P2) where both P1 and P2 are BGPs. Evaluating a CONSTRUCT
query over G in a graph where P1 is instantiated with all the mappings in
JP2KG.
      </p>
      <p>Let us now formalize no-value information. We rst de ne no-value
statements to capture which information is non-existent.</p>
      <p>De nition 1 (No-Value Statement). A no-value statement N is de ned as
No(P ) where P is a BGP. To N , we associate the CONSTRUCT query QN =
(CONSTRUCT P P ).</p>
      <p>
        We use BGP to have a exibility to represent complex no-values which need more
than one triple patterns. We then de ne an incomplete data source to model the
OWA of RDF graphs. As in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], an incomplete data source G = (Ga; Gi) is a pair
of an available graph Ga and an ideal graph Gi such that Ga Gi. Here, an
available graph is the graph that we have, whereas an ideal graph is a possible
extension over the available graph, which represents a version of ideal, complete
information.
      </p>
      <p>Having no-value statements restricts the possibilities of ideal graphs since
they must not contain any instantiations of the information denoted by the
statements. Over a graph G, we de ne the transfer operator TN (G) = SN2N JQN KG.
We de ne the semantics of no-value statements as follows.</p>
      <p>De nition 2 (Satisfaction of No-Value Statements). An incomplete data
source G = (Ga; Gi) satis es a set N of no-value statements, written as G j= N ,
if and only if TN (Gi) = ;.</p>
      <p>Note that since Ga Gi holds by the de nition of an incomplete data source,
TN (Gi) = ; implies TN (Ga) = ;. Next, we de ne the emptiness of a query over
an incomplete data source.</p>
      <p>De nition 3 (Query Emptiness). Let G = (Ga; Gi) be an incomplete data
source and Q a query. To express that Q is empty, we write Empty(Q). It is
the case that G j= Empty(Q) if and only if JQKGi = ;.</p>
      <p>Query emptiness over one incomplete data source does not always mean
that it always holds also over other incomplete data sources. For this reason, we
de ne that the entailment of a set N of no-value statements and query emptiness
Empty(Q) holds, written as N j= Empty(Q), if for any incomplete data source
G j= N , we have that G j= Empty(Q). If the entailment holds, we can guarantee
that the query will always return an empty answer no matter which possible
extensions of a graph are considered. We have the following theorem to check if
a set of no-value statements can guarantee the emptiness of queries.
Theorem 1 (Query Emptiness Entailment from No-Value Statements).
Let N be a set of no-value statements, Q be a query, and P~ be the prototypical
graph of Q. It is the case that N j= Empty(Q) if and only if TN (P~) 6= ;.
Example 1. Let N = No(f (obama; child; ?c); (?c; gender; male) g) be a no-value
statement about Obama having no sons. Consider the query Q = (f?c; ?sg;
f (obama; child; ?c); (?c; gender; male); (?c; school; ?s) g) asking for the schools of
Obama's sons. We have that TfNg(P~) 6= ;. Thus, from Theorem 1, it holds that
fN g j= Empty(Q). This means that Q returns an empty answer because of
nonexistence of the information that is asked, not by the incompleteness of the data
source. In contrast, suppose the constant male in the query Q were a variable
?g. If Q returns an empty answer over the data source, that may be due to the
incompleteness of the data source.</p>
      <p>As seen in the above example, if there is some part of the query that cannot
return any answer due to no-value information, then the whole query does not
return any answer. Now, we can distinguish between empty query answers from
possibly incomplete information, and empty query answers from non-existent
information.</p>
    </sec>
    <sec id="sec-3">
      <title>3 RDF Representation of No-Value Statements</title>
      <p>To concretely represent no-value statements in RDF, we use the rei cation
technique. Given a no-value statement No(f (s1; p1; o1); : : : ; (sn; pn; on) g), we
represent the statement as a resource of the class NoValStatement, while for each
triple pattern, we use a blank node with the properties subject, predicate,
and object. Variables are also represented via blank nodes with the property
varName. Each of the patterns' blank nodes is linked to the statement's resource
via the property hasPattern. For instance, we represent the no-value statement
\Obama has no sons" as follows:</p>
    </sec>
    <sec id="sec-4">
      <title>4 Discussion</title>
      <p>
        In this paper, we present a technique for representing no-value information in
RDF and checking the emptiness of queries based on such information. The
novalue information restricts possible extensions of an RDF graph wrt. OWA. As
a consequence, queries always return an empty answer if they try to capture
novalue information. No-value information can also be seen as a stronger version of
completeness information, since it also enforces that all possible extensions must
contain no corresponding information. Furthermore, when a query is ensured to
always return an empty answer, it is obvious that the query is also complete.
Hence, this work can complement the completeness reasoning framework for
RDF data sources as described in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Another use of no-value information is for
data cleaning. If we assume that no-value statements are correct, then we can
detect dirty data sources by checking if they contain information that has been
stated to be non-existent by the statements. For future work, we will study the
relation of our approach to OWL and to more expressive queries.
Acknowledgments The research was supported by the projects \MAGIC:
Managing Completeness of Data" funded by the Bolzano province and \CANDy:
Completeness-Aware Querying and Navigation on the Web of Data" by UniBZ.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Fredo</given-names>
            <surname>Erxleben</surname>
          </string-name>
          , Michael Gunther, Markus Krotzsch, Julian Mendez, and
          <string-name>
            <given-names>Denny</given-names>
            <surname>Vrandecic</surname>
          </string-name>
          .
          <article-title>Introducing Wikidata to the Linked Data Web</article-title>
          .
          <source>In ISWC</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Serge</surname>
            <given-names>Abiteboul</given-names>
          </string-name>
          , Richard Hull, and
          <string-name>
            <given-names>Victor</given-names>
            <surname>Vianu</surname>
          </string-name>
          .
          <source>Foundations of Databases. Addison-Wesley</source>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Steve</given-names>
            <surname>Harris</surname>
          </string-name>
          and Andy Seaborne, editors.
          <source>SPARQL 1</source>
          .
          <article-title>1 Query Language</article-title>
          .
          <source>W3C Recommendation, 21 March</source>
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Fariz</given-names>
            <surname>Darari</surname>
          </string-name>
          , Werner Nutt, Giuseppe Pirro, and
          <string-name>
            <given-names>Simon</given-names>
            <surname>Razniewski</surname>
          </string-name>
          .
          <article-title>Completeness Statements about RDF Data Sources and Their Use for Query Answering</article-title>
          .
          <source>In ISWC</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>