<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A General Approach to Query the Web of Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Xin Liu</string-name>
          <email>liu@disi.unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Science and Engineering, University of Trento</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>With the development of the Semantic Web, an increasing amount of data with semantics has been published on the web according to the Linked Data principles and become ubiquitous. The requirement of utilizing the entire web of data to answer a query has arised since the desire for the application of such ubiquitous semantic data sources. However, limited research work has been made on the query above the web of data and there is not a formal way to describe the query processing procedure. In this paper, we propose a general query processing method on the web of data, which contains three steps: data inference configuration, data discovery, result generation and ranking. Finally, we briefly represent the work already done and the future work.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>advance, the available approaches or theories of querying upon the web of data
is still missing .</p>
      <p>In this paper, we try to fill the research gap by proposing a general approach
to execute SPARQL queries on the entire web of data. The task of semantic web
search engines (like Sindice, Falcns, etc) belongs to the information retrieval
approaches while the work in this paper mainly derived from the execution
of (SPARQL) queries on the semantic web. The proposed mechanism mainly
includes three steps: data inference configuration, data discovery and extraction,
result generation and ranking. Data inference configuration is used to configure a
query processing procedure, that is whether allowed to use implicit information
got from data reasoning to answer a query or not. The task of data discovery
and extraction is to discover new data sources those are unknown in advance
for the query processing according to the data inference configuration. The final
results will be get and ranked in the result generation and ranking step.</p>
      <p>The remainder of the paper is structured as follows: in section 2, I briefly
introduce the state of the art of the data querying methods. In section 3, I
propose a general approach to process a query on the web of data and finally, I
summerize this proposal and briefly represent the future work in section 4.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>With the development of the semantic web, the query process on the semantic
data is moving forward from one single data source to various, distributed data
sources and even the entire web of data.</p>
      <p>DARQ4 is a query engine for federated SPARQL queries. It provides a
transparent query access to multiple, distributed SPARQL endpoints like querying a
single RDF graph. SEWimQ[7] is similar to the DARQ approach. It provides
a middleware that enables the virtual access to distributed RDF data sources.
It set up an RDF endpoint for each non-RDF format data source, make use of
query parser, optimizer and engine to rewrite a SPARQL and federate the results.
Virtuoso5, a comprehensive data integration software developed by OpenLink
Software, combines the functionality of a traditional RDBMS, virtual database,
RDF, XML, free-text, web application server and file server functionality in a
single system. It enables the integration of numerous heterogeneous data from
distributed data sources. However, all of the methods mentioned above have the
precondition that the data sources to be used are already known before the query
process.</p>
      <p>The Sematic Web Client Library6 presents the complete Semantic Web
as a single RDF graph and support to query on the semantic web. It makes use
of a pipeline based approach to dynamically retrieve information on the web
by dereferencing HTTP URIs or query semantic web search engines during the
process of a query. Paolo et al. present a formal description of query the web of</p>
      <sec id="sec-2-1">
        <title>4 http://darq.sourceforge.net/</title>
      </sec>
      <sec id="sec-2-2">
        <title>5 http://virtuoso.openlinksw.com/</title>
      </sec>
      <sec id="sec-2-3">
        <title>6 http://www4.wiwiss.fu-berlin.de/bizer/ng4j/semwebclient/</title>
        <p>data in [8]. They propose a model of open collection of RDF graphs on the web
and three different ways in which a query can be answered.</p>
        <p>The approaches adopted in the two papers try to solve the problems to some
extent on the data discovery during the query processing but they do not make
a further research on the data relevance analysis and result ranking algorithms.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Proposed Methods and Preliminary Results</title>
      <p>The semantic web languages such as OWL, RDF make open world assumption
and assume incomplete information by default. One single ontology may not
answer a query but a set of ontologies can do. So the query processing on the
web of data requires data integration. On the current semantic web, there are
significant amount of distributed ontologies. One single RDF graph may not
include enough data to get answers for a query but may contribute to part of
them, thus if one query is answerable on the web of data, then the integration
of possible relevant RDF graphs can complete the answers.</p>
      <p>The methodology of the query processing approach is based on the analysis of
the query language. On the semantic web, SPARQL is a W3C recommendation
language to query RDF data. A relational algebra for SPARQL7 can be used to
analyze and express a SPARQL query in relational algebra. In order to answer a
SPARQL query, it is required to find all solutions [4] for each triple pattern in the
query. The set of solutions for one triple pattern can be treated as an “relation”
in relational algebra, called “RDF relation”. And then the relational algebra
operators, such as selection, projection, rename, inner join, etc. can be executed
on such RDF relations to get the final answer. Therefore, the following discussion
focus on the solutions for one triple pattern. As long as we get solutions for each
triple pattern, we can get the final answers through relational operators.</p>
      <p>
        The goal of global query processing is to use the entire (or as much as
possible) web of data to execute SPARQL queries and generate answers. Many issues
have arised due to the features of “openness” and “incompleteness” of the RDF
graphs. (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ). The data can be published anywhere, we cannot find all the data
to answer a query; (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ). People don’t know the schema of each data source so
that we cannot send a precise query to a specific RDF data source as we use
SQL to query relational databases; (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ). The answer should include not only the
explicit information represented in RDF data but also the implicit information
which can be got through data inference. Based on the issues above, I propose a
way with three steps to query the entire web of data using semantic web query
language.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Data Inference Configuration</title>
        <p>The first step to execute a query is to configure the data inference. The answer of
a query may be various according to the function of data inference. Current
available RDFS/OWL reasoners mainly include Jena, Pallet, FaCT++, etc. Some</p>
        <sec id="sec-3-1-1">
          <title>7 http://www.hpl.hp.com/techreports/2005/HPL-2005-170.html</title>
          <p>rules in RDFS/OWL vocabularies used to generate implicit statements includes:
rdfs:subClassOf, rdfs:subPropertyOf, rdfs:domain, rdfs:range, owl:equivalentClass,
owl:equivalentProperty, owl:inverseOf, owl:tran
sitiveProperty, owl:sameAs, etc. Besides, user-defined rules can also be used in
Jena reasoner, like “uncle rule”: (?x &lt;http://foo.com/rel/botherOf&gt;?y) (?y
&lt;http://foo.com/rel/fatherOf&gt;?z) = (?x &lt;http://foo.com/rel/uncleOf&gt;?z).
One or multiple inference rules can be used to complete the execution of a query.
The strategy of data discovery will depend on the configuration result. As long
as the inference configuration is fixed, we can further make a decision on the
data discovery strategy.
3.2</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Data Discovery</title>
        <p>One of the main idea of the semantic web is to use information from anywhere
on the web. When a query is executed, it is infeasible to locate every data source
in advance. Therefore, in order to get answers from the web of data, the first
step should be the data discovery. The task of data discovery is to collect as
many data sources as possible those may be relevant to answer a query. Based
on different features of current RDF data sources on the web, there are two
ways to locate globally distributed sources, one is relying on the links between
different graphs and the other is depending on the semantic web search engines.</p>
        <p>In terms of the semantic data according to the Linked Data principle, we
can rely on RDF links to locate other data sources. RDF links take the form of
RDF triples, in which subject is a URI reference with the namespace of one data
source and object is a URI reference with the namespace of the other one. When
defining a global query, we should simultaneously set a starting point (an RDF
graph or a merge of a set of graphs) from which we can crawl the web of data
to collect more data sources. The staring point can be either dereferenced from
the URIs in the query or provided by users. The process of collecting interlinked
graph will follow the principle of web page collection of a web crawler.</p>
        <p>Besides, graphs may be discovered through the semantic web search engines
(e.g. Sindice8, Sig.ma9, Swoogle10, etc.). They can be used to collect the graphs
those include the same keywords or URIs. Due to the introduction of data
inference, the data discovery through search engines is an iteration process. Different
inference rules will generate different data discovery strategy. In general, it will
stop until there is no new statements found. Data discovery task through search
engines will collect two kinds of statements: one is the set of statements with the
same triple pattern as the query, the other is the set of statements with available
data inference rules.</p>
        <sec id="sec-3-2-1">
          <title>8 http://sindice.com/</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>9 http://sig.ma/ 10 http://swoogle.umbc.edu/</title>
          <p>Through the last two steps, potentially relevant statements have already been
discovered and collected. It is feasible to execute a query on such statements
relying on available reasoners. There may be many answers found for a query,
so it is necessary to return a ranking result according to the importance of
each answer. Evaluating the weight of one answer mainly lies on two factors:
the number of occurrences of one answer and the number of triples which can
get this answer. The more number of occurrences of one answer is and the less
number of triples is, the more weight of the answer is.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future work</title>
      <p>In this paper, we propose a general approach to complete the query processing
upon the web of data, which mainly includes data infererce configuration, data
discovery, result generation and ranking.</p>
      <p>
        For the future, my research work will focus on the following points: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ).
The design and development of efficient and precise data discovery strategy.
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ). The design of result ranking algorithm. (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ). The implementation of this
general approach. We can use some real world queries to evaluate this approach
on performance and quality. (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ). The definitions of relevant sub-graph (part of
an RDF graph) for a query. Currently, the granularity of relevance is a whole
RDF graph, but the fact is that not every triple in a graph is relevant to a query.
So the collection of relevant sub-graphs will make the query more efficient.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Tim</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          .
          <article-title>Design Issues: Linked Data</article-title>
          . Online, last change:
          <source>June</source>
          <year>2009</year>
          , http://www.w3.org/DesignIssues/LinkedData.html.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Jeen</given-names>
            <surname>Broekstra</surname>
          </string-name>
          and
          <string-name>
            <given-names>Arjohn</given-names>
            <surname>Kampman</surname>
          </string-name>
          .
          <article-title>Serql: An rdf query and transformation language</article-title>
          .
          <source>August</source>
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Hewlett-Packard Laboratories Bristol Eric</surname>
          </string-name>
          <article-title>Prud'hommeaux, Andy Seaborne</article-title>
          .
          <article-title>SPARQL Query Language for RDF</article-title>
          ,
          <year>January 2008</year>
          . http://www.w3.org/TR/rdfsparql-query/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Marcelo</given-names>
            <surname>Arenas Jorge Perez</surname>
          </string-name>
          and
          <string-name>
            <given-names>Claudio</given-names>
            <surname>Gutierrez</surname>
          </string-name>
          .
          <article-title>Semantics of sparql</article-title>
          .
          <source>Technical report</source>
          , May
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Greg</given-names>
            <surname>Karvounarakis</surname>
          </string-name>
          , Sofia Alexaki, Vassilis Christophides, Dimitris Plexousakis, and
          <string-name>
            <given-names>Michel</given-names>
            <surname>Scholl</surname>
          </string-name>
          .
          <article-title>Rql: A declarative query language for rdf</article-title>
          . pages
          <fpage>592</fpage>
          -
          <lpage>603</lpage>
          . ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Graham</given-names>
            <surname>Klyne</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jeremy J.</given-names>
            <surname>Carroll</surname>
          </string-name>
          .
          <article-title>Resource Description Framework (RDF): Concepts and Abstract Syntax</article-title>
          . http://www.w3.org/TR/rdf-concepts/.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Langegger</surname>
          </string-name>
          , Wolfram W¨oß, and
          <article-title>Martin Bl¨ochl. A Semantic Web Middleware for Virtual Data Integration on the Web</article-title>
          .
          <source>In ESWC</source>
          , pages
          <fpage>493</fpage>
          -
          <lpage>507</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Chiara</given-names>
            <surname>Ghidini Paolo Bouquet</surname>
          </string-name>
          and
          <string-name>
            <given-names>Luciano</given-names>
            <surname>Serafini</surname>
          </string-name>
          .
          <article-title>Query the web of data: A formal approach</article-title>
          .
          <source>In Inproceedings of Asian Semantic Web Conference</source>
          <year>2009</year>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>