<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Web-Scale Querying through Linked Data Fragments</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ruben Verborgh Sam Coppens</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miel Vander Sande Erik Mannens</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Multimedia Lab - Ghent University - iMinds Gaston Crommenlaan 8 bus 201 B-9050 Ledeberg-Ghent</institution>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Pieter Colpaert Rik Van de Walle</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>8</volume>
      <issue>2014</issue>
      <abstract>
        <p>To unlock the full potential of Linked Data sources, we need flexible ways to query them. Public sparql endpoints aim to fulfill that need, but their availability is notoriously problematic. We therefore introduce Linked Data Fragments, a publishing method that allows e cient o oading of query execution from servers to clients through a lightweight partitioning strategy. It enables servers to maintain availability rates as high as any regular http server, allowing querying to scale reliably to much larger numbers of clients. This paper explains the core concepts behind Linked Data Fragments and experimentally verifies their Web-level scalability, at the cost of increased query times. We show how trading server-side query execution for inexpensive data resources with relevant a ordances enables a new generation of intelligent clients.</p>
      </abstract>
      <kwd-group>
        <kwd>Linked Data</kwd>
        <kwd>querying</kwd>
        <kwd>availability</kwd>
        <kwd>scalability</kwd>
        <kwd>sparql</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Whenever there is a large amount of data, people will want to
query it—and nothing is more intriguing to query than the vast
amounts of Linked Data published over the last few years [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
With over 800 million triples in the widely used dbpedia [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], only
one of the many datasets in a large ecosystem, the need for various
specialized information searches has never been this high before.
sparql has been specifically designed [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] to fulfill this requirement
for reliable and standardized access to data in the rdf triple format.
Consisting of a query language [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and a protocol [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], sparql is the
de facto choice to publish rdf data in a flexible way, and allows to
select with high precision the data that interests us.
      </p>
      <p>
        There is one issue: it appears to be very hard to make a sparql
endpoint available reliably. A recent survey examining 427 public
endpoints concluded that only one third of them have an
availability rate above 99%; not even half of all endpoints reach 95% [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
To put this into perspective: 95% availability means the server is
unavailable for one and a half days every month. These figures are
quite disturbing given the fact that availability is usually measured
in “number of nines” [
        <xref ref-type="bibr" rid="ref25 ref5">5, 25</xref>
        ], counting the number of leading nines
in the availability percentage. In comparison, the fairly common
three nines (99.9%) amounts to 8.8 hours of downtime per year.
The disappointingly low availability of public sparql endpoints is
the Semantic Web community’s very own “Inconvenient Truth”.
      </p>
      <p>More precisely, practice reveals that the following three
characteristics are irreconcilable for sparql endpoints:
a) being publicly available;
b) o ering unrestricted queries;
c) having many concurrent users.</p>
      <p>
        This is because the load of a server is proportional to the product
of the variety, complexity, and amount of requests, the first two
of which remain virtually unbounded for sparql. Any endpoint’s
availability can be considerably improved by sacrificing one of
these three characteristics: private sparql endpoints perform well
because the server load can be predicted more reliably, limiting
query possibilities eliminates slow queries that can bring down the
server, and low demand of course contributes positively to
availability. http servers on the other hand have no problem combining these
characteristics, as the complexity of each request can be limited
because the server restricts what “queries” a client can execute by
determining the o ered http resources [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. We do not claim by
any means this comparison is fair, as sparql servers have to
perform significantly more work per request. On the contrary, it is
exactly because sparql requests require more processing that sparql
endpoints do not scale well compared to http servers.
      </p>
      <p>This paper challenges the idea that servers should spend their cpu
cycles on expensive queries, and proposes a model in which the
client solves a complex query by only asking the server for simple
data retrieval operations. Instead of answering a complex sparql
query, the server sends a Linked Data Fragment that corresponds to
a specific triple pattern. This fragment then contains metadata that
allows the client itself to execute the complex query. While this leads
to an increased number of http requests between clients and servers,
each request is answered easily and also fully cacheable. Therefore,
this is the scalable and sustainable approach to Web querying: with
sparql, each new client requires additional processing power from
the server, whereas with Linked Data Fragments, clients take care
of their own processing. We e ectively trade fast answers but
low scalability for increased (yet manageable) query times with
Web-level scalability. Most importantly, this makes it possible to
fully query datasets of publishers who cannot invest in hosting and
maintaining an expensive sparql endpoint—which is most of us.
Client</p>
      <p>Client
(a) sparql endpoints perform all processing on the server, leading to fast
query execution with low data bandwidth, and a rapidly overloaded server.
(b) ldf servers only support simple requests and can thus handle far higher
loads. Clients perform the querying, so they need more (cacheable) data.</p>
      <p>In the next section, we critically examine the scalability problems
of sparql. Next, we introduce Linked Data Fragments, followed by
the implementation of a server (Section 4) and a client (Section 5).
Section 6 evaluates the improved scalability. We then discuss our
method and its context in Section 7, and conclude in Section 8.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        At its core, the sparql query language [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] allows clients to find
triples based on basic graph patterns. For instance, consider the
following sparql query:
      </p>
      <p>SELECT ?p ?c WHERE {
?p a &lt;http://dbpedia.org/ontology/Artist&gt;.
?p &lt;http://dbpedia.org/ontology/birthPlace&gt; ?c.
?c &lt;http://xmlns.com/foaf/0.1/name&gt; "York"@en.
}</p>
      <p>Listing 1: Search for artists born in places named “York”.
Such queries facilitate searching relevant information in datasets
that can contain hundreds of millions of triples. Most triple stores,
such as Virtuoso and Sesame, o er a sparql interface, which is
referred to as a “sparql endpoint” when exposed through http.</p>
      <p>
        Before we dive into the details of sparql endpoints, let us first
briefly recapitulate the architectural properties of the Web and why
they enable the Web to scale the way it does. The Web is a
distributed hypermedia application that conforms to the constraints of
the Representational State Transfer architectural style (rest, [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]).
The main building blocks of the Web are resources, which are
identified by urls through the uniform interface o ered by http [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
Resources can be represented in hypermedia formats, which can
link to other resources. These links remove the need for the server to
maintain the application state between di erent interactions, as each
representation (and not the server) retains the next steps a client can
take [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The combination of the uniform interface and
statelessness makes it possible for intermediaries to cache server responses,
which significantly improves scalability.
      </p>
      <p>
        sparql endpoints essentially implement a protocol on top of http
through a strictly standardized set of constraints [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The client
sends a sparql query to a server, which executes it and sends back
the results. Given the amount of data involved and the arbitrary
complexity of queries, the server possibly needs to execute a
significant amount of work to obtain the results of each query. In
contrast to regular http servers, a sparql endpoint does not expose
resources on an application-specific level, but rather one “endpoint”
resource that acts as a data handling process [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], and an unlimited
set of “query answer” resources that correspond to all queries [
        <xref ref-type="bibr" rid="ref11 ref9">9, 11</xref>
        ].
Therefore, regular http caching strategies for resources below query
level cannot be applied; each unique query still needs full execution.
This query-based partitioning of resources gives sparql poor scaling
properties, as illustrated in Figure 1a. The inherent problem with
such an endpoint architecture is that the required time to generate
each query answer resource is potentially very high, and that all
processing needs to happen at the server side. While this makes
querying rather convenient for clients, it puts an enormous burden
on providers of sparql endpoints, as sparql engines can strain cpu
and ram intensively even for common queries [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. It should thus
not surprise us that maintaining high availability rates for public
sparql endpoints is exceptionally challenging [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This problem
seldom occurs with regular http servers, as the granularity of
offered resources can be adjusted such that each individual resource
does not require excessive processing time. Additionally, this finer
granularity allows those resources to be cached e ciently [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        The performance of sparql has been the subject of multiple
benchmarks [
        <xref ref-type="bibr" rid="ref29 ref4">4, 29</xref>
        ]. Several caching strategies have been proposed on
various levels, for instance, by placing a proxy in front of a sparql
endpoint [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], or by integrating caching information into the triple
store itself to allow http caching [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ]. However, these techniques
consider caching of results for entire queries, which means related
but non-identical queries do not benefit. Syntax-agnostic approaches
can cache based on the algebraic representation and allow subquery
caching [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ], also enabled by other specific techniques [
        <xref ref-type="bibr" rid="ref22 ref23 ref31">22, 23, 31</xref>
        ].
      </p>
      <p>
        A category of approaches for executing sparql queries over
Linked Data [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] is based on link traversal [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and relies on the
principle of dereferencing [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Link traversal strongly benefits from
caching [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] because the granularity is refined to the level of data
needed for queries—as opposed to the full result set of a single query.
This technique resembles the querying method we will introduce
in this paper, because of the active role clients play in fetching and
evaluating data, as well as the potential of pipelining through
nonblocking iterators [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. However, our method does not rely on one
primary data source per uri (a consequence of dereferencing) and
we use additional information to reduce the execution time of
typical queries by more than an order of magnitude. While optimizing
planning heuristics exist [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], our planning strategy employs more
reliable indicators. Furthermore, the present initial paper focuses
on vastly improving the scalability of individual endpoints, even
though the method is generalizable to distributed querying.
      </p>
      <p>
        Closely tied to the publication of Linked Data is the specification
of a standard read/write interface, which is the goal of the Linked
Data Platform (ldp, [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]). While the definitions in Section 3 will
seemingly demand a comparison with ldp, it is crucial to note that
ldp and Linked Data Fragments are orthogonal, i.e., a server can
choose to support either or both of them independently. More
specifically, ldp proposes a subject-centric read/write interface, while the
goal of Linked Data Fragments is to o er scalable query execution.
Our design permits any resource to additionally implement ldp.
3.
3.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>LINKED DATA FRAGMENTS</title>
    </sec>
    <sec id="sec-4">
      <title>Motivation</title>
      <p>
        As indicated above, the concept of querying through endpoints
entails serious availability issues [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the root cause of which is
the non-scalability of the expensive component, the sparql server
(Figure 1a). While all client–server interactions on the Web can
lead to server overloading, sparql is especially vulnerable to this
because of its partitioning in (potentially expensive) query answer
resources. For instance, compare dbpedia access through its sparql
endpoint1 versus its subject pages2. The former provides access to
the unlimited set of query answers, whereas the latter provides the
same data through a limited set of subject resources listing all triples
per subject. It it straightforward to understand that, regardless of
the used technology, the latter demands less server usage because
the underlying queries are answered easily by simple index lookups.
In fact, the second case does not even require an on-demand query
processor: because the subject set is finite, a static file server could
serve pre-generated subject pages, which are updated periodically
by another process. Furthermore, such a finite set can be cached
e ciently by regular http caches as several clients reuse the same
pages, whereas a large amount of sparql queries are client-specific.
      </p>
      <p>Admittedly, even though the same data is exposed in both cases,
the sparql endpoint is more powerful when available, because it
o ers custom client-centric views on specific parts of the data. In
contrast, the server-driven partitioning in subjects might or might
not be helpful for a specific client’s goal. Yet this is exactly the
reason why hosting a sparql endpoint is such a risky endeavor: the
scalability of http and thus the whole Web are based on e ective
partitioning of resources. It is only natural that a server goes down if
it commits itself to serving an unlimited set of expensive resources.
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Definitions and examples</title>
      <p>To solve these availability and scalability issues, we need to
create a compromise between o ering a very limited partitioning and
allowing unrestricted sparql queries. Each of the o ered resources
should additionally contain the necessary information for clients to
execute sparql queries e ciently themselves. To that end, we
introduce Linked Data Fragments, o ering a hybrid solution between
limited subject-based Linked Data dereferencing and the di cultly
scalable server-side sparql execution.</p>
      <p>Definition 1. A Linked Data Fragment (ldf) of a Linked Data
dataset is a resource consisting of those elements of this dataset
that match a specific selector, together with their metadata and the
controls to retrieve related Linked Data Fragments.</p>
      <p>
        We will first discuss the selector aspect, before we detail the
metadata and control constraints. The concept is not unlike that of
a media fragment [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ], which selects a part of a media resource.
We instead select parts of Linked Data resources, without a priori
restricting the kind of selector. Therefore, the data of an ldf could
for instance be that of a subject page (which would have the triple
pattern selector { &lt;s&gt; ?p ?o } for a specific &lt;s&gt;) or even a sparql
result resource (which would have the sparql query as a selector).
However, we are primarily interested in those ldfs that a) are
useful for client-side query answering b) only require a low server
processing cost. Therefore, we define the following:
Definition 2. A basic Linked Data Fragment (basic ldf) is a Linked
Data Fragment with a triple pattern as selector, count metadata,
and the controls to retrieve any other basic ldf of the same dataset,
in particular other fragments the matching elements belong to.
1http:// dbpedia.org/ sparql
2e.g., http:// dbpedia.org/ page/ Pete_Townshend (via Pubby [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ])
Linked Data subject pages o er only a subset of all basic ldfs,
namely those triple patterns with a fixed subject and variable
predicates and objects. However, to avoid exhaustive searches when
solving queries with variable predicates or objects, a partitioning
into basic ldfs includes all combinations of { ?s ?p ?o } with each
component either a variable or a specific uri or literal. For instance,
the basic ldfs for dbpedia include “triples with Pete Townshend as
subject”, “triples with The Who as object”, as well as “triples with
Pete Townshend as subject and birth place as predicate”. In other
words, as each component can either be variable or fixed, each triple
in a dataset belongs to exactly 23 = 8 basic ldfs.
      </p>
      <p>
        This data partitioning is only one of the aspects that sets ldfs apart
from alternatives. Definition 1 also mentions metadata and controls,
which are defined in the same open way as the possible selectors.
Together, they transform the ldf into an a ordance [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] that enables the
client to perform actions, in particular sparql query execution. Each
ldf thereby provides the client with context on how this fragment
relates to the dataset and other fragments. The metadata on the one
hand includes information such as the fragment’s selector, since the
client might have received this fragment from a third party, unaware
of the precise selector used. Controls on the other hand include
links to other fragments, allowing clients to discover more data.
Providing a ordance is a required part of any rest interface [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], as
it allows statelessness and reduces client–server coupling [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
      </p>
      <p>
        Which metadata and controls we need is constrained by
Definition 2. Since ldfs can be very long (for instance, dbpedia counts
more than 60 million matches for { ?s rdf:type ?o }), they
sometimes need pagination [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. To compensate, each basic ldf should
provide the (estimated) total number of triples that match the pattern.
As we will see in Section 5, this is crucial for e cient querying.
      </p>
      <p>
        Furthermore, each basic ldf should provide the (hypermedia)
controls to access other basic ldfs. A concrete implementation
could be that the basic ldf for the pattern { ?s rdf:type ?o } links
to the ldf for { ?s ?p foaf:Person } (if this fragment contains
foaf:Person triples). Additionally, each basic ldf representation
must contain a form or similar control that allows to retrieve any
basic ldf with a triple pattern selector of choice. This is necessary
for the independent evolution of ldf clients and servers: a client
should not need to know how a server exposes its ldf resources.
Concretely, servers are free to choose the urls of the ldf resources
they o er. This makes ldf compatible with a partitioning that does
require a specific url structure such as the sparql protocol, which
demands a “?query=” parameter [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>The next section will discuss the design and implementation of
ldf servers, followed by a section on the design of ldf clients that
can execute sparql queries through ldfs.
4.
4.1</p>
    </sec>
    <sec id="sec-6">
      <title>SERVER</title>
    </sec>
    <sec id="sec-7">
      <title>Architecture</title>
      <p>The desired architectural characteristics for a server of ldfs are
availability, scalability, and performance, in that order. This means
that, at any time, its top priority is to ensure clients can reach the
server and retrieve a response within a time interval that is similar
to other http servers. As typical http response time is of the order
of a few hundreds of milliseconds, this is what we aim for (and
preferably less). In addition, the infrastructure should scale with the
number of clients. We define an ldf server as follows.
Definition 3. A Linked Data Fragments server (ldf server) is an
http server that o ers Linked Data Fragments covering one or more
datasets in at least one triple-based representation.
Relational Database</p>
      <p>LDF Server</p>
      <p>HTTP Cache
Other Data Source</p>
      <p>Note that under this definition, sparql endpoints, subject page
servers, and http servers with data dumps are also ldf servers.
While servers choose which specific ldf partitionings they use, those
servers that o er basic ldfs strike an optimal balance between low
server-side complexity and e cient client-side querying (Section 5);
we call these basic ldf servers. In addition to triple-based
representations (such as Turtle or html with rdfa), servers can o er others
(such as text or regular html) and/or more ldfs with more complex
selectors. However, these features are optional, as complex selectors
might negatively impact server availability and performance.</p>
      <p>Figure 2 shows a schematic display of how an ldf server interacts
with its environment. On the back-end side, the server fetches
data from an underlying data source to construct ldf fragments.
Such a data source could be a triple store, perhaps through a sparql
endpoint, but even a relational database or an rdf source file. The
fact that we would still use sparql endpoints after criticizing them
for low availability seems a contradiction, but it is not: as stated in
Section 1, when the complexity of queries can be limited, endpoints
can perform very well. Since ldf servers are only required to obtain
results for single triple patterns, the endpoint is not stressed in any
way. Alternatively, regular relational databases can also perform
well because of the simplicity of the lookup patterns.</p>
      <p>
        The main task of the ldf server is o ering a rest interface [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
to its ldfs by providing and maintaining a url space. It translates
each request into a specific query for the appropriate data server that
collects the needed data and metadata. This is then combined into
an ldf and represented in a media type the client understands (such
as Turtle or html) by adding the data as well as the metadata and
controls that provide the a ordance towards next steps [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        On the front-end side, the ldf server can be proxied through
a regular http cache [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], restricting the load on the ldf server.
Furthermore, the ldf server can restrict load on the underlying data
sources by caching responses as well, for instance, using existing
sparql http caching mechanisms [
        <xref ref-type="bibr" rid="ref31 ref35">31, 35</xref>
        ].
      </p>
      <p>
        This architecture maximizes availability and performance by two
key decisions. First, the o ered resources consist of ldfs that are
simple to generate, minimizing processing time for each resource.
In contrast to endpoints o ering an unlimited sparql interface, this
places an upper bound on the execution time of each request; and
lower server loads directly lead to higher availability. Second, when
partitioning in basic ldfs, the entire dataset is exposed in a way
that maximizes reuse across clients, and hence enables e cient
caching. Furthermore, since the set of basic ldfs of a dataset is finite,
substantial parts can be pre-generated and pre-cached, leading to
lower server load and faster response times. The scalability is then
guaranteed through the properties induced by the rest architectural
style [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Caching at the front-end can happen hierarchically, and
load-balancing between multiple ldf servers is possible. Back-end
caching and load-balancing can happen as well, but synchronization
might be required to ensure consistency between di erent servers.
However, as the load caused by the front-end server is predictable,
a su cient infrastructure for synchronization can be planned.
      </p>
      <p>We have implemented an example ldf server, the source code
of which is available at http:// linkeddatafragments.org/ software/ .
A public instance of this server with several datasets is running at
http:// data.linkeddatafragments.org/ . We will demonstrate the
discussed features of ldfs through this public instance. Note that every
step in the following discovery process happens entirely through
the a ordance supplied by the ldfs, i.e., by using links and forms,
indicating the client’s decoupling from any server’s url structure.</p>
      <p>When you open http:// data.linkeddatafragments.org/ in a browser,
you will see links to di erent datasets. This start resource is in
fact an ldf that allows to browse all datasets on the server. One of
them is dbpedia, which is located at /dbpedia. This initial dbpedia
ldf lists some triples of the dataset to allow browsing. Using the
provided links, we can click through to see related fragments. For
example, when we click an rdf:type link of a triple, we arrive
at the basic ldf of all triples with the rdf:type predicate, located
at /dbpedia?predicate=rdf%3Atype. We can also use the form to
navigate to a specific basic ldf. For example, Figure 3 shows the ldf
for the pattern { ?s rdf:type dbpedia-owl:Person }, located at
/dbpedia?predicate=rdf%3Atype&amp;object=dbpedia-owl%3APerson.
While both urls follow a convention adopted by this particular
server, they remain opaque identifiers that servers can assign freely
to ldf resources as long as they provide the necessary controls.</p>
      <p>
        The html representation of ldfs generated by this server contains
rdfa markup to enable interpretation by automated clients. All
triples and metadata are annotated. However, parsing html involves
an overhead that can be avoided by directly parsing Turtle. For
that reason, our implementation also o ers Turtle representations of
each ldf through http content negotiation [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. For example:
curl http://data.linkeddatafragments.org/dbpedia \
-H "Accept: text/turtle"
This results in a Turtle representation of the ldf we retrieved
earlier. In contrast to html, which has &lt;a&gt; and &lt;form&gt; elements, rdf
o ers no native support for hypermedia controls [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]—except for
the uris of its triple components, which only allow dereferencing
(cf. dbpedia subject pages). Since basic ldfs must contain controls
towards all other basic ldfs of the dataset, we have to describe them
declaratively. This happens in three ways. First, for each of the
non-variable parts of the basic ldf’s triple pattern, rdfs:seeAlso
links are provided to the ldfs that have these parts in subject or
object position. Second, the representation provides an alternative
possible variable bindings B fg;
split pattern Q in connected subpatterns S = fS1; : : : ; Sng;
foreach subpattern Si 2 S do
foreach triple pattern tj in subpattern Si do
      </p>
      <p>Fj GET first page of basic ldf for tj through F;
end
return ? if any fragment Fj has 0 matching triples;
if all fragments Fj have exactly 1 matching triple then
B[Si][bj] nil 8tj where bj ..= binding of tj to Fj;
return ? if S bj is inconsistent;
Fm ..= Fj with minimal total number of matches;
Fm0 Fm [ f GET remaining pages of Fm g;
foreach binding bk of pattern tm in fragment Fm0 do</p>
      <p>
        Si0 apply binding bk to subpattern Si;
B[Si][bk] FindVariableBindings(Si0, Fm0 );
to the html form using the Hydra hypermedia api vocabulary [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ],
allowing the client to query any basic ldf of the dataset. Finally,
it o ers a dataset description using the void vocabulary [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which
defines properties such as triple count. These annotations give the
Turtle representation the same a ordance as its html counterpart.
4.3
      </p>
    </sec>
    <sec id="sec-8">
      <title>Dereferencing</title>
      <p>
        At first sight, it might appear that ldf voids the Linked Data
principles that enable dereferencing [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. After all, the fact that
the identifier of a concept (uri) also serves as its address (url)
forms the foundation of Linked Data. For example, not only does
http:// dbpedia.org/ resource/ Pete_Townshend uniquely identify the
musician Pete Townshend, it also a ords retrieving information
about him. In contrast, this information on the previous ldf server
is located at /dbpedia?subject=dbpedia%3APete_Townshend and
/dbpedia?object=dbpedia%3APete_Townshend. Yet, dereferencing
and ldfs actually play complementary roles, as indicated below.
      </p>
      <p>First, the use of ldfs does not break dereferencing. Since ldf
servers are not bound by url constraints, they can choose to serve
the ldf about the resource at its own url. In fact, dbpedia.org is
an ldf server: it could host the ldf with Pete Townshend as subject at
http:// dbpedia.org/ resource/ Pete_Townshend, and could in principle
also o er support for basic ldfs. This shows that dereferencing and
ldf can work in conjunction seamlessly.</p>
      <p>Second, the Web is founded on the idea that “anyone can say
anything about anything”. While dereferencing is fast and easy, it only
leads to the source that happens to host the identifier, which does
not mean this source also has the information we are looking for.
Compare this to regular Web browsing: the best source for objective
information about a certain company is likely not that company’s
homepage. It would be unpractical to assign a new identifier every
time another party wants to add statements about a resource.
Furthermore, no single representation can contain all facts; for instance,
http:// dbpedia.org/ ontology/ Person does not contain a list of all
people on the Web. A basic ldf server instead lets us indicate what
triples we want to obtain about a certain resource, di erentiating
between subject, predicate, and object positions. While a basic ldf
server would also not represent millions of people on a single page,
it allows to retrieve a list of them page by page through the paginated
basic ldf resource for the pattern { ?s a dbpedia-owl:Person }.
That way, we can ask to obtain all dbpedia-owl:Person instances
from any dataset, even when not hosted on the dbpedia url space.
Additionally, dereferencing only works with urls, whereas ldf
affordances also function with generic uris.</p>
      <p>
        Third, the fourth Linked Data principle demands to include links
to other resources [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This means representations of resources such
as http:// dbpedia.org/ resource/ Pete_Townshend could link to ldfs
that contain more data. This closes the circle, as ldfs themselves
link to a) the concept’s url (through the data) and b) related ldfs
(through the metadata). The main di erence between dereferencing
and ldfs is that the former uses the implicit a ordance in the url,
whereas ldfs explicitly provide multiple a ordances.
      </p>
    </sec>
    <sec id="sec-9">
      <title>CLIENT</title>
    </sec>
    <sec id="sec-10">
      <title>Querying basic Linked Data Fragments</title>
      <p>The scalability of ldfs as depicted in Figure 1b can be achieved
when the server o ers a partitioning that is inexpensive to generate
but still allows e cient client-side querying. When using basic ldfs,
leading to a partitioning in basic triple patterns, clients can solve
queries for basic graph patterns autonomously and e ciently. Each
resource operation requires only minimal cost from the server, is
fully cacheable, and likely to be reused.
1 Function FindVariableBindings(Q, F)</p>
      <p>Input: basic graph pattern query Q and start fragment F
Output: possible variable bindings (nil if none needed)</p>
      <p>B where B[si][bk] 6= ?;
Algorithm 1: An ldf client can e ciently find the possible variable
bindings of any basic graph pattern through basic ldfs.</p>
      <p>The main task of an ldf client is to find possible variable bindings
of queries such as the one in Listing 1. Algorithm 1 details this
process for basic ldfs. Count metadata is used to ensure an e cient
solution path (line 13). FindVariableBindings has three possible
types of output for queries and their subqueries:
the empty set ? if no valid binding exists. For instance, given
a query { ?person foaf:name "Fake Name"@en }, no value
for ?person exists, so the set of possible bindings is empty.
nil if no binding is necessary to satisfy the query. For instance,
{ dbpedia:Keith_Moon foaf:name "Keith Moon"@en } is
satisfied if (and only if) the corresponding triple exists.
a hierarchical list of bindings in all other cases. For instance,
a solution to the query in Listing 1 could have bindings for
?c, each of which can have one or multiple bindings for ?p.</p>
      <p>We will now run through a possible execution of the algorithm for
the query in Listing 1, assuming a basic ldf server with the dbpedia
dataset. We invoke FindVariableBindings with Q Listing 1 and
F the basic ldf at http:// data.linkeddatafragments.org/ dbpedia.
As the entire query pattern is connected (i.e., there exists a path from
any triple to any other by following shared variables), there is only
one subpattern: S = fS1g = fQg. There are three triple patterns in
S1, so we use the controls in F to GET the corresponding basic ldfs.
For each of them, we receive the first 100 matches and metadata:
fragment selector matches
F1 { ?p a dbpedia-owl:Artist } 68,237
F2 { ?p dbpedia-owl:birthPlace ?c } 469,849</p>
      <p>F3 { ?c foaf:name "York"@en } 12
As each fragment has more than 1 match, lines 8 to 11 of the
algorithm are skipped and we go straight to line 13 where F3 is
selected as smallest fragment Fm. Since the 12 triples fit on one
result page, Fm0 = Fm. The possible bindings of ?c in Fm0 include
dbpedia:York_(album), dbpedia:York, dbpedia:York,_Ontario,
dbpedia:York,_New_York, dbpedia:28220_York, and seven others.
Each of those is in turn bound to S1 (= Q), which gets stored in S10.
We now follow the recursive invocation of FindVariableBindings
with Q S10 = S1 bound to ?c = dbpedia:York and F Fm0 . The
graph pattern query Q thus becomes:
t1 ?p a dbpedia-owl:Artist.
t2 ?p dbpedia-owl:birthPlace dbpedia:York.
t3 dbpedia:York foaf:name "York"@en.</p>
      <p>This time, there are two connected subpatterns: S1 = ft1;t2g with
the first two triple patterns containing variable ?p, and S2 = ft3g
with the last triple without variables. For S2, only 1 matching triple
exists, so this results in nil as no binding is necessary. For each
triple pattern in S1, basic ldfs are retrieved (F1 was cached):
fragment selector
F1 { ?p a dbpedia-owl:Artist }
F2 { ?p dbpedia-owl:birthPlace dbpedia:York }
matches
68,237
75</p>
      <p>Since F2 has the lowest number of matches, it is used for Fm0 = Fm.
Possible bindings for ?p include dbpedia:Paul_Banks_(musician),
dbpedia:Eddie_Robson, dbpedia:Thomas_Turton, and 72 others.</p>
      <p>For each of them, FindVariableBindings is executed again. We
will follow the execution with the following Q parameter:
t1 dbpedia:Eddie_Robson a dbpedia-owl:Artist.</p>
      <p>t2 dbpedia:Eddie_Robson ...:birthPlace dbpedia:York.
Both corresponding basic ldfs are retrieved (F2 can be generated
from cache). Essentially, we verify whether the triple t1 exists, i.e.,
whether Eddie Robson is an artist according to the dbpedia dataset.
F1 and F2 both have 1 matching triple, so the check at line 9 is
successful; no further bindings are necessary.</p>
      <p>By contrast, if we would follow the execution for the binding
?p = Thomas_Turton, who was a mathematician, F1 would have
0 matches, which results in the empty binding on line 8.</p>
      <p>When all of the 75 possibilities have been scanned for artists,
control is returned to the earlier FindVariableBindings invocation,
which is now at line 21. Out of 75 matches for people with York as
birthplace, 12 are artists. They are returned as part of B, which also
contains the binding to dbpedia:York.</p>
      <p>Execution continues similarly for the 11 other matches for ?c at
the highest level, most of which have no matches for birthPlace.
Finally, the returned bindings B from this level are the following:
?c = dbpedia:York
– ?p = dbpedia:Eddie_Robson
– ?p = dbpedia:Dustin_Gee
– ?p = dbpedia:Paul_Banks_(musician)
– ?p = dbpedia:Johnny_Leeze
– ?p = dbpedia:Joe_Van_Moyland
– ?p = dbpedia:Mark_Simpson_(journalist)
– ?p = dbpedia:David_Reed_(comedian)
– ?p = dbpedia:Andrew_Martin_(novelist)
– ?p = dbpedia:Sam_Forrest
– ?p = dbpedia:Seebohm_Rowntree
– ?p = dbpedia:Peter_John_Allan
– ?p = dbpedia:John_Barry_(composer)
?c = dbpedia:York,_Ontario</p>
      <p>– ?p = dbpedia:Dawn_Langstroth
From these bindings, the result set can be generated unambiguously.
They are the same results we get when executing the query on the
dbpedia sparql endpoint (given the same version of the dataset).</p>
      <p>This algorithm has been implemented in the ldf client, which
has been made available at http:// linkeddatafragments.org/ software/ .
Caching is added where possible, so the same ldf is only retrieved
once—even though the algorithm might need it multiple times.
5.2</p>
    </sec>
    <sec id="sec-11">
      <title>Querying other Linked Data Fragments</title>
      <p>
        In Section 5.1, we explained how any basic graph pattern can
be solved at the client side by retrieving basic ldfs. However, not
all Linked Data servers will be partitioned (only) in basic ldfs;
some will support more detailed ldfs (e.g., sparql endpoints), others
will merely support less detailed fragments (e.g., Pubby subject
pages [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]). While we believe that basic ldfs strike a fair balance
between server e ort and client e ort, Algorithm 1 can be extended
to optimally query servers with any ldf partitioning.
      </p>
      <p>The main di erence would be how a subpattern Si is divided in
fragment selectors (line 3). Since each basic ldf corresponds to
a single triple pattern selector tj, the original algorithm retrieves
fragments Fj for each triple pattern. If the partitioning in ldfs is
di erent, the subpattern can be divided in other fragment selectors
to minimize the number of needed requests. However, each of those
requests might be more (or less) expensive to a server, so the server
should carefully consider which partitioning it o ers.</p>
      <p>
        Below are examples of possible ldf partitionings.
a (limited) sparql endpoint – If the server o ers sparql, each sparql
query corresponds to an ldf with that query as selector. While
a full sparql endpoint would be able to solve any basic graph
pattern directly, it would su er from the aforementioned
scalability issues. It is therefore beneficial to limit the possible
query forms. For instance, if an endpoint would only
allow sparql queries containing up to two triple patterns, the
query discussed in Section 5.1 could be solved faster, since
the Artist/birthPlace subpattern could be retrieved in one
request instead of having to test for Artist 75 times.
basic ldfs with extra data – Any basic ldf server is free to send
extra triples along that might be helpful to a client. For
instance, a server could decide to always send the rdfs:label
of any triple pattern component. That way, if the query in
Listing 1 would additionally ask for artists’ labels, no extra
requests would be necessary.
only subject pages – When the server o ers only subject-based
dereferencing (such as dbpedia subject pages), triple patterns
with variable subjects cannot be retrieved easily. In that case,
a lot more requests are needed; the algorithm in fact becomes
regular Linked Data querying with link traversal [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ].
Before a client can decide how a subpattern can be divided, it
must know the available partitioning(s) of the server. They can be
advertised in rdf using void and/or other vocabularies.
6.
6.1
      </p>
    </sec>
    <sec id="sec-12">
      <title>EVALUATION</title>
    </sec>
    <sec id="sec-13">
      <title>Experimental design</title>
      <p>The main characteristic of basic ldfs is that they allow a much
higher availability and scalability than other ldf partitionings such
as sparql result sets. The primary purpose of this evaluation is thus
to verify whether the availability and scalability of ldf client/server
setups is significantly higher than that of sparql endpoints. To this
end, we will execute a series of sparql queries against a sparql
endpoint and through clients of a basic ldf server.</p>
      <p>
        We built a prototype implementation of a basic ldf client that can
execute sparql queries consisting of basic graph patterns against
a basic ldf server. Since existing benchmark suites [
        <xref ref-type="bibr" rid="ref29 ref4">4, 29</xref>
        ] use
additional features such as filters, we could not meaningfully reuse
them here. An examination of logs from popular endpoints such
as dbpedia revealed that these are presently not the best sources of
varied, non-trivial queries consisting of only basic graph patterns.
We therefore developed a generator of basic graph pattern queries,
available at https:// github.com/ LinkedDataFragments/ Benchmarks,
which can provide us with varied queries for any given dataset. The
algorithm generates basic graph pattern queries Q = fq1; : : : ; qng,
where each query qi consists of triple patterns, using the following:
1. Select a random type &lt;t&gt; from the dataset ({_:s a ?t}) and
add the triple pattern {?s1 a &lt;t&gt;} to the query.
2. Select a random subject &lt;s&gt; with this type ({?s a &lt;t&gt;}).
3. Select a random property &lt;p&gt; of this subject ({&lt;s&gt; ?p _:o})
and add the pattern {?s1 &lt;p&gt; ?o1}.
4. Select matching objects &lt;o1&gt; and &lt;o2&gt; ({&lt;s&gt; &lt;p&gt; ?o}).
5. For non-literal &lt;o1&gt; and &lt;o2&gt;, find triples {&lt;o&gt; ?p3 ?o3}
and possibly {&lt;o3&gt; ?p4 ?o4}, using the results to augment
the query with further triple patterns.
      </p>
      <p>Below is an example query for dbpedia generated by this algorithm:
SELECT * WHERE {
?s1 a dbpedia-owl:Agent.
?s1 dbpedia-owl:associatedMusicalArtist ?o1.
?o1 dbpedia-owl:genre ?o2.
?o1 dbpedia-owl:recordLabel ?o3.
?o2 a dbpedia-owl:Genre.</p>
      <p>?o3 rdfs:label "Paramount Records"@en.</p>
      <p>}
We can see this is representative for the kind of queries we would be
interested in on dbpedia. The algorithm generated 275 such queries.
6.2</p>
    </sec>
    <sec id="sec-14">
      <title>Experimental setup</title>
      <p>For this experiment, we installed Virtuoso 7 and our basic Linked
Data Fragments server on a Ubuntu Linux machine (four 6-core
processors at 2.4 ghz, 24 gb ram). Virtuoso was configured with the
recommended optimal settings, but result caching was disabled to
ensure the results were served from the database and not from memory.
The English dbpedia 3.8 was then ingested (427,670,470 triples).</p>
      <p>
        Two di erent data sources were configured on the basic ldf server
(as in Figure 2). The first one is the Virtuoso 7 server described
above. The basic ldf server will execute two types of queries:
1) CONSTRUCT queries for basic graph patterns; 2) COUNT queries for
the same. While Virtuoso can execute the former really fast, counts
for large result sets are inherently slow. Therefore, we configured
a second data source with dbpedia in the hdt format (Header,
Dictionary, Triples) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. hdt is a compressed format for rdf that allows
fast triple patterns queries and fast (approximate) counts. The fact
that sparql only allows exact counts—even though approximations
are su cient for basic ldfs—is a major advantage for hdt.
server type
data source
sparql endpoint
Virtuoso 7
basic ldf server
Virtuoso 7 back-end
basic ldf server
hdt back-end
number
of clients
1 client
2 clients
4 clients
1 client
2 clients
4 clients
1 client
2 clients
4 clients
average
cpu usage
121.62%
241.51%
477.96%
66.58%
82.35%
116.30%
0.60%
0.67%
0.48%
average
ram usage
3.81 gb
4.52 gb
5.18 gb
3.32 gb
3.32 gb
3.41 gb
0.98 gb
1.09 gb
0.88 gb
      </p>
      <p>The http load testing tool JMeter was used to test the throughput
of queries with a distributed setup, alternating between 1, 2, and
4 physical client machines that each attempted to execute 10 queries
per second. If no response was received within 60 seconds, this was
considered a timeout. A monitor on the server sampled the cpu and
ram usage of the data source processes every second.
6.3</p>
    </sec>
    <sec id="sec-15">
      <title>Availability/scalability results</title>
      <p>Table 1 shows the averages of the measured cpu and ram usage
for the Virtuoso process and the hdt process. When we execute the
queries against the sparql endpoint, the cpu load on the Virtuoso
process is high and increases linearly with the number of clients;
ram usage also increases steadily. Extrapolating the cpu usage, our
24-core server could handle at most 20 clients at the same time.
Beyond that, availability would become compromised.</p>
      <p>The basic ldf server with Virtuoso as back-end handles an
increasing number of clients with less cpu load; cpu deltas are also lower.
Remarkably, ram usage remains constant, likely due to the fact that
no join operations need to be performed but only basic selections
and counts. Extrapolation reveals the server handles 46 clients.</p>
      <p>However, if we choose a data source that is optimized for basic
triple patterns and counts, such as hdt, we see that the scalability and
resulting availability could be improved drastically. hdt is not
cpubound or ram-bound, as it basically streams the needed segments
from disk. We hardly see the influence with a low number of clients.
Note however that these numbers only include the data access part
and not the cost of handling the http interactions; they will probably
be the first bottleneck in most scenarios.</p>
      <p>The above indicates basic ldf servers scale better than sparql
endpoints and thus can guarantee a much higher availability, certainly
with data sources optimized for triple pattern access and counts.
6.4</p>
    </sec>
    <sec id="sec-16">
      <title>Performance results</title>
      <p>Table 2 summarizes the media and average query times for our
test set, and the percentage of queries that time out (time &gt; 60 s).
Note how the median is in all cases far lower than the average,
indicating that there are outliers with a high query time.</p>
      <p>Without any doubt, a sparql endpoint such as Virtuoso solves
sparql queries much faster under availability. However, solutions
generated by basic ldf clients do not require excessive time: results
generally arrive in a matter of seconds. The query time for the basic
ldf server with 4 clients increased in our tests, yet this was not due
to the data process, as Table 1 reveals, but due to the http server
process, to which more cpu cycles could be allocated. Additionally,
regular http caching would allow major performance improvements
for the ldf server—and unlike sparql, even across di erent queries.
server type
data source
sparql endpoint
Virtuoso 7
basic ldf server 1 client
Virtuoso 7 back-end 2 clients</p>
      <p>4 clients
basic ldf server
hdt back-end
number
of clients
1 client
2 clients
4 clients
1 client
2 clients
4 clients
median
time
753 ms
837 ms
902 ms
1,539 ms
1,551 ms
1,743 ms
907 ms
922 ms
1,333 ms
average</p>
      <p>time
2,338 ms
2,544 ms
2,623 ms
6,136 ms
6,275 ms
6,214 ms
3,460 ms
3,520 ms
5,044 ms
timeouts
1.09%
1.45%
1.82%
4.73%
5.09%
3.73%
2.18%
2.18%
2.55%
7.1</p>
    </sec>
    <sec id="sec-17">
      <title>Linked Data Fragments in the Semantic</title>
    </sec>
    <sec id="sec-18">
      <title>Web context</title>
      <p>
        The sparql language and protocol have always been important
to the Semantic Web’s infrastructure, and we do not see a
necessity for this to change. However, we do question the scalability of
public endpoints that aim to o er unrestricted queries to a large
number of users. The main strength of the endpoint philosophy is
also its Achilles’ heel: the fact that one server accepts the
responsibility of answering arbitrarily complex requests inevitably leads
to availability problems, as evidenced by recent statistics [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. It is
important to understand that this cannot be solved by building more
e cient sparql servers—the problem is inherent to the concept of
such a powerful endpoint. The resource partitioning of regular http
servers on the Web can be chosen by its developers in such a way
that each resource can be delivered within acceptable bounds; the
resources of sparql endpoints are query results that can be
unpredictably complex. After all, for every 100 distinct queries a certain
sparql server can answer in one second, there exists at least one
query it cannot: the union of those queries.
      </p>
      <p>We do see several important roles for sparql endpoints. First, as
a private or internal data source, since the load is predictable; for
instance, as a back-end of Web or desktop applications, similar to
how relational databases are used. Second, when the query forms are
somehow constrained; for instance, by limiting the allowed number
of triple patterns or the execution time. Third, in environments where
the number of users is limited; for instance, for highly specialized
datasets. In those cases, the product of query variety, complexity,
and access rate, which correlates with endpoint load, is minimized
because one of its factors is controlled. For public sparql endpoints
with a high number of users, the only option to guarantee high
availability is to limit query complexity, but this often conflicts with
the motivations for o ering a queryable endpoint in the first place.</p>
      <p>
        This paper pleads to move the intelligence that enables
querying from the server to the client side. As clients have become
increasingly powerful compared to servers—even mobile devices
now exceed older laptops’ capabilities—a model in which the client
performs most of the work is realistic. This results in a significant
increase in scalability, as depicted in Figure 1b. Even though clients
have to issue many more requests, each of those requests 1) requires
minimal server processing cost—and the server decides how much
e ort it is willing to spend; 2) can be cached and reused across di
erent clients, as the granularity of responses is much finer. As the Web
has been designed with per-resource access and caching [
        <xref ref-type="bibr" rid="ref11 ref13">11, 13</xref>
        ],
ensuring that each resource can be rapidly generated and subsequently
reused contributes more to availability and scalability than o ering
highly specific and expensive resources.
      </p>
      <p>
        Performance-wise, the ldf querying approach cannot outperform
sparql; availability-wise, it certainly can. This has a considerable
impact on average query times, as shown in Table 3. If we look
at the sparql a scenario in which the server has 99% availability
(which is only the case with one third of sparql endpoints [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]), and
assuming the average episode of downtime lasts 15 minutes, then
an average query time of 0.2 seconds under availability comes down
to an average query time of 4.70 seconds in general:
0:2 s + (1
0:99)
      </p>
      <p>
        (15 min = 2) = 4:70 s
If we look at those endpoints with 95% availability (sparql b, less
than half of all endpoints [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]), the generalized average query time
increases to 45 seconds. In contrast, maintaining a 99.9% availability
level for a basic ldf server is reasonable; with increased query times
of 5 and 10 seconds, the generalized average query times become
respectively 5.15 and 10.15 seconds (Table 3).
use case
basic ldf a
basic ldf b
sparql a
sparql b
      </p>
      <p>average server average adjusted
query time availability downtime query time</p>
      <p>So while sparql is certainly an order of magnitude faster under
availability, actual availability percentages are su ciently low that,
when considering them in the average query time calculation, the
di erence with ldf querying becomes much smaller. The trade-o
is the increased usage of bandwidth, which might be acceptable for
desktop devices but perhaps di cult for mobile devices on slow
connections. The improved caching can partly compensate for this.</p>
    </sec>
    <sec id="sec-19">
      <title>Linked Data Fragments and Linked Data</title>
      <p>
        Above all, Linked Data Fragments are a publishing strategy for
Linked Data, with basic ldfs o ering a partitioning that allows
clientside querying at low server-side cost. Implementing basic ldfs can
be seen as adding additional constraints to the Linked Data
principles [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]: each basic ldf with a fixed subject ({ &lt;s&gt; ?p ?o }) has its
own http uri, represents triples about a certain subject, and includes
links to other documents that allow to discover more things (all
related basic ldfs). As discussed in Section 4.3, the uri-based
dereferencing concept is retained, and actually augmented with
hypermedia controls that allow to retrieve di erent fragments about a topic.
For instance, while dereferencing http:// dbpedia.org/ ontology/ Artist
only leads to dbpedia’s metadata of Artist, the same uri allows
basic ldf servers to show 1) their own metadata of Artist; 2) all
resources that have type Artist. Dereferencing a topic’s uri on a
basic ldf server might lead to the fragment of those triples that have
the topic as subject, and this fragment contains controls towards all
other basic ldfs of that dataset.
      </p>
      <p>Additionally, the way we have defined ldfs in Definition 1
allows to consider all existing published Linked Data sets as Linked
Data Fragments; an ldf is literally any “fragment” of a Linked Data
source. All of the following are ldf partitionings, from coarse- to
fine-grained: a single-file data dump in Turtle format, a dataset
exposed as subject pages, a collection of basic ldfs, a sparql
endpoint. Furthermore, the algorithm discussed in Section 5 and its
generalization allow to query any ldf partitioning, thereby providing
a means to evaluate which partitioning is best for e cient client-side
querying—while still guaranteeing server availability.</p>
      <p>Basic ldfs are not the only way of partitioning, but they set
an example for novel ways to publish Linked Data, with a focus
on enabling more intelligent clients through added metadata and
hypermedia controls. It would be interesting to see which other ldf
partitionings emerge and how they influence client capabilities.</p>
      <p>
        Of crucial importance is the independence of clients and servers.
While sparql is an expressive language, its use in a contract between
a client and a server determines to a certain extent the way a client
operates and behaves. Basic ldf servers impose a much less strict
contract. The resources they o er can be used to solve sparql
queries, but not that is not their only purpose. They can be used
for browsing, to solve queries in other languages or even without
a specific query language, to solve sparql queries partially (if not all
results are needed), and for several other purposes. In that way, ldfs
can enable a more serendipitous reuse [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] of Linked Data that is
able to transcend individual data silos. The key to such an approach
is that the client is in control of recombining individual pieces of
data, and inter-fragment links aid this combination process.
7.3
      </p>
    </sec>
    <sec id="sec-20">
      <title>Towards a new querying paradigm</title>
      <p>In addition to improving the availability of queryable Linked
Data sources, we believe that client-side querying can contribute
to a new querying paradigm. sparql approaches querying in the
traditional, non-Web-specific way: a client asks a question, the
server computes the answer while the client waits, and finally, the
client receives the whole answer at once. However, we should
ask ourselves how realistic and desirable such a single delineated
answer is in the context of an open and unpredictable Web. sparql
endpoints of course never pretend to o er complete answers (they
cannot, because no data source is ever complete); but each query is
answered with a finite-length response, and the entire query needs
to be asked again to check whether there are any changes.</p>
      <p>
        Therefore, when we say “Web-scale querying”, we do not only
mean our method of querying can technically scale with an
increasing number of clients; we also mean that ldfs are able to embrace
the open nature of the Web. Even though we presented the querying
algorithm in Section 5 in a synchronous way, its steps can actually
be completed asynchronously and iteratively, streaming
intermediate results as soon as they become available. This is not unlike other
Linked Data querying strategies [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], but on a smaller time scale
because of the more e cient partitioning of the data source.
      </p>
      <p>Concretely, partial results can be communicated as soon as they
are known. Revising the artists query example discussed in
Section 5, the uri of a person born in York could be sent directly after
it has been determined that he/she is an artist, without having to
wait until all other 74 York inhabitants have been checked. This
improves the latency of applications on behalf of which the requests
are made, as the results can either be shown iteratively as they arrive
(for instance, visualized on a map), or the first incoming results
might already be su cient to make a decision. It makes sense to
trade the idea that delineated queries demand delineated answers
for a more fluid way of querying answering. In theory, the artist
query could even run indefinitely, returning new answers as dbpedia
(or any other data source) gets updated. In some cases, tentative
answers might also be useful, e.g., “this person is a potential match
because she lives in a city named York, but the verification whether
she is an artist is pending”.</p>
      <p>
        Another aspect of being Web-scale is the use of more than one
data source. While sparql federated query [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] enables querying
data from multiple sparql endpoints, low sparql endpoint
availability makes the mechanism brittle. After all, if two endpoints
each have an availability of 95%, the a priori probability of both
endpoints being available decreases to 95% 95% 90%.
      </p>
      <p>ldfs allow querying of distributed sources in a transparent way.
Since the mechanism of basic ldfs is based on hyperlinks, each
ldf can link to ldfs on the same or another server. At no point in
Sections 4 and 5 have we used any knowledge about the server’s uri
structure, because only links and forms were followed. If, for any
reason, those links lead to another server, the querying algorithm
can be completed as usual. Furthermore, the client can decide to
have multiple starting fragments; for instance, it might ask birth
place information from dbpedia and use bbc MusicBrainz to verify
whether somebody is an artist. Interestingly, in sharp contrast to
sparql federation, ldf querying actually becomes faster when using
di erent data sources, because the http requests are distributed
across di erent servers. The use of di erent data sources also fits
well with iterative results: dbpedia might not contain the birthplace
of a certain person, while Freebase does.</p>
      <p>We end up with an information-gathering process that bears more
similarities with the way human consumers would answer questions.
Instead of posing the question to an omniscient oracle, we consult
targeted data sources to refine an answer iteratively.
8.</p>
    </sec>
    <sec id="sec-21">
      <title>CONCLUSIONS AND FUTURE WORK</title>
      <p>In this paper, we introduced the concept of Linked Data
Fragments, discussed the development of ldf servers and clients, and
made example implementations of a server and a client available at
http:// linkeddatafragments.org/ software/ . We thereby aim to
facilitate further experiments with more intelligent clients, starting with
a di erent resource publication strategy at the server side. Below
are various directions for future work.</p>
      <p>
        Above all, this paper strives to encourage research into o ering
datasets as fragments in addition to traditional partitionings such
as data dumps, subject pages, and sparql endpoints. Even though
basic ldfs already improve scalability and illustrate the powerful
architectural properties of fragments, they are likely not the final
destination of the quest for scalability. Other specifically designed
partitionings could reduce bandwidth, which would significantly
improve performance. The example query in Section 5 required
75 artist type checks; if they could somehow be bundled into fewer
requests, the entire query can execute much faster. One way to do
this is at the protocol level, for instance using http 2.0 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which
allows to send multiple requests to a single server more e ciently.
Another way is a more granular partitioning than basic ldfs, so that
multiple similar triple patterns can be queried at once. For instance:
dbpedia:{Dustin_Gee,Thomas_Turton} a dbpedia-owl:Artist.
This would decrease the number of needed requests, but each
individual request would become more expensive. Furthermore, caching
e ciency would be reduced. It is again up to the server to decide
how much processing time it is willing to spend on each resource.
      </p>
      <p>This brings us to another important research topic, namely how
servers can indicate what kind of resource partitioning they support.
A straightforward approach would be to create a vocabulary for
di erent types, such as “single-file data dump”, “subject pages”,
“basic ldfs”, and “limited/full sparql”. However, we envision that
di erent kinds of ldf partitionings will emerge, and that these might
even vary dynamically depending on server load. Perhaps a semantic
way to express the data, metadata, and hypermedia controls of each
fragment will be necessary.</p>
      <p>A next technological step is the implementation of a streaming
client. At the moment, the current algorithm and implementation
follow a bottom-up approach, where each iteration downloads all
pages from the smallest fragment. A top-down approach with a data
pipeline would read fragment data one page at a time. This would
make partial results available earlier, and thus allow faster decisions.</p>
      <p>This paper has focused on querying basic graph patterns. In time,
the full expressivity of the sparql query language could be supported
e ciently as well. This would involve support for filters; one way
to implement them is to o er ldfs with regular expression selectors.
Such features would then also be indicated by a server.</p>
      <p>In order to enable ldf querying in an uniform way, we should
look at standardizing basic ldfs and related technologies. A first
e ort is our website http:// linkeddatafragments.org/ , which o ers
documentation and example source code, as well as ldf sources.</p>
      <p>
        Finally, we are eager to explore links between ldfs and other
technologies and standards. In particular, we see an important role
for provenance [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] to explain how a client obtained an answer and
what data sources were used in the process.
      </p>
      <p>With Linked Data Fragments, we have introduced a novel way
to look at Linked Data querying. By adjusting the granularity of
information and equipping each fragment with metadata and the
controls needed to find others, clients become able to consume
Linked Data in more flexible ways. We believe the best way to
make intelligent clients happen is to stop creating intelligent servers.
The ultimate objective of Linked Data Fragments is therefore to
build servers that foster intelligent clients.</p>
    </sec>
    <sec id="sec-22">
      <title>ACKNOWLEDGMENTS</title>
      <p>Ruben wishes to thank Richard Cyganiak for insightful
discussions, Mario Arias for his help with hdt, and Johannes Lorey for
suggestions on related work.</p>
      <p>The described research activities were funded by Ghent
University, the Institute for the Promotion of Innovation by Science and
Technology in Flanders (iwt), the Fund for Scientific Research
Flanders (fwo Flanders), and the European Union.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Belshe</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peon</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Melnikov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <source>Hypertext Transfer Protocol version 2</source>
          .0.
          <string-name>
            <surname>Internet</surname>
            <given-names>draft</given-names>
          </string-name>
          ,
          <source>Internet Engineering Task Force (Dec</source>
          <year>2013</year>
          ), http:// tools.ietf.org/ html/ draft-ietf
          <source>-httpbis-http2-09</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Linked Data - the story so far</article-title>
          .
          <source>International Journal on Semantic Web and Information Systems</source>
          <volume>5</volume>
          (
          <issue>3</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          (
          <year>Mar 2009</year>
          ), http:// tomheath.com/ papers/ bizer
          <article-title>-heath-bernerslee-ijswis-linked-data</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>: dbpedia - a crystallization point for the Web of Data</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>7</volume>
          (
          <issue>3</issue>
          ),
          <fpage>154</fpage>
          -
          <lpage>165</lpage>
          (
          <year>2009</year>
          ), http:// www.websemanticsjournal.org/ index.php/ ps/ article/ view/ 164
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schultz</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The Berlin sparql benchmark</article-title>
          .
          <source>International Journal on Semantic Web and Information Systems</source>
          <volume>5</volume>
          (
          <issue>2</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          (
          <year>2009</year>
          ), http:// wifo5-
          <fpage>03</fpage>
          .informatik.uni-mannheim.de/ bizer/ pub/ Bizer-SchultzBerlin
          <string-name>
            <surname>-SPARQL-</surname>
          </string-name>
          Benchmark-IJSWIS.pdf
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Bottomley</surname>
            ,
            <given-names>J.E.J.:</given-names>
          </string-name>
          <article-title>Implementing clusters for high availability</article-title>
          .
          <source>In: Proceedings of the Annual Conference on usenix Annual Technical Conference. usenix Association</source>
          (
          <year>2004</year>
          ), http:// dl.acm.org/ citation.cfm?id=
          <volume>1247415</volume>
          .
          <fpage>1247459</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Buil-Aranda</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hogan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Umbrich</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vandenbussche</surname>
          </string-name>
          , P.Y.:
          <article-title>sparqlWeb-querying infrastructure: Ready for action?</article-title>
          <source>In: Proceedings of the 12th International Semantic Web Conference (Nov</source>
          <year>2013</year>
          ), http:// link.springer.com/ chapter/ 10.1007/ 978-3-
          <fpage>642</fpage>
          -41338-4_
          <fpage>18</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Pubby - a Linked Data frontend for sparql endpoints</article-title>
          , http:// wifo5-
          <fpage>03</fpage>
          .informatik.uni-mannheim.de/ pubby/
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alexander</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hausenblas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Vocabulary of Interlinked Datasets (void)</article-title>
          . Interest group note,
          <source>World Wide Web Consortium (Mar</source>
          <year>2011</year>
          ), http:// www.w3.org/ TR/ media-frags/
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Feigenbaum</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>G.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>K.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torres</surname>
          </string-name>
          , E.:
          <article-title>sparql 1.1 protocol</article-title>
          . Recommendation,
          <source>World Wide Web Consortium (Mar</source>
          <year>2013</year>
          ), http:// www.w3.org/ TR/ sparql11-protocol/
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Fernández</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martínez-Prieto</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gutiérrez</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polleres</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arias</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Binary rdf representation for publication and exchange (hdt)</article-title>
          .
          <source>Journal of Web Semantics</source>
          <volume>19</volume>
          ,
          <fpage>22</fpage>
          -
          <lpage>41</lpage>
          (
          <year>Mar 2013</year>
          ), http:// dx.doi.org/ 10.1016/ j.websem.
          <year>2013</year>
          .
          <volume>01</volume>
          .002
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Fielding</surname>
          </string-name>
          , R.T.:
          <article-title>Architectural Styles and the Design of Network-based Software Architectures</article-title>
          .
          <source>Ph.D. thesis</source>
          , University of California (
          <year>2000</year>
          ), http:// www.ics.uci.edu/ ~fielding/ pubs/ dissertation/ top.htm
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Fielding</surname>
          </string-name>
          , R.T.:
          <article-title>rest apis must be hypertext-driven</article-title>
          . Untangled - Musings of Roy T. Fielding (Oct
          <year>2008</year>
          ), http: // roy.gbiv.com/ untangled/ 2008/ rest
          <article-title>-apis-must-be-hypertext-driven</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Fielding</surname>
            ,
            <given-names>R.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gettys</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mogul</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frystyk</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Masinter</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leach</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Hypertext Transfer Protocol (http)</article-title>
          .
          <source>Request For Comments</source>
          <volume>2616</volume>
          ,
          <string-name>
            <surname>Internet</surname>
            <given-names>Engineering Task Force</given-names>
          </string-name>
          (
          <year>Jun 1999</year>
          ), http:// tools.ietf.org/ html/ rfc2616
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moreau</surname>
          </string-name>
          , L.:
          <article-title>prov overview</article-title>
          . Working group note,
          <source>World Wide Web Consortium (Apr</source>
          <year>2013</year>
          ), http:// www.w3.org/ TR/ prov-overview/
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seaborne</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>: sparql 1.1 query language</article-title>
          .
          <source>Recommendation, World Wide Web Consortium (Mar</source>
          <year>2013</year>
          ), http:// www.w3.org/ TR/ sparql11-query/
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Hartig</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>How caching improves e ciency and result completeness for querying Linked Data</article-title>
          .
          <source>In: Proceedings of the 4th Workshop on Linked Data on the Web (Mar</source>
          <year>2011</year>
          ), http:// ceur-ws.
          <source>org/</source>
          Vol-
          <volume>813</volume>
          /
          <fpage>ldow2011</fpage>
          -
          <lpage>paper05</lpage>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Hartig</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Zero-knowledge query planning for an iterator implementation of link traversal based query execution</article-title>
          .
          <source>In: Proceedings of the 8th Extended Semantic Web Conference on The Semantic Web</source>
          . pp.
          <fpage>154</fpage>
          -
          <lpage>169</lpage>
          . Springer (
          <year>2011</year>
          ), http:// dl.acm.org/ citation.cfm?id=
          <volume>2008892</volume>
          .
          <fpage>2008906</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Hartig</surname>
            ,
            <given-names>O.:</given-names>
          </string-name>
          <article-title>An overview on execution strategies for linked data queries</article-title>
          .
          <source>Datenbank-Spektrum</source>
          <volume>13</volume>
          (
          <issue>2</issue>
          ),
          <fpage>89</fpage>
          -
          <lpage>99</lpage>
          (
          <year>2013</year>
          ), http:// dx.doi.org/ 10.1007/ s13222-013-0122-1
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Hartig</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freytag</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          :
          <article-title>Executing sparql queries over the Web of Linked Data</article-title>
          .
          <source>In: Proceedings of the 8th International Semantic Web Conference</source>
          . pp.
          <fpage>293</fpage>
          -
          <lpage>309</lpage>
          . Springer (
          <year>2009</year>
          ), http:// www2.informatik.hu-berlin.de/ ~hartig/ files/ HartigEtAl_ QueryTheWeb_ISWC09_Preprint.pdf
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Kjernsmo</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>The necessity of hypermedia rdf and an approach to achieve it</article-title>
          .
          <source>In: Proceedings of the Workshop on Linked apis for the Semantic Web (May</source>
          <year>2012</year>
          ), http:// lapis2012.linkedservices.org/ papers/ 1.pdf
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Lanthaler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gütl</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Hydra: A vocabulary for hypermedia-driven Web apis</article-title>
          .
          <source>In: Proceedings of the 6th Workshop on Linked Data on the Web (May</source>
          <year>2013</year>
          ), http:// ceur-ws.
          <source>org/</source>
          Vol-
          <volume>996</volume>
          / papers/ ldow2013-paper-03.pdf
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Lorey</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naumann</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Caching and prefetching strategies for sparql queries</article-title>
          .
          <source>In: Proceedings of the 3rd International Workshop on Usage Analysis and the Web of Data (May</source>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Lorey</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naumann</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Detecting sparql query templates for data prefetching</article-title>
          .
          <source>In: Proceedings of the 10th Extended Semantic Web Conference (May</source>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Unbehauen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Improving the performance of Semantic Web applications with sparql query caching</article-title>
          .
          <source>In: The Semantic Web: Research and Applications, Lecture Notes in Computer Science</source>
          , vol.
          <volume>6089</volume>
          , pp.
          <fpage>304</fpage>
          -
          <lpage>318</lpage>
          . Springer (
          <year>2010</year>
          ), http:// dx.doi.org/ 10.1007/ 978-3-
          <fpage>642</fpage>
          -13489-0_
          <fpage>21</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Marwah</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maciel</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christian</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Almeida</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Araújo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Souza</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Callou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Galdino</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pires</surname>
          </string-name>
          , J.:
          <article-title>Quantifying the sustainability impact of data center availability</article-title>
          .
          <source>sigmetrics Performance Evaluation Review</source>
          <volume>37</volume>
          (
          <issue>4</issue>
          ),
          <fpage>64</fpage>
          -
          <lpage>68</lpage>
          (
          <year>Mar 2010</year>
          ), http:// doi.acm.
          <source>org/ 10</source>
          .1145/ 1773394.1773405
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Nottingham</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Feed paging and archiving</article-title>
          .
          <source>Request For Comments</source>
          <volume>5005</volume>
          ,
          <string-name>
            <surname>Internet</surname>
            <given-names>Engineering Task Force</given-names>
          </string-name>
          (
          <year>Sep 2007</year>
          ), http:// tools.ietf.org/ html/ rfc5005
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Pautasso</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilde</surname>
          </string-name>
          , E.:
          <article-title>Why is the Web loosely coupled? - A multi-faceted metric for service design</article-title>
          .
          <source>In: Proceedings of the 18th International Conference on World Wide Web</source>
          . pp.
          <fpage>911</fpage>
          -
          <lpage>920</lpage>
          . acm, New York (
          <year>2009</year>
          ), http:// www2009.eprints.org/ 92/ 1/ p911.pdf
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Prud'hommeaux</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Buil-Aranda</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>: sparql 1.1 federated query</article-title>
          .
          <source>Recommendation, World Wide Web Consortium (Mar</source>
          <year>2013</year>
          ), http:// www.w3.org/ TR/ sparql11-federated-query/
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Schmidt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hornung</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pinkel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lausen</surname>
          </string-name>
          , G.:
          <article-title>SP2Bench: A sparql performance benchmark</article-title>
          .
          <source>In: Semantic Web Information Management</source>
          , pp.
          <fpage>371</fpage>
          -
          <lpage>393</lpage>
          . Springer (
          <year>2010</year>
          ), http:// dx.doi.org/ 10.1007/ 978-3-
          <fpage>642</fpage>
          -04329-1_
          <fpage>16</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Shadbolt</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>The Semantic Web revisited</article-title>
          .
          <source>Intelligent Systems</source>
          <volume>21</volume>
          (
          <issue>3</issue>
          ),
          <fpage>96</fpage>
          -
          <lpage>101</lpage>
          (
          <year>Jul 2006</year>
          ), http:// eprints.soton.ac.uk/ 262614/
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Shu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Compton</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Müller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , K.:
          <article-title>Towards content-aware sparql query caching for Semantic Web applications</article-title>
          .
          <source>In: Web Information Systems Engineering, Lecture Notes in Computer Science</source>
          , vol.
          <volume>8180</volume>
          , pp.
          <fpage>320</fpage>
          -
          <lpage>329</lpage>
          . Springer (
          <year>2013</year>
          ), http:// dx.doi.org/ 10.1007/ 978-3-
          <fpage>642</fpage>
          -41230-1_
          <fpage>27</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Speicher</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arwe</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malhotra</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Linked Data Platform 1.0</article-title>
          . Working draft,
          <source>World Wide Web Consortium (Jul</source>
          <year>2013</year>
          ), http:// www.w3.org/ TR/ 2013/ WD-ldp-
          <volume>20130730</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Troncy</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mannens</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , Pfei er, S.,
          <string-name>
            <surname>Van Deursen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Media fragments uri 1.0 (basic</article-title>
          ).
          <source>Recommendation, World Wide Web Consortium (Sep</source>
          <year>2012</year>
          ), http:// www.w3.org/ TR/ media-frags/
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>Vinoski</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Serendipitous reuse</article-title>
          .
          <source>Internet Computing</source>
          <volume>12</volume>
          (
          <issue>1</issue>
          ),
          <fpage>84</fpage>
          -
          <lpage>87</lpage>
          (
          <year>Jan 2008</year>
          ), http:// steve.vinoski.net/ pdf/ IEEE-Serendipitous_Reuse.pdf
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>G.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weaver</surname>
          </string-name>
          , J.:
          <article-title>Enabling fine-grained http caching of sparql query results</article-title>
          .
          <source>In: Proceedings of the 10th International Conference on The Semantic Web</source>
          . pp.
          <fpage>762</fpage>
          -
          <lpage>777</lpage>
          . Springer (
          <year>2011</year>
          ), http:// www.cs.rpi.edu/ ~weavej3/ papers/ iswc2011.pdf
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>d</year>
          .:
          <article-title>Improving sparql query performance with algebraic expression tree based caching and entity caching</article-title>
          .
          <source>Journal of Zhejiang University science c 13(4)</source>
          ,
          <fpage>281</fpage>
          -
          <lpage>294</lpage>
          (
          <year>2012</year>
          ), http:// dx.doi.org/ 10.1631/ jzus.C1101009
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>