<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LDCache - a cache for Linked Data-driven Web applications</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hannes Ebner</string-name>
          <email>hannes@metasolutions.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthias Palmér</string-name>
          <email>matthias@metasolutions.se</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>KTH Royal Institute of Technology</institution>
          ,
          <country country="SE">Sweden</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>MetaSolutions AB</institution>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>67</fpage>
      <lpage>73</lpage>
      <abstract>
        <p>Modern Web applications that make use of external data sources, be it Linked Data (LD) or not, usually run into the same problems; if a data source performs badly or is offline, the user experience is affected negatively. In some cases this may result in long response times, whereas in more extreme cases the application becomes unusable. This paper presents LDCache [4], a caching service that ensures that Linked Data-driven Web applications remain functional with good user experience regardless of the status of the external data sources they eventually integrate. First, some requirements are defined that typically occur when data from disparate sources are integrated. Then the features of the implemented solution are summarized together with documentation on how the service can be taken advantage of. This is followed by a description of the architecture. The authors conclude with a road map for future development and a short summary of the work carried out.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Background</title>
      <p>When implementing Web applications that take advantage of multiple distributed
data sources a range of problems have to be solved to maintain an acceptable
user experience. Performance, scalability and reliability are important aspects
that must not be ignored during the design and development process.</p>
      <p>
        Because of the use of HTTP it is possible to cover a majority of the
requirements (see following section for details) by taking full advantage of the
mature HTTP tooling ecosystem; reverse proxies, HTTP caches, Content
Delivery Networks (CDN), mirroring of datasets by importing RDF dumps into local
instances, etc., are ways to gain control and increase reliability and performance.
There has also been some work on solving slightly related client-side problems
such as the Semantic Web Client Library [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The Linked Data Fragments
initiative [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is based on subsets of Linked Data collections which are queried on the
client side. However, nothing really allows for the dynamic optimization of the
usage of a diverse range of only minor parts of possibly large datasets. Datasets
such as DBpedia [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] bring a large overhead at a high cost if they are to be
self-hosted while most probably only using a subset of the available data.
      </p>
      <p>LDCache aims to be a light-weight solution that implements a set of features
that may be used by typical LD-driven applications. It assists in traversing
(parts of) the Linked Data graph by following links and provides a pre-caching
and filtering mechanism to manage data that is relevant for particular client-side
applications. Such an approach encourages to couple an LD-driven application
with an LDCache configuration that exactly matches the application’s
requirements. The technical requirements of the necessary infrastructure can be kept
at a minimum. To clarify LDCache’s behavior in comparison with e.g. caching
HTTP proxies it has to be emphasized that no requests to the source servers
are performed during normal cache operation. Due to the currently applied
precaching mechanism LDCache only returns data that has been cached in advance,
which is something to take into consideration if frequently updated data is to be
cached.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Requirements of typical LD-driven applications</title>
      <p>The following requirements have been formalized during the development of
several interactive Web applications that take advantage of Linked Data from
third parties. The requirements are listed and described individually, but typical
scenarios require combinations of one or more of them.
2.1</p>
      <sec id="sec-2-1">
        <title>Reliability issues of third party data sources</title>
        <p>It is common that third party data sources have regular down times or
inconsistent performance. Especially big data hubs with a high load such as DBpedia
are prone to being unreliable. Applications in which information from different
sources is mashed up can be prepared in various ways for the unavailability of
one or more data sources: (1) the application may degrade gracefully and still
work, showing less information, (2) the application breaks and stops working,
or (3) the application uses a local copy of (parts of) the data source and is not
affected by any reliability issues. LDCache provides an easy to use solution for
(3).
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Streamlining multiple dereferenciation steps</title>
        <p>Dereferenciation introduces complexity; redirects have to be followed which causes
multiple requests that most probably cause delays that affect interactive
applications negatively. With LDCache only a single request is necessary as the
dereferenciation steps are streamlined on the server-side into one response.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>RDF format alignment</title>
        <p>If an application integrates different data sources then the support of various
RDF formats may become an issue as more complex content negotiation
scenarios apply and JavaScript-libraries have to be added to support additional
formats. With LDCache as an intermediate layer a Web application can use the
same content type for all RDF resources, even if the desired format is not
directly supported by data source. E.g., an application can rely on JSON-LD, even
if the data source only supports RDF/XML.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Offloading of Web resources to a CDN</title>
        <p>It is common to make use of a CDN for the most requested content in high
load scenarios. This requires that the DNS of the server can be controlled, which
usually is not the case for third party data sources. LDCache makes it possible
to provide URLs for cached information and allows for using a CDN where the
URLs and their URL-parameters usually are changed into canonical
representations.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Features and usage</title>
      <sec id="sec-3-1">
        <title>Graph traversal</title>
        <p>The configuration allows to define databundles consisting of lists of resources that
are to be used as starting point for the caching process which is initiated during
the startup of the service, i.e., the cache is pre-populated. Initially all resources
are fetched from the provided URIs; if the graph traversal’s depth parameter is
greater than zero, a depth-first search (DFS) is performed (loops are detected).
If follow and/or followTuples parameters are provided then resources in object
or subject positions, respectively, are followed. This process has implications for
the definition of an LDCache databundle; because links to other sources are
followed, an LDCache databundle is not limited to a single data source and may
contain resources from multiple LD datasets.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Cached resources</title>
        <p>
          The REST resource that accesses the cache requires that the requested RDF
resources (and subsequent resources if links are to be followed) are cached in
the local LDCache repository. See the configuration section of the LDCache
documentation [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] for instructions on how to configure LDCache to prefetch
RDF resources. The request URL for cached resources is at the root of the
LDCache: http://{ldc-base}/.
        </p>
        <p>The following parameters are available when accessing cached resources:
– url: The RDF resource to be fetched.
– follow: A comma-separated list of predicates. The objects of matching triples
are followed during graph traversal and if these objects correspond to cached
resources their RDF graphs will be merged into the response graph.
– followTuples: A comma-separated list of predicate-object tuples in the format
predicate|object, i.e., the predicate is separated from the object by the
pipe symbol. The subjects of matching triples are followed.
– followDepth: The maximum distance from the root-resource that should be
followed. Default is 0, i.e., only the resource identified by the url parameter
will be fetched and no links are followed.
– includeDestinations: A comma-separated whitelist with prefixes of
destinations to include when traversing the graph.
– includeLiteralLanguages: A comma-separated whitelist with language tags
to specify which literals to include.
– callback: Name of the callback method to be used for a JSONP response.
– format: A valid RDF MIME type. If an Accept header is supplied the format
parameter takes precedence.</p>
        <p>The parameters follow and followTuples may be used in combination.
Namespace expansion is applied to all parameters, refer to the LDCache documentation
which contains a list of supported prefixes.</p>
        <p>It has to be emphasized that the configuration parameters match the
available request parameters of the API. In consequence the cache does not return
satisfying results for requests that are out of scope of the cached data. However,
the proxy resource may be used for such out-of-scope requests.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Uncached (proxied) resources</title>
        <p>The proxy service bypasses the cache and allows arbitrary RDF resources to be
fetched. Format conversions are handled transparently through normal content
negotiation between proxy and data source and proxy and client. The proxy
resource is quite simple as it only proxies one resource without the smartness that
the cache provides. The proxy has to be explicitly activated in the configuration
and is, like the cache, restricted to RDF resources.</p>
        <p>The request URL of the proxy is: http://{ldc-base}/proxy.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4 Implicit dereferenciation</title>
        <p>A common issue in LD-based applications are dereferenciations. In
combination with following links to other resources and consecutive redirects this can
cause unacceptable delays in user interfaces. LDCache flattens redirects
automatically, e.g., when requesting http://dbpedia.org/resource/Marie_Curie
in format N3 the response from LDCache corresponds to http://dbpedia.org/
data/Marie_Curie.n3. No redirects are necessary and the client receives data
that can be used immediately. This works for the whole traversal path in case
any links to other resources are to be followed in the context of the request.
3.5</p>
      </sec>
      <sec id="sec-3-5">
        <title>Rate limitation of HTTP requests</title>
        <p>LDCache supports the limitation of requests per host to avoid overloading servers
which may lead to being blacklisted. This feature is configurable.
3.6</p>
      </sec>
      <sec id="sec-3-6">
        <title>Transparent handling of RDF formats</title>
        <p>
          All content negotiation is handled transparently. During the caching process
LDCache negotiates with the RDF source; when a client requests a cached resource
from LDCache any RDF format supported by Sesame Rio [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] plus JSON-LD
[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] is supported. This also means that LDCache may be used as transparent
RDF converter in case an RDF source does not support the desired RDF format
directly.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Example requests</title>
      <p>The requests below illustrate how resources may be fetched, including related
resources that are pointed out by outgoing links. The requests require that all
data are prefetched and cached as there will be no on-demand requests to any
data source. All matching resources are merged into a single RDF graph which
then is returned to the requesting client.
4.1</p>
      <sec id="sec-4-1">
        <title>Marie Curie, her doctoral students and their doctoral students</title>
        <p>Outgoing links to resources in object position are followed, but restricted to
doctoral students that have are identified by a URI that starts with dbpedia.org.
The graph is traversed two levels down, the resource graphs are merged and the
result is returned in JSON-LD.</p>
        <p>GET http://{ldc-base}/?
url=http://dbpedia.org/resource/Marie_Curie&amp;
follow=http://dbpedia.org/ontology/doctoralStudent&amp;
followDepth=2&amp;
includeDestinations=http://dbpedia.org/&amp;
format=application/ld+json
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>All Nobel laureates by following matching resources in subject position</title>
        <p>All outgoing links are followed because the follow parameter (acting as whitelist) is
omitted. In addition to following all resources in object position, the information from
the followTuples parameter (which is a tuple consisting of predicate and object) is used
to follow matching resources in subject position. The graph is traversed one level down,
the graphs are merged and the result is returned in Turtle format.</p>
        <p>GET http://{ldc-base}/?
url=http://data.nobelprize.org/all/laureate&amp;
followTuples=rdf:type|nobel:Laureate&amp;
followDepth=1&amp;
format=text/turtle
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Returning only partial data</title>
        <p>In the next milestone support for a filter parameter will be added. This allows to reduce
the size of responses by only returning a set of triples as defined by the filter parameter.
This is useful for cases where e.g. only labels are relevant for the requesting
application. In such situations it will be possible to only request e.g. rdfs:label, skos:prefLabel,
dc:title, and dcterms:title; all other triples will be omitted.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Architecture</title>
      <p>
        The architecture is kept simple and consists of few components:
– A Restlet-supported API [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] is wrapped around a Sesame SAIL [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
– The application can be started either stand-alone (supported by the Simple
framework [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]) or as web app in a container using Tomcat, Jetty, etc.
– Every cached RDF resource is stored as one named graph [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
– Administrative information about datasets is stored in separate named graphs.
– The API is intentionally kept simple and currently contains only two REST
resources: one for the cache and one for the proxy.
      </p>
      <p>– The configuration format is JSON.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Road map for future development</title>
      <p>
        The development road map contains some major features and improvements:
– Continuous background updates of databundles and their cached resources.
– Enhanced property filtering to be smart about storing “interesting” values
– On-the-fly caching of proxied resources; proxied RDF resources are currently
discarded after a proxy request.
– Support for SPARQL queries to find a list of resources that should be cached.
– Explore whether support for Linked Data Fragments [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] would make sense.
– Investigate how SPARQL requests could be cached.
7
      </p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>
        LDCache has been designed as a light-weight and easy to use Linked Data cache that
supports the development of Linked Data-driven Web applications. The current
version is still in its early stages with room for improvements, however, it is considered
stable and safe to use in production environments. A showcase where LDCache is used
in production are Nobel Media’s [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] profile pages of Nobel laureates that consist of
information from various data sources, e.g., Nobel Prize Linked Data [
        <xref ref-type="bibr" rid="ref6 ref7">6,7</xref>
        ] and DBpedia
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The goal of fulfilling the stated requirements has been met and the authors feel
that LDCache improves the lead time for developing Linked Data-driven applications.
Also, it improves the performance, scalability, and reliability of applications that rely
on third party data sources and brings back control of the data to the application
developer.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>
        The work presented in this paper has been partially carried out with financial support
from the VINNOVA-funded project “Vidareutveckling av länkade öppna data för
Nobelpris” [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and the EC-funded project “Open Discovery Space” [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] which the authors
gratefully acknowledge.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Carroll</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stickler</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Named graphs</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>3</volume>
          (
          <issue>4</issue>
          ),
          <fpage>247</fpage>
          -
          <lpage>267</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>DBpedia</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://dbpedia.org
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>JSON-LD - JSON for Linking Data</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://json-ld.org, accessed
          <year>2014</year>
          -
          <volume>07</volume>
          -25
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. LDCache documentation (
          <year>2014</year>
          ), http://entrystore.org/ldcache/
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Linked</given-names>
            <surname>Data Fragments</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://linkeddatafragments.org/, accessed 2014-
          <volume>07</volume>
          -24
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Nobel</given-names>
            <surname>Prize Linked Open Data</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://data.nobelprize.org, accessed
          <year>2014</year>
          -
          <volume>07</volume>
          -24
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <article-title>Specification of Nobel Prize Linked Data (</article-title>
          <year>2013</year>
          ), http://www.nobelprize. org/nobel_organizations/nobelmedia/nobelprize_org/developer/ manual-linkeddata/terms.html, accessed
          <year>2014</year>
          -
          <volume>07</volume>
          -24
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Nobel</given-names>
            <surname>Media</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://www.nobelprize.org/nobel_organizations/ nobelmedia/, accessed 2014-
          <volume>07</volume>
          -24
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <article-title>Vidareutveckling av länkade öppna data för Nobelpris (</article-title>
          <year>2014</year>
          ), http://www.vinnova.se/sv/Resultat/Projekt/Effekta/2012-00741/
          <article-title>Vidareutveckling-av-lankade-oppna-data-for-</article-title>
          <string-name>
            <surname>Nobelpris</surname>
          </string-name>
          /, accessed 2014-
          <volume>07</volume>
          - 24
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. Open Discovery Space (
          <year>2014</year>
          ), http://opendiscoveryspace.eu
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>OpenRDF Sesame</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://www.openrdf.org/, accessed 2014-
          <volume>07</volume>
          -24
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <article-title>Restlet - REST framework for Java (</article-title>
          <year>2014</year>
          ), http://restlet.com/, accessed 2014-
          <volume>07</volume>
          -24
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Simple Framework</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://www.simpleframework.org/, accessed 2014-
          <volume>07</volume>
          - 24
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <article-title>Semantic Web Client Library - Querying the complete Semantic Web with SPARQL (</article-title>
          <year>2009</year>
          ), http://wifo5-
          <fpage>03</fpage>
          .informatik.uni-mannheim.de/bizer/ng4j/ semwebclient/, accessed 2014-
          <volume>07</volume>
          -24
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>