<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LDSpider: An open-source crawling framework for the Web of Linked Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Robert Isele</string-name>
          <email>robertisele@googlemail.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jurgen Umbrich</string-name>
          <email>juergen.umbrich@deri.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Bizer</string-name>
          <email>chris@bizer.de</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Harth</string-name>
          <email>harth@kit.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AIFB, Karlsruhe Institute of Technology</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Digital Enterprise Research Institute, National University of Ireland</institution>
          ,
          <addr-line>Galway</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Web of Linked Data is growing and currently consists of several hundred interconnected data sources altogether serving over 25 billion RDF triples to the Web. What has hampered the exploitation of this global dataspace up till now is the lack of an open-source Linked Data crawler which can be employed by Linked Data applications to localize (parts of) the dataspace for further processing. With LDSpider, we are closing this gap in the landscape of publicly available Linked Data tools. LDSpider traverses the Web of Linked Data by following RDF links between data items, it supports di erent crawling strategies and allows crawled data to be stored either in les or in an RDF store.</p>
      </abstract>
      <kwd-group>
        <kwd>Linked Data</kwd>
        <kwd>Crawler</kwd>
        <kwd>Spider</kwd>
        <kwd>Linked Data tools</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        As of September 2010, the Web of Linked Data contains more than 200
interconnected data sources totaling in over 25 billion RDF triples4. Applications
that need to localize data from the Web of Linked Data [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for further
processing currently either need to implement their own crawling code or rely on
pre-crawled data provided by Linked Data search engines or in the form of data
dumps, for example the Billion Triples Challenge dataset5. With LDSpider, we
are closing this gap in the Linked Data tool landscape. LDSpider is an extensible
Linked Data crawling framework, enabling client applications to traverse and to
consume the Web of Linked Data.
      </p>
      <p>The main features of LDSpider are:
LDSpider can process a variety of Web data formats including RDF/XML,
Turtle, Notation 3, RDFa and many microformats by providing a plugin
architecture to support Any236.
4 http://lod-cloud.net
5 http://challenge.semanticweb.org/
6 http://any23.org/
Crawled data can be stored together with provenance meta-information
either in a le or via SPARQL/Update in an RDF store.</p>
      <p>LDSpider o ers di erent crawling strategies, such as breadth- rst traversal
and load-balancing, for following RDF links between data items.
Besides of being usable as a command line application, LDSpider also o ers
a simple API which allows applications to con gure and control the details
of the crawling process.</p>
      <p>The framework is delivered as a small and compact jar with a minimum of
external dependencies.</p>
      <p>The crawler is high-performing by employing a multi-threaded architecture.</p>
      <p>LDSpider can be downloaded from Google Code7 under the terms of the
GNU General Public License v3. In the following, we will give an overview of
the LDSpider crawling framework and report about several use cases in which
we employed the framework.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Using LDSpider</title>
      <p>LDSpider has been developed to provide a exible Linked Data crawling
framework, which can be customized and extended by client applications. The
framework is implemented in Java and can be used through a command line application
as well as a exible API.
2.1</p>
      <sec id="sec-2-1">
        <title>Using the command line application</title>
        <p>The crawling process starts with a set of seed URIs. The order how LDSpider
traverses the graph starting from these seed URIs is speci ed by the crawling
strategy. LDSpider provides two di erent round-based crawling strategies:</p>
      </sec>
      <sec id="sec-2-2">
        <title>The breadth- rst strategy takes three parameters: &lt;depth&gt; &lt;uri-limit&gt;</title>
        <p>&lt;pld-limit&gt;. In each round, LDSpider fetches all URIs extracted from the
content of the URIs of the previous round, before advancing to the next
round. The depth of the breadth- rst traversal, the maximum number of
URIs crawled per round and per pay-level8 domain as well as the maximum
number of crawled pay-level domains can be speci ed. This strategy can be
used in situations where only a limited graph around the seed URIs should
be retrieved.</p>
        <p>The load-balancing strategy takes a single parameter: &lt;max-uris&gt;. This
strategy tries to fetch the speci ed number of URIs as quickly as possible while
adhering to a minimum and maximum delay between two successive requests
to the same pay-level domain. The load-balancing strategy is useful in
situations where the fetched documents should be distributed between domains
without overloading a speci c server.</p>
        <sec id="sec-2-2-1">
          <title>7 http://code.google.com/p/ldspider/</title>
          <p>
            8 \A pay-level domain (PLD) is any domain that requires payment at a TLD or
ccTLD registrar."[
            <xref ref-type="bibr" rid="ref3">3</xref>
            ]
          </p>
          <p>LDSpider: A crawling framework for the Web of Linked Data</p>
          <p>LDSpider will fetch URIs in parallel employing multiple threads. The strategy
can be requested to stay on the domains of the seed URIs.</p>
          <p>Crawled data can be written to di erent sinks: File output writes the crawled
statements to les using the N-Quads format. Triple store output writes the
crawled statements to endpoints that support SPARQL/Update.
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Using the API</title>
        <p>LDSpider o ers a exible API to be used in client applications. Each component
in the fetching pipeline can be con gured by either using one of the
implementations already included in LDSpider or by providing a custom implementation.
The fetching pipeline consists of the following components:
The Fetch Filter determines whether a particular page should be fetched by
the crawler. Typically, this is used to restrict the MIME types of the pages
which are crawled (e.g. to RDF/XML).</p>
        <p>The Content Handler receives the document and tries to extract RDF data
from it. LDSpider includes a content handler for documents formatted in
RDF/XML and a general content handler, which forwards the documents
to an Any23 server to handle other types of documents including Turtle,
Notation 3, RDFa and many microformats.</p>
        <p>The Sink receives the extracted statements from the content handler and
processes them usually by writing them to some output. LDSpider includes
sinks for writing various formats including N-Quads and RDF/XML as well
as to write directly to a triple store using SPARQL/Update. Both sinks can
be con gured to write metadata containing the provenance of the extracted
statements. When writing to a triple store, the sink can be con gured to
include the provenance using a Named Graph layout.</p>
        <p>The Link Filter receives the parsed statements from the content handler and
extracts all links which should be fetched in the next round. A common use
of a link lter is to restrict crawling to a speci c domain. Each Link Filter
can be con gured to follow only ABox and/or TBox links. This can be used
for example to con gure the crawler to get the schema together with the
primary data.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.3 Implementation</title>
        <p>LDSpider is implemented in Java and uses 3 external libraries: The parsing of
RDF/XML, N-Triples and N-Quads is provided by the NxParser library9. The
HTTP functionality is provided by the Apache HttpClient Library10, while the
Robot Exclusion Standard is repected through the use of the Norbert11 library.</p>
        <sec id="sec-2-4-1">
          <title>9 http://sw.deri.org/2006/08/nxparser/ 10 http://hc.apache.org/ 11 http://www.osjava.org/norbert/</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Usage examples</title>
      <p>
        We have employed LDSpider for the following crawling tasks:
{ We employed LDSpider to crawl interlinked FOAF pro les and write them
to a triple store. For that purpose, we crawled the graph around a single seed
pro le (http://www.wiwiss.fu-berlin.de/suhl/bizer/foaf.rdf) and
compared the number of traversed FOAF pro les for di erent number of rounds:
rounds
pro les
As the number of pro les grows faster than in the previsous use case, we
can conclude that the interlinked Twitter pro les build a much denser graph
than the FOAF web.
{ LDSpider is used in an online service which executes live SPARQL queries
over the LOD Web12
{ We used LDSpider to gather datasets for various research projects; e.g. the
study of link dynamics [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] or the evaluation of SPARQL queries with data
summaries over Web data [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
      </p>
      <p>In summary, LDSpider can be used to collect small to medium-sized Linked
Data corpora up to hundreds of millions of triples.
12 http://swse.deri.org/lodq</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bizer</surname>
          </string-name>
          , Tom Heath, and
          <string-name>
            <surname>Tim</surname>
          </string-name>
          Berners-Lee.
          <article-title>Linked data - the story so far</article-title>
          .
          <source>Int. J. Semantic Web Inf. Syst.</source>
          ,
          <volume>5</volume>
          (
          <issue>3</issue>
          ):1{
          <fpage>22</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Harth</surname>
          </string-name>
          , Katja Hose, Marcel Karnstedt, Axel Polleres,
          <string-name>
            <surname>Kai-Uwe Sattler</surname>
          </string-name>
          , and
          <article-title>Jurgen Umbrich. Data summaries for on-demand queries over linked data</article-title>
          .
          <source>In WWW '10: Proceedings of the 19th international conference on World wide web</source>
          , pages
          <volume>411</volume>
          {
          <fpage>420</fpage>
          , New York, NY, USA,
          <year>2010</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hsin-Tsang</surname>
            <given-names>Lee</given-names>
          </string-name>
          , Derek Leonard,
          <string-name>
            <given-names>Xiaoming</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Dmitri</given-names>
            <surname>Loguinov</surname>
          </string-name>
          .
          <article-title>Irlbot: scaling to 6 billion pages and beyond</article-title>
          .
          <source>In WWW '08: Proceeding of the 17th international conference on World Wide Web</source>
          , pages
          <volume>427</volume>
          {
          <fpage>436</fpage>
          , New York, NY, USA,
          <year>2008</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Michael; Hogan Aidan; Polleres Axel; Decker Stefan Umbrich, Jurgen; Hausenblas.
          <article-title>Towards dataset dynamics: Change frequency of linked open data sources</article-title>
          .
          <source>3rd International Workshop on Linked Data on the Web (LDOW2010)</source>
          ,
          <source>in conjunction with 19th International World Wide Web Conference</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>