<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Live SPARQL Auto-Completion</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Insight Centre for Data Analytics, National University of Ireland</institution>
          ,
          <addr-line>Galway</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The amount of Linked Data has been growing increasingly. However, the e cient use of that knowledge is hindered by the lack of information about the data structure. This is re ected by the di culty of writing SPARQL queries. In order to improve the user experience, we propose an auto-completion library1 for SPARQL that suggests possible RDF terms. In this work, we investigate the feasibility of providing recommendations by only querying the SPARQL endpoint directly.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Over the years, many contributions have been done towards facilitating the use of
SPARQL, either visually [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], or by completely hiding SPARQL from the user [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In
this work, we aim to help users with a knowledge of SPARQL by providing an
autocompletion feature. Several systems have been proposed in this direction. Although
1 Gosparqled: https://github.com/scampi/gosparqled
the focus in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is the visual interface, it can provide recommendations of terms such
as predicates and classes. In [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] possible recommendations are taken from query
logs. The system proposed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] provides recommendations based on the data itself,
with a focus on SPARQL federation. Instead, we aim to make available an
easy-touse library which core feature is to provide data-based recommendations. In [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] an
editor with auto-completion was developed that leverage a data-generated schema
(i.e., a graph summary ). We investigate in this work the practicability of bypassing
the graph summary by relying only on the data.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Live Auto-Completion</title>
      <p>
        We propose a data-based auto-completion which retrieves possible items with
regards to the current state of the query. Recommended items can be predicates,
classes, or even named graphs. Firstly, we indicate the position in the SPARQL
query that is to be auto-completed, i.e., the Point Of Focus (POF), by inserting the
character `&lt;'. Secondly, we reduce the query down to its recommendation scope [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Finally, we transform the POF into the SPARQL variable \?POF" which is used
for retrieving recommendations. The retrieved recommendations are then ranked,
e.g., by the number of occurrences of an item.
      </p>
      <p>
        Recommendation Scope. While building a SPARQL query, not all triple patterns are
relevant for the recommendation. Therefore, we de ne the scope as the connected
component that contains the POF. Figure 1a depicts a SPARQL query where the
POF is associated with the variable \?s": it seeks possible predicates that occur with
a \:Person" having the predicate \:name". Figure 1b depicts the previous SPARQL
query reduced to its recommendation scope. Indeed, the pattern on line 4 is removed
since it is not part of the connected component containing the POF.
Recommendation Capabilities. The scope may include content-speci c terms, e.g,
resources and lters, unlike to [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] since the graph summary is an abstraction that
captures only the structure of the data. Recommendations about predicates, classes
and named graphs are possible as in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In addition, the use of the data directly
allows to provide recommendations for speci c resources.
4
      </p>
    </sec>
    <sec id="sec-3">
      <title>Evaluation</title>
      <p>
        Systems. In this section, we evaluate the recommendations returned by the proposed
system, that we refer to as \S1", against the ones provided by the approach in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
which we refer to as \S2".
      </p>
      <p>
        Settings. We compare the recommendations with regards to (1) the response-time,
i.e., the time spent on retrieving the recommendations via a SPARQL query; and
(2) the quality of the recommendations. A run of the evaluation consists of the
following steps. First, we vary the amount of information retrieved via the \LIMIT"
clause. Then, we compare the ranked TOP-10 recommendations against a gold
standard. The ranking is based on the number of occurrences of a recommendation. The
gold standard consists in retrieving recommendations directly from the data without
the LIMIT clause, and retaining only the 10 most occurring terms. The TOP-10 of
the gold standard and the system are compared using the Jaccard similarity. We
consider that the higher the similarity, the higher the quality of recommendations.
Queries. We used the query logs of the DBpedia endpoint version 3.3 available
from the USEWOD20132 dataset. The queries3 were stripped of any pattern about
speci c resources, in order to keep only the structure of the query. In addition,
we removed queries that contain more than one connected component. Queries are
grouped according to their complexity, which depends on the number of triple
patterns and on the number of star graphs. A group is identi ed by a string that has
as many numbers as there are stars, with numbers separated by a dash '-' and
representing the number of triple patterns in a star. For example, a query with two
stars and one triple pattern each is then identi ed with 1-1. This de nition of query
complexity exhibits the potential errors, i.e., a recommendation having zero-result,
that a graph summary can have, as described in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>Graphs. We loaded into an endpoint the English part of the Dbpedia3.34 dataset,
which consists of 167 199 852 triples. The graph summary consists of 29 706 051
triples, generated by grouping resources sharing the same set of classes.
Endpoint. We used a Virtuoso5 SPARQL endpoint. The endpoint is deployed on a
server with 32GB of RAM and with SSD drives.</p>
      <p>
        Comparison. For each group of query complexity QC, we report in Table 1 the
results of the evaluation, with J 1 (resp., J 2) the average Jaccard similarity for the
system S1 (resp., S2); and T 1 (resp., T 2) the average response-time in ms for the
system S1 (resp., S2). The reported values are the averages over 5 runs. We can see
that as the LIMIT gets larger, the higher the Jaccard similarity becomes.Since the
graph summary used in S2 is a concise representation of the graph structure, the
data sample at a certain LIMIT value contains more terms than in S1. However,
this impacts negatively on the quality of S2 as re ected by the values of J2. This
shows the graph summary is subject to errors [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], i.e., zero-result recommendations.
Nonetheless, it is interesting to remark that in S1 the recommendations can lead
the query to an \isolated" part of the graph, from which the way out is through
the use of \OPTIONAL" clauses. In S2, the graph summary allows to reduce this
e ect. The response-times for either system is similar, with S2 being slightly faster
than S1. This indicates that directly querying the endpoint for recommendations
is feasible. However, the signi cant di erence in sizes between the graph summary
and the original graph would become increasingly pre-dominant as the data grows.
2 http://usewod.org/
3 https://github.com/scampi/gosparqled/tree/master/eval/data
4 http://wiki.dbpedia.org/Downloads33
5 Virtuoso v7.1.0 at https://github.com/openlink/virtuoso-opensource
      </p>
      <p>T 1 T 2</p>
      <p>2
107 81
108 81
141 91</p>
      <p>1-3
101 108
103 105
126 105</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgement</title>
      <p>This material is based upon works supported by the European FP7 projects LOD2
(257943).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Ambrus</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mller</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Handschuh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Konduit vqb: a visual query builder for sparql on the social semantic desktop</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Campinas</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delbru</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tummarello</surname>
          </string-name>
          , G.:
          <article-title>E ciency and precision trade-o s in graph summary algorithms</article-title>
          .
          <source>In: Proceedings of the 17th International Database Engineering &amp; Applications Symposium</source>
          . pp.
          <volume>38</volume>
          {
          <fpage>47</fpage>
          . IDEAS '13,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Campinas</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perry</surname>
            ,
            <given-names>T.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceccarelli</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delbru</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tummarello</surname>
          </string-name>
          , G.:
          <article-title>Introducing rdf graph summary with application to assisted sparql formulation</article-title>
          .
          <source>In: Proceedings of the 2012 23rd International Workshop on Database and Expert Systems Applications</source>
          . pp.
          <volume>261</volume>
          {
          <fpage>266</fpage>
          . DEXA '12, IEEE Computer Society, Washington, DC, USA (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Sparql views: A visual sparql query builder for drupal</article-title>
          .
          <source>In: International Semantic Web Conference</source>
          . pp. {
          <volume>1</volume>
          {
          <issue>1</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Gombos</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiss</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Sparql query writing with recommendations based on datasets</article-title>
          . In: Yamamoto,
          <string-name>
            <surname>S</surname>
          </string-name>
          . (ed.)
          <source>Human Interface and the Management of Information. Information and Knowledge Design and Evaluation, Lecture Notes in Computer Science</source>
          , vol.
          <volume>8521</volume>
          , pp.
          <volume>310</volume>
          {
          <fpage>319</fpage>
          . Springer International Publishing (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Kramer</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dividino</surname>
            ,
            <given-names>R.Q.</given-names>
          </string-name>
          , Groner, G.:
          <article-title>Space: Sparql index for e cient autocompletion</article-title>
          . In: Blomqvist,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Groza</surname>
          </string-name>
          , T. (eds.)
          <source>International Semantic Web Conference. CEUR Workshop Proceedings</source>
          , vol.
          <volume>1035</volume>
          , pp.
          <volume>157</volume>
          {
          <fpage>160</fpage>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Buhmann, L.:
          <article-title>Autosparql: Let users query your knowledge base</article-title>
          .
          <source>In: Proceedings of the 8th Extended Semantic Web Conference on The Semantic Web: Research and Applications - Volume Part I</source>
          . pp.
          <volume>63</volume>
          {
          <fpage>79</fpage>
          . ESWC'
          <volume>11</volume>
          , Springer-Verlag, Berlin, Heidelberg (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Rietveld</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoekstra</surname>
          </string-name>
          , R.: Yasgui:
          <article-title>Not just another sparql client</article-title>
          .
          <source>In: SALAD@ESWC</source>
          . pp.
          <volume>1</volume>
          {
          <issue>9</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>