<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ephedra: SPARQL federation over RDF data and services</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andriy Nikolov</string-name>
          <email>an@metaphacts.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Haase</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Johannes Trame</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Artem Kozlov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>metaphacts GmbH</institution>
          ,
          <addr-line>Walldorf</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Knowledge graph management use cases often require addressing hybrid information needs that involve a multitude of data sources, a multitude of data modalities (e.g., structured, keyword, geospatial search), and availability of computation services (e.g., machine learning and graph analytics algorithms). Although SPARQL queries provide a convenient way of expressing data requests over RDF knowledge graphs, the level of support for hybrid information needs is limited: existing query engines usually focus on retrieving RDF data and only support a set of hard-coded built-in services accessible via SPARQL 1.1 queries. To deal with this problem, we present Ephedra: a SPARQL federation engine aimed at processing hybrid queries, which provides a flexible declarative mechanism for including hybrid services into a SPARQL federation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>In many practical knowledge graph management use cases there is a need to address
hybrid information needs. Such needs can be characterized by the following dimensions:
queries. SPARQL 1.1 with its SERVICE clauses provides a convenient data retrieval
formalism: a complex information request over several data sources can be expressed
using a single query. With Ephedra we adopt the SPARQL 1.1 federation mechanism,
but we broaden its usage to include custom services as data sources and optimize such
hybrid queries to be executed efficiently.</p>
      <p>Expressing a complex hybrid information request using a SPARQL query can be
non-trivial due to the variety of potential types of services and the limitations of the
SPARQL syntax: e.g., a service can take as input a single set of parameters or a list of
arbitrary length; it can return as output one value, several values or a table of multiple
records, etc. With Ephedra we take these factors into account to support practical hybrid
federation use cases of the metaphactory platform and address hybrid information needs
in an efficient way. In this paper, we propose a reusable architecture in which hybrid
services can be easily plugged in, described in a declarative way, and invoked using
federated SPARQL queries.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Hybrid SPARQL federation</title>
      <p>To support the hybrid querying functionality, we developed the Ephedra query
processing engine as a part of the metaphactory platform. A crucial requirement is the ability
to plug in additional services with minimal effort and reference them from SPARQL.</p>
      <p>Figure 1 shows the generic architecture of the metaphactory platform. Ephedra is
used as a hybrid query federation layer to access the data repositories and services. In
our use case scenarios, the user’s information need is captured interactively: the user can
define search clauses, explore partial results, incrementally add new clauses, while the
system provides relevant suggestions. These interactions generate information requests
that are expressed as SPARQL 1.1 queries and given to Ephedra to process them.</p>
      <p>As the basis for Ephedra implementation, we used the RDF4J Federation SAIL API2
reusing the common functions such as query parsing and accessing remote SPARQL
endpoints. However, Ephedra extends the RDF4J object model and overrides the static
optimization and query execution strategies to deal with hybrid queries. The Ephedra
query evaluation strategy sends the sub-clauses of the query to the corresponding data
sources and invokes the relevant processing services, then gathers the partial results,
combines them using the union and join operations, and produces the final result set.
In this way, processing becomes transparent: hybrid information needs are processed in
the same way as ordinary SPARQL queries to an RDF triple store without the need to
integrate related processing services at the UI level.</p>
      <p>In order to configure the services as federation members, the system requires
relevant information about the service type as well as service instances. Ephedra includes
two types of hybrid services: extension services and aggregate services. Extension
services take as input a partial query solution (binding set) and extend it with additional
variable bindings. Extension services are called in the query via a SPARQL SERVICE
clause. On the contrary, aggregate services operate over a set of multiple query solutions
as the SPARQL aggregate functions (e.g., AVG, MIN, MAX) do: they take as input a
list of records and produce one or more resulting binding sets. As with the SPARQL
aggregates, aggregate services are referenced as function calls in the SELECT clause.</p>
      <p>Relevant meta-level information about the hybrid service types is summarized using
the service descriptors structured according to the service description ontology. The
ontology expands the well-known SPIN3 ontology for SPARQL query engines to capture
the relevant parameters of services.</p>
      <p>A service descriptor contains the following information:
– Input parameters and their expected datatypes. An input parameter is described
using the SPIN ontology vocabulary as a spl:Argument resource.
– Output parameters and their expected datatypes. An output parameter is described
as a spin:Column resource in the SPIN ontology.
– Expected graph pattern. The special triple patterns expected by the service are
expressed using the SPIN SPARQL syntax4. The placeholders for input/output
parameters are expressed as resources which are referenced from the input/output
parameter descriptors.
– Input and output cardinalities of a service call (optional).
:WikidataTextSearch a eph:Service ;
rdfs:label "A wrapper for the Wikidata test search." ;
eph:hasSPARQLPattern ([
2 http://docs.rdf4j.org/sail/
3 http://spinrdf.org/
4 http://spinrdf.org/sp.html#sp-TriplePattern
sp:subject :_uri ;
sp:predicate wikidata:search ;
sp:object :_token ]) ;
spin:constraint [
a spl:Argument ;
rdfs:comment "Input token" ;
spl:predicate :_token ;
spl:valueType xsd:string ] ;
spin:column [
a spin:Column ;
rdfs:comment "URI of the Wikidata resource" ;
spl:predicate :_uri ;
spl:valueType rdf:Resource ] .</p>
      <p>A descriptor for an aggregation service declares the input and output parameters
in a similar way, but instead of the list of triple patterns it defines a custom aggregate
function which will be referenced by its URI.
2.1</p>
      <sec id="sec-2-1">
        <title>Implementing service extensions</title>
        <p>To simplify the integration of new hybrid services into the framework, the architecture
provides a generic API to wrap arbitrary services and include them as SPARQL
federation members. To this end, Ephedra reuses and extends the RDF4J SAIL API. A service
is represented as a SAIL module which is responsible for extracting the values of input
parameters from a given SPARQL tuple expression, executing the actual service call,
and returning the results by binding resulting values to the output variables. Ephedra
provides abstract implementations for a generic service SAIL as well as a specific
wrapper for HTTP services. The common routines, such as extracting the input values and
output variables and wrapping the results as binding sets do not depend on the actual
service and are performed in a generic way using the declarative service descriptor.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Outlook</title>
      <p>The main directions for the future work concern further minimizing the adaptation effort
needed to deploy metaphactory in a new use case. This involves, for example,
building a library of reusable data analytics services (e.g., for common machine learning
algorithms). There are also several promising directions for improving the query
performance, in particular, exploiting more detailed meta-data about other types of federation
members: e.g., summary of the content for RDF triple stores and sets of R2RML
mappings for relational databases.</p>
      <sec id="sec-3-1">
        <title>Acknowledgements</title>
        <p>This work has been supported by the Eurostars project DIESEL (E!9367) and by the
German BMWI Project GEISER (project no. 01MD16014).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . (
          <year>2013</year>
          )
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>