<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Including Co-referent URIs in a SPARQL Query</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christian Y A Brenninkmeijer</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carole Goble</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alasdair J G Gray</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul Groth</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonis Loizou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Steve Pettifer</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, VU University of Amsterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computer Science, University of Manchester</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Linked data relies on instance level links between potentially di ering representations of concepts in multiple datasets. However, in large complex domains, such as pharmacology, the inter-relationship of data instances needs to consider the context (e.g. task, role) of the user and the assumptions they want to apply to the data. Such context is not taken into account in most linked data integration procedures. In this paper we argue that dataset links should be stored in a stand-o fashion, thus enabling di erent assumptions to be applied to the data links during query execution. We present the infrastructure developed for the Open PHACTS Discovery Platform to enable this and show through evaluation that the incurred performance cost is below the threshold of user perception.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        A key mechanism of linked data is the use of equality links between resources
across di erent datasets. However, the semantics of such links are often not
trivial: as Halpin et al. have shown sameAs, is not always sameAs [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]; and equality
is often context- and even task-speci c, especially in complex domains. For
example, consider the drug \Gleevec". A search over di erent chemical databases
return records about two di erent chemical compounds { \Imatinib" and
\Imatinib Mesylate" { which di er in chemical weight and other characteristics. When
a user requires data to perform an analysis of the compound, they would require
that these compounds are kept distinct. However, if they were investigating the
interactions of \Gleevec" with targets (e.g. proteins) then they would like these
records related. This requires two di erent sets of links between the records: one
that links data instances based on their chemical structure and another where
the data instances are linked based on their drug name. Users require the ability
to switch between these alternative scienti c assumptions [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        To support this requirement, we have developed a new approach that
applies context-dependent sets of data instance equality links at query time. The
links are stored in a stand-o fashion so that they are not intermingled with
the datasets, and are accessible through services such as BridgeDB [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] or
sameas.org3 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This allows for multiple, context-dependent linksets that can
      </p>
      <sec id="sec-1-1">
        <title>3 http://sameas.org/ accessed July 2013</title>
        <p>evolve without impacting the underlying datasets and can be activated
depending upon the user's context. This exibility is in contrast to both linked data and
traditional data integration approaches, that expose a single precon gured view
of the data. However, the use of stand-o mappings should not impact query
evaluation times from the user's perspective.</p>
        <p>
          In this paper, we present a query infrastructure to expand the instance URIs
in a given query with equivalent URIs drawn from linksets stored in a
stando fashion. It is assumed that such a query speci es which properties are to
be taken from each of the data sources, i.e. a global-as-view query [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], and as
such we do not consider ontology alignment issues [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. This query
infrastructure allows the linkset store to decide about which links are activated under a
speci c context; this will be explored in future work. We present a performance
evaluation of the query infrastructure in comparison with a typical linked data
approach (Section 4), which shows that query execution times are within the
required performance characteristics (i.e. interactive speed) while allowing for
contextual views. After the evaluation, we discuss related work and conclude.
2
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Motivation and Context</title>
      <p>
        This work is motivated by the needs of pharmacology researchers within the
context of the Open PHACTS project4; a public-private partnership aimed at
addressing the problem of public domain data integration for both academia
and major pharmaceutical companies [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Data integration is a prerequisite for
modern drug discovery as researchers need to make use of multiple
information sources to nd and validate potential drug candidates. The integration of
knowledge from these disparate sources presents a signi cant problem to
scientists, where intellectual and scienti c challenges are often overshadowed by
the need to repeatedly perform error-prone and tedious mechanical integration
tasks.
      </p>
      <p>A key challenge in the pharmacological domain is that the complexity of
the domain makes developing a single view on the pharmacological data space
di cult. For example, in some areas of the life sciences it is common to refer
to a gene using the name of the protein it encodes (a geneticist would interpret
\Rhodopsin" as meaning `the gene RHO that encodes the protein Rhodopsin')
whereas in other areas the terms are treated as referring to distinct, non
interchangeable entities. Clearly \RHO" and \Rhodopsin" cannot in all cases be
considered as synonyms, since they refer to di erent types of biological concept.
Another challenge is that the domain scientists require di erent data record links
depending on the context of their task. For example, when searching for
information about \Protein Kinase C Alpha", the scientist may want information
returned for that protein as it exists in humans5, mice6 or both. Additionally,
when connecting across databases one may want to treat these proteins as equal</p>
      <sec id="sec-2-1">
        <title>4 http://www.openphacts.org/ accessed July 2013</title>
      </sec>
      <sec id="sec-2-2">
        <title>5 http://www.uniprot.org/uniprot/P17252 accessed July 2013</title>
      </sec>
      <sec id="sec-2-3">
        <title>6 http://www.uniprot.org/uniprot/P20444 accessed July 2013</title>
        <p>in order to bring back all possible information whereas in other cases (e.g. when
the researcher is focused on humans) only information on the particular
protein should be retrieved. In both of the above cases, a single normalized view is
impossible to produce.</p>
        <p>
          To address the need for such multiple views as well as other data
integration issues, the Open PHACTS Discovery Platform has been developed using
semantic technologies to integrate the data. The overall architecture, and the
design decisions for it, are detailed in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Here, we brie y give an overview of
a typical interaction with the Discovery Platform and motivate the need for the
query infrastructure presented in this paper.
        </p>
        <p>The Open PHACTS Discovery Platform exposes a domain speci c web
service API to a variety of end user applications. Two key groups of methods are
provided: (i) resolution methods that resolve a user entry, e.g. text for the name
of a compound, to a URI for the concept; and (ii) data retrieval methods that
extract the data from the underlying datasets which have been cached into a
single triplestore. A typical interaction rst uses a resolution method to obtain
a URI for the user input and then uses that URI to retrieve the integrated data.</p>
        <p>Each data retrieval method corresponds to a SPARQL query that determines
which properties are selected from each of the underlying data sources and the
alignment between the source models. However, the query is parameterized with
a single URI that is passed in the method call. Prior to execution over the
triplestore this single URI needs to be expanded to the equivalent identi ers for
each of the datasets involved in the query. The next section provides details of the
query infrastructure that enables this functionality through stand-o mappings
that can be varied depending upon the context of the user.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Identity Mapping and Query Expansion Services</title>
      <p>To support the Open PHACTS Discovery Platform, we have developed a Query
Expansion service that replaces data instance URIs in a query with \equivalent"
URIs. The equivalent URIs are decided by the Identity Mapping Service (IMS).
Note that equivalence is assumed to be with respect to some context. Obtaining
and managing contextual links, and their e ects, are not discussed in this
paper. Instead we focus on ensuring the query infrastructure can return results in
interactive time, even when there are many equivalent URIs.
3.1</p>
      <sec id="sec-3-1">
        <title>Identity Mapping Service</title>
        <p>
          The IMS extends the BridgeDB database identi er cross-reference service [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]
for use in a linked data system. Given a URI, the IMS returns a list of equivalent
URIs drawn from the loaded VoID linksets. A VoID linkset provides a set of links
that relate two datasets together with associated provenance information [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>Multiple namespaces can be used for a dataset, e.g. UniProt is available at
http://www.uniprot.org/uniprot/ and http://purl.uniprot.org/uniprot/.
1 PREFIX chemspider: &lt;http://rdf.chemspider.com/#&gt;
2 PREFIX sio: &lt;http://semanticscience.org/resource/&gt;
3 SELECT DISTINCT ?inchikey ?molformula ?molweight
4 WHERE {
5 GRAPH &lt;http://rdf.chemspider.com&gt; {
6 &lt;http://rdf.chemspider.com/2157&gt; chemspider:inchikey ?inchikey .
7
8
9
10
11
12
13
14 } } }
}
GRAPH &lt;http://linkedchemistry.info/chembl&gt; {
&lt;http://rdf.chemspider.com/2157&gt; sio:CHEMINF_000200 _:node1 .
_:node1 a sio:CHEMINF_000042; sio:SIO_000300 ?molformula .</p>
        <p>OPTIONAL {
&lt;http://rdf.chemspider.com/2157&gt; sio:CHEMINF_000200 _:node2 .
_:node2 a sio:CHEMINF_000198; sio:SIO_000300 ?molweight .</p>
        <p>Rather than duplicate the mappings across each of these equivalent namespaces,
we support the mapping of namespaces at the dataset level.</p>
        <p>The IMS exposes a web service API on top of a MySQL database. The
database contains the mappings together with the metadata available about the
linksets from which they were read. The source code is available from http://
github.com/openphacts/BridgeDB and the service is available through http:
//dev.openphacts.org/.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Query Expansion Service</title>
        <p>The Query Expansion service takes as input a SPARQL query and expands
all of the data instance URIs appearing in that query with equivalent URIs
drawn from the IMS. It is assumed that the query has been written with
regard to the schema of each of the datasets involved but does not have
appropriate mappings for the resource URIs. For example, consider a query such as
the one given in Fig. 1; a simpli ed version of the Open PHACTS Discovery
Platform compound information query which retrieves data from two datasets.
The query has been instantiated with the URI from the ChemSpider database
for \Aspirin", viz. http://rdf.chemspider.com/2157, which was returned by
the resolution step (Section 2). The query expander is responsible for replacing
each instance URI with \equivalent" URIs retrieved from the IMS, in this case
http://linkedchemistry.info/chembl/molecule/m1280.</p>
        <p>The query expander has implemented two equivalent expansion strategies
which can be selected between by providing a parameter. The rst replaces each
instance URI by a variable and introduces a FILTER statement to consider all
equivalent URIs. For the example query, line 6 would be replaced with
?uri1 chemspider:inchikey ?inchikey .</p>
        <p>FILTER(?uri1 = &lt;http://rdf.chemspider.com/2157&gt; || ?uri1 =
&lt;http://linkedchemistry.info/chembl/molecule/m1280&gt;)
The second approach introduces UNION clauses for each of the equivalent URIs.
For the example query, line 6 would be replaced with
{ &lt;http://rdf.chemspider.com/2157&gt; chemspider:inchikey ?inchikey . }
UNION { &lt;http://linkedchemistry.info/chembl/molecule/m1280&gt;
chemspider:inchikey ?inchikey . }
Note that the information contained in the results of queries constructed based
on these expansion strategies is the same, even though strictly speaking the
queries will produce di erent bindings sets.</p>
        <p>
          Data integration queries consist of a collection of statements over multiple
data sources. The localise SPARQL subpatterns heuristic, where statements are
grouped into graph blocks according to their data source, results in more
performant queries [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. For instance URI expansion, if the heuristic is not followed
then each statement is expanded independently generating an exponential
number of alternatives which generally will not yield query answers. Additionally, if
there is a one-to-many instance mapping, then the source URI will be expanded
to two separate target URIs in the same dataset in two di erent statements. This
is not the evaluation semantics that the user desires. Queries written according
to the heuristic can avoid this shortcoming by the expansion service employing
two optimisations. First, given a mapping between the graph block name and
the underlying data source, the query expander can limit the set of expanded
URIs to those that occur in that dataset. Second, when there are many URIs
for a given graph block, these are bound once for the block rather than on a
per statement basis, i.e. if there are y equivalent URIs you will get y bindings
for the graph block rather than xy, where x is the number of statements in the
graph block. For instance, when expanding the statement on line 6 of Fig. 1 we
only need to insert the ChemSpider URI, not the equivalent ChEMBL one as it
does not appear in the ChemSpider dataset. Note that since there is only one
ChemSpider URI we do not need to employ a lter or union block and directly
insert the corresponding URI. As shown in [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], eliminating FILTER and UNION
blocks results in more performant queries. These optimisations are essential for
the complex queries used in the Open PHACTS Discovery Platform.
        </p>
        <p>The code for the query expansion service is available from https://github.
com/openphacts/QueryExpander and a deployment is accessible from http:
//openphacts.cs.man.ac.uk:9090/QueryExpander/.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>This section describes our experimental evaluation. We compare the query
evaluation time of using stand-o mappings through the query expansion service
with two baseline approaches. The evaluation investigates whether the use of a
query expansion service introduces undue performance costs.
4.1</p>
      <sec id="sec-4-1">
        <title>Experimental Design</title>
        <p>The two URI instance expansion strategies presented in this paper (
lter-bygraph and union-by-graph) are compared with the performance of the
corresponding queries when all the URIs are directly inserted in the query
(perfectURIs) and the linked data form of the query where link statements are included
in the query (linked-data). That is, for the example query in Fig. 1 the statement
&lt;http://rdf.chemspider.com/2157&gt; skos:exactMatch ?chemblid .
is added to the where clause and the instance URIs in the ChEMBL graph
replaced with the variable ?chemblid. We compare the performance of
calling the query expansion service and executing the resulting query with
directly running the baseline queries against the same triplestore. As a
consequence we anticipate that the query expander queries will be slower since they
have to call the query expansion service before being executed over the
triplestore. The queries, datasets and scripts used in the evaluation are available from
http://openphacts.cs.man.ac.uk/qePaper/.</p>
        <p>Queries and Datasets. The queries used are adapted from the Open PHACTS
Discovery Platform API7 methods \Compound Information", \Compound
Pharmacology Paginated", \Target Information", and \Target Pharmacology
Paginated". The only di erences between the queries used in this paper and the
corresponding Discovery Platform API methods are that pagination is not
considered, and CONSTRUCT blocks are replaced with SELECT statements. The queries
are characterised by their use of graph blocks to group statements by source (i.e.
following the localise SPARQL subpatterns heuristic); involving a large number
of optional statements (between 2 and 15); and returning a large number of
properties (between 12 and 21). For comparison purposes, we also used
simplied forms of the \compound information" and \target information" queries that
returned between 3 and 6 variables and at most one optional statement.</p>
        <p>
          The data used corresponds to RDF dumps of ChEMBL version 13 [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ],
ChemSpider [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], ConceptWiki8, and DrugBank [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The datasets mainly describe two
types of resource: chemical compounds and targets (e.g. proteins). Additionally,
a third type of resource is the interaction between a compound and a target
which is used to answer the two pharmacology queries. In total, the data
contains 168,783,592 triples, 290 predicates and are loaded in 4 separate named
graphs (one per dataset).
        </p>
        <p>Linksets have been separated out from the datasets and relate the concepts
across the datasets. In total, there are ve linksets providing 2,114,584 links.
Note that the linksets were loaded into the evaluation triplestore to enable the
linked-data evaluation, and represent one equivalence context.</p>
        <p>Computational Environment. The experiments were conducted on a
machine with 2 Intel 6 Core Xeon E5645 2.4GHz, 96GB RAM 1333Mhz, and
4.3TB RAID 6 (7 1TB 7200rpm) hard drives. The Virtuoso Enterprise
triplestore9, version 07.00.3202 was used for the experiments.</p>
        <sec id="sec-4-1-1">
          <title>7 https://dev.openphacts.org/ accessed July 2013</title>
        </sec>
        <sec id="sec-4-1-2">
          <title>8 http://ops.conceptwiki.org/ accessed July 2013</title>
        </sec>
        <sec id="sec-4-1-3">
          <title>9 http://virtuoso.openlinksw.com/ accessed July 2013</title>
          <p>
            Evaluation Framework. The experimental evaluation was conducted using
the Manchester University Multi-Benchmarking framework (MUM-Benchmark)10
[
            <xref ref-type="bibr" rid="ref2">2</xref>
            ], which extends the Berlin SPARQL Benchmark [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] by allowing a query
workload to be de ned and run over an arbitrary dataset. The benchmark consists
of two phases; a generation phase and an evaluation phase. Only the evaluation
phase was timed. We ran 35 repetitions of the benchmark.
          </p>
          <p>During the generation phase three versions of the queries were instantiated
from the template queries with a randomly selected ConceptWiki URI of the
correct type { compound or protein { for the query. This corresponds to the
resolution step of the interaction with the Open PHACTS Discovery Platform
(Section 2). During this phase, the perfect-URIs queries used the IMS to discover
the equivalent URI for each graph block and these were inserted in the query.
The other queries only contained the ConceptWiki URI.</p>
          <p>During the evaluation phase the queries were executed either directly over the
triplestore endpoint (perfect-URIs and linked-data) or through an endpoint that
rst called the query expansion service and then ran the resulting query (
lterby-graph and union-by-graph). Each run consisted of 10 warm-up evaluations
of the query before timing the execution of 50 evaluations and reporting the
average evaluation time.
4.2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Results</title>
        <p>
          Due to space limitations, we are unable to present all of the results from the
experiments. However, the raw experimental results and the generated graphs
are available from [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
        <p>In each of the graphs we use red to present the perfect-URIs baseline, blue
for the linked-data baseline, green for the lter-by-graph expansion strategy, and
orange for the union-by-graph strategy. The black dashes depict the number of
query results returned. Note that the compound and protein information queries
(queries 1, 2, 4 and 5) are expected to return a single result while the
pharmacology queries (queries 3 and 6) have a variable number of results dependent
upon the number of target/compound interactions.</p>
        <p>
          Fig. 2 presents the average query execution time for each strategy over 35
random seed values focusing on the rst ve queries. The results show that the
query expansion strategies are generally slower than our baselines as expected.
However, in the worst case on query 2 (complete compound information) the
query execution time for the lter-by-graph approach is 0.030 seconds, which is
below human perception [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], and is compared to 0.006 seconds and 0.011 seconds
for the perfect-URIs and linked-data baselines respectively. For query 6 (protein
pharmacology, not shown due to clarity of presentation) the linked-data baseline
performs signi cantly worse than all of the others; an average query execution
time of 124 seconds compared to 6 seconds for each of the others.
        </p>
        <p>In the compound information (query 2) and target information (query 5)
queries, the linked data baseline produced two answers instead of the expected
10 https://code.google.com/p/mum-benchmark/ (revision 29) accessed July 2013.
single answer in ten and 22 cases respectively11. The cause for the two answers
results from the DrugBank graph block being contained in an OPTIONAL clause.
Due to the inclusion of the linking statements of the form</p>
        <p>&lt;http://rdf.chemspider.com/2157&gt; skos:exactMatch ?chemblid .
in the linked-data query and the evaluation semantics of SPARQL this means
that two results are obtained when the DrugBank URI is bound. This binding
does not take place in the other forms of the query as they do not contain the link
statement. However, this extra result does not signi cantly a ect the evaluation
time of the linked-data baseline. For the two simpli ed forms of these queries,
queries 1 and 4, the linked-data baseline has more variation in its performance
and was slower than the expansion strategies in runs 17 and 5 respectively.</p>
        <p>Fig. 3 presents the query execution times for the compound pharmacology
query, ordered by the number of results returned (shown by the black dashes).
For all of the approaches, as the number of results increases the query execution
time increases. In the majority of cases the linked-data baseline performed almost
as well as the perfect-URIs baseline.</p>
        <p>Fig. 4 presents the results for the protein pharmacology query (query 6),
ordered by result size. The linked-data results as well as the results for runs 23
and 35 are ommitted for clarity of presentation. The linked-data approach
performs exceptionally poorly for this query (taking at least 100 seconds). We
suspect this is due to linking to the compound name for each interaction with the
11 Not re ected in Fig. 2.
seed protein. On runs 23 and 35, all four approaches took about 106 seconds to
return 847 and 746 answers respectively. The expansion strategies have similar
performance to the perfect-URIs baseline. Overall the results show a correlation
between the result set size and the evaluation time.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Discussion</title>
        <p>
          The performance measures show that in general the query expansion strategies
are slower than both the perfect-URIs and linked-data baselines. This is due to
the query expansion strategies rst calling the query expansion service to
expand the queries and then executing the resulting queries, whereas the baseline
approaches directly executed the queries over the RDF store. The average
difference to the perfect-URIs baseline for both query expansion approaches is 0.02
seconds. This is particularly evident in the results from the two pharmacology
queries (Figs. 3 and 4) where large results sets are returned. In these queries,
the time taken to return the answers dominates the execution time and thus the
cost of URI expansion is mitigated. We note that the overall execution time of
the query expansion queries is below the levels of human perception which are
0.05 to 0.2 seconds [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          Of the two query expansion strategies, the union-by-graph strategy
outperforms the lter-by-graph strategy in almost all cases. However, as shown in [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]
the di erence between processing FILTER and UNION clauses is dependent upon
the triplestore used and it was shown that Virtuoso performs better with UNION.
These experiments corroborate that result.
The ultimate goal of our work is to integrate data from multiple data sources.
Data integration has been widely studied both in the relational database
community and in the semantic web community [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Integration systems expose a
single view of the world to users and require the work of a domain expert to
interrelate the datasets to be integrated. Dataspaces [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] aim to lower the up
front cost by starting with rough relationships that can be re ned
automatically through user feedback. Our approach is similar in that the integration is
achieved through queries and the relationships between datasets is captured in
our global-as-view queries. However we enable di erent views of the data to be
shown based on the equivalence relationships for the instance URIs in the query.
        </p>
        <p>
          Linked data [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] is another approach to integrate data which uses instance
level relationships. Links are embedded in the data that relate the entities in
one dataset with those in another. The owl:sameAs predicate is widely used as
the linking relationship, even though the intended interpretation does not match
the OWL semantics. In fact, Halpin et al. showed that it is hard to correctly
characterise the semantic relationship between two entities [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Moreover, the
embedding of equivalence links in the data imposes a single view of the linked
data web. In our work, we enable di erent views to be con gured at query time,
by requesting di erent sets of equivalent URIs from a mapping service. In this
way we avoid the need to misuse owl:sameAs, or other equivalence predicates,
and support the ability to adaptively turn linksets on and o at query time.
        </p>
        <p>
          Correndo et al [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] proposed an approach for rewriting a SPARQL query
expressed over one ontology to a query over a second ontology which also rewrote
the instance URIs in the query. This is achieved through ontology alignments
and the sameas.org service. They identi ed the problem that a rewriting does
not hold in all contexts, which is the motivation for our query expansion
infrastructure. The principal di erence between their approach and ours is that they
completely rewrite the query into a query over the second source whereas we
leave the global-as-view query unchanged except for the inclusion of equivalent
instance URIs. This is because we focus on instance level integration rather than
schema integration. An important di erence is that this work provides
performance details on queries produced from real-world datasets.
6
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>
        In this paper, we presented a query infrastructure to enable the use of co-referent
URIs drawn from stand-o mappings. This allows the equivalence of URIs to be
decided by a mapping service such as BridgeDB or sameas.org, rather than
embedded in the data. Such a service is required to support contextualised
mappings that provide multiple views over linked data at query execution time, thus
enabling the use of scienti c lenses [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The performance of two strategies,
implemented as a query expansion service on top of a BridgeDB identity mapping
service, were compared with two baselines. While the query expansion and
execution performance times were slower (in general) than the baselines, the di erence
in response times was still below the threshold of human perception. The extra
time can be accounted for by the need to call the expansion service to expand
the query with equivalent URIs. As future work we will be enabling the IMS to
support di erent views over the data through the application of scienti c lenses
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. We will also investigate the generation of context dependent linksets from
information available in existing datasets.
      </p>
      <p>Acknowledgements The research has received support from the Innovative Medicines
Initiative Joint Undertaking under grant agreement number 115191, resources of which
are composed of nancial contribution from the European Union's Seventh Framework
Programme (FP7/2007- 2013) and EFPIA companies' in kind contribution, and the
UK EPSRC myGrid platform grant (EP/G026238/1). We would like to thank Bijan
Parsia for the discussions on SPARQL semantics and Samantha Bail for her help with
the MUM-Benchmark.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alexander</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hausenblas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Describing linked datasets with the void vocabulary</article-title>
          . Note,
          <source>W3C (March</source>
          <year>2011</year>
          ), http://www.w3.org/TR/void/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bail</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alkiviadous</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parsia</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Workman</surname>
            , D., van Harmelen,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goncalves</surname>
            ,
            <given-names>R.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garilao</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Fishmark: A linked data application benchmark</article-title>
          .
          <source>In: SSWS+HPCSW</source>
          . pp.
          <volume>1</volume>
          {
          <fpage>15</fpage>
          . CEUR Workshop Proceedings (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schultz</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The berlin sparql benchmark</article-title>
          .
          <source>International Journal on Semantic Web and Information Systems</source>
          <volume>5</volume>
          (
          <issue>2</issue>
          ),
          <volume>1</volume>
          {
          <fpage>24</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Brenninkmeijer</surname>
            ,
            <given-names>C.Y.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evelo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>A.J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stevens</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Willighagen</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          :
          <article-title>Scienti c lenses over linked data: An approach to support task speci c views of the data. a vision</article-title>
          .
          <source>In: LISC2012. CEUR</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Brenninkmeijer</surname>
            ,
            <given-names>C.Y.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>A.J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loizou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Complete results for including co-referent uris in a sparql query</article-title>
          . http://dx. doi.org/10.6084/m9.figshare.
          <volume>701203</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Card</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moran</surname>
            ,
            <given-names>T.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Newell</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The Psychology of Human-Computer Interaction</article-title>
          . Lawrence Erlbaum Associates, Inc (
          <year>1983</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Correndo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salvadores</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Millard</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glaser</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shadbolt</surname>
          </string-name>
          , N.:
          <article-title>Sparql query rewriting for implementing data integration over linked data</article-title>
          . In: EDBT/ICDT Workshops (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Doan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Halevy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ives</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Principles of data integration</article-title>
          . Morgan Kaufmann (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Euzenat</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shvaiko</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Ontology matching. Springer, rst edn. (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>A.J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loizou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Askjaer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brenninkmeijer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burger</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chichester</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evelo</surname>
            ,
            <given-names>C.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harland</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waagmeester</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.J.:</given-names>
          </string-name>
          <article-title>Applying linked data approaches to pharmacology: Architectural decisions and implementation</article-title>
          . Semantic Web Journal To appear. http://www.semantic
          <article-title>-web-journal</article-title>
          .net/system/files/swj258_1.pdf
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Halevy</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franklin</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maier</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Principles of dataspace systems</article-title>
          .
          <source>In: PODS 2006</source>
          . pp.
          <volume>1</volume>
          {
          <issue>9</issue>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Halpin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCusker</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuinness</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>H.S.:</given-names>
          </string-name>
          <article-title>When owl: sameas isn't the same: An analysis of identity in linked data</article-title>
          .
          <source>In: ISWC (1)</source>
          . pp.
          <volume>305</volume>
          {
          <issue>320</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Linked Data: Evolving the Web into a Global Data Space</article-title>
          . Morgan &amp;
          <string-name>
            <surname>Claypool</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. van Iersel,
          <string-name>
            <given-names>M.P.</given-names>
            ,
            <surname>Pico</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.R.</given-names>
            ,
            <surname>Kelder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Hanspers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Conklin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.R.</given-names>
            ,
            <surname>Evelo</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.T.A.</surname>
          </string-name>
          :
          <article-title>The bridgedb framework: standardized access to gene, protein and metabolite identi er mapping services</article-title>
          .
          <source>BMC Bioinformatics 11</source>
          ,
          <issue>5</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Loizou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>On the formulation of performant sparql queries Submitted for publication</article-title>
          . http://arxiv.org/abs/1304.0567
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Pence</surname>
            ,
            <given-names>H.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>ChemSpider: An Online Chemical Information Resource</article-title>
          .
          <source>Journal of Chemical Education</source>
          <volume>87</volume>
          (
          <issue>11</issue>
          ),
          <volume>1123</volume>
          {
          <fpage>1124</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Samwald</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jentzsch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouton</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kallesoe</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Willighagen</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hajagos</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marshall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prud</surname>
          </string-name>
          'hommeaux, E.,
          <string-name>
            <surname>Hassanzadeh</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pichler</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stephens</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Linked open drug data for pharmaceutical research and development</article-title>
          .
          <source>Journal of Cheminformatics</source>
          <volume>3</volume>
          (
          <issue>1</issue>
          ),
          <volume>19</volume>
          + (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harland</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chichester</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Willighagen</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evelo</surname>
            ,
            <given-names>C.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blomberg</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ecker</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mons</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Open</surname>
            <given-names>PHACTS</given-names>
          </string-name>
          :
          <article-title>Semantic interoperability for drug discovery</article-title>
          .
          <source>Drug Discovery Today (21-22)</source>
          ,
          <volume>1188</volume>
          {
          <fpage>1198</fpage>
          (
          <year>2012</year>
          ),
          <volume>10</volume>
          .1016/j.drudis.
          <year>2012</year>
          .
          <volume>05</volume>
          .016
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Willighagen</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waagmeester</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spjuth</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ansell</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tkachenko</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hastings</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wild</surname>
            ,
            <given-names>D.J.:</given-names>
          </string-name>
          <article-title>The chembl database as linked open data</article-title>
          .
          <source>Journal of Cheminformatics</source>
          <volume>5</volume>
          (
          <issue>23</issue>
          ) (
          <year>2013</year>
          ),
          <volume>10</volume>
          .1186/
          <fpage>1758</fpage>
          -2946-5-23
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>