<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Answering SPARQL Queries using Views</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Universite de Nantes</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Views are used to optimize queries and to integrate data in Databases. The data integration schema is composed of terms, they are used to pose queries to the integration system, and to describe sources data. When the data descriptions are SPARQL conjunctive queries, their number and the complexity of answering queries using them may be very high. In order to keep query answering cost low, and increase the usability of views to answer queries, we make the assumption of the simplest form of replication among services, and use triple pattern views as the Linked Data Fragments and Col-graph approaches have recently done. We propose two approaches: the SemLAV approach to integrate heterogeneous data using Local-as-View mappings to describe Linked Data and Deep Web sources, and the FEDRA approach to process queries against federations of SPARQL endpoints with replicated fragments.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>own SPARQL endpoints. These endpoints are mostly mirrors, and their use is
restricted to local users. This replication strategy is limited by the user resources,
and a user with a modest amount of resources may not set up endpoints for all
datasets relevant for all the queries she may want to answer.</p>
      <p>
        Linked Data Fragments (LDF) [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] has been proposed to exploit clients
resources to relieve servers resources and improve availability, and Col-graph [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]
has been proposed to use clients resources to replicate datasets fragments and
improve data quality. We think Linked Data users may also bene t from
replicating datasets fragments to execute federated queries; but in order to achieve
this, they need to put their own replicated fragments at other users disposal.
      </p>
      <p>
        Solving the availability limitations of Linked Data using data consumer
resources, to replicate data and to make these data available to others, may lead
to query processing performance concerns. For example, given popular datasets
with many data consumers willing to set up endpoints with datasets subsets, how
are these endpoints going to be used to execute a federated query? A very simple
solution may be to declare all of these endpoints as part of the federation used by
a federated query engine, but this simple solution may incur in very expensive
execution times. For example, in [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], we give an example where two replicas
of the relevant fragments leads to increase the execution time in two orders of
magnitude for federated query engines FedX [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] and ANAPSID [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In order
to properly exploit the bene ts of replicated datasets fragments, we propose a
source selection strategy that is aware of fragment replication able to enhance
federated query processing engines. This idea of sharing replicated fragments
among Linked Data consumers is new, and so there are no existing techniques
that can perform a source selection aware of data replication. Even if techniques
to detect data overlapping based on data summaries exists [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], applying them
to scenarios with data replication will incur in expensive computations that are
unnecessary in a replication scenario, will have accuracy limitations that may
be overcome by tailored approaches, and will make greedy decisions that choose
the best sources for a triple pattern without considering that the triple pattern
would be executed in combination with other triple patterns.
      </p>
      <p>In this thesis, our aim is to process SPARQL queries using views in two
contexts. First, in the context of data integration, answer queries using views
that describe sources as conjunctive SPARQL queries. And second, in the context
of SPARQL federations of endpoints, answer queries using queries that describe
the data fragments that have been replicated across SPARQL endpoints.</p>
      <p>The research objective is to improve the query execution
performance. This improvement may be measured in terms of amount of transferred
data or answer throughput.</p>
      <p>In particular, we want to answer the following research questions: I) When
SPARQL is used as language for LAV global schema, can view loading
outperform traditional LAV query rewriting techniques in query answering? II) How
can the knowledge about the fragment replication be used to safely reduce the
number of selected sources by federated query engines while keeping the same
answer completeness? III) Does considering BGPs instead of just triple patterns
at source selection time produce better source selections that lead to less
intermediate results?</p>
      <p>The main challenges of this work are: i) Choosing the order in which views
are loaded into the local graph that represents a partial instance of the global
schema. ii) Keeping the size of transferred data low. iii) Keeping low the
complexity of the containment computation among replicated fragments.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Proposed Approach</title>
      <p>
        We propose two approaches to solve our problem: the SemLAV approach, and
the FEDRA approach. SemLAV integrates heterogeneous data from Linked Data
and Deep Web, following the Local-as-View (LAV) paradigm [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to describe the
setup, it uses a graph instance that integrates views instances, built during query
execution, to answer users queries. FEDRA selects the sources to be contacted
to execute each subquery in order to produce the query answer.
2.1
      </p>
      <p>Formal De nition and Properties of the Approach
The SemLAV approach is based on the LAV paradigm, wrappers are used to
transform data from Deep Web sources into the global schema. Given a SPARQL
query Q on a set M of LAV views, SemLAV selects relevant views for Q and ranks
them in order to maximize query results. Next, data collected from selected
views are included into a partial instance of the global schema, where Q can
be executed whenever new data is included. The SemLAV approach has the
following properties: 1) If the global schema instance includes the data from
all the relevant views ranked by SemLAV, the query execution produces the
same answer as traditional rewriting-based query processing approaches. 2) The
e ectiveness of SemLAV is proportional to the number of covered rewritings.
3) The view loading and query execution time linearly depends on the number
of views loaded in the global schema instance. 4) Answers may be produced
incrementally if all the relevant views do not t in memory.</p>
      <p>The FEDRA approach is based on query containment and equivalence to
prune relevant sources to retrieve data from, and a reduction to the set
covering problem to produce as few subqueries as possible, sending to the endpoints
subqueries composed by as many triple patterns as possible, and hopefully
reducing the number of intermediate results to transfer from endpoints to the
query engine. Given the set of the endpoints descriptions that constitute the
federation, and a query, nd a function D that for each query triple pattern
returns the set of endpoints that need to be contacted in order to obtain a query
answer as complete as possible while intermediate results are reduced. The
FEDRA approach has the following properties: 1) FEDRA source selection used
in combination with query engines, produces at least as many answers as the
query engines alone. 2) FEDRA selects as few sources as possible for each triple
pattern. 3) FEDRA source selection aims to reduce the number of transferred
tuples from endpoints to query engines.
2.2</p>
      <p>
        Relationship between your approach and state-of-art approaches
Several approaches have been proposed for querying the Web of Data [
        <xref ref-type="bibr" rid="ref12 ref13 ref19 ref2 ref6">2, 6, 12,
13, 19</xref>
        ]. All these approaches assume that queries are expressed in terms of RDF
vocabularies used to describe the data in the RDF sources; thus, their main
challenge is to e ectively select the sources, and e ciently execute the queries on the
data retrieved from the selected sources. In contrast, SemLAV attempts to
integrate data sources, and relies on a global schema to describe data sources and to
provide a uni ed interface to the users. As a consequence, in addition to
collecting and processing data transferred from the selected sources, SemLAV decides
which of these sources need to be contacted rst, to quickly answer the query.
The Local-As-View paradigm for data integration allows to easily integrate new
data sources; further, data sources that publish entities of several concepts in
the global schema, can be naturally de ned as LAV views. Query rewriters allow
to answer queries against sources described using the LAV paradigm, some
examples are: MCD-SAT [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], GQR [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], Bucket Algorithm [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], and MiniCon [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        Recently, DAW, a source selection duplication aware approach for Linked
Data has been recently proposed [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. DAW relies on indexes with data
summaries, that allow to measure the overlapping among sources and properly reduce
the number of selected sources. Nevertheless, keeping up-to-date endpoints data
summaries may be very expensive, and they cannot guarantee the correct
overlapping detection. Moreover, DAW is a triple pattern wise approach, and data
localities are not used to choose sources that may reduce the number of
transferred tuples from endpoints to the query engines. FEDRA endpoints description
are perfect data summaries, with no accuracy issues with a size independent of
the data size, and that requires less updates that data summaries. FEDRA does
take into account the data localities to produce as few subqueries as possible
and reduce the size of intermediate results.
      </p>
      <p>
        Linked Data Fragments (LDF) [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], a new publishing strategy that reduces
the data publisher resources usage, and proposes a more active role for the data
consumers. Like LDF, our approach seeks to use data consumers computational
resources, and hopefully it may contribute to improve data availability of data
publishers. Contrarily to LDF, FEDRA approach seeks to share the query
execution among data consumers, and does not intend that each client communicates
with the data publisher, and creates a cache of her queries datasets.
      </p>
      <p>
        FedX [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] and ANAPSID [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] are state-of-art query engines for federations of
SPARQL endpoints. Both, FedX and ANAPSID, send queries for triple patterns
that should be evaluated in more than one endpoint, individually to the
endpoint; and group into exclusive groups (FedX) or star-shaped groups
(ANAPSID) the triple patterns that can be solely evaluated using one endpoint. We
have extended FedX and ANAPSID query engines with FEDRA source
selection strategy, and study FEDRA impact during query execution.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Implementation of the Proposed Approach</title>
      <p>
        We use SPARQL conjunctive queries as views that describe sources contents. As
in the Bucket Algorithm [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], relevant sources are selected per query subgoal, i.e.,
triple pattern. And the selected sources are ranked according to the number of
query subgoals that are covered by their views. Loading rst the views that cover
more subgoals aims to produce answers as soon as possible. We expect to observe
that using LAV paradigm to answer SPARQL queries allows to integrate
heterogeneous data sources, and that the SemLAV approach produce more answers
sooner than traditional query rewriting-based approaches. We will use variable
substitution [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] to detect containment among fragments, and containment to
detect equivalence of fragments [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Endpoints that have replicated equivalent
fragments are considered as equivalent sources for that fragment. Fragment
containment will be used to select the set of non-redundant relevant fragments for
each triple pattern. Then, a second selection will be done at BGP level to reduce
the number of subqueries to be sent to the endpoints, and in consequence, the
number of transferred tuples. To further prune the selected sources, a set
covering heuristic [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] will be used to determine the minimal set of endpoints that can
be used to execute the subqueries. The approach implementation may be done
stand-alone, having as input a query without SERVICE clauses, it produces
a query with SERVICE clauses that delegate subqueries execution to remote
endpoints; or existing federated query engines may be extended with FEDRA
source selection strategy, to be used before the engines optimizations. We expect
to observe that query engines enhanced with FEDRA incur in less intermediate
results, and produce answers sooner.
      </p>
      <p>
        The SemLAV approach has been formalized [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], and empirical tests have
been performed [
        <xref ref-type="bibr" rid="ref21 ref9">21, 9</xref>
        ]. Some limitations of SemLAV implementation are that it
lacks of strategies to overcome the memory limitations, and that views are loaded
sequentially because the data store used did not allow parallel loading of data.
The FEDRA source selection strategy for federations with replicated fragments
has been proposed in [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], and some empirical tests have been performed [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
Currently FedX and ANAPSID query engines have been extended to include
the FEDRA source selection strategy. A limitation of the FedX extension is that
it may increase the number of intermediate results in some cases, when one
endpoint is selected for triple patterns that do not share variables.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Empirical Evaluation Methodology</title>
      <p>The approaches will be used to execute queries in di erent setups, and they will
be compared with alternative approaches.</p>
      <p>
        For the SemLAV approach, the Berlin Benchmark [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] dataset generator will
be used to generate a ten million triples dataset, and queries and views based
on the ones proposed in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] will be used. To stress the scalability of SemLAV
and query rewriting, 476 views are used. SemLAV performance is compared to
the performance of three query rewriters: GQR [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], MCD-SAT [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and
MiniCon [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Query execution performance will be measured in terms of answer size,
query execution time, amount of memory used, and answer throughput.
      </p>
      <p>
        Our research hypothesis is that the SemLAV approach will produce a higher
answer throughput than traditional query rewriting-based approaches, for LAV
data integration in the context of Linked Data and the Deep Web. Currently,
the SemLAV evaluation has been performed, and results are available at [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>For the FEDRA approach, we will study two federated query engines: FedX
and ANAPSID. As baseline we will consider the federated query engines source
selection strategy. We will consider real and synthetic datasets of varying sizes
from 72 thousand to 10 millions triples. Query generators will be used to produce
more than 10,000 queries for each dataset. And we will use each dataset to setup
a federation of ten endpoints, and simulate the execution of federated queries.
The federations will contain endpoints with heterogeneous replicated fragments,
and a sample of 100 random queries will be taken for each federation. For each
query, the number of selected sources per triple pattern, source selection time,
execution time, answer completeness, and number of transferred tuples will be
measured. The obtained results will be statistically analyzed using the Wilcoxon
signed rank test for paired non-uniform data.</p>
      <p>In particular, we have the following research hypotheses: 1) The FEDRA
approach will select signi cantly less sources than federated query engines like
FedX and ANAPSID. 2) The FEDRA approach enhances federated query like
FedX and ANAPSID, and reduces the number of intermediate results during
query execution. 3) The reductions that the FEDRA approach achieves in terms
of number of selected sources and intermediate results, do not reduce the number
of obtained answers.</p>
      <p>
        Experiments with six federations and extended version of the query engines
have already been performed, and they support our research hypotheses [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Lessons Learned, Open Issues and Future Directions</title>
      <p>Query rewriting approaches are too stressed by the large number of triple
patterns in SPARQL queries and the high number of sources in the Web, these
characteristics prevents query rewriters to o er a practical query evaluation strategy
for Linked Data. SemLAV uses query rewriters most basic information, buckets,
to select relevant views, and ranks sources in a way that when k views have been
loaded, they cover the maximal number of rewriting that can be covered with
k views. The SemLAV approach implementation can be improved by loading
views in parallel, and considering memory limits.</p>
      <p>In federations of SPARQL endpoints with replicated fragments, federated
query engines need to be enhanced with a source selection approach like FEDRA
that allows to select sources for each triple pattern in a way that data transferred
from endpoints to query engines is reduced. The strategy used by FEDRA gets
excellent results when used with ANAPSID, but results are less good when
used with FedX. ANAPSID does not send subqueries with Cartesian products
to the endpoints, but FedX may do it when triple patterns that do not share
variables need to be executed in the same endpoint. Cartesian products may
signi cantly increase the number of transferred tuples when compared to the
evaluation of each triple pattern individually. These limitations may be overcome
with a stand-alone implementation of the FEDRA approach that transforms
a plain query into a query with query decomposition and source localization
represented as SERVICE clauses that avoids Cartesian products. Unfortunately,
such implementation has faced with query engines limited support for SERVICE
clauses. In this direction, we are currently working in a version that do rewrite
the query using SERVICE clauses, and use some heuristics to overcome current
federated query engines limitations during execution.</p>
      <p>
        Other extensions of FEDRA include the usage of cost functions, and the
consideration of divergent fragments. Cost functions may be used to select the
endpoints that satisfy the user criteria, e.g., in [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] public endpoint usage was
reduced. Removing the assumption that all the endpoints fragments descriptions
are up-to-date, endpoints may o er data with di erent levels of divergence with
respect the current dataset, in the same direction as in [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>Acknowledgments. This thesis is supervised by Pascal Molli and Hala
SkafMolli. It has received contributions from Maria-Esther Vidal.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>S.</given-names>
            <surname>Abiteboul</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Manolescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rigaux</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-C. Rousset</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Senellart</surname>
          </string-name>
          .
          <source>Web Data Management</source>
          . Cambridge University Press, New York, NY, USA,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Acosta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vidal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lampo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Castillo</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. Ruckhaus. ANAPSID:</surname>
          </string-name>
          <article-title>An Adaptive Query Processing Engine for SPARQL Endpoints</article-title>
          .
          <source>In Aroyo et al. [4]</source>
          , pages
          <fpage>18</fpage>
          {
          <fpage>34</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>G.</given-names>
            <surname>Antoniou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Grobelnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. P. B.</given-names>
            <surname>Simperl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Parsia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Plexousakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. D.</given-names>
            <surname>Leenheer</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. Z</surname>
          </string-name>
          . Pan, editors.
          <source>The Semantic Web: Research and Applications - 8th Extended Semantic Web Conference, ESWC</source>
          <year>2011</year>
          , Heraklion, Crete, Greece, May 29-June 2,
          <year>2011</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>I</given-names>
          </string-name>
          , volume
          <volume>6643</volume>
          of Lecture Notes in Computer Science. Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>L.</given-names>
            <surname>Aroyo</surname>
          </string-name>
          et al., editors.
          <source>ISWC</source>
          <year>2011</year>
          ,
          <string-name>
            <surname>Part</surname>
            <given-names>I</given-names>
          </string-name>
          , volume
          <volume>7031</volume>
          <source>of LNCS</source>
          . Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Arvelo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bonet</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.-E.</given-names>
            <surname>Vidal</surname>
          </string-name>
          .
          <article-title>Compilation of query-rewriting problems into tractable fragments of propositional logic</article-title>
          .
          <source>In AAAI</source>
          , pages
          <volume>225</volume>
          {
          <fpage>230</fpage>
          . AAAI Press,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>C.</given-names>
            <surname>Basca</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernstein</surname>
          </string-name>
          .
          <article-title>Avalanche: Putting the Spirit of the Web back into Semantic Web Querying</article-title>
          . In A. Polleres and H. Chen, editors,
          <source>ISWC Posters&amp;Demos</source>
          , volume
          <volume>658</volume>
          <source>of CEUR Workshop Proceedings. CEUR-WS.org</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          and
          <string-name>
            <surname>A. Schultz.</surname>
          </string-name>
          <article-title>The berlin sparql benchmark</article-title>
          .
          <source>Int. J. Semantic Web Inf. Syst.</source>
          ,
          <volume>5</volume>
          (
          <issue>2</issue>
          ):1{
          <fpage>24</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>R.</given-names>
            <surname>Castillo-Espinola</surname>
          </string-name>
          .
          <article-title>Indexing RDF data using materialized SPARQL queries</article-title>
          .
          <source>PhD thesis</source>
          , Humboldt-Universitat zu Berlin,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>P.</given-names>
            <surname>Folz</surname>
          </string-name>
          , G. Montoya,
          <string-name>
            <given-names>H.</given-names>
            <surname>Skaf-Molli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Molli</surname>
          </string-name>
          , and
          <string-name>
            <surname>M.-E. Vidal.</surname>
          </string-name>
          <article-title>SemLAV: Querying Deep Web and Linked Open Data with SPARQL</article-title>
          .
          <source>In ESWC: Extended Semantic Web Conference</source>
          , volume
          <volume>476</volume>
          of The Semantic Web:
          <article-title>ESWC 2014 Satellite Events</article-title>
          , pages
          <volume>332</volume>
          {
          <fpage>337</fpage>
          , Anissaras/Hersonissou, Greece, May
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>C. Gutierrez</surname>
            ,
            <given-names>C. A.</given-names>
          </string-name>
          <string-name>
            <surname>Hurtado</surname>
            ,
            <given-names>A. O.</given-names>
          </string-name>
          <string-name>
            <surname>Mendelzon</surname>
            , and
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Perez</surname>
          </string-name>
          .
          <article-title>Foundations of Semantic Web databases</article-title>
          .
          <source>J. Comput. Syst. Sci.</source>
          ,
          <volume>77</volume>
          (
          <issue>3</issue>
          ):
          <volume>520</volume>
          {
          <fpage>541</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>A. Y.</given-names>
            <surname>Halevy</surname>
          </string-name>
          .
          <article-title>Answering queries using views: A survey</article-title>
          .
          <source>VLDB J</source>
          .,
          <volume>10</volume>
          (
          <issue>4</issue>
          ):
          <volume>270</volume>
          {
          <fpage>294</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>A.</given-names>
            <surname>Harth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Karnstedt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          , K.-U. Sattler, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Umbrich</surname>
          </string-name>
          .
          <article-title>Data summaries for on-demand queries over linked data</article-title>
          . In M. Rappa,
          <string-name>
            <given-names>P.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Freire</surname>
          </string-name>
          , and S. Chakrabarti, editors,
          <source>WWW</source>
          , pages
          <volume>411</volume>
          {
          <fpage>420</fpage>
          . ACM,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>O.</given-names>
            <surname>Hartig</surname>
          </string-name>
          .
          <article-title>Zero-knowledge query planning for an iterator implementation of link traversal based query execution</article-title>
          .
          <source>In Antoniou et al. [3]</source>
          , pages
          <fpage>154</fpage>
          {
          <fpage>169</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>B.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , and
          <string-name>
            <surname>K. C.-C.</surname>
          </string-name>
          <article-title>Chang. Accessing the Deep Web</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>50</volume>
          (
          <issue>5</issue>
          ):
          <volume>94</volume>
          {
          <fpage>101</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. L. D. Iban~ez, H.
          <string-name>
            <surname>Skaf-Molli</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Molli</surname>
            , and
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Corby.</surname>
          </string-name>
          Col-Graph:
          <article-title>Towards Writable and Scalable Linked Open Data</article-title>
          . In P. Mika et al., editors,
          <source>ISWC</source>
          <year>2014</year>
          ,
          <string-name>
            <surname>Part</surname>
            <given-names>I</given-names>
          </string-name>
          , volume
          <volume>8796</volume>
          <source>of LNCS</source>
          , pages
          <volume>325</volume>
          {
          <fpage>340</fpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Johnson</surname>
          </string-name>
          .
          <article-title>Approximation Algorithms for Combinatorial Problems</article-title>
          . In A. V.
          <string-name>
            <surname>Aho</surname>
          </string-name>
          et al., editors,
          <source>ACM Symposium on Theory of Computing</source>
          , pages
          <volume>38</volume>
          {
          <fpage>49</fpage>
          . ACM,
          <year>1973</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17. G. Konstantinidis and
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Ambite</surname>
          </string-name>
          .
          <article-title>Scalable query rewriting: a graph-based approach</article-title>
          . In T. K. Sellis,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kementsietsidis</surname>
          </string-name>
          , and Y. Velegrakis, editors,
          <source>SIGMOD Conference</source>
          , pages
          <volume>97</volume>
          {
          <fpage>108</fpage>
          . ACM,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>D.</given-names>
            <surname>Kossmann</surname>
          </string-name>
          .
          <article-title>The state of the art in distributed query processing</article-title>
          .
          <source>ACM Computer Survey</source>
          ,
          <volume>32</volume>
          (
          <issue>4</issue>
          ):
          <volume>422</volume>
          {
          <fpage>469</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. G. Ladwig and
          <string-name>
            <given-names>T.</given-names>
            <surname>Tran</surname>
          </string-name>
          . Sihjoin:
          <article-title>Querying remote and local linked data</article-title>
          .
          <source>In Antoniou et al. [3]</source>
          , pages
          <fpage>139</fpage>
          {
          <fpage>153</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>A. Y.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rajaraman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Ordille</surname>
          </string-name>
          .
          <article-title>Querying heterogeneous information sources using source descriptions</article-title>
          . In T. M.
          <string-name>
            <surname>Vijayaraman</surname>
            ,
            <given-names>A. P.</given-names>
          </string-name>
          <string-name>
            <surname>Buchmann</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Mohan</surname>
          </string-name>
          , and N. L. Sarda, editors,
          <source>VLDB</source>
          , pages
          <volume>251</volume>
          {
          <fpage>262</fpage>
          . Morgan Kaufmann,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21. G. Montoya, L. D. Iban~ez, H.
          <string-name>
            <surname>Skaf-Molli</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Molli</surname>
            , and
            <given-names>M.-E. Vidal.</given-names>
          </string-name>
          <article-title>SemLAV: Local-As-View Mediation for SPARQL Queries</article-title>
          . Transactions on
          <string-name>
            <surname>Large-Scale Dataand Knowledge-Centered Systems</surname>
            <given-names>XIII</given-names>
          </string-name>
          , pages
          <volume>33</volume>
          {
          <fpage>58</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. G. Montoya,
          <string-name>
            <given-names>H.</given-names>
            <surname>Skaf-Molli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Molli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.-E.</given-names>
            <surname>Vidal</surname>
          </string-name>
          . Fedra:
          <article-title>Query Processing for SPARQL Federations with Divergence</article-title>
          .
          <source>Technical report</source>
          , Universite de Nantes, May
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. G. Montoya,
          <string-name>
            <given-names>H.</given-names>
            <surname>Skaf-Molli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Molli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.-E.</given-names>
            <surname>Vidal</surname>
          </string-name>
          .
          <article-title>E cient Query Processing for SPARQL Federations with Replicated Fragments</article-title>
          . Jan.
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24. G. Montoya,
          <string-name>
            <given-names>H.</given-names>
            <surname>Skaf-Molli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Molli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.-E.</given-names>
            <surname>Vidal</surname>
          </string-name>
          .
          <source>Federated SPARQL Queries Processing with Replicated Fragments. In The Semantic Web - ISWC 2015 - 14th International Semantic Web Conference</source>
          , Bethlehem, United States, Oct.
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>M. Saleem</surname>
            ,
            <given-names>A.-C. N.</given-names>
          </string-name>
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>J. X.</given-names>
          </string-name>
          <string-name>
            <surname>Parreira</surname>
            ,
            <given-names>H. F.</given-names>
          </string-name>
          <string-name>
            <surname>Deus</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Hauswirth</surname>
          </string-name>
          . DAW:
          <string-name>
            <surname>Duplicate-AWare Federated Query</surname>
          </string-name>
          <article-title>Processing over the Web of Data</article-title>
          . In H. Alani et al., editors,
          <source>ISWC</source>
          <year>2013</year>
          ,
          <string-name>
            <surname>Part</surname>
            <given-names>I</given-names>
          </string-name>
          , volume
          <volume>8218</volume>
          <source>of LNCS</source>
          , pages
          <volume>574</volume>
          {
          <fpage>590</fpage>
          . Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <given-names>A.</given-names>
            <surname>Schwarte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Haase</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schenkel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          .
          <source>FedX: Optimization Techniques for Federated Query Processing on Linked Data. In Aroyo et al. [4]</source>
          , pages
          <fpage>601</fpage>
          {
          <fpage>616</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <given-names>R.</given-names>
            <surname>Verborgh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. V.</given-names>
            <surname>Sande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Colpaert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Coppens</surname>
          </string-name>
          , E. Mannens, and R. V. de Walle.
          <article-title>Web-Scale Querying through Linked Data Fragments</article-title>
          . In C. Bizer et al., editors,
          <source>WWW Workshop on LDOW</source>
          <year>2014</year>
          , volume
          <volume>1184</volume>
          <source>of CEUR Workshop Proceedings. CEUR-WS.org</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>