<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Deducing Federated SPARQL queries from RDF Mappings</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benjamin Moreau</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manoé Kieffer</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patricia Serrano-Alvarado</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Nantes University</institution>
          ,
          <addr-line>LS2N, CNRS, UMR6004, 44000 Nantes</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>OpenDataSoft</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Datasets can be virtually integrated into the Linked Data space through RDF mappings. An RDF mapping consists of rules that map data from an input dataset to RDF triples. Defining SPARQL queries is an arduous task. As RDF mappings define semantic schemas, it is possible to use them to deduce federated queries. In this demonstration, we propose MaRQ, a tool that deduces federated queries from a set of RDF mappings. The goal is to provide a clear idea of the conjunction possibilities of a dataset with other virtually integrated datasets.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction and motivation</title>
      <p>
        Having data as linked data enables vast amounts of datasets to be interconnected
creating new and innovative applications. To avoid expensive investments in
terms of storage and time, some data providers use RDF mappings to integrate
virtually and on-demand, non-RDF data into the Linked Data [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. Making
mappings is not easy, but some tools help to generate them [
        <xref ref-type="bibr" rid="ref2 ref8">2, 8</xref>
        ]. For data
providers, it would be interesting to know how a particular dataset might benefit
from existing semantic datasets. That is, to which extent, their dataset can be
combined with datasets of the Linked Data, i.e., which conjunctive queries could
be processed by a federation that includes their dataset?
      </p>
      <p>
        In the state of the art, [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] allows to generate federated queries based on RDF
datasets. This approach provides flexible parameterization of realistic
conjunctive benchmark queries. Query parameterization includes structure, complexity,
and cardinality constraints. Similarly, [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] can generate a variety of federated
queries over a given set of RDF datasets to facilitate the process of
benchmarking for federated query processing. Generated queries are conjunctive and use
the OPTIONAL and FILTER keywords. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] proposes a solution for generating
federated queries from query logs executed in the past. Generated queries are
conjunctive queries and queries using UNION, OPTIONAL, FILTER, etc.
      </p>
      <p>This work provides a low-cost solution that uses RDF mappings instead of
RDF datasets or query logs. We consider, that the particular dataset that a
? Copyright c 2021 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
provider wants to integrate into the Linked Data virtually is not materialized in
the RDF format, and a log of executed SPARQL queries does not exist.</p>
      <p>Our tool, named MaRQ, analyses RDF mappings to deduced conjunctive
BGPs. An RDF mapping can be seen as an RDF summary of the dataset it
describes. Thus, its analysis can provide helpful information with limited overhead.
MaRQ identifies the conjunction capabilities between a particular dataset and
a set of datasets of the Linked Data that are virtually integrated through RDF
mappings. It provides a ranked list of RDF datasets according to their degree of
conjunction with a particular dataset.
2</p>
      <p>MaRQ: a tool to detect joins from RDF mappings
Consider Figure 1 that shows two mappings. Mapping 1 concerns a doctors’
directory dataset. Mapping 2 concerns an Airbnb accommodations dataset. In
bold, we highlight templates or references existing in the RDF mappings. We
want to know how corresponding datasets could enrich one another. We recall
that to deduce federated SPARQL queries, MaRQ uses nothing but RDF
mappings. So deduced queries are not verified against instances of corresponding
datasets.</p>
      <p>MaRQ deduces queries of type subject-subject, object-object, and
subjectobject. The matching of terms is based on the pairwise Jaccard similarity of
the types that describe them.</p>
      <p>In MaRQ, the Jaccard similarity is the number of types in common divided by
the total number of types of two subjects (templates) described in two mappings.
Two terms from different mappings are joinable if their similarity is greater than
a configurable threshold between 0 and 1. For instance, the Jaccard similarity
between Address and Home is 0:20, i.e., 1=5. This is because these templates
have only one type in common (schema:Place), and the total number of types
of both templates is 5. If the threshold of MaRQ is equal to or greater than 0.2,
then a join between these templates will be deduced.</p>
      <p>Subject-subject queries. To identify two joinable subjects, MaRQ calculates
the similarity of all subjects. Considering a similarity threshold 0:2, in our
example, Doctor is similar to Host, and Address is similar to Home and
Neighbourhood. Thus MaRQ deduces three queries. Each query contains one join
BGP. BGPs will contain all types (rdf:type and corresponding classes) and all
predicates with the same variable in subjects and, for not rdf:type predicates,
different variables in objects. Listing 1.1 shows one of these queries. It asks for
Doctors that are also BusinessEntities.</p>
      <p>Object-object queries. To identify two joinable objects, MaRQ uses objects
that are described semantically in the mappings. In our example, three objects
are described: Address, Host and Neighbourhood. If there is a pairwise similarity
a query is deduced. Only Address is similar to Neighbourhood. BGP will
contain different variables in all subjects, the same variable in the object, and the
predicates where the objects are used. Listing 1.2 shows this query. It asks for
Doctors’ locations that are neighbourhood of Airbnb homes.</p>
      <p>It is possible also to deduce object-object queries only based on types. In our
example, MaRQ deduces two queries of this type, one with the BGP {?s1 rdf:type
schema:Person} and another with the BGP {?s1 rdf:type schema:Place}. These
queries ask for all persons and all places of both datasets.</p>
      <p>Mapping 1
ex:Directory/$Doctor</p>
      <p>Mapping 2
ex:AirBnB/$Home
rdf:type
rdf:type
rdf:type
rdf:type
schema:Person
lgdo:Doctor
vcard:Home
schema:Place
schema:location</p>
      <p>ex:Directory/$Address
dbo:speciality</p>
      <p>ex:Directory/$codeProfession
rdf:type schema:Residence
dbo:owner</p>
      <p>ex:AirBnB/$Host
schema:containedInPlace
ex:AirBnB/$Neighbourhood
juso:fullAddress</p>
      <p>ex:Directory/$address
schema:postalCode</p>
      <p>ex:Directory/$code_postal
rdf:type
rdf:type
rdf:type
rdf:type
rdf:type
rdf:type
juso:Address
schema:Place
schema:Person
gr:BusinessEntity
schema:Place
dbo:PopulatedPlace</p>
      <p>Subject-object queries. To identify subject-object joins, MaRQ identifies
objects that are described semantically in the second mapping. In our example,
objects Host and Neighbourhood are described in the second mapping. Then, it
identifies subjects in the first mapping that are joinable with these objects. In
our example, the subject Doctor is similar to the object Host, and the subject
Address is similar to the object Neighbourhood. Thus, two subject-object queries
are deduced. BGPs will contain all types (rdf:type and corresponding classes)
and properties of the joinable subject with the same variable in the subject, and
the predicates where the joinable object is used with the same subject’ variable.
Listing 1.3 shows one of these queries. It asks for doctors that are also owners
of Airbnb homes.</p>
      <p>Deduction of object-subject queries follows the algorithm of subject-object
ones, but mappings are taken in inverse order. In our example, the object Address
is similar to the subjects Home and Neighbourhood. Thus, two more queries are
deduced, and they are considered subject-object queries. Listing 1.4 shows one
of these queries. It asks for Airbnb places that are also doctors’ locations.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Demonstration</title>
      <p>
        The MaRQ implementation uses YARRRML mappings [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. You can test it in
command line https://github.com/Manoe-K/MaRQ. It is also available as a Web
application available at https://marq-priloo.univ-nantes.fr/.
      </p>
      <p>During the demonstration, attendees will be able to choose some RDF
mappings. They will then be able to choose one mapping in order to compare it to all
the others. A graph representing the number of triple patterns that MaRQ
generates by query type will be shown for each mapping. The similarity threshold
used by the Web application is 0:2 (this threshold is configurable in the command
line version). Deduced queries will be shown, and they are also downloadable.</p>
      <p>Figure 2 shows a screenshot of the MaRQ web application. In this
example, the dataset evenements-publics-cibul is compared other three datasets. The
graph shows the number of potential join queries and their types
(subjectsubject, subject-object, object-subject).</p>
      <p>Final remarks. MaRQ identifies possible conjunctive queries from RDF
mappings. The pertinence of deduced queries depends on the quality of mappings
and the similarity threshold. In addition, as each BGP contains all the possible
triple patterns of joinable terms, it is unlikely that these queries, with such a
big number of constraints, give results. However, they provide a valuable set of
BGPs whose triple patterns can be analyzed to identify joins that may return
instances.</p>
      <p>
        In addition, join queries to integrate different datasets face the problem of
entity-matching. For instance, URIs can be of different domains but referring
to the same entity. This problem is out of the scope of this work, but several
solutions exist, see [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] for a survey.
      </p>
      <p>Acknowledgment. Authors thank Fatim Touré (Master student of the University
of Nantes) for her participation in the early stages of this work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Görlitz</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thimm</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Staab</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : SPLODGE:
          <article-title>Systematic Generation of SPARQL Benchmark Queries for Linked Open Data</article-title>
          .
          <source>In: International Semantic Web Conference (ISWC)</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szekely</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knoblock</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taheriyan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muslea</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Karma: A System for Mapping Structured Sources into the Semantic Web</article-title>
          .
          <source>In: Extended Semantic Web Conference (ESWC)</source>
          ,
          <source>Poster&amp;Demo</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hacques</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skaf-Molli</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Molli</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hassad</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          :
          <article-title>PFed: Recommending Plausible Federated SPARQL Queries</article-title>
          .
          <source>In: International Conference on Database and Expert Systems Applications (DEXA)</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Heyvaert</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Meester</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dimou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verborgh</surname>
          </string-name>
          , R.:
          <article-title>Declarative Rules for Linked Data Generation at Your Fingertips</article-title>
          ! In: Extended Semantic Web Conference (ESWC),
          <source>Poster&amp;Demo</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Köpcke</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahm</surname>
          </string-name>
          , E.:
          <article-title>Frameworks for entity matching: A comparison</article-title>
          .
          <source>Data &amp; Knowledge Engineering</source>
          <volume>69</volume>
          (
          <issue>2</issue>
          ),
          <fpage>197</fpage>
          -
          <lpage>210</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faron Zucker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montagnat</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A Mapping-based Method to Query MongoDB Documents with SPARQL</article-title>
          .
          <source>In: International Conference on Database and Expert Systems Applications (DEXA) (Sep</source>
          <year>2016</year>
          ), https://hal. archives-ouvertes.fr/hal-01330146
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Moreau</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serrano</surname>
            <given-names>Alvarado</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Desmontils</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Thoumas</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>Querying non-RDF Datasets using Triple Patterns</article-title>
          . In: International Semantic Web Conference (ISWC),
          <source>Poster&amp;Demo (Oct</source>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Moreau</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Terpolilli</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serrano-Alvarado</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>A Semi-Automatic Tool for Linked Data Integration</article-title>
          . In: International Semantic Web Conference (ISWC),
          <source>Poster&amp;Demo (Oct</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Rakhmawati</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saleem</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lalithsena</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Decker</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>QFed: Query Set for Federated SPARQL Query Benchmark</article-title>
          .
          <source>In: International Conference on Information Integration and Web-based Applications &amp; Services</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>