<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extending R2RML to a source-independent mapping language for RDF</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anastasia Dimou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miel Vander Sande</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pieter Colpaert</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Erik Mannens</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rik Van de Walle</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ghent University - iMinds - Multimedia Lab Gaston Crommenlaan 8</institution>
          ,
          <addr-line>bus 201, B-9050 Ledeberg-Ghent</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Although reaching the fth star of the Open Data deployment scheme demands the data to be represented in RDF and linked, a generic and standard mapping procedure to deploy raw data in RDF was not established so far. Only the R2RML mapping language was standardized but its applicability is limited to mappings from relational databases to RDF. We propose the extension of R2RML to also support mappings of data sources in other structured formats (indicatively CSV, TSV, XML, JSON). Broadening further its scope, the focus is put on the mappings and their optimal reuse. The language becomes sourceagnostic, and resources are integrated and interlinked at a primary stage.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>State of the art</title>
      <p>
        Beyond R2RML which has already several implementations3, other RDB2RDF
mapping languages were de ned [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In the same context, there are corresponding
languages to support CSV-to-RDF mappings (CSV2RDF), e.g., the XLWrap's
mapping language [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], the Mapping Master's M2 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and Vertere4. On the other
hand, in the case of mappings from XML to RDF (XML2RDF), the di erent
tools rely mostly on existing XML solutions. To be more precise, XSLT-based
approaches were explored, as the Krextor [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and the AstroGrid-D5 mapping
tools, while other implementations deploy mappings using XPath and XQuery,
e.g., the Tripliser6 and the XSPARQL [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. These solutions for XML sources lead
to mappings on the syntactic level rather than on the semantic level or fail to
provide a solution applicable to a broader domain. Beyond the standard
ExtractMap-Load (EML) mappings, dynamic query translation was also explored, e,g,
in the case of Tarql7 (CSV2RDF) and XSPARQL (mapping and integration
of XML, RDB and RDF resources).
      </p>
      <p>
        In general, most tools deploy mappings from a certain source format to RDF
(source-centric approaches ). There are only a few tools that provide mappings
from various source formats to RDF {DataLift [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the DataTank [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Karma [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ],
Open Re ne8 and Virtuoso Sponger9 are the most well known{ but only the
DataTank uses a mapping language. For the latter's needs, Vertere was extended
not only to cover CSV2RDF mappings but mappings from other structured data
sources as well, namely databases, XML and JSON. Since R2RML became a
W3C standard and due to its analogous nature to Vertere, the extension of
R2RML is considered a prominent solution and its applicability veri ed.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Extending R2RML for a more generic use</title>
      <p>An extension of the R2RML language is proposed, aiming to broaden its scope
beyond RDB2RDF mappings, to cover every structured data format (a
GlobalAs-View approach), and to address the limitations of existing languages. The
R2RML's RDF graphs are used to express mappings independently of the
source format. Therefore, the same custom mappings are reused whether the
source les are in the same format or not, only by redetermining the references
to the source values to be mapped, as the expected custom mapping de
nitions remain the same. The vocabulary extending the R2RML is available at
http://mmlab.be/users/andimou/rml.ttl. The expansion is achieved as follows:
3 http://www.w3.org/2001/sw/rdb2rdf/wiki/Implementations
4 https://github.com/knudmoeller/Vertere-RDF
5 http://www.gac-grid.de/project-products/Software/XML2RDF.html
6 http://daverog.github.io/tripliser/
7 https://github.com/cygri/tarql
8 http://openre ne.org/
9 http://virtuoso.openlinksw.com/dataspace/doc/dav/wiki/Main/VirtSponger
Extending RDF triples mapping. Triples map is extended not only to
map each row in the logical table, but each resource in the logical source.
To this end, the rr:logicalTable and rr:tableName become a sub-property of
rml:logicalSource and rml:sourceName respectively, while rr:elementName for
XML sources and rr:objectName for JSON sources are introduced. In the
example, books is a logical table's, a JSON object's or an XML element's name.
&lt;# RDB_CSV_map &gt; rml : logicalSource [ rr: tableName " BOOKS " ];
rr: subjectMap [ rr: template " http :// data . example . com / books /{ ISBN }" ];
rr: predicateObjectMap [ rr: predicate ex:id; rr: objectMap [ rml : resource "ID" ] ].
&lt;# XML_map &gt; rml : logicalSource [ rml : elementName "/ books " ];
rr: subjectMap [ rr: template " http :// data . example . com / books /{ book / ISBN }"];
rr: predicateObjectMap [ rr: predicate ex:id; rr: objectMap [ rml : resource " book / ISBN@id " ] ].
&lt;# JSON_map &gt; rml : logicalSource [ rml : objectName " books " ];
rr: subjectMap [ rr: template " http :// data . example . com / books /{ book . ISBN }"];
rr: predicateObjectMap [ rr: predicate ex:id; rr: objectMap [ rml : resource " book .id" ] ].
Extending resources' mapping. In the same context, term maps are
extended to generate RDF terms from any logical resource, either this is a table
row, an XML element or a JSON object. The column-valued term map is
extended to cover every resource term map. Therefore, the R2RML's rr:column
property becomes a sub-property of the rml:resource which is a valid column
name for relational databases and CSV les, a valid XPath expression for an
XML node's or attribute's absolute path and a valid path pattern in JavaScript
syntax for objects in JSON source les, as in the aforementioned example.
Multiple entities per row. Most of the mapping languages (including R2RML)
follow the entity-per-row model and consider that each row's RDF triples are
mapped to the same subject. In its extended version, R2RML can map sets of
columns to di erent subjects, which are then related among each other with a
predicate-object triples map. For example, a row may have several columns with
information about an event and a few of them refer to its location, e.g., latitude
and longitude. Using this single row a triples map may be de ned for the event
while another triples map may be de ned for the location where the event takes
place (this mapping de nition might be reused for other locations' mapping)
and the two of them are related with a predicate-object triples map, as in the
following example:
&lt;# Event_map &gt; rml : logicalSource [ rml : elementName "/ events " ];
rr: subjectMap [ rr: template " http :// data . example . com / events /{ event /id }" ];
rr: predicateObjectMap
[ rr: predicate ex: location ; rr: objectMap [ rr: parentTriplesMap &lt;# Location_map &gt; ] ] ,
[ rr: predicate ex: transport ; rr: objectMap [ rr: parentTriplesMap &lt;# Transport_map &gt; ;
rr: joinCondition [ rr: child " event / bus / num "; rr: parent " BUS_NUM " ] ] ].
&lt;# Location_map &gt; rml : logicalSource [ rml : elementName "/ events " ];
rr: subjectMap [
rr: template " http :// data . example . com / location /{ event / location / lat } ,{ event / location / long }"].
&lt;# Transport_map &gt; rml : logicalSource [ rr: tableName " TRANSPORTATIONS " ];
rr: subjectMap [ rr: template " http :// data . example . com / transport /{ TYPE }/{ BUS_NUM }"];
rr: predicateObjectMap [ rr: predicate ex: name ; rr: objectMap [ rml : resource " BUS_NAME " ] ].
Extended the logical sources. According to the rr:sqlQuery, the rml:xmlQuery
is adapted and both are sub-properties of rml:query to serve a query against a
source le. In the same context the rml:queryLanguage is de ned to determine
which language is used (indicatively, a W3C standard in the case of XML).
Integrated mapping. Extending the reference object map, one can use the
subjects of another triples map as the objects generated by a predicate-object map.
Since the triples maps may be based on di erent logical sources, the potential to
create triples based on integrated sources emerges. At the aforementioned event
example, an element node may refer to the number of the bus going to the event
location, but the bus names are associated to the bus numbers at a separate
table which is mapped by another triples map. The mappings of both of them
are de ned and a predicate-object terms map may be used to relate them.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <p>A generic mapping language is proposed to handle the mappings from di erent
source formats to RDF. The uppermost goal of such an extension is to keep the
focus on the mappings to be expressed rather than on the data and their original
structure. With this work, we bring into discussion its feasibility, possible barriers
and aspects that should be taken into consideration. In the future the arising
generic mapping language will be used at the DataTank, instead of Vertere, to
cover mappings from di erent source formats to RDF and, in the same time, to
confront with the standard mapping language for the RDB2RDF mappings.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Hert</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reif</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gall</surname>
          </string-name>
          , H.C.
          <article-title>: A comparison of RDB-to-RDF mapping languages</article-title>
          .
          <source>In: Proceedings of the 7th International Conference on Semantic Systems. I-Semantics '11</source>
          , New York, NY, USA, ACM (
          <year>2011</year>
          )
          <volume>25</volume>
          {
          <fpage>32</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Langegger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Wo , W.:
          <article-title>XLWrap { Querying and Integrating Arbitrary Spreadsheets with SPARQL</article-title>
          .
          <source>In: Proceedings of the 8th International Semantic Web Conference. ISWC '09</source>
          , Berlin, Heidelberg, Springer-Verlag (
          <year>2009</year>
          )
          <volume>359</volume>
          {
          <fpage>374</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>O</given-names>
            <surname>'Connor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.J.</given-names>
            ,
            <surname>Halaschek-Wiener</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Musen</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.A.</surname>
          </string-name>
          :
          <article-title>Mapping Master: a exible approach for mapping spreadsheets to OWL</article-title>
          .
          <source>In: Proceedings of the 9th International Semantic Web Conference on The Semantic Web - Volume Part II. ISWC'10</source>
          , Berlin, Heidelberg, Springer-Verlag (
          <year>2010</year>
          )
          <volume>194</volume>
          {
          <fpage>208</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Lange</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Krextor - an extensible framework for contributing content math to the Web of Data</article-title>
          .
          <source>In: Proceedings of the 18th Calculemus and 10th international conference on Intelligent computer mathematics. MKM'11</source>
          , Berlin, Heidelberg, SpringerVerlag (
          <year>2011</year>
          )
          <volume>304</volume>
          {
          <fpage>306</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bischof</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Decker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krennwallner</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopes</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polleres</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Mapping between rdf and xml with xsparql</article-title>
          .
          <source>Journal on Data Semantics</source>
          <volume>1</volume>
          (
          <issue>3</issue>
          ) (
          <year>2012</year>
          )
          <volume>147</volume>
          {
          <fpage>185</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Schar e, F.,
          <string-name>
            <surname>Atemezing</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Troncy</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gandon</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villata</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bucher</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamdi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bihanic</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kepeklian</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cotton</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Euzenat</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vandenbussche</surname>
            ,
            <given-names>P.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vatant</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Enabling Linked Data publication with the Datalift platform</article-title>
          .
          <source>In: Proc. AAAI workshop on semantic cities</source>
          , Toronto, Canada (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Vander</given-names>
            <surname>Sande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Colpaert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Van Deursen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Mannens</surname>
          </string-name>
          , E., Van de Walle, R.:
          <article-title>The DataTank: an open data adapter with semantic output</article-title>
          .
          <source>In: 21st International Conference on World Wide Web, Proceedings</source>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szekely</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knoblock</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taheriyan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muslea</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Karma: A system for mapping structured sources into the Semantic Web</article-title>
          .
          <source>In: 9th Extended Semantic Web Conference (ESWC2012)</source>
          . (May
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>