<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Instance Matching Benchmark for Spatial Data: A Challenge Proposal to OAEI</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Irini Fundulaki</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Axel-Cyrille Ngonga-Ngomo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute for Applied Informatics, University of Leipzig</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Computer Science - FORTH</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>A benchmark for linking geo-spatial entities So far, only a limited number of link discovery benchmarks target the problem of linking geo-spatial entities e.g., PABench [11]. However, some of the largest knowledge bases on the Linked Open Data Web are geo-spatial knowledge bases (e.g., LinkedGeoData with more than 30 billion triples). Linking spatial resources requires techniques that differ from the classical mostly string-based approaches. In particular, considering the topology of the spatial resources and the topological relations between them is of central importance to systems driven by spatial data. We believe that due to the large amount of available geo-spatial datasets employed in Linked Data and in several domains, it is critical that benchmarks for geo-spatial link discovery are developed. For OAEI 2017, we propose the introduction of a new challenge for such systems. The benchmark that the systems will use will be based upon the Lance [9] scalable, schema-agnostic benchmark generator extended with appropriate transformations to tackle geo-spatial link</p>
      </abstract>
      <kwd-group>
        <kwd>Instance Matching</kwd>
        <kwd>Benchmarks</kwd>
        <kwd>Spatial Datasets</kwd>
        <kwd>Linked Data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The number of datasets published in the Web of Data as part of the Linked
Data Cloud is constantly increasing. The Linked Data paradigm is based on the
unconstrained publication of information by different publishers, and the
interlinking of Web resources across knowledge bases. In most cases, the cross-dataset
links are not explicit in the dataset and must be automatically determined using
Instance Matching (IM) tools (also known as record linkage [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], duplicate
detection [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and, entity resolution [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]) amongst others. The large variety of techniques
requires their comparative evaluation to determine which one is best suited for
a given context. Performing such an assessment generally requires well-defined
and widely accepted benchmarks to determine the weak and strong points of the
proposed techniques and/or tools.
      </p>
      <p>
        A number of real and synthetic benchmarks that address different data
linking challenges have been proposed for evaluating the performance of such
systems. Those include, but are not limited to, IIMB 2012 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Sandbox 2012 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
RDFT 2013 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], ID-REC 2014 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], ONTOBI 2010 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Author - Task 2015 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and
Lance 2015 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] to mention few. A more complete survey can be found in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
discovery tasks. Lance is able to produce, from an initial ontology, datasets of
arbitrary size and complexity. The configuration of the transformations will be
derived from real data. More specifically, we will use the widely accepted Linked
Data datasets such as GeoNames, LinkedGeoData, and DBpedia for this task.
The tasks proposed will focus on the different types of spatial object
representations and will be provided with different severity levels for the applied
transformations. In these transformations, objects may keep their representation, they
may change their geometry, type or attributes, merge with other objects, or
can completely disappear. This is a scenario that stems from the heterogeneous
datasets (in structure and semantics) used to describe geo-spatial entities. The
produced tasks will be used by IM tools that implement string-based as well as
topological approaches for identifying matching entities. The IM frameworks will
be evaluated for both accuracy (precision, recall and f-measure) and scalability.
Furthermore, the results will be made available in both human and
machinereadable form for further processing. Since Lance is schema-agnostic, contrary
to PABench, it will be used to produce benchmarks for different (source)
ontologies to accommodate the different requirements that stem from a variety of
applications.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Acknowledgments</title>
      <p>This work was supported by a grant from the EU H2020 Framework Programme
provided for the project HOBBIT (GA no. 688227).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Mehrotra</surname>
          </string-name>
          .
          <article-title>Supporting efficient record linkage for large data sets using mapping techniques</article-title>
          .
          <source>In WWW</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>A. K. Elmagarmid</surname>
            ,
            <given-names>P.G.</given-names>
          </string-name>
          <string-name>
            <surname>Ipeirotis</surname>
            , and
            <given-names>V.S.</given-names>
          </string-name>
          <string-name>
            <surname>Verykios</surname>
            . Duplicate Record Detection:
            <given-names>A</given-names>
          </string-name>
          <string-name>
            <surname>Survey. TKDE</surname>
          </string-name>
          ,
          <volume>19</volume>
          (
          <issue>1</issue>
          ),
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>I.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Getoor</surname>
          </string-name>
          .
          <article-title>Entity resolution in graphs</article-title>
          .
          <source>Mining Graph Data</source>
          . Wiley and Sons,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>J.</given-names>
            <surname>Aguirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Eckert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Euzenat</surname>
          </string-name>
          , et al.
          <article-title>Results of the ontology alignment evaluation initiative 2012</article-title>
          .
          <string-name>
            <surname>In</surname>
            <given-names>OM</given-names>
          </string-name>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>B.</given-names>
            <surname>Cuenca Grau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dragisic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Eckert</surname>
          </string-name>
          , et al.
          <article-title>Results of the ontology alignment evaluation initiative 2013</article-title>
          .
          <string-name>
            <surname>In</surname>
            <given-names>OM</given-names>
          </string-name>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dragisic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Eckert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Euzenat</surname>
          </string-name>
          , et al.
          <article-title>Results of the ontology alignment evaluation initiative 2014</article-title>
          .
          <string-name>
            <surname>In</surname>
            <given-names>OM</given-names>
          </string-name>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>K.</given-names>
            <surname>Zaiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Conrad</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Vater</surname>
          </string-name>
          .
          <article-title>A Benchmark for Testing Instance-Based Ontology Matching Methods</article-title>
          . In KMIS,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>M.</given-names>
            <surname>Cheatham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dragisic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Euzenat</surname>
          </string-name>
          , et al.
          <article-title>Results of the ontology alignment evaluation initiative 2015</article-title>
          .
          <string-name>
            <surname>In</surname>
            <given-names>OM</given-names>
          </string-name>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>T.</given-names>
            <surname>Saveta</surname>
          </string-name>
          , E. Daskalaki,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Flouris, I. Fundulaki, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Ngonga-Ngomo</surname>
          </string-name>
          .
          <article-title>LANCE: Piercing to the Heart of Instance Matching Tools</article-title>
          .
          <source>In ISWC</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. E. Daskalaki, G. Flouris,
          <string-name>
            <surname>I. Fundulaki</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Saveta</surname>
          </string-name>
          .
          <article-title>Instance matching benchmarks in the era of linked data</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>B.</given-names>
            <surname>Berjawi</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Duchateau</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Favetta</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Miquel</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Laurini</surname>
          </string-name>
          . Pabench:
          <article-title>Designing a taxonomy and implementing a benchmark for spatial entity matching</article-title>
          .
          <source>In GeoProcessing</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>