<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Enhancing Dataset Quality Using Keys</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tommaso Soru</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edgard Marx</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Axel-Cyrille Ngonga Ngomo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>tsoru</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ngongag@informatik.uni-leipzig.de</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AKSW, Department of Computer Science, University of Leipzig</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Linked Data principles provide a decentral approach for publishing structured data in RDF on the Web. A consequence of this architectural choice is a high variance in the quality of the RDF datasets which constitute the Linked Data cloud. In this demo paper, we address a particular aspect of quality, i.e., the discriminability of resources. During our demo, we will present our simple three-step approach and interface, which allows data publishers to detect the resources in their dataset that are indistinguishable with respect to a given set of properties. Our approach is highly scalable as it relies on ROCKER, a novel algorithm for key discovery. Our evaluation on DBpedia suggests that even very commonly-used data sources are still in need to significant improvement to abide by the discriminability criterion.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The quality of RDF datasets on the Web varies significantly [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. While several
approaches have been developed to check the quality of datasets (see [
        <xref ref-type="bibr" rid="ref2 ref5">2,5</xref>
        ] for an overview),
the discriminability of resources (see Section 2 for a formal definition) has been paid
little attention to. However, improving the discriminability of resources has been shown
to be beneficiary for the quality of the links created by automatic linking processes [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
Hence, this process promises datasets to improve their abiding by the fourth Linked
Data principle.
      </p>
      <p>
        In this demo, we present a tool that addresses exactly this gap. Our framework1
makes use of ROCKER [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], a state-of-the-art algorithm for key discovery, to detect
indistinguishable resources and facilitate the curation of these resources. The following
section introduces some preliminaries; we then describe the approach implemented on
the demo and illustrate a use case; thereafter, we conclude.
1 Available online at http://rocker.aksw.org/.
2 For the definition of CBD, see http://www.w3.org/Submission/CBD/.
two resources r1; r2 2 R distinguishable w.r.t. a set of properties P = fp1; : : : png iff
9p 2 fp1; : : : png9o : ((r1; p; o) ^ :(r2; p; o)) _ (:(r1; p; o) ^ (r2; p; o)).
      </p>
      <p>Given a knowledge base K, the idea behind key discovery is to find one or all sets
of properties which make their respective subjects distinguishable in K. We call a set
of properties P P a key for a knowledge base K (short: key, denoted key(P; K)) if
all resources in K are distinguishable w.r.t. P . Let us consider the smallest set S0 S
such that every pair of distinct elements in S n S0 are distinguishable from each other.
The discriminability score of S w.r.t. P is then calculated as scoreP (S) = jSnS0j . P is
jSj
called a k-almost-key if S0 has cardinality k.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Enhancing Dataset Quality Using Keys</title>
      <p>The intuition behind our framework is based on the following principle. Since keys are
defined as unique descriptions of resources, any description collision can be considered
as a potential error. We then assume that the error rate for a key is no lower than a
threshold value . Thereafter, we ask ROCKER to find any P such that scoreP (S)
= 1 jSkj . We expect that for each k-almost-key P , curating the k resources in S0
would lead to a dataset with higher accuracy, easier to link and thus fitter for use in
applications which rely, e.g., on federated data sources.</p>
      <p>
        A similar approach was introduced in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], where the authors discover owl:sameAs
links among resources which do not obey by almost-keys3. However, the definition of
key adopted in the paper presented some imperfections, as discussed in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>The contributions of our work can be listed as follows: (i) We devise a method for
the detection of data issues using keys; (ii) To the best of our knowledge, we implement
the first interface for dataset repair using keys; (iii) We release a vocabulary4 for key
discovery which relies on a more correct definition of keys; (iv) We provide a RESTful
API endpoint for our algorithm ROCKER.</p>
      <p>
        Our framework implements a three-step approach composed by threshold selection,
key selection, and issue visualization and export. During the demo, we will show all of
the steps presented below:
Threshold Selection. The first step, i.e. the threshold selection, is depicted in
Figure 1a. As can be seen, users choose the dataset and the class on which the key
discovery will be carried out. A threshold value of discriminability is also required. This
parameter represents the discriminability factor, i.e. the minimum score for a set of
properties to be considered an almost-key. Previous research has shown that there is no
direct correlation among and the amount of retrieved almost-keys, nor the resources
(i.e., runtime and memory) needed for the computation [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>Key Selection. The key selection step is depicted in Figure 1c. In this step, the result of
the key discovery task is shown to the user. Each composite almost-key can be selected
to show the respective resources with issues, if any.
3 Henceforth, we will call them almost-keys to generalize for any k-almost-key.
4 Described at http://rocker.aksw.org/vocab.xhtml.
(a) First step, threshold selection.</p>
      <p>(b) Third step, issue visualization.</p>
      <p>(c) Second step, key selection.</p>
      <p>Issue Visualization and Export. In the last step, users are shown the subgraph
containing the resources having issues (red circles), their types and common object values (blue
circles). Figure 1b shows how the subgraph is rendered to the user. Finally, almost-keys
and resources with issues can be exported in a human- and machine-readable format.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Use Case</title>
      <p>In this demo, we show how our approach can be applied on one real-world and two
synthetic datasets. The first dataset was generated from DBpedia 3.95 and contains 360
instances of class dbpedia-owl:Monument and their CBD. The remaining two
datasets belong to the 2010 Instance Matching track of the Ontology Alignment
Evaluation Initiative6. The former describes persons, while the latter describes restaurants.
All datasets are available online on the respective websites.</p>
      <p>In order to understand the workflow, we now illustrate a use case in detail, which
will be shown during the demo session. Figure 1 shows the three stages applied to
5 http://oldwiki.dbpedia.org/Downloads39/
6 http://oaei.ontologymatching.org/2010/im/
our use case, which refers to the OAEI Restaurant1 dataset. As can be seen, class
oaei:Restaurant and a threshold value of = 0:99 are selected. As soon as
the user presses the Discover Keys button, ROCKER is launched via a
RESTful API call, whereas a loading screen is presented to the user. In the following step,
the user is asked to select almost-keys and issues. Since the first two almost-keys
oaei:has address and oaei:name are perfect keys (i.e., they achieve a score
of 1), they present no issues. On the other hand, the third key – composed only by
property oaei:phone number – presents one issue, which is then selected by the
user.</p>
      <p>Please note that some issues may be critical only for certain almost-keys,
depending on the specific domain. While common sense would suggest that two restaurants
having the same address or name may coexist, two restaurants having the same phone
number are unlikely to find. Therefore, resources related to some almost-keys – oaei:
phone number in our example – have more probability to contain errors than others.</p>
      <p>The subgraph displaying the issue is then rendered. In the graph, all values for each
property (in addition to rdf:type) are shown. Should the graph display no object
values, the resources are not distinguishable by the fact of sharing only null values.
Here, two restaurants (displayed in red color) share the same oaei:phone number,
therefore they are not distinguishable w.r.t. the almost-key foaei:phone numberg.
The user can thus export the data to a single file. In case the user thinks that the two
restaurants refer to the same real-world entity, then an owl:sameAs link should
subsist among them. Otherwise, one of the two phone numbers is probably incorrect.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Summary</title>
      <p>This demo paper presents a method to enhance the quality of Linked Datasets using
keys. The web interface shows how ROCKER, a state-of-the-art algorithm for key
discovery, can help finding inaccuracies in the data. Furthermore, we provide a RESTful
API endpoint for key discovery, yielding results in JSON-LD format. In future work,
we expect to open the possibility for users to upload their own datasets and we will
provide an online editor in order to speed up the data repair task.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Atencia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>David</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Scharffe</surname>
          </string-name>
          .
          <article-title>Keys and pseudo-keys detection for web datasets cleansing and interlinking</article-title>
          .
          <source>In Knowledge Engineering and Knowledge Management</source>
          .
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>D.</given-names>
            <surname>Kontokostas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Westphal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hellmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lehmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cornelissen</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Zaveri</surname>
          </string-name>
          .
          <article-title>Test-driven evaluation of linked data quality</article-title>
          .
          <source>In Proceedings of the 23rd International Conference on World Wide Web</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>T.</given-names>
            <surname>Soru</surname>
          </string-name>
          , E. Marx, and A.
          <string-name>
            <surname>-C. Ngonga</surname>
          </string-name>
          <article-title>Ngomo. ROCKER: A refinement operator for key discovery</article-title>
          .
          <source>In Proceedings of the 24th International Conference on World Wide Web</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>D.</given-names>
            <surname>Symeonidou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Armant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pernelle</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Sa</surname>
          </string-name>
          <article-title>¨ıs. SAKey: Scalable almost key discovery in RDF data</article-title>
          .
          <source>In International Semantic Web Conference</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>A.</given-names>
            <surname>Zaveri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maurino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pietrobon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lehmann</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          .
          <article-title>Quality assessment for linked data: A survey</article-title>
          .
          <source>Semantic Web Journal</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>