<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>M. Mountantonakis);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Demonstration of L O D C h a i n : How to Tackle the Problem of Low Connectivity for your RDF Dataset</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michalis Mountantonakis</string-name>
          <email>mountant@ics.forth.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yannis Tzitzikas</string-name>
          <email>tzitzik@ics.forth.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Connectivity, Data Integration, Visualizations, Data Discovery</institution>
          ,
          <addr-line>Data Enrichment, owl:sameAs</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science, University of Crete</institution>
          ,
          <addr-line>Heraklion</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Computer Science</institution>
          ,
          <addr-line>FORTH, Heraklion</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1951</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This paper demonstrates LODChain, which is an online service for tackling the problem of low connectivity for a given RDF dataset, by strengthening its connections to the rest of LOD Cloud. The demo focuses on presenting the connectivity analytics, visualizations and enrichment services for the user/publisher, which are ofered for several parts of the dataset, e.g., for owl:sameAs mappings, entities, schema elements and triples. Moreover, we show how these analytics can be exploited for improving the connectivity of a dataset, and thus its discoverability and reusability, by using a scenario with a real RDF dataset.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Current approaches for dataset search and discovery are mainly metadata-based and ignore
the content of the datasets [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], and existing RDF approaches do not favor their discoverability
and reusability [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. As a result, “Linked Data Cloud consists of loosely inter-linked individual
subgraphs” [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For tackling the problem of low connectivity, we present L O D C h a i n ; an online
web service, which is available at https://demos.isl.ics.forth.gr/LODChain, that strengthens
the connectivity of an RDF dataset, by computing the transitive and symmetric closure of
equivalence relationships, such as owl:sameAs, between the given dataset and hundreds of RDF
datasets indexed by LODsyndesis [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Moreover, it finds common entities, schema elements and
triples, and produces connectivity analytics, visualizations and enrichment services. The target
is to aid the publishers a) to evaluate the connectivity of their dataset, e.g., before publishing it
to the LOD Cloud, b) to strengthen its connectivity, i.e, by discovering new connections, and c)
to verify and enrich its content by finding common and complementary data.
      </p>
      <p>
        This demo paper is a complement to an accepted paper of ISWC’22 Resource Track [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The
accepted resource paper presents all the methods and algorithms for computing the connectivity
analytics of L O D C h a i n , and an evaluation with real datasets. On the contrary, this demo paper
focuses on presenting in more details the ofered analytics, visualizations, and data enrichment
services, by providing a scenario of how the services can be used for a real dataset.
https://users.ics.forth.gr/~mountant/ (M. Mountantonakis); https://users.ics.forth.gr/~tzitzik/ (Y. Tzitzikas)
      </p>
      <p>© 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>The rest of this demo paper is organized as follows: §2 introduces the related work, §3
describes in brief the process of L O D C h a i n , §4 presents all its services, and §5 concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        There are several approaches for aiding the interlinking, discoverability and reusability of RDF
datasets. Concerning interlinking, applications like WIMU [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], MetaLink [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and LODsyndesis
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] can be used for finding all the datasets and equivalent URIs of a given entity. Regarding
discoverability and reusability, services such as LOD Cloud (http://lod-cloud.net), Google Dataset
Search [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], LODatio [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and LODAtlas [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] ofer metadata-based search for discovering and
reusing RDF datasets. Comparing to these approaches, L O D C h a i n is the first online service
for strengthening the connectivity of an RDF dataset (of any domain), even before its actual
publishing, by exploiting the content of hundreds of RDF datasets.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. The Process and Architecture of L O D C h a i n</title>
    </sec>
    <sec id="sec-4">
      <title>4. Connectivity Analytics, Visualizations &amp; Enrichment</title>
      <p>
        We present the ofered services, divided in six categories by using a real dataset from publications
domain, i.e., WW1LOD [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], including 47,616 triples, 547 owl:sameAs mappings and connections
with 5 RDF datasets. The target is to explain how the results of each category can be used for
evaluating/improving the connectivity of a given dataset. For more details, a video is available
in https://youtu.be/Kh9751p32tM, and screenshots including analytics for more datasets are
presented in https://zenodo.org/record/6467419. In the demonstration, we will showcase the
ofered analytics and visualizations, by showing results over such real datasets.
      </p>
      <p>Category 1. Analytics over owl:sameAs Mappings. The objective is to evaluate for the
given dataset if L O D C h a i n a) inferred new mappings through the computation of owl:sameAs
closure, and b) detected possible owl:sameAs errors, i.e., L O D C h a i n checks if an entity is equivalent
with two or more real entities in LODsyndesis. L O D C h a i n ofers a chart analyzing the results
of the computation of closure, including the number of a) possible errors, and b) owl:sameAs
relationships before and after the computation of closure. In case of detecting errors, they can
be downloaded from the user for aiding the process of fixing them. As Fig. 2(a) shows, we
inferred thousands of owl:sameAs mappings for WW1LOD, without detecting errors, which is
very positive for its connectivity, i.e., we expect that more connections have been discovered.</p>
      <p>Category 2. Analytics over Entities. The objective is to evaluate how connected is the
given dataset to the rest of LOD Cloud, and to quantify the gain of the computation of owl:sameAs
closure. L O D C h a i n provides charts for the number of unique/common entities, the connections
before and after the closure, the top-10 connected datasets, and its connectivity compared to
datasets of several domains. For instance, Fig. 2(c-d) shows that WW1LOD shares 825 entities
with other datasets (such as DBpedia and Wikidata), whereas 25 new inferred connections were
discovered through L O D C h a i n . Moreover, a graph visualizes all the connections of the dataset
(Fig. 2(f)), where the red nodes depict the old connections and the green nodes the new ones
(due to closure), and the labels of edges the number of common entities. From these charts, we
can see that the connectivity of WW1LOD was highly increased. On the contrary, in a scenario
where a dataset has either zero or few connections, even after the computation of closure, it
could be an indication of bad/low connectivity, e.g., few or/and incorrect links to other datasets.
Finally, charts like Fig. 2(e) can be used for quantifying the connectivity of a dataset comparing
to other domains, e.g., WW1LOD is connected with datasets from 5 diferent domains.</p>
      <p>
        Category 3. Analytics over Schema Elements. The goal is to evaluate if existing
ontologies are reused, since it is important for a) comparing the values of the same fact for an
entity [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], e.g., for data verification and/or for detecting conflicts, and b) for enabling
schemabased integration, e.g., for creating a mediator or a data warehouse [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. L O D C h a i n ofers charts
showing the number of common/unique properties or classes, and the top-10 most connected
datasets for these URIs. For example, WW1LOD shares many common properties (Fig. 2(g)),
which can be important for schema-based integration. On the contrary, if a dataset has a low
number of common properties/classes, existing schemas are not reused and the creation of
further mappings will be probably needed, e.g., for achieving schema-based integration.
      </p>
      <p>Category 4. Analytics over Triples. The objective is to evaluate whether a dataset shares
common triples (or facts) with others, and if there are available complementary facts for its
entities. L O D C h a i n ofers charts showing the number of common/unique facts, and the top-10
datasets ofering the most common and complementary facts for the entities of the given
dataset. Fig. 2(i) shows that WW1LOD contains 368 common facts with at least one RDF dataset.
Certainly, it is not required for a dataset to contain common facts with other ones for being
included in the LOD Cloud, however, they can be of primary importance for data verification.
On the contrary, the existence of complementary facts can be useful for enriching the content
of the given dataset, e.g., we found 362,339 complementary facts for the entities of WW1LOD.</p>
      <p>
        Category 5. Dataset Discovery and Selection. In many cases, the publishers desire to
select and integrate their dataset with only K datasets, i.e., since it can be time-consuming to
export and to integrate with any available relevant dataset [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. L O D C h a i n ofers measurements
for finding the K most relevant datasets for a given dataset according to the number of common
entities, properties, classes and facts, and of complementary facts. Fig. 2(i) shows the top-3
datasets according to the number of a) common and b) complementary facts for our running
example. The most relevant datasets difer in these two cases, which shows that for diferent
needs (verification versus data enrichment), diferent combinations of datasets can be used.
      </p>
      <p>Category 6. Data Enrichment. L O D C h a i n gives the option to the publisher to export the
results of the connectivity analytics and to enrich the given dataset with additional data for
creating an enriched version of their dataset. In particular, one can export in RDF format all the
common entities, schema elements and facts (and their provenance), the inferred owl:sameAs
relationships, complementary facts, and rich metadata including the results of the analytics. All
these data (or any subset of them) can be used from a publisher for improving his/her dataset.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Concluding Remarks</title>
      <p>We presented the connectivity analytics, visualizations and enrichment services, ofered by
L O D C h a i n ; a web application for strengthening the connectivity of an RDF dataset over hundreds
of datasets. As a future work, we plan to extend L O D C h a i n for ofering more services, e.g., finding
connections for almost or totally disconnected datasets (by performing instance matching), and
services for further aiding the process of fixing connectivity errors, e.g., owl:sameAs errors.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments References</title>
      <p>This work has received funding from the European Union’s Horizon 2020 coordination and
support action 4CH (Grant agreement No 101004468).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chapman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Simperl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Koesten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Konstantinidis</surname>
          </string-name>
          , L.
          <string-name>
            <surname>-D. Ibáñez</surname>
            , E. Kacprzak,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Groth</surname>
          </string-name>
          ,
          <article-title>Dataset search: a survey</article-title>
          ,
          <source>The VLDB Journal</source>
          <volume>29</volume>
          (
          <year>2020</year>
          )
          <fpage>251</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Debattista</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Attard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brennan</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>O'Sullivan, Is the LOD Cloud at risk of becoming a museum for datasets? looking ahead towards a fully collaborative and sustainable lod cloud</article-title>
          ,
          <source>in: Proceedings of WWW Conference</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>850</fpage>
          -
          <lpage>858</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Hitzler</surname>
          </string-name>
          ,
          <article-title>A review of the semantic web field</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>64</volume>
          (
          <year>2021</year>
          )
          <fpage>76</fpage>
          -
          <lpage>83</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mountantonakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tzitzikas</surname>
          </string-name>
          ,
          <article-title>Content-based union and complement metrics for dataset search over RDF knowledge graphs</article-title>
          ,
          <source>ACM JDIQ 12</source>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mountantonakis</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Tzitzikas,</surname>
          </string-name>
          <article-title>LODChain: Strengthen the connectivity of your RDF dataset to the rest LOD Cloud, in: ISWC 2022 (Accepted in Resource Track</article-title>
          ),
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Valdestilhas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Soru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nentwig</surname>
          </string-name>
          , E. Marx,
          <string-name>
            <given-names>M.</given-names>
            <surname>Saleem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-C. N.</given-names>
            <surname>Ngomo</surname>
          </string-name>
          ,
          <article-title>Where is my URI?</article-title>
          , in: European Semantic Web Conference, Springer,
          <year>2018</year>
          , pp.
          <fpage>671</fpage>
          -
          <lpage>681</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>W.</given-names>
            <surname>Beek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Raad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Acar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. v.</given-names>
            <surname>Harmelen</surname>
          </string-name>
          ,
          <article-title>Metalink: A travel guide to the LOD cloud</article-title>
          ,
          <source>in: European Semantic Web Conference</source>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>481</fpage>
          -
          <lpage>496</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Gottron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Scherp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Krayer</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Peters,
          <article-title>LODatio: A schema-based retrieval system for linked open data at web-scale</article-title>
          ,
          <source>in: ESWC</source>
          , Springer,
          <year>2013</year>
          , pp.
          <fpage>142</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Pietriga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gözükan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Appert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Destandau</surname>
          </string-name>
          , Š. Čebirić,
          <string-name>
            <given-names>F.</given-names>
            <surname>Goasdoué</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Manolescu</surname>
          </string-name>
          ,
          <article-title>Browsing linked data catalogs with LODAtlas</article-title>
          , in: ISWC, Springer,
          <year>2018</year>
          , pp.
          <fpage>137</fpage>
          -
          <lpage>153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Mäkelä</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Törnroos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lindquist</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Hyvönen,</surname>
          </string-name>
          <article-title>WW1LOD: an application of CIDOC-CRM to World War 1 linked data</article-title>
          ,
          <source>IJDL</source>
          <volume>18</volume>
          (
          <year>2017</year>
          )
          <fpage>333</fpage>
          -
          <lpage>343</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mountantonakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tzitzikas</surname>
          </string-name>
          ,
          <article-title>Large-scale semantic integration of linked data: A survey</article-title>
          ,
          <source>ACM CSUR 52</source>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>