<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>for the Biodiversity Domain</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Samira Babalou</string-name>
          <email>samira.babalou@uni.jena.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Erik Kleinsteuber</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Badr El Haouni</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Franziska Zander</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Schellenberger Costa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jens Kattge</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Birgitta König-Ries</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Semantic Web, Knowledge Graph Platforms, Biodiversity</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>German Center for Integrative Biodiversity Research (iDiv)</institution>
          ,
          <addr-line>Halle-Jena-Leipzig</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Computer Science, Friedrich Schiller University Jena</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Max Planck Institute for Biogeochemistry</institution>
          ,
          <addr-line>Jena</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>creation process. We present iKNOW, a platform for building and managing biodiversity Knowledge Graphs (KG), currently under development. We show the architecture of iKNOW and look at the planned workflow of the KG 1. Introduction &amp; Literature Review In the biodiversity domain, the potential benefits of Knowledge Graphs (KGs) have been recognized for quite some while [1] and first graphs on specific sub topics exist. Still, uptake is disappointingly slow. Furthermore, most of the already proposed KGs in this area focus on data from natural history collections, only [2, 3, 4]. In particular, KGs leveraging the wealth of tabular data available in the biodiversity domain are still lacking. We believe, that - as in other domains - one major roadblock to wider adoption is the large efort and high semantic web expertise still needed to create and manage KGs. Addressing this problem, in our ongoing project, iKNOW [5], we aim to build a semantic-based toolbox for KG generation in the biodiversity domain. iKNOW focuses on the (semi-) automatic, reproducible transformation of tabular biodiversity data into So far, in biodiversity as in many other domains, the few existing KGs have been created largely manually in one-of eforts. While over the last few years several KG platforms have been proposed, none meets the requirements of the biodiversity domain for both generic (e.g., ingest of data in diferent formats, provenance management) and discipline-specific functionality (e.g., resolution of species names). If the potential for KGs is to be leveraged for this important ISWC-Posters-Demos-Industry 2022 (International Semantic Web Conference (ISWC) 2022: Posters, Demos, and Industry ∗Corresponding author.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Tracks)</p>
    </sec>
    <sec id="sec-2">
      <title>Access control and security checker</title>
      <sec id="sec-2-1">
        <title>Search+SPARQL</title>
      </sec>
      <sec id="sec-2-2">
        <title>Visualization</title>
      </sec>
      <sec id="sec-2-3">
        <title>KG Creation</title>
      </sec>
      <sec id="sec-2-4">
        <title>KG Update</title>
        <p>Update KG
Data Quality
Visualization
Keyword Search
Data Management
Access Control / Security
Query Catalog
Knowledge Augmentation
User Profile
Provenance Management</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data Access Infrastructure</title>
      <p>?
Tool Matcher
Data Cleaning
Data Linking
Data Authoring
RDF Generation</p>
      <sec id="sec-3-1">
        <title>User Administration</title>
      </sec>
      <sec id="sec-3-2">
        <title>Web-based UI</title>
      </sec>
      <sec id="sec-3-3">
        <title>Platform Services</title>
        <sec id="sec-3-3-1">
          <title>Non-curated</title>
        </sec>
        <sec id="sec-3-3-2">
          <title>Repository</title>
        </sec>
        <sec id="sec-3-3-3">
          <title>Curated</title>
        </sec>
        <sec id="sec-3-3-4">
          <title>Repository</title>
        </sec>
        <sec id="sec-3-3-5">
          <title>User</title>
        </sec>
        <sec id="sec-3-3-6">
          <title>Information</title>
        </sec>
        <sec id="sec-3-3-7">
          <title>Provenance</title>
        </sec>
        <sec id="sec-3-3-8">
          <title>Information</title>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>Data Storage</title>
        <p>domain, it is our conviction, that a KG management platform providing both generic and
discipline-specific functionality is needed for developing, maintaining, and using KGs. Such a
platform can reduce the barriers for non-semantic web experts to use and finally benefit from
KGs to explore new exciting findings. In this paper, we show the architecture of iKNOW along
with its planned functionalities.
2. iKNOW: The Proposed Platform
The iKNOW project is a joined efort by computer scientists and domain experts from the
German Centre for Integrative Biodiversity Research (iDiv) (www.idiv.de). The work benefits
from the wealth of well-curated data sources and expert knowledge on their creation, cleaning,
and harmonization available at iDiv. Thus, for now, iKNOW focuses on the (semi-)automatic,
reproducible transformation of tabular biodiversity data into RDF statements. It also includes
provenance tracking to ensure reproducibility and update ability. Further, options for
visualization, search, and query are planned. Once established, this platform will be open-source
and available to the biodiversity community. Thus, it can significantly contribute to making
biodiversity data widely available, easily discoverable, and integrable. In this section, we present
shortly the architecture and workflow of KG generation at iKNOW.</p>
        <p>Datasets uploading</p>
        <p>Data
Cleaning</p>
        <p>Entity
Extraction</p>
        <p>Relation
Extraction</p>
        <p>Schema
Generation</p>
        <p>Data
Authoring</p>
        <p>Triple
Generation
Legend
Optional
Mandatory</p>
        <p>Dataflow</p>
        <p>Query
Building</p>
        <p>Saving
/Pushing</p>
        <p>Quality</p>
        <p>Checking
2.1. iKNOW Architecture
2.2. Workflow in the KG Creation Scenario
Figure 2 shows the planned iKNOW workflow for the KG creation scenario. The workflow
shows the data flow between the steps towards KG generation. Not all steps are mandatory;
some optional processes in each step can add further value to the KG based on the user’s needs.
For every uploaded dataset, we build a sub-KG. It will be the subgraph of the main KG in iKNOW.
In the first step, users go through the authentication process. The verified users can upload their
datasets. If required, the Data Cleaning process will take place. We plan to ofer diferent tools
for this step, which users can select and adjust based on their needs. In the Entity Extraction step,
we map the entities of the dataset to the corresponding concepts in the real world (which build
instances of sub-KGs). This mapping is the basis for interlinking entities with external KGs like
Wikidata or domain-specific ones. Each mapped entity is a node in the KG. For this process,
we will embedded diferent tools at iKNOW, in which users can select the desired tool along
with the desired external KGs. In the Relation Extraction step, the relations between the KG’s
nodes will be extracted via the user-selected tool. Note that in the entity and relation extraction
steps, the tools return the extracted entities and relations to the user. Through our GUI, the
user can edit them (Data Authoring step). Each column from the relational dataset refers to a
category in the world. We consider the types of the column as classes in the KG. Along with the
extracted relations in the previous step, the schema of this sub-KG will be created in the Schema
Generation step. In the Triple Generation step, (subject, predicate, object)-triples based on the
extracted information from the previous steps will be created. Nodes in the KG are subjects and
objects, and relationships are predicates. The triples are generated for classes and instances in
the sub-KG.</p>
        <p>After these processes, the generated sub-KG can be used directly. However, one can take
further steps such as: Triple Augmentation (generate new triples and extra relations to ease
KG completion), Schema Refinement (refine the schema, e.g., via logical reasoning for the KG
completion and correctness), Quality Checking (check the quality of the generated sub-KG), and
Query Building (create customized SPARQL queries for the generated sub-KG). In the Pushing
step of our platform, the generated KGs are saved first at a temporal repository (shown by
“non-curated repository” in Figure 2). After a manual data curation by domain experts in the
Curation step, the KG will be published in the main repository of our platform. With this step,
we aim to increase the trust and correctness of the information on the KG.</p>
        <p>All information regarding the user-selected tools with parameters and settings along with
the initial dataset and intermediate results will be saved in every step of our platform. With
the help of this, users can redo the previous steps (which shows by arrows in both directions).
Moreover, this enables us to track the provenance of created sub-KG. In each step mentioned
above, we plan to have a tool-recommendation service to help the user select the right tool for
every process. For that, we will consider diferent parameters, such as the characteristics of the
dataset and tools.
3. Implementation
The iKNOW platform is currently under development (https://planthub.idiv.de/iknow) and is
distributed under an open-source license in github.com/fusion-jena/iKNOW. The Python web
framework Django (www.djangoproject.com) is used for the backend with a PostgreSQL (www.
postgresql.org/) database to maintain users, services, tools, datasets, and the KG generation
parameters in the iKNOW platform (used in provenance tracking). We use the compiler Svelte
(https://svelte.dev/) with SvelteKit as a framework for building web applications to create
a user-friendly web interface. For security, maintenance, and provenance reasons, all tools
from external providers used within the workflow will be executed in a sandbox using Docker
(www.docker.com/). For managing the triplestore, we are using the graph database Blazegraph
(https://blazegraph.com/). Any sub-KG created by an end-user, first, will be placed at the
noncurated triplestore. After curation by domain experts, the new sub-KG will be added to the
curated triplestore. The curated triplestore also serves as the base for SPARQL queries and the
keyword search via search engine Elasticsearch (www.elastic.co/elasticsearch/).</p>
        <p>iKNOW is a modular platform, which increases the flexibility of our platform and allows
adding new tools. Our ultimate goal is to provide a large set of tool choices for the end-user.
Although only a few tools are embedded so far, we plan to add more tools for each functionality
in the platform. Then users have a variety of choices with respect to diferent needs and use
cases. Our open-source code and modular designs of our platform make both the front and
backend of our platform easily extendable. We encourage users (new developers) to use or
extend our reusable UI components to speed up their development.</p>
        <p>Acknowledgments
The work described in this paper is conducted in the iKNOW Flexpool project of iDiv, the German
Centre for Integrative Biodiversity Research, funded by DFG (Project number 202548816). We
thank our colleague Sven Thiel for comments on the manuscript and the iKNOW PIs Helge
Bruelheide, Christine Römermann and Christian Wirth for their insights into biodiversity
research and data integration needs.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Sachs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Page</surname>
          </string-name>
          , et al.,
          <article-title>Training and hackathon on building biodiversity knowledge graphs</article-title>
          ,
          <source>Research Ideas and Outcomes</source>
          <volume>5</volume>
          (
          <year>2019</year>
          )
          <article-title>e36152</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R. D.</given-names>
            <surname>Page</surname>
          </string-name>
          ,
          <article-title>Ozymandias: a biodiversity knowledge graph</article-title>
          ,
          <source>PeerJ</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Penev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dimitrova</surname>
          </string-name>
          , et al.,
          <article-title>Openbiodiv: a knowledge graph for literature-extracted linked open data in biodiversity science</article-title>
          ,
          <source>Publications</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Stocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heger</surname>
          </string-name>
          , et al.,
          <article-title>Skg4eosc-scholarly knowledge graphs for eosc: Establishing a backbone of knowledge graphs for fair scholarly information in eosc</article-title>
          ,
          <source>Research Ideas and Outcomes</source>
          <volume>8</volume>
          (
          <year>2022</year>
          )
          <article-title>e83789</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Babalou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Schellenberger</given-names>
            <surname>Costa</surname>
          </string-name>
          , et al.,
          <article-title>Towards a semantic toolbox for reproducible knowledge graph generation in the biodiversity domain-how to make the most out of biodiversity data</article-title>
          ,
          <source>INFORMATIK</source>
          <year>2021</year>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>