<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Strati ed Data Integration</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>usto Giun</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>higli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ssio Z</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>yukh B</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simon</string-name>
          <email>simone.boccag@unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering and Computer Science (DISI), University of Trento</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We propose a novel approach to the problem of semantic heterogeneity where data are organized into a set of strati ed and independent representation layers, namely: conceptual (where a set of unique alinguistic identi ers are connected inside a graph codifying their meaning), language (where sets of synonyms, possibly from multiple languages, annotate concepts), knowledge (in the form of a graph where nodes are entity types and links are properties), and data (in the form of a graph of entities populating the previous knowledge graph). This allows us to state the problem of semantic heterogeneity as a problem of Representation Diversity where the di erent types of heterogeneity, viz. Conceptual, Language, Knowledge, and Data, are uniformly dealt within each single layer, independently from the others. In this paper we describe the proposed strati ed representation of data and the process by which data are rst transformed into the target representation, then suitably integrated and then, nally, presented to the user in her preferred format. The proposed framework has been evaluated in various pilot case studies and in a number of industrial data integration problems.</p>
      </abstract>
      <kwd-group>
        <kwd>Semantic Heterogeneity</kwd>
        <kwd>Knowledge Graph Construction</kwd>
        <kwd>Strati ed Data Integration</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Semantic Heterogeneity [
        <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
        ], namely the existence of variance in the
representation of the same real-world phenomenon, has long been a major impediment
for e ective large scale data integration implementations. It is inescapable in
nature and is rooted in the diversity which is inherent in di erent means of
representation (which itself is rooted in world diversity) [
        <xref ref-type="bibr" rid="ref3 ref4">3,4</xref>
        ]. The pervasiveness of
semantic heterogeneity in its various forms, has long been an overarching
concern in the data management research landscape [
        <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
        ]. In particular, its obtrusive
rami cations with reference to data integration scenarios have been widely
acknowledged, and partial approaches towards their resolution at the schema and
data level have been proposed, see for instance [
        <xref ref-type="bibr" rid="ref1 ref5 ref6 ref7">1,5,6,7</xref>
        ].
      </p>
      <p>Copyright c 2021 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>
        In this paper we propose a novel approach which is based on the key
assumption that data should be represented as a set of strati ed and independent
representation layers, namely:
1. conceptual, where concepts, codi ed as a set of unique alinguistic identi ers,
are connected inside a graph codifying their meaning;
2. language, where sets of synonyms, i.e., synsets [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], possibly from multiple
languages, annotate concepts;
3. knowledge, in the form of a graph where nodes are entity types and links are
properties; and
4. data, in the form of a graph of entities populating the previous knowledge
graph.
      </p>
      <p>The representation of data according to these four layers allows us to state
the problem of semantic heterogeneity as a problem of Representation Diversity,
where heterogeneity distributes itself over these layers thus being strati ed into
four di erent types of diversity, viz. Conceptual, Language, Knowledge, and Data,
which can then be are uniformly dealt within each a single layer, independently
from the others. The proposed approach has three major advantages. The rst
is that the combinatorial explosion deriving from the interaction of the four
di erent types of diversity is avoided and the complexity of the data integration
problem reduces to the sum of the complexity of each layer. Each layer can be
dealt with as if all the other layers presented no heterogeneity at all. The second
is that the techniques developed for each layer can be composed with the ones
developed in the other layers irrespectively of how heterogeneity appears in the
current problem. The third, which is a direct consequence of the second, is that,
within each layer, it is possible to exploit the large body of work which can be
found in the literature (see the related work section for a detailed discussion on
this point).</p>
      <p>In this paper, after restating the semantic heterogeneity problem as a
representation diversity problem, we describe the proposed strati ed representation
of data and the process by which data are rst transformed into the target
representation, then suitably integrated and then, nally, presented to the user in
her preferred format. The main contributions of this paper can be articulated as
follows:1. a novel articulation of the problem of semantic heterogeneity as strati ed
into the above four layers of representation diversity (Section 2);
2. a viable solution to the issues above in the form of an end-to-end logical data
management architecture grounded in our four layered strati cation of
representation diversity, wherein we focus on accommodating the heterogeneity
resident in each layer, independently from all the other layers (Section 3);
3. an illustration of an implemented semi-automated data integration pipeline
which exploits the representation of diversity presented in Section 3 (Section
4).</p>
      <p>The pipeline presented in Section 4 has been extensively evaluated within many
data integration pilot studies, which have spanned three years and has then
applied in various real world problems. Section 5 presents a short highlight of these
activities. Towards the end of the paper, Section 6 contextualizes the
contribution of our work by surveying the state of the art, while Section 7 outlines the
future research ventures.
There are three datasets fCar, Vettura, Vehicleg which refer to the same real
world entity, namely a car, which we assume has plate `FP372MK'. The rst
dataset describes FP372MK as a car, having ve attributes, four of which are
expressed using the automotive extension of schema.org.1 The second dataset also
considers the entity as a car but its description is provided in Italian. The third
dataset encodes the entity as a vehicle having four attributes expressed employing
the Vehicle Sales Ontology2 namespace `vso:'.</p>
      <p>By a close look at the above example, it is easy to notice four di erent types of
diversity which we can brie y describe as follows:
1 https://schema.org/docs/automotive.html
2 http://www.heppnetz.de/ontologies/vso/ns
{ Conceptual diversity (L1):3 the same real world object is mentioned in two
datasets using the concept denoted by the word car while in the third data
set is called vehicle, namely using a more general term (because, for instance,
in this latter case there is no need to distinguish among the various types of
vehicles as the issue is that of counting the number of free parking lots).
{ Language diversity (L2): the same real object is described using three
different lexicons, namely, that of a natural language, i.e., Italian, and two
namespaces, i.e., from the automobile extension of schema.org and from the
Vehicle Sales Ontology, both of which use a di erent natural language, i.e.,
English, as base language. Notice that these are three di erent lexicons,
independently developed where, therefore, the meaning of the terms used are
intuitively similar but formally unrelated.
{ Knowledge diversity (L4): the same real world object is described using
different properties, the motivation being most likely in the di erent focus of
the three databases. Thus, for instance, the rst could be the description
used in an online car rental which codi es its data using schema.org, the
second could be the description used by the Italian Automobile Club, while
the third could be the description used in an online sales portal.
{ Data diversity (L5): the same real world object is described in a way that,
even when associated with the same properties, the corresponding values are
di erent. There can be many reasons for this last source of heterogeneity,
for instance, di erent approximations, di erent formats, di erent units of
measure, di erent reference standards (e.g., date standards for dates) and
so on.</p>
      <p>Let us consider these four layers in detail.</p>
      <p>
        Conceptual Diversity. The notion of concept is well known in the Philosophy
of Language literature, see, e.g. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and in Computational Linguistics, see, e.g.
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].. Here we follow our own work, as described in [
        <xref ref-type="bibr" rid="ref11 ref12">11,12</xref>
        ], and take concepts to be
unique alinguistic identi ers. Concepts are organized in multiple hierarchies, one
per syntactic type (i.e., noun, verb, adjective) wherein a child (father) concept
is taken to be more specialized (more general) than the father (child) [
        <xref ref-type="bibr" rid="ref12 ref8">8,12</xref>
        ]. For
instance noun hierarchies are organized in terms of `hypernym-hyponym' links
where, in the example in Table 1, the term car is a direct hyponym of the term
vehicle.
      </p>
      <p>
        Language Diversity. Languages, taken here in a very broad sense to include,
e.g., natural languages, namespaces and formal languages. Language diversity
occurs both across and within languages. Thus, there are multiple languages
available for the purpose of representing the same concept, but also, even within
the same language, linguistic phenomena like polysemy and synonymy allow for
multiple diverse representations of entities. As a result there is a many-to-many
3 L1, L2, L3, L4, L5 are the ve layers into which we organize the representation of
data. L3, the layer used to represent causality [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], is not discussed here because it is
irrelevant to the goals of the paper.
mapping between words and concepts, both within the same language and across
languages [
        <xref ref-type="bibr" rid="ref11 ref12">11,12</xref>
        ].
      </p>
      <p>
        Knowledge Diversity. We model knowledge as a set of entity types, also called
etypes, meaning by this, classes of entities with associated properties. Knowledge
diversity arises from the many-to-many mapping between etypes and the
properties employed to describe them [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], and can appear in one of two di erent forms.
The rst appears when there are `n' representations of di erent etypes described
in terms of the same set of properties. The second manifests itself when there
are `n' representations of the same etype with di erent sets of properties. As an
example of the latter situation, in Table 1 two datasets describe the same etype
car, but the two etypes are associated to three di erent groups of properties.
Data Diversity. We model data, meaning by this the concrete, ground knowledge,
that we have about objects in the world, as entities each associated with property
values, where properties are inherited from the etype of the entity. Data diversity
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] exists because of the fact that the mapping between entities and the property
values used to describe them is many-to-many. Data diversity appears as well in
two di erent forms, wherein the same real world entity, associated with the same
properties, is described using di erent data values, while dually, two di erent
real world entities, still associated with the same properties, can be described
using the same data values. As an example of the latter situation, there can be
two identical cars which are both described by a set of attributes which do not
contain their plate or any other identifying attribute. The example in Table 1
provides an example of the former situation. Here, the three datasets refer to
the same entity, a car with plate `FP372MK', which shares a common attribute
which is the car speed, but this property has three di erent values.
      </p>
      <p>Notice how the strati cation of diversity presented above has the following
crucial characteristics. The rst is that each layer models a di erent type of
phenomenon and the corresponding type of diversity. The second is that how
diversity appears in one layer is completely independent of how it appears in any
other layer. The third is that L1, the conceptual layer, provides the grounding
for unifying the various types of diversity as they appear in the other layers. In
fact, in all respects, L1 is a logical theory, which can be codi ed in a logical
language, e.g., Description Logics, where the semantics of all terms (the alinguistic
identi ers) is univocally de ned by the links of the hierarchy. The fourth and
last observation is that the diversity mappings which appear in L2, L4 and L5
are all many-to-many and this generates the type of combinatorial complexity
which makes it so di cult to handle the problem of semantic heterogeneity. As
already hinted to in the introduction, the strati cation of semantic heterogeneity
provides a major advantage in that it allows to structure it in four independent
and much smaller problems, where each problem can be treated uniformly by
developing methods and techniques which are specialized just for that layer.
This last observation is the basis for the work presented in the remaining of this
paper.
Fig. 1 depicts the proposed data management architecture instantiated (partially,
for lack of space) to the example in Table 1. Here the arrows represent the
functional dependencies which must be enforced among the di erent layers and,
therefore, implicitly de ne the order of execution which must be followed during
a data integration task, starting from the user input and concluding with the
fully integrated data. Fig.3 in Section 4 will later depict the process, tools and
algorithms which exploit the di erent representation layers in Fig.1 towards
producing, in the end, the target data integration.</p>
      <p>The language representation layer (L2) appears rst and last in the
architecture in Fig.1. L2 enforces the input and the output dependence of the
representation of data on the user language. In fact, language is the key enabler of the
bidirectional interaction between users and the platform. In the rst phase, the
L2 input language is translated into the system internal L1 conceptual language
and the input language is only resumed during the last step, when the results
of the data integration steps are presented back to the user. To this extent,
notice how, in Fig.1 the language used in the LEG (L1), ETG (L4) and EG (L5)
are just ids, while the conversion table of the rst phase in the repository is
the mapping from internal and external language(s). In this process, L2 is key
in keeping completely distinct the multilingual user-de ned data representation
and the alinguistic system-level data representation. It is also important to
notice how the proposed data management architecture is natively multilingual
as a result of the L1 alinguistic concepts being the convergence of semantically
equivalent words in di erent languages. A very important case which can be
dealt by this architecture is the heterogeneity of namespaces, as also re ected in
the running example. Any number of namespaces and natural languages can in
fact be seamlessly integrated following the same uniform process.</p>
      <p>
        The management of conceptual diversity (L1), which functionally comes next
in sequence, involves the organization of the L1 alinguistic concepts, as
identied in the rst step, into a Language Entity Graph (LEG) which codi es the
semantic relations across concepts (and, therefore, among, the corresponding L2
input words). In order to achieve this goal we exploit, as a-priori knowledge, a
multilingual lexico-semantic resource, called Universal Knowledge Core (UKC)
[
        <xref ref-type="bibr" rid="ref11 ref12">11,12</xref>
        ] which represents words, synonyms, hyponyms and hypernyms quite
similarly to the Princeton Wordnet [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], still with important di erences [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The net
result of this phase is an LEG with the following properties:
{ the concepts identi ed during the rst phase are all and only the nodes in
this graph;
{ these nodes are annotated with the input L2 terms, across languages;
{ these nodes are organized into a hierarchy which preserves the ordering,
across the links of the UKC (in the case of nouns, the synonym/ hyponym/
hypernym relations).
      </p>
      <p>The third representation layer (L4), dedicated to the management of
knowledge diversity, involves the construction of a (alinguistic) Entity Type Graph
(ETG) encoded using only concepts occurring in the LEG constructed during
the previous two phases. In this phase, the rst step is to distinguish concepts
into etypes and properties (both object properties and datatype properties) while
the second step is to organize them into a subsumption hierarchy. The key
observation here, which constitutes a major departure from the previous work is
that etype subsumption, as encoded in the ETG, is enforced to be coherent with
the concept hierarchy encoded in the LEG. Thus for instance, as from the above
example, the object with plate FP372MK, being a car, can be encoded to be
an entity of etype vehicle, but not of, e.g., etype organism. This fact, which
is natively enforced by the lexico-semantic hierarchy of the UKC for what
concerns natural languages (in the above example, Italian), is extended to cover also
the terms belonging to namespaces (in the above example that of schema.org
and that of the Vehicle Sales Ontology). This alignment of meanings across
languages and namespaces, which absorbs a major source of heterogeneity present
in the (Semantic) Web is a natural consequence of the language and conceptual
alignment performed during the rst two phases.</p>
      <p>
        In the fourth representation layer (L5), we tackle the heterogeneity of data
values by employing an Entity Graph (EG), namely, a data-level knowledge graph
populating the ETG with the entities extracted from the input datasets. Fig.2
reports the EG resulting from the rst four phases. As you can notice this graph is
constituted of a backbone of L1 alinguistic ids, each annotated with the input L2
terms where, for each L2 term, the system remembers the dataset it comes from.
This information is crucial in case of iterated (multi-phased) data integration,
as it is usually the case, as the system needs to remember which new terms and
values substitute which old terms and values. This mechanism is implemented via
a provenance mechanism, not represented in Fig.1, which applies to all the input
dataset elements, both at the schema and at the data level. A last observation is
that in Fig.1 the unique id identifying the car with plate FP72MK is #589625,
a new identi er which never appeared before. In fact, any time a new entity is
identi ed, it is associated a unique id which is managed internally by the system
and that the user can see and also query, but not modify.
The pipeline in Fig.3 describes the process used to manage and integrate the
diversity as it appears in the four layers described in the previous section. This
partially automated pipeline is highly exible, independent of the input data
and domain, and largely customizable. As detailed in the next section, it has
been applied to the modeling of events, facilities, personal data, medical data,
university data [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and many other domains. Its intended target users are Data
Scientists with no programming knowledge but with an understanding of the
domain and data integration problem to be solved. All tasks are managed via
a exible user-friendly user interface and the process is assumed to start with
the input datasets being in a repository from where they can be uploaded. The
only required programming e ort, which we assume should not be performed by
the data scientist, is that needed to extract the input datasets from, e.g., legacy
systems or open data repositories, and to pre-process them.
      </p>
      <p>
        The pipeline is composed of four main phases, as depicted in Fig.3, where
the rst phase produces the integrated L1 and L2 representation of the LEG,
the second phase produces the ETG and the third phase produces the EG. The
fourth phase, the phase of EG Presentation, depicted in Fig.3 for completeness,
is representative of an external client application exploiting the LEG, ETG and
EG produced by the pipeline. During this phase, not further discussed, the data
scientist usually selects the language in which she wants the EG to be presented
to her, e.g., using the words of the input datasets or any other language supported
by the UKC.4
LEG Construction: This rst activity takes in input all data to be integrated (in
the example above, this consists of all the three tables represented in Table 1),
and it extracts each word and multi-word occurring in the input tables (in the
case of namespaces, the namespace itself and the word are treated as a single
concept). The output of this phase is the L1 and L2 representation of the input
data organized in a LEG. During this phase the prior knowledge codi ed in the
UKC is heavily exploited (see the previous section for details). This activity is
performed with the help of the Word Sense Disambiguation (WSD) component
SCROLL, a multilingual NLP library and pipeline which is specialised to handle
the Language of Data, as de ned in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], namely the type of NLP sentences that
are usually found in data. At the moment SCROLL supports seven languages
(including various European languages but also Mongolian and Chinese) but, as
we have found out, because of the similarity of many languages, of the relatively
simple structure of the language of data, and of the fact that the system
processing is in control of the Data Scientist which validates each step, SCROLL can
also be useful in various other similar languages (where similar here means not
diverse, with language diversity being de ned as in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]). The main features of
SCROLL which make it quite suitable for multilingual data integration are:
{ it has been developed to be highly modular and with a clear split between
the modules which are language independent from those which are language
dependent;
{ in SCROLL, the tasks that are language dependent and must therefore be
implemented for each new language (e.g., word segmentation in English and
Italian is very di erent from word segmentation in Chinese) are executed
4 In the current state of implementation, this phase can only perform a word by word
translation without being able to reconstruct the overall meaning of a sequence of
words.
      </p>
      <p>as soon as possible. The net advantage is that the more semantics
dependent tasks, e.g., Entity Recognition (ER) and Word Sense Disambiguation
(WSD) work on a conceptual representation of the data and are therefore
implemented once for all;
{ SCROLL's NLP pipeline is highly optimized and fully integrated with the
UKC, mainly with the goal of implementing a highly optimized and highly
e ective language agnostic WSD task which is also domain aware;
{ it is often the case SCROLL encounters new words which are not in the UKC
which, in turn may or may not contain the corresponding concept id. These
situations are dealt with by suitably enriching the UKC according to a
dedicated mechanism.</p>
      <p>The output of this phase is a spreadsheet, including the structured de nitions
of new concepts and their relations, which, suitably interpreted on the basis of
the UKC hierarchy, codi es the LEG.</p>
      <p>
        ETG Construction: This activity takes in input the schemas of the input datasets,
where all the words are now annotated with the LEG concepts and it constructs
the ETG which integrates them. This phase is performed via a Knowledge
Editor, similar in spirit to Protege,5 but highly integrated in the pipeline in Fig.3,
which is used interactively by the data scientist. Two are the main operations
which are performed during this phase with the help of the knowledge editor:
{ perform schema matching. This activity is performed manually but it
bene ts from the suggestions provided by a multilingual schema matcher [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]
(where, however, an e ective way to integrate this matcher in the knowledge
editor is yet to be found);
{ build the resulting integrated ETG and possibly align it with a reference
ontology. The goal of this step is to produce a clean and highly reusable
ETG. This step at the moment is performed manually, but the plan is to
integrate this functionality inside the Knowledge Editor, exploiting the results
described in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>The output of this phase is an OWL le codifying the ETG where all the terms
which are used are annotated with the LEG's concept ids.</p>
      <p>
        EG Construction: This activity takes the ETG and the input datasets,
annotated with the LEG concepts, and constructs a mapping between the data values
within each dataset and the ETG built during the previous step. The EG
Construction is an iterative activity which considers one dataset at the time. The
mapping operations are implemented through the usage of a speci c tool, called
KarmaLinker, which consists of the integration of the Karma data integration
tool [
        <xref ref-type="bibr" rid="ref5 ref6">5,6</xref>
        ], which does not do any NLP, with SCROLL. Within each activity
iteration, KarmaLinker maps the data values in the input dataset to the etypes
and properties of the ETG. Some of the most important and non trivial
operations performed here are that of recognizing the entities which are implicitly
5 https://protege.stanford.edu/
mentioned in the input datasets (entity detection), of recognizing their entity
types (etype recognition) and, nally, of recognizing whether there are multiple
occurrences of the same real world entities, possibly described using di erent
properties and property values (entity mapping, see the example of data
heterogeneity presented in Section 2). In the example of Fig.1, all the three input
datasets are recognized to describe the same car which, as from Fig.2, is then
assigned the unique id #589625. The nal output is an EG encoding the
information in the initial datasets, at all the four di erent levels of diversity, stored using
a language agnostic JSON-LD 6 format. See Fig.2 for a partial representation of
the EG constructed from the datasets in Table.1.
5
      </p>
    </sec>
    <sec id="sec-2">
      <title>Evaluation</title>
      <p>The representation architecture and pipeline described above, have been
experimented and evaluated during the past three years (2018, 2019, 2020), the last
year still being ongoing, as part of the Knowledge and Data Integration (KDI)
class, a six credit course of the Master Degree in Computer Science of the
University of Trento.7 Table 2 reports the information regarding the population
involved (excluding PhD students which are not counted) and number of projects.
During this class students, 2-5 people per group, must generate an EG using the
pipeline above starting from a high level problem speci cation. Part of the task
is also to identify the most suitable datasets and pre-process them. Datasets are
usually found in open data repositories but some of them are also scraped from
the Web. The overall project has an elapsed time of 14 weeks during which
students have to work intensely, even not full time. We estimate the overall e ort
that each group puts into building an EG in around 4-8 man-months, depending
on the case. At the end of the course, after the nal exam, students are asked to
evaluate the methodology they have used (as partially described in this paper).
This evaluation involves various aspects including application scenarios, datasets
used, ETG and EG generation, language management and LEG generation, and
evaluation of the overall pipeline. As of to day we have piloted 24 projects and
75 evaluations.</p>
      <p>A rst question in the evaluation is about L2 and the management of language
diversity, with a speci c emphasis on the use of the UKC. The percentage of
students which believe that the explicit management of language diversity, as a
dedicated independent phase, is worthwhile was 79.4% in 2018 and 80% in 2019.
A second question is about L4 and the management of knowledge diversity. More
speci cally, here the purpose of the question is to understand if students nd
it helpful to de ne the etypes of the input data, and if the pipeline properly
supports this task. The 69% of the students in 2018, and 95% in 2019, provided
6 https://json-ld.org/
7 See https://unitn-kdi-2020.github.io/unitn-kdi-2020/ for more details. This site
contains the material used during the 2020 edition of the course and it consists of
theoretical and practical lectures, as well as demos of the tools to be used, some of which
have been mentioned above.
a positive answer. Moreover the 72.4% (2018), and 60% (2019), stated that
grounding the types with the reference ontology, usually an upper ontology,
facilitates the construction of the ETG and in particular the positioning of the
entity types. A last set of questions are asked about the overall usability of the
methodology, where each question aims at the evaluation of a speci c usability
aspect. The answers are reported in Fig.4 for both 2018 and 2019, over a 1 to 7
scale (1 maximally positive, 7 maximally negative). Observing the gures below
we can notice an overall positive trend. Furthermore we can see that, with respect
to 2018, in 2019 students found the methodology easier to use and also easier
to learn. Moreover, we noticed that the level of e ciency in the accomplishment
of the data integration objectives increased in 2019. This is part of an overall
positive trend that, we believe, will also be con rmed in 2020 and that, we
believe, relates to the continuous adjustments to the methodology and to the
tools that we perform every year, also based on the feedback collected during
the KDI course.
As from the introduction, our approach to representing and managing
semantic heterogeneity as the strati cation of four independent problems, was never
proposed before. However, the very same fact that we are stratifying semantic
heterogeneity into the four problems of conceptual diversity, language diversity,
knowledge diversity and data diversity, one at the time, allows us to refer and
exploit the huge amount of work which has been independently done in these
areas. The following of this section concentrates on this work, across the four
layers, including also earlier work from the authors.</p>
      <p>
        The notion of language diversity, as described here, was rst introduced in
[
        <xref ref-type="bibr" rid="ref11 ref12">11,12</xref>
        ] which also provide a detailed description of the UKC. However, this
work builds upon decades of the work in the development of multilingual lexical
resources, see, e.g. [
        <xref ref-type="bibr" rid="ref18 ref8">8,18</xref>
        ]. The main innovation with respect with this earlier work
is that LEGs, and the UKC in particular, have a separated and independent
conceptual layer while, in the previous lexico-semantic resources, L1 and L2 are
collapsed. The strati cation between L1 and L2 on one side and L4 on the other
side, and the exploitation of lexical resources in order to do schema integration is
a direct application of the ideas and technology developed in the eld of ontology
matching, see, e.g., [
        <xref ref-type="bibr" rid="ref19 ref20">19,20</xref>
        ]. The work in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], used during the ETG construction,
builds upon this previous work by proposing NuSM - a multilingual ontology
matching framework which heavily exploits the UKC.
      </p>
      <p>
        Our proposal of using knowledge graphs is very much in sync with the huge
amount of work now being developed in this area [
        <xref ref-type="bibr" rid="ref21 ref22">21,22</xref>
        ]. Di erently from the
previous work, we uniquely stratify Knowledge Graphs into four layers,
however, see [
        <xref ref-type="bibr" rid="ref23 ref24 ref25">23,24,25</xref>
        ] for an approach which is quite similar in spirit to ours, in
particular in the distinction between L4 and L5. Furthermore, the validity of
an ontology guided, knowledge graph backed approach towards data integration
and presentation has been favourably discussed in [
        <xref ref-type="bibr" rid="ref26 ref27">26,27</xref>
        ].
      </p>
      <p>
        In the context of semantic data integration [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], the survey in [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] noted the
prevailing di culties as non-standardized identity management, multilinguality
management, data and schema heterogeneity, namely all issues which our work
addresses. The work in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] also mentions architectural, structural, syntactic and
semantic heterogeneity in data integration frameworks, all issues that our
proposed approach tackles. The work in [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] and [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] combined, highlights diverse
parametric aspects of six major openly available knowledge graphs, with [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]
calling for newer approaches in knowledge modeling and new forms of knowledge
graphs. More speci cally, Wikidata [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] emerges as a feature-rich, cross-domain
openly available knowledge graph [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. Still, due to its very nature, with respect
to the work proposed here, the work on Wikidata lacks an adaptive schema
customizable to di erent data integration scenarios, and an overall explicit,
stratied data management architecture.
7
      </p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>In this paper we have presented an innovative organization of data management
strati ed across four layers of heterogeneity - namely concept, language,
knowledge and data. This has allowed the re-interpretation of semantic heterogeneity
as a problem of representation diversity and the proposal of a strati ed logical
architecture which deals with this problem. The future work will consist of a
generalization of the pipeline presented in this paper into a full- edged knowledge
graph based methodology for data integration.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>The research conducted by Fausto Giunchiglia, Mayukh Bagchi and Simone
Bocca has received funding from the \DELPhi - DiscovEring Life Patterns"
project funded by the MIUR Progetti di Ricerca di Rilevante Interesse Nazionale
(PRIN) 2017 { DD n. 1062 del 31.05.2019. The research conducted by Alessio
Zamboni was supported by the InteropEHRate project, co-funded by the
European Union (EU) Horizon 2020 programme under grant number 826106.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Halevy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Why your data won't mix: Semantic Heterogeneity</article-title>
          .
          <source>Queue</source>
          <volume>3</volume>
          (
          <issue>8</issue>
          ),
          <volume>50</volume>
          {
          <fpage>58</fpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hull</surname>
          </string-name>
          , R.:
          <article-title>Managing semantic heterogeneity in databases: a theoretical perspective</article-title>
          .
          <source>In: Proceedings of the sixteenth ACM SIGACT-SIGMOD-SIGART symposium on Principles of database systems</source>
          , pp.
          <volume>51</volume>
          {
          <fpage>61</fpage>
          . (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bouquet</surname>
          </string-name>
          ,
          <article-title>Paolo and Giunchiglia, Fausto: Reasoning about theory adequacy. A new solution to the quali cation problem, Fundamenta Informaticae</article-title>
          , IOS Press.
          <volume>23</volume>
          (
          <issue>2</issue>
          ,
          <issue>3</issue>
          ,
          <issue>4</issue>
          ),
          <volume>247</volume>
          {
          <fpage>262</fpage>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Giunchiglia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Maltese</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Dutta</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Domains and context: rst steps towards managing diversity in knowledge, Journal of Web Semantics, Special Issue on Reasoning with Context in the Semantic Web, (</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Knoblock</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szekely</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ambite</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lerman</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muslea</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taheriyan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mallick</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Semi-automatically mapping structured sources into the semantic web</article-title>
          .
          <source>In: Extended Semantic Web Conference</source>
          , pp.
          <volume>375</volume>
          {
          <fpage>390</fpage>
          . Springer, Heidelberg (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Knoblock</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szekely</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Exploiting semantics for big data integration</article-title>
          .
          <source>AI</source>
          Magazine
          <volume>36</volume>
          (
          <issue>1</issue>
          ),
          <volume>25</volume>
          {
          <fpage>38</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Leida</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gusmini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davies</surname>
          </string-name>
          , J.:
          <article-title>Semantics-aware data integration for heterogeneous data sources</article-title>
          .
          <source>Journal of Ambient Intelligence and Humanized Computing</source>
          <volume>4</volume>
          (
          <issue>4</issue>
          ),
          <volume>471</volume>
          {
          <fpage>491</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beckwith</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gross</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>K.J.</given-names>
          </string-name>
          :
          <article-title>Introduction to WordNet: An on-line lexical database</article-title>
          .
          <source>International journal of lexicography 3(4)</source>
          ,
          <volume>235</volume>
          {
          <fpage>244</fpage>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Giunchiglia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fumagalli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Teleologies: Objects, actions and functions</article-title>
          .
          <source>In: International Conference on Conceptual Modeling</source>
          , pp.
          <volume>520</volume>
          {
          <issue>534</issue>
          (ER
          <year>2017</year>
          ). Springer, Cham (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Millikan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          : Language, thought, and
          <article-title>other biological categories: New foundations for realism</article-title>
          . MIT press. (
          <year>1984</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Giunchiglia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batsuren</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freihat</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          :
          <article-title>One world{seven thousand languages</article-title>
          .
          <source>In: Proceedings 19th International Conference on Computational Linguistics and Intelligent Text Processing, CiCling2018</source>
          , (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Giunchiglia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batsuren</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bella</surname>
          </string-name>
          , G.:
          <article-title>Understanding and exploiting language diversity</article-title>
          .
          <source>In: Proceedings of the 26th International Joint Conference on Arti cial Intelligence</source>
          , pp.
          <volume>4009</volume>
          {
          <fpage>4017</fpage>
          . (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Giunchiglia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fumagalli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Entity Type Recognition - dealing with the Diversity of Knowledge</article-title>
          .
          <source>In: Seventeenth International Conference on Principles of Knowledge Representation and Reasoning</source>
          . (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Bella</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gremes</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giunchiglia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Exploring the Language of Data</article-title>
          .
          <source>In: Proceedings of the 28th International Conference on Computational Linguistics</source>
          , pp.
          <fpage>6638</fpage>
          -
          <lpage>6648</lpage>
          . (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Maltese</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giunchiglia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Foundations of Digital Universities</article-title>
          .
          <source>Cataloging &amp; Classi cation Quarterly</source>
          <volume>55</volume>
          (
          <issue>1</issue>
          ),
          <volume>26</volume>
          {
          <fpage>50</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Bella</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zamboni</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Giunchiglia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Domain-Based Sense Disambiguation in Multilingual Structured Data</article-title>
          . DIVERSITY workshop, ECAI,
          <volume>53</volume>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Bella</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giunchiglia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McNeill</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Language and domain aware lightweight ontology matching</article-title>
          .
          <source>Journal of Web Semantics</source>
          <volume>43</volume>
          ,
          <issue>1</issue>
          {
          <fpage>17</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S.P.:</given-names>
          </string-name>
          <article-title>BabelNet: Building a very large multilingual semantic network. In: Proceedings of the 48th annual meeting of the association for computational linguistics</article-title>
          , pp.
          <volume>216</volume>
          {
          <fpage>225</fpage>
          . (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Euzenat</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shvaiko</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Ontology Matching. Vol.
          <volume>18</volume>
          . Springer, Heidelberg (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Jimenez-Ruiz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hassanzadeh</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Efthymiou</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srinivas</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <source>SemTab</source>
          <year>2019</year>
          :
          <article-title>Resources to Benchmark Tabular Data to Knowledge Graph Matching Systems</article-title>
          . In: European Semantic Web Conference, pp.
          <volume>514</volume>
          {
          <fpage>530</fpage>
          . Springer, Cham (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Iglesias</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jozashoori</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaves-Fraga</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Collarana</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Vidal</surname>
          </string-name>
          , M.E.:
          <article-title>SDMRDFizer: An RML interpreter for the e cient creation of rdf knowledge graphs</article-title>
          .
          <source>In: Proceedings of the 29th ACM International Conference on Information Knowledge Management</source>
          , pp.
          <fpage>3039</fpage>
          -
          <lpage>3046</lpage>
          . (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Jozashoori</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaves-Fraga</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iglesias</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vidal</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Corcho</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>FunMap: E cient Execution of Functional Mappings for Knowledge Graph Creation</article-title>
          . In: International Semantic Web Conference, pp.
          <fpage>276</fpage>
          -
          <lpage>293</lpage>
          . Springer, Cham (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Fensel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simsek</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angele</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huaman</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , Karle, E.,
          <string-name>
            <surname>Panasiuk</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toma</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Umbrich</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wahler</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <source>Knowledge Graphs. 1st edn</source>
          . Springer International Publishing,
          <string-name>
            <surname>Switzerland</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Kejriwal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Domain-Speci c Knowledge Graph Construction. 1st edn</article-title>
          . Springer, Heidelberg (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Bonatti</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Decker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polleres</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Presutti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Knowledge graphs: New directions for knowledge representation on the semantic web (dagstuhl seminar 18371)</article-title>
          .
          <source>Dagstuhl Reports</source>
          <volume>8</volume>
          (
          <issue>9</issue>
          ) (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Gagnon</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Ontology-based integration of data sources</article-title>
          .
          <source>In: 10th International Conference on Information Fusion</source>
          , pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          . IEEE. (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Gawriljuk</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knoblock</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Szekely</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A scalable approach to incrementally building knowledge graphs</article-title>
          .
          <source>In: International Conference on Theory and Practice of Digital Libraries</source>
          , pp.
          <volume>189</volume>
          {
          <fpage>199</fpage>
          . Springer, Cham (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Cheatham</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pesquita</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Semantic data integration</article-title>
          .
          <source>Handbook of big data technologies (Springer)</source>
          .
          <volume>263</volume>
          {
          <issue>305</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Mountantonakis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tzitzikas</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Large-scale semantic integration of linked data: A survey</article-title>
          .
          <source>ACM Computing Surveys (CSUR) 52(5)</source>
          ,
          <volume>1</volume>
          {
          <fpage>40</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30. Farber,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Ell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Menne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Rettinger</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>A comparative survey of dbpedia, freebase, opencyc, wikidata, and yago</article-title>
          .
          <source>Semantic Web Journal</source>
          <volume>1</volume>
          (
          <issue>1</issue>
          ), 1{
          <issue>5</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Ringler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>One knowledge graph to rule them all? Analyzing the di erences between DBpedia, YAGO, Wikidata &amp; co</article-title>
          .
          <source>In: Joint German/Austrian Conference on Arti cial Intelligence (Kunstliche Intelligenz)</source>
          , pp.
          <volume>366</volume>
          {
          <fpage>372</fpage>
          . Springer, Cham (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Vrandecic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Krotzsch, M.:
          <article-title>Wikidata: a free collaborative knowledge base</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>57</volume>
          (
          <issue>10</issue>
          ),
          <volume>78</volume>
          {
          <fpage>85</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>