<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>June</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>An approach to extracting thematic views from highly heterogeneous sources of a data lake</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudia Diamantini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Lo Giudice</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lorenzo Musarella</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Domenico Potena</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emanuele Storti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Domenico Ursino</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DII, Polytechnic University of Marche</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>DIIES, University \Mediterranea" of Reggio Calabria</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>2</volume>
      <fpage>4</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>In the last years, data lakes are emerging as an e ective and e cient support for information and knowledge extraction from a huge amount of highly heterogeneous and quickly changing data sources. Data lake management requires the de nition of new techniques, very di erent from the ones adopted for data warehouses in the past. One of the main issues to address in this scenario consists in the extraction of thematic views from the (very heterogeneous and generally unstructured) data sources of a data lake. In this paper, we propose a new network-based model to uniformly represent structured, semi-structured and unstructured sources of a data lake. Then, we present a new approach to, at least partially, \structure" unstructured data. Finally, we de ne a technique to extract thematic views from the sources of a data lake, based on similarity and other semantic relations among the metadata of data sources.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In the last years, data lakes are emerging as an e ective and e cient answer
to the problem of extracting information and knowledge from a huge amount
of highly heterogeneous and quickly changing data sources [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Data lake
management requires the de nition of new techniques, very di erent from the ones
adopted for data warehouses in the past. These techniques may exploit the large
set of metadata always supplied with data lakes, which represent their core and
the main tool allowing them to be a very competitive framework in the big data
era. One of the main issues to address in a scenario comprising many data sources
extremely heterogeneous in their format, structure and semantics, consists in the
extraction of thematic views from data sources [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], i.e., the construction of views
concerning one or more topics of interest for the user, obtained by extracting
and merging data coming from di erent sources. This problem has been largely
investigated in the past for structured and semi-structured data sources stored in
a data warehouse [
        <xref ref-type="bibr" rid="ref22 ref26 ref8">26, 8, 22</xref>
        ], and this witnesses its extreme relevance. However,
it is esteemed that, currently, more than 80% of data sources are unstructured
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. As a consequence, it is just this type of source that represents the main actor
of the big data scenario and, consequently, of data lakes.
      </p>
      <p>
        In this paper, we aim at providing a contribution in this setting. Indeed, we
propose a supervised approach to extracting thematic views from highly
heterogeneous sources of a data lake. Our approach represents all the data lake sources
by means of a suitable network. Indeed, networks are very exible structures that
allow the modeling of almost all phenomena that researchers aim at investigating
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Thanks to this uniform representation of the data lake sources, the
extraction of thematic views from them can be performed by exploiting graph-based
tools. We de ne \supervised" our approach because it requires the user to
specify the set of topics T = fT1; T2; : : : ; Tng that should be present in the thematic
view(s) to extract. Our approach consists of two steps. The former is mainly
based on the structure of involved sources. It exploits several notions typical
of (social) network analysis, such as the notion of ego network, which actually
represents the core of the proposed approach. The latter exploits a knowledge
repository, which is used to discover new relationships, other than synonymies,
among metadata, with the purpose to re ne the integration of di erent thematic
views obtained after the rst step. In this step, our approach relies on DBpedia.
      </p>
      <p>This paper is organized as follows: Section 2 illustrates related literature. In
Section 3, we present the proposed approach. In particular, rst we describe a
unifying model for data lake representation; then, we present our approach to
partially structuring unstructured sources; nally, we discuss the two steps of
our approach for thematic view extraction. In Section 4, we present our example
case, whereas, in Section 5, we draw our conclusions and discuss future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Literature</title>
      <p>
        The new data lake scenario is characterized by several peculiarities that make
it very di erent from the data warehouse paradigm. Hence, it is necessary to
adapt (if possible) old algorithms conceived for data warehouses or to de ne
new approaches capable of handling and taking advantage of the speci cities of
this new paradigm. However, most approaches proposed in the literature for data
integration, query answering and view extraction do not completely t the data
lake paradigm. For instance, [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] proposes some techniques for building views on
semi-structured data sources based on some expected queries. Other researchers
focus on materialized views and, speci cally, on throughput and execution time;
therefore, they a-priori de ne a set of well-known views and, then, materialize
them. Two surveys on this issue can be found in [
        <xref ref-type="bibr" rid="ref1 ref16">16, 1</xref>
        ]. The authors of [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]
investigate the same problem but they focus on XML sources. The approach
of [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] addresses the same issue by means of query rewriting; speci cally, it
transforms a query Q into a set of new queries, evaluates them, and, then,
merges the corresponding answers to construct the materialized answer to Q.
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] proposes an approach to constructing materialized views for heterogeneous
      </p>
      <p>DBpedia: http://dbpedia.org
databases; it requires the presence of a static context and the pre-computation
of some queries.</p>
      <p>
        Another family of approaches exploits materialized views to perform tree
pattern querying [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] and graph pattern queries [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Unfortunately, all these
approaches are well-suited for structured and semi-structured data, whereas they
are not scalable and lightweight enough to be used in a dynamic context or with
unstructured data. An interesting advance in this area can be found in [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Here,
the authors propose an incremental approach to address the graph pattern query
problem on both static and dynamic real-life data graphs. Other kinds of views
are investigated in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In particular, this last paper uses virtual views to
access heterogeneous data sources without knowing many details of them. For
this purpose, it creates virtual views of the data sources.
      </p>
      <p>
        Finally, semantic-based approaches have long been used to drive data
integration in databases and data warehouses. More recently, in the context of big
data, formal semantics has been speci cally exploited to address issues
concerning data variety/heterogeneity, data inconsistency and data quality in such a
way as to increase understandability [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In the data lake scenario, semantic
techniques have been successfully applied to more e ciently integrate and
handle both structured and unstructured data sources by aligning data silos and
better managing evolving data models. For instance, in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], the authors discuss
a data lake system with a semantic metadata matching component for
ontology modeling, attribute annotation, record linkage, and semantic enrichment.
Furthermore, [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] presents a system to discover and enforce expressive integrity
constraints from data lakes. Similarly to what happens in our approach,
knowledge graphs in RDF are used to drive integration. To reach their objectives,
these techniques usually rely on information extraction tools (e.g., Open Calais)
that may assist in linking metadata to uniform vocabularies (e.g., ontologies or
knowledge repositories, such as DBpedia).
3
3.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Description of the proposed approach</title>
      <sec id="sec-3-1">
        <title>A unifying model for data lake representation</title>
        <p>In this section, we illustrate our network-based model to represent and handle a
data lake, which we will use in the rest of this paper. In our model, a data lake
DL is represented as a set of m data sources: DL = fD1; D2; ; Dmg. A data
source Dk 2 DL is provided with a rich set Mk of metadata. We denote with
MDL the repository of the metadata of all the data sources of DL: MDL =
fM1; M2; : : : ; Mmg.</p>
        <p>
          According to [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], our model represents Mk by means of a triplet: Mk =
hMkT ; MkO; MkBi. Here: (i) MkT denotes technical metadata. It represents the
type, the format, the structure and the schema of the corresponding data. It
RDF Concepts and abstract
REC-rdf-concepts-20040210/
http://www.opencalais.com
        </p>
        <p>Syntax:
http://www.w3.org/TR/2004/
is commonly provided by the source catalogue. (ii) MkO represents operational
metadata. It includes the source and target locations of the corresponding data,
the associated le size, the number of their records, and so on. Usually, it is
autoB
matically generated by the technical framework handling the data lake. (iii) Mk
indicates business metadata. It comprises the business names and descriptions
assigned to data elds. It also covers business rules, which can become integrity
constraints for the corresponding data source.</p>
        <p>In this paper, we consider only business metadata. Indeed, they denote, at
the intensional level, the information content stored in the data lake sources and
are those of interest for supporting the extraction of thematic views from a data
lake, which is our ultimate goal. In order to represent MkB, our model adopts
a notation typical of XML, JSON and many other semi-structured models.
According to this notation, Objk indicates the set of all the objects stored in MkB.
It consists of the union of three subsets: Objk = Attk [ Smpk [ Cmpk. Here: (i)
Attk indicates the set of the attributes of MkB; (ii) Smpk represents the set of
the simple elements of MkB; (iii) Cmpk denotes the set of the complex elements
of MkB. In this context, the meaning of the terms \attribute", \simple element"
and \complex element" is the one typical of semi-structured data models.</p>
        <p>MkB can be also represented as a graph: MkB = hNk; Aki. Nk is the set of
the nodes of MkB. There exists a node nkj 2 Nk for each object okj 2 Objk.
According to the structure of Objk, Nk consists of the union of three subsets:
Nk = NkAtt [ NkSmp [ NkCmp. Here, NkAtt (resp., NkSmp, NkCmp) indicates the
set of the nodes corresponding to Attk (resp., Smpk, Cmpk). There is a
oneto-one correspondence between a node of Nk and an object of Objk. Therefore,
in the following, we will use the two terms interchangeably. Let x be a complex
element of MkB. Objkx indicates the set of the objects directly contained in x,
whereas NkOxbj denotes the set of the corresponding nodes. Finally, let y be a
simple element of MkB. Attky represents the set of the attributes of y, whereas
NkAytt denotes the set of the corresponding nodes.</p>
        <p>Ak indicates the set of the arcs of MkB. It consists of two subsets: Ak =
A0k [ A00. Here: (i) A0k = f(nx; ny)jnx 2 N Cmp; ny 2 NnOxbj g, i.e., there is an
k k
arc from a complex element of MkB to each object directly contained in it. (ii)
A0k0 = f(nx; ny)jnx 2 NkSmp; ny 2 NnAxttg, i.e., there is an arc from a simple
element of MkB to each of its attributes.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>An approach to partially structuring unstructured sources</title>
        <p>Our network-based model for representing and handling a data lake is perfectly
tted for representing and managing semi-structured data because it has been
designed having XML and JSON in mind. Clearly, it is su ciently powerful
to represent structured data. The highest di culty regards unstructured data
because it is worth avoiding a at representation consisting of a simple element
for each keyword provided to denote the source content. As a matter of fact, this
kind of representation would make the reconciliation, and the next integration,
of an unstructured source with the other (semi-structured and structured) ones
of the data lake very di cult. Therefore, it is necessary to (at least partially)
\structure" unstructured data.</p>
        <p>Our approach to addressing this issue consists of four phases, namely: (i)
creation of nodes; (ii) derivation and management of part-of relationships; (iii)
derivation of lexical and string similarities; (iv) management of lexical and string
similarities.</p>
        <p>Phase 1. During this phase, our approach creates a complex element
representing the source as a whole, and a simple element for each keyword.
Furthermore, it adds an arc from the source to each of the simple elements. Initially,
there is no arc between two simple elements. To determine the arcs to add, the
next phases are necessary.</p>
        <p>
          Phase 2. During this phase, our approach adds an arc from the node nk1 ,
corresponding to the keyword k1, to the node nk2 , corresponding to the keyword
k2, if k2 is registered as a lemma of k1 in a suitable thesaurus. Taking the
current trends into account, this thesaurus should be a multimedia one; for this
purpose, in our experiments, we have adopted BabelNet [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. When this arc has
been added, nk1 must be considered a complex element, instead of a simple one.
        </p>
        <p>
          Phase 3. During this phase, our approach starts by deriving lexical
similarities. In particular, it states that there exists a similarity between the nodes nk1 ,
corresponding to the keyword k1, and nk2 , corresponding to the keyword k2, if k1
and k2 have at least one common lemma in a suitable thesaurus. Also in this case,
we have adopted BabelNet. After having found lexical similarities, our approach
derives string similarities and states that there exists a similarity between nk1
and nk2 if the string similarity degree kd(k1; k2), computed by applying a suitable
string similarity metric on k1 and k2, is \su ciently high" (see below). We have
chosen N-Grams [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] as string similarity metric because we have experimentally
seen that it provides the best results in our context. Now, we illustrate in detail
what \su ciently high" means and how our approach operates. Let KeySim
be the set of the string similarities for each pair of keywords of the source into
consideration. Each record in KeySim has the form hki; kj ; kd(ki; kj )i. Our
approach rst computes the maximum keyword similarity degree kdmax present
in KeySim. Then, it examines each keyword similarity registered therein. Let
hk1; k2; kd(k1; k2)i be one of these similarities. If ((kd(k1; k2) thk kdmax) and
(kd(k1; k2) thkmin)), which implies that the keyword similarity degree
between k1 and k2 is among the highest ones in KeySim and that, in any case,
it is higher than or equal to a minimum threshold, then it concludes that there
exists a similarity between nk1 and nk2 . We have experimentally set thk = 0:70
and thkmin = 0:50. At the end of this phase, our approach has found some
(lexical and/or string) similarities, each stating that a node nki is similar to a node
nkj .
        </p>
        <p>
          Phase 4. During this phase, our approach aims at managing the similarities
found during Phase 3. In particular, if there exists a (lexical and/or string)
simIn this paper, we use the term \lemma" according to the meaning it has in
BabelNet [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. Here, given a term, its lemmas are other objects (terms, emoticons, etc.)
contributing to specify its meaning.
ilarity between two nodes nki and nkj , it merges them into one node nkij , which
inherits all the incoming and outgoing arcs of nki and nkj . After all similarities
have been considered, it could happen that there exist two or more arcs from a
node nki to a node nkj . In this case, our approach merges them into one arc.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>An approach to extracting thematic views</title>
        <p>Our approach to extracting thematic views operates on a data lake DL whose
data sources are represented by means of the model described in Section 3.1.
It is called \supervised" because it requires the user to specify the set of topics
T = fT1; T2; ; Tlg, which should be present in the thematic view(s) to extract.
It consists of two steps, the former mainly based on the structure of the sources
at hand, the latter mainly focusing on the corresponding semantics. We describe
these two steps in the next subsections.</p>
        <p>
          Step 1. The rst step of our approach receives a data lake DL, a set of topics
T = fT1; T2; ; Tlg, representing the themes of interest for the user, and a
dictionary Syn of synonymies involving the objects stored in the sources of DL.
This dictionary could be a generic thesaurus, such as BabelNet [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], a
domainspeci c thesaurus, or a dictionary obtained by taking into account the structure
and the semantics of the sources, which the corresponding objects refer to (such
as the dictionaries produced by XIKE [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], MOMIS [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] or Cupid [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]).
        </p>
        <p>
          In this step, the concept of ego network [
          <xref ref-type="bibr" rid="ref11 ref2">2, 11</xref>
          ] plays a key role. We recall that
an ego network consists of a focal node (the ego) and the nodes it is directly
connected to (the \alters"), plus the ties, if any, between the alters.
        </p>
        <p>Let Ti be a topic of T . Let Obji = foi1 ; oi2 ; ; oiq g be the set of the
objects synonymous of Ti in DL. Let Ni = fni1 ; ni2 ; ; niq g be the
corresponding nodes. First, Step 1 constructs the ego networks Ei1 ; Ei2 ; ; Eiq
having ni1 ; ni2 ; ; niq as the corresponding egos. Then, it merges all the egos
into a unique node ni. In this way, it obtains a unique ego network Ei from
Ei1 ; Ei2 ; ; Eiq . If a synonymy exists between two alters belonging to di erent
ego networks, then these are merged into a unique node and the corresponding
arcs linking them to the ego ni are merged into a unique arc. At the end of this
task, we have a unique ego network Ei corresponding to Ti.</p>
        <p>After having performed the previous task for each topic of T , we have a
set E = fE1; E2; ; Elg of l ego networks. At this point, Step 1 nds all the
synonymies of Syn involving objects of the ego networks of E and merges the
corresponding nodes. After all the possible synonymies involving objects of the
ego network of E have been considered and the corresponding nodes have been
merged, a set V = fV1; ; Vgg, 1 g l, of networks representing potential
views is obtained.</p>
        <p>If g = 1, then it is possible to conclude that Step 1 has been capable of
extracting a unique thematic view comprising all the topics required by the
user. Otherwise, there exist more views each comprising some (but not all) of
the topics of interest for the user. If g = 1, Step 2 is performed to make more
precise and complete the unique view representing all the topics of T . If g &gt; 1,
Step 2 aims at nding other relationships, di erent from synonymies, among
the objects of the views of V in such a way as to try to obtain a unique view
embracing all the topics of interest for the user.</p>
        <p>Step 2. This step starts by enriching each view Vi 2 V . For this purpose, it
connects each of its elements to all the semantically related concepts taken from
a reference knowledge repository.</p>
        <p>In this work, we rely on DBpedia, one of the largest knowledge graphs in the
Linked Data context, including more than 4.58 million entities in RDF. To this
aim, rst each element of Vi (including its synonyms) is mapped to the
corresponding entry in DBpedia. In many cases, such a mapping is already provided by
BabelNet. Then, for each DBpedia entry, all the related concepts are retrieved.
In DBpedia, knowledge is structured according to the Linked Data principles,
i.e. as an RDF graph built by triples. Each triple hs(ubject); p(roperty); o(bject)i
states that a subject s has a property p, whose value is an object o. Both
subjects and properties are resources (i.e., nodes in DBpedia's knowledge graph),
whereas objects may be either resources or literals (i.e., values of some primitive
data types, such as strings or numbers). Each triple represents the minimal
component of the knowledge graph. This last is built by merging triples together.
Therefore, retrieving the related concepts for a given element x implies nding
all the triples where x is either the subject or the object.</p>
        <p>For each view Vi 2 V , the procedure to extend it consists of the following
three substeps:
1. Mapping : for each node n 2 Vi, its corresponding DBpedia entry d is found.
2. Triple extraction: all the related triples hd; p; oi and hs; p; di, i.e., all the
triples in which d is either the subject or the object, are retrieved.
3. View extension: for each retrieved triple hd; p; oi (resp., hs; p; di), Vi is
extended by de ning a node for the object o (resp., s), if not already existing,
linked to n through an arc labeled as p.</p>
        <p>These three tasks are repeated for all the views of V . As previously pointed
out, this enrichment procedure is particularly important if jV j &gt; 1 because the
new derived relationships could help to merge the thematic views that was not
possible to merge during Step 1. In particular, let Vi 2 V and Vj 2 V be two
views of V , and let Vi0 and Vj0 be the extended views corresponding to them. If
there exist two nodes nih 2 Vi0 ad njk 2 Vj0 such that nih = njk , then they can
be merged in one node; if this happens, Vi0 and Vj0 become connected. After all
equal nodes of the views of V have been merged, all the views of V could be
either merged in one view or not. In the former case, the process terminates with
success. Otherwise, it is possible to conclude that no thematic views comprising
Whenever this does not happen, the mapping can be automatically provided by the
DBpedia Lookup Service (http://wiki.dbpedia.org/projects/dbpedia-lookup).</p>
        <p>Here, two nodes are equal if the corresponding name coincide.
all the topics speci ed by the user can be found. In this last case, our approach
still returns the enriched views of V and leaves the user the choice to accept of
reject them.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>An example case</title>
      <p>In this section, we present an example case aiming at illustrating the various
tasks of our approach. Here, we consider: (i) a structured source, called Weather
Conditions (W , in short), whose corresponding E/R schema is not reported for
space limitations; (ii) two semi-structured sources, called Climate (C, in short)
and Environment (E, in short), whose corresponding XML Schemas are not
reported for space limitations; (iii) an unstructured source, called Environment
Video (V , in short), consisting of a YouTube video and whose corresponding
keywords are: garden, f lower, rain, save, earth, tips, recycle, aurora, planet,
garbage, pollution, region, lif e, plastic, metropolis, environment, nature, wave,
eco, weather, simple, f ineparticle, climate, ocean, environmentawareness,
educational, reduce, power, bike.</p>
      <p>By applying the approaches mentioned in Section 3.2, we obtain the
corresponding representations in our network-based model, shown in Figure 1.</p>
      <p>Assume, now, that a user speci es the following set T of topics of her
interest: T = fOcean; Areag. First, our approach determines the terms (and,
then, the objects) in the ve sources that are synonyms of Ocean and Area.
As for Ocean, the only synonym present in the sources is Sea; as a
consequence, Obj1 comprises the node Ocean of the source V (V:Ocean) and the
node Sea of the source C (C:Sea). An analogous activity is performed for
Area. At the end of this task we have that Obj1 = fV:Ocean; C:Seag and
Obj2 = fW:P lace; C:P lace; V:Region; E:Locationg.</p>
      <p>Step 1 of our approach proceeds by constructing the ego networks
corresponding to the objects of Obj1 and Obj2. They are reported in Figure 2.</p>
      <p>Now, consider the ego networks corresponding to V:Ocean and C:Sea. Our
approach merges the two egos into a unique node. Then, it veri es whether
further synonyms exist between the alters. Since none of these synonyms exists,
it returns the ego network shown in Figure 3(a). The same task is performed to
the ego networks corresponding to W:P lace, C:P lace, V:Region and E:Location.
In particular, rst the four egos are merged. Then, synonyms between the alters
W:City and C:City and the alters W:Altitude and C:Altitude are retrieved.
Based on this, W:City and C:City are merged in one node, W:Altitude and
C:Altitude in another node, the arcs linking the ego to W:City and C:City are
merged in one arc and the ones linking the ego to W:Altitude and C:Altitude in
another arc. In this way, the ego network shown in Figure 3(b) is returned. At
this point, there are two ego networks, EOcean and EArea, each corresponding
to one of the terms speci ed by the user.</p>
      <p>Here and in the following, we use the notation S:o to indicate the object o of the
source S.</p>
      <p>Now, Step 1 veri es if there are any synonyms between a node of EOcean and
a node of EArea. Since this does not happen, it terminates and returns the set
V = fVOcean; VAreag, where VOcean (resp., VArea) coincides with EOcean (resp.,
EArea).</p>
      <p>Fig. 2. Ego networks corresponding to V:Ocean, C:Sea, W:P lace, C:P lace, V:Region
and E:Location.</p>
      <p>At this point, Step 2 is executed. As shown in Figure 4, rst each term
(synonyms included) is semantically aligned to the corresponding DBpedia entry
(e.g., Ocean is linked to dbo:Sea, Area is linked to dbo:Location and dbo:Place,
while Country to dbo:Country , respectively). After a single iteration, the triples
hdbo:sea rdfs:range dbo:Seai and hdbo:sea rdfs:domain dbo:Placei are retrieved.
They correspond to the DBpedia property sea, which relates a country to the
sea from which it is lapped. Other connections can be found by moving to
speci c instances of the mentioned resources. In this way, the following triples are
retrieved: hinstance rdf:type dbo:Seai, hinstance rdf:type dbo:Locationi, hinstance
rdf:type dbo:Placei, meaning that there are resource instances having types Sea,
Location and Place simultaneously (e.g., dbr:Mediterranean Sea). Furthermore,
a triple hinstance dbo:country dbo:Countryi can be retrieved, meaning that those
instances being a Sea, a Location or a Place specify in which dbo:Country they
are located through the dbo:country property. In this example case, Step 2
succeeded in merging the two views that Step 1 had maintained separated.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper, we have presented a new network-based model to uniformly
represent the structured, semi-structured and unstructured sources of a data lake.
Then, we have proposed a new approach to, at least partially, \structuring"
Pre xes dbo and dbr stand for http://dbpedia.org/ontology/ and http://
dbpedia.org/resource/
unstructured data. Finally, based on these two tools, we have de ned a new
approach to extracting thematic views from the sources of a data lake consisting of
two steps, based on ego networks (Step 1) and semantic relationships (Step 2).
This paper is not to be intended as an ending point. Actually, it could be the
starting point of a new family of approaches aiming at handling information
systems in the new big data-oriented scenario. By proceeding in this direction, rst
we plan to de ne an unsupervised approach to extracting thematic views from
a data lake. Then, we plan to de ne new approaches to supporting a exible
and lightweight querying of the sources of a data lake, as well as approaches to
schema matching, schema mapping, data reconciliation and integration strongly
oriented to data lakes based mainly on unstructured data sources.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>S.</given-names>
            <surname>Abiteboul</surname>
          </string-name>
          and
          <string-name>
            <given-names>O.M.</given-names>
            <surname>Duschka</surname>
          </string-name>
          .
          <article-title>Complexity of answering queries using materialized views</article-title>
          .
          <source>In Proc. of the International Symposium on Principles of database systems (SIGMOD/PODS'98)</source>
          , pages
          <fpage>254</fpage>
          {
          <fpage>263</fpage>
          , Seattle, WA, USA,
          <year>1998</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>V.</given-names>
            <surname>Arnaboldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Conti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passarella</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Pezzoni</surname>
          </string-name>
          .
          <article-title>Analysis of ego network structure in online social networks</article-title>
          .
          <source>In Proc. of the International Conference on Privacy (PASSAT'12)</source>
          , pages
          <fpage>31</fpage>
          {
          <fpage>40</fpage>
          , Amsterdam, Netherlands,
          <year>2012</year>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>L.</given-names>
            <surname>Aversano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Intonti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Quattrocchi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Tortorella</surname>
          </string-name>
          .
          <article-title>Building a virtual view of heterogeneous data source views</article-title>
          .
          <source>In Proc. of the International Conference on Software and Data Technologies (ICSOFT'10)</source>
          , pages
          <fpage>266</fpage>
          {
          <fpage>275</fpage>
          , Athens, Greece,
          <year>2010</year>
          . INSTICC Press.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>C.</given-names>
            <surname>Bachtarzi</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Bachtarzi</surname>
          </string-name>
          .
          <article-title>A model-driven approach for materialized views de nition over heterogeneous databases</article-title>
          .
          <source>In Proc. of the International Conference on New Technologies of Information and Communication (NTIC'15)</source>
          , pages
          <fpage>1</fpage>
          <lpage>{</lpage>
          5,
          <string-name>
            <surname>Mila</surname>
          </string-name>
          , Algeria,
          <year>2015</year>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Castano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vincini</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Beneventano</surname>
          </string-name>
          .
          <article-title>Semantic integration and query of heterogeneous information sources</article-title>
          .
          <source>Data &amp; Knowledge Engineering</source>
          ,
          <volume>36</volume>
          (
          <issue>3</issue>
          ):
          <volume>215</volume>
          {
          <fpage>249</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>J.</given-names>
            <surname>Biskup</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Embley</surname>
          </string-name>
          .
          <article-title>Extracting information from heterogeneous information sources using ontologically speci ed target views</article-title>
          .
          <source>Information Systems</source>
          ,
          <volume>28</volume>
          (
          <issue>3</issue>
          ):
          <volume>169</volume>
          {
          <fpage>212</fpage>
          ,
          <year>2003</year>
          . Elsevier.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Bouadjenek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hacid</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Bouzeghoub</surname>
          </string-name>
          .
          <article-title>Social networks and information retrieval, how are they converging? A survey, a taxonomy and an analysis of social information retrieval approaches and platforms</article-title>
          .
          <source>Information Systems</source>
          ,
          <volume>56</volume>
          :1{
          <fpage>18</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>S.</given-names>
            <surname>Castano</surname>
          </string-name>
          and V. De Antonellis.
          <article-title>Building views over semistructured data sources</article-title>
          .
          <source>In Proc. of the International Conference on Conceptual Modeling (ER'99)</source>
          , pages
          <fpage>146</fpage>
          {
          <fpage>160</fpage>
          , Paris, France,
          <year>1999</year>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>A.</given-names>
            <surname>Corbellini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mateos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zunino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Godoy</surname>
          </string-name>
          , and
          <string-name>
            <surname>S.N.</surname>
          </string-name>
          <article-title>Schia no. Persisting big-data: The NoSQL landscape</article-title>
          .
          <source>Information Systems</source>
          ,
          <volume>63</volume>
          :1{
          <fpage>23</fpage>
          ,
          <year>2017</year>
          . Elsevier.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>P. De Meo</surname>
            , G. Quattrone, G. Terracina, and
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Ursino</surname>
          </string-name>
          .
          <article-title>Integration of XML Schemas at various \severity" levels</article-title>
          .
          <source>Information Systems</source>
          ,
          <volume>31</volume>
          (
          <issue>6</issue>
          ):
          <volume>397</volume>
          {
          <fpage>434</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>R. DeJordy</surname>
            and
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Halgin</surname>
          </string-name>
          .
          <article-title>Introduction to ego network analysis</article-title>
          . Boston MA: Boston College and the Winston Center for Leadership &amp; Ethics,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>W.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <article-title>Answering pattern queries using views</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <volume>28</volume>
          (
          <issue>2</issue>
          ):
          <volume>326</volume>
          {
          <fpage>341</fpage>
          ,
          <year>2016</year>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>H.</given-names>
            <surname>Fang</surname>
          </string-name>
          .
          <article-title>Managing data lakes in big data era: What's a data lake and why has it became popular in data management ecosystem</article-title>
          .
          <source>In Proc. of the International Conference on Cyber Technology in Automation (CYBER'15)</source>
          , pages
          <fpage>820</fpage>
          {
          <fpage>824</fpage>
          ,
          <string-name>
            <surname>Shenyang</surname>
          </string-name>
          , China,
          <year>2015</year>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>M. Farid</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Roatis</surname>
            ,
            <given-names>I.F.</given-names>
          </string-name>
          <string-name>
            <surname>Ilyas</surname>
            , H. Ho mann, and
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Chu</surname>
          </string-name>
          .
          <article-title>CLAMS: bringing quality to Data Lakes</article-title>
          .
          <source>In Proc. of the International Conference on Management of Data (SIGMOD/PODS'16)</source>
          , pages
          <year>2089</year>
          {
          <year>2092</year>
          , San Francisco, CA, USA,
          <year>2016</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>R.</given-names>
            <surname>Hai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Geisler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Quix</surname>
          </string-name>
          .
          <article-title>Constance: An intelligent data lake system</article-title>
          .
          <source>In Proc. of the International Conference on Management of Data (SIGMOD/PODS'16)</source>
          , pages
          <year>2097</year>
          {
          <volume>2100</volume>
          , San Francisco, CA, USA,
          <year>2016</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>A.</given-names>
            <surname>Halevy</surname>
          </string-name>
          .
          <article-title>Answering queries using views: A survey</article-title>
          .
          <source>The VLDB Journal</source>
          ,
          <volume>10</volume>
          (
          <issue>4</issue>
          ):
          <volume>270</volume>
          {
          <fpage>294</fpage>
          ,
          <year>2001</year>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>P.</given-names>
            <surname>Hitzler</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Janowicz. Linked Data</surname>
          </string-name>
          ,
          <source>Big Data, and the 4th Paradigm. Semantic Web</source>
          ,
          <volume>4</volume>
          (
          <issue>3</issue>
          ):
          <volume>233</volume>
          {
          <fpage>235</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>G.</given-names>
            <surname>Kondrak</surname>
          </string-name>
          .
          <article-title>N-gram similarity and distance</article-title>
          .
          <source>In String processing and information retrieval</source>
          , pages
          <volume>115</volume>
          {
          <fpage>126</fpage>
          ,
          <year>2005</year>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>J. Madhavan</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          <string-name>
            <surname>Bernstein</surname>
            , and
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Rahm</surname>
          </string-name>
          .
          <article-title>Generic schema matching with Cupid</article-title>
          .
          <source>In Proc. of the International Conference on Very Large Data Bases (VLDB</source>
          <year>2001</year>
          ), pages
          <fpage>49</fpage>
          {
          <fpage>58</fpage>
          , Rome, Italy,
          <year>2001</year>
          . Morgan Kaufmann.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>R.</given-names>
            <surname>Navigli</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.P.</given-names>
            <surname>Ponzetto. BabelNet:</surname>
          </string-name>
          <article-title>The automatic construction, evaluation and application of a wide-coverage multilingual semantic network</article-title>
          .
          <source>Arti cial Intelligence</source>
          ,
          <volume>193</volume>
          :
          <fpage>217</fpage>
          {
          <fpage>250</fpage>
          ,
          <year>2012</year>
          . Elsevier.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>A.</given-names>
            <surname>Oram</surname>
          </string-name>
          .
          <article-title>Managing the Data Lake</article-title>
          . Sebastopol, CA, USA,
          <year>2015</year>
          .
          <string-name>
            <given-names>O</given-names>
            <surname>'Reilly.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. L.
          <string-name>
            <surname>Palopoli</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Pontieri</surname>
            , G. Terracina, and
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Ursino</surname>
          </string-name>
          .
          <article-title>Intensional and extensional integration and abstraction of heterogeneous databases</article-title>
          .
          <source>Data &amp; Knowledge Engineering</source>
          ,
          <volume>35</volume>
          (
          <issue>3</issue>
          ):
          <volume>201</volume>
          {
          <fpage>237</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>K.</given-names>
            <surname>Singh</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Singh</surname>
          </string-name>
          .
          <article-title>Answering graph pattern query using incremental views</article-title>
          .
          <source>In Proc. of the International Conference on Computing (ICCCA'16)</source>
          , pages
          <fpage>54</fpage>
          {
          <fpage>59</fpage>
          ,
          <string-name>
            <surname>Greater</surname>
            <given-names>Noida</given-names>
          </string-name>
          , India,
          <year>2016</year>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24. J.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>and J.X.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
          </string-name>
          .
          <article-title>Answering tree pattern queries using views: a revisit</article-title>
          .
          <source>In Proc. of the International Conference on Extending Database Technology (EDBT/ICDT'11)</source>
          , pages
          <fpage>153</fpage>
          {
          <fpage>164</fpage>
          ,
          <string-name>
            <surname>Uppsala</surname>
          </string-name>
          , Sweden,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.X.</given-names>
            <surname>Yu</surname>
          </string-name>
          .
          <article-title>Revisiting answering tree pattern queries using views</article-title>
          .
          <source>ACM Transactions on Database Systems</source>
          ,
          <volume>37</volume>
          (
          <issue>3</issue>
          ):
          <fpage>18</fpage>
          ,
          <year>2012</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>X. Wu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Theodoratos</surname>
            , and
            <given-names>W.H.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Answering XML queries using materialized views revisited</article-title>
          .
          <source>In Proc. of the International Conference on Information and Knowledge Management (CIKM '09)</source>
          , pages
          <fpage>475</fpage>
          {
          <fpage>484</fpage>
          ,
          <string-name>
            <surname>Hong</surname>
            <given-names>Kong</given-names>
          </string-name>
          , China,
          <year>2009</year>
          . ACM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>