<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Joint Conference (March</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Value-driven Approach for Designing Extended Data Warehouses</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nabila Berkani</string-name>
          <email>n-berkani@esi.dz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Selma Khouri</string-name>
          <email>s-khouri@esi.dz</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ladjel Bellatreche</string-name>
          <email>ladjel.bellatreche@ensma.fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carlos Ordonez</string-name>
          <email>carlos@central.uh.edu</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ecole nationale Supérieure d'Informatique</institution>
          ,
          <addr-line>BP 68M, 16309, Oued-Smar, Alger, Algérie.</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ecole nationale Supérieure d'Informatique</institution>
          ,
          <addr-line>BP 68M, 16309, Oued-Smar, Alger, Algérie.</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ISAE-ENSMA</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Houston</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>26</volume>
      <issue>2019</issue>
      <abstract>
        <p>In a very short time, the data warehouse (DW) technology has gone through all the phases of a technological product's life: introduction on the market, growth, maturity and decline. Maturity means there is a clearly identified design life cycle plus a race and competition between companies to increase their decision-making power. Decline was signaled by the appearance of Big Data. It is therefore essential to find other challenges that will contribute to the revival of DW while taking advantage of the V's of Big Data. The arrival of Linked Open Data (LOD) era is an excellent opportunity for both the DW academia and industry communities. LOD may bring an additional Value that the sources feeding a DW typically do not usually succeed to yield. Ofering the added value of a DW is related to a high Variety of sources. In this paper, first, we conceptualize the variety of internal and external sources and study its impact on the ETL phase to ease the value capturing. Secondly, three scenarios for integrating LOD in the DW are given. Finally, experiments are conducted to show the efectiveness of our approach.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>In contrast to traditional database applications, the process of
building DW is a complex, expensive, and time-consuming task.
Knowing this risk, companies willing to conduct a DW project
should never start unless managers are convinced that its benefits
outweigh the cost, known as Return Of Investment (ROI). Survey
studies conducted by analytical companies such as International
Data Corporation and SAS clearly conclude that DW technology
provides a good payback, in the sense that the average ROI for
a DW is far above the industry average, confirming the
addedvalue of DW technology.</p>
      <p>With the arrival of Big Data, companies owning a DW had
to change their BI strategy and align it. This alignment comes
from facing the V’s brought by Big Data (Volume, Variety,
Velocity, Veracity). This situation pushes these companies to enhance
their DW environment with Big Data technology, including
distributed programming, cloud computing, parallel processing
and so on. These technologies mainly focus on managing
Volume and Velocity of data, leaving Variety as a second priority.
More recently, considering value-requirements fixed by a
company such as money invested, awed customers, increased sales,
etc. has spawned another V: Value.</p>
      <p>
        The DW value has to be evaluated considering risks [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
One of the most important risks is the lack of satisfaction of
user requirements [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This is due to the fact that data sources
participating in the DW construction are not rich enough in
terms of concepts and instances. This limitation can decrease the
value of the target DW and consequently its ROI, where
decisions will be made perhaps without data. This scenario is deeply
related to the trade-of between Closed World (CWA) and Open
World (OWA) Assumptions. The CWA assumption states what
is not known to be true must be false, whereas OWA is the
opposite. Therefore, building a DW only from database sources
may penalize its value. To deal with the Value risk, recent
studies propose augmenting traditional DW sources with external
Web sources such as LOD and knowledge graphs [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Usually,
LOD store data with high quality, since important eforts in
curation, cleaning, entity resolution, etc. are deployed. If a company
succeeds in selecting LOD, their integration into DW will not
afect its quality. This is because the traditional ETL processes
will be augmented to deal with this new source of data. The price
to pay by designers when considering LOD is managing data
Variety. They bring a new format of data usually incompatible
with traditional ones. To augment a DW with LOD, important
eforts in conceptualizing and managing data Variety have to be
made. This variety concerns these main aspects: (a) the universe
of discourse (UOD) of sources, (b) their conceptual formalisms,
and (c) their physical implementations.
      </p>
      <p>In the context of LOD, their schema has the potential for
being the chosen model that federates external sources, as it is open,
standardized, visualizable, and associated with graph database
tools, while maintaining interoperability with semantic databases
and ontologies. Moreover, it allows conceptualizing variety at
two levels: vocabulary (using ontologies) and formalisms (RDF
schema). In this paper, we defend this LOD-augmented scenario
and we show that it significantly impacts the ETL environment
including its tasks, workflows and operators. With this
motivation in mind, we propose three comprehensive scenarios for a
company to integrate LOD into the DW design, while meeting
the Variety and added Value requirements at the conceptual level.
These scenarios are distinguished based on ordering events
stipulating the time when the company decides to build its DW and
the time when it decides to connect its DW to a relevant LOD.
More precisely, these scenarios follow two main schedules: (a)
the DW meets LOD and (b) the DW was designed before
LOD. In the first schedule, the DW is built from scratch by a
simultaneous integration of internal and external data sources,
whereas in (b), we assume that the DW was constructed well
before the company decided to integrate LOD. Therefore, such
DW needs to continuously integrate data from local sources
and LOD. Thanks to conceptual modeling, variety is managed
at ETL level. Value is handled by introducing metrics related to
requirement satisfaction.</p>
      <p>This paper is structured as follows: Section 2 positions LOD
in the DW landscape. Section 3 is related to the
conceptualization of the variety and its impact on ETL. Section 4 details our
scenarios integrating LOD in the DW design to add value. An
experimental study is conducted in Section 5. Section 6 concludes
our paper and outlines future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        A couple of recent studies consider LOD in the process of DW
construction, without emphasizing on value. These studies can
be projected on the conventional DW life cycle design that
includes: requirements definition, conceptual design, ETL, logical
and physical design. Even though the experience has shown that
bring requirements in forefront ensure the DW to be tightly
tailored to the users requirements [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], in most studies considering
LOD in DW design, users’ requirements are either ignored or
assumed defined. Our work is motivated by the importance of
requirements on the DW system incorporating LOD in
identifying missing concepts and instances required for DW value
augmentation.
      </p>
      <p>
        LOD works have come up with new approaches for
managing variety of sources, covering only some parts of the life
cycle phases, namely : (i) Conceptual level: the unification of the
universe of discourse is either ignored, or handled relying on
ad-hoc structures such as correspondence tables (using
similarity measures) [
        <xref ref-type="bibr" rid="ref15 ref6">6, 15</xref>
        ] or a shared ontology [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. (ii) Logical level:
most studies highlight multidimensional models and map LOD
sources to this logical format. These models are either generic
multidimensional models or ad-hoc models. Other studies
privilege LOD format based on the graph representation, for the
target DW [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5–7</xref>
        ]. (iii) ETL : variety management process was
handled using conventional ETL process that integrate and load
external sources to the DW [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. Other studies propose an
incremental fetching and storing of external sources on-demand,
i.e., as they are needed during the analysis process [
        <xref ref-type="bibr" rid="ref2 ref5">2, 5</xref>
        ]. In these
studies, a single scenario for integrating LOD is assumed, thus
variety is treated partially, in the sense that it has already been
dealt with in the initial DW. (iv) Physical level: proposed
scenario obliges designers to manage variety of LOD according
the physical implementations of DW: at unification formalism
level [
        <xref ref-type="bibr" rid="ref10 ref2 ref6">2, 6, 10</xref>
        ] or at querying level ([
        <xref ref-type="bibr" rid="ref11 ref13 ref16">11, 13, 16</xref>
        ].
      </p>
      <p>Contrary to existing studies, our approach proposes three main
contributions: (i) a conceptualization of variety in the presence
of internal and external LOD sources, (ii) it proposes diferent
scenarios inspired from the organizational level of a company
that decide to incorporate LOD within its internal sources. (iii)
Our approach analyzes then the impact of these variety scenarios
on the Value.
3
3.1</p>
    </sec>
    <sec id="sec-3">
      <title>VARIETY AND VALUE MANAGEMENT</title>
    </sec>
    <sec id="sec-4">
      <title>First V: Variety</title>
      <p>
        In this section, we provide a conceptualization of the ETL
environment covering the three scenarios proposed. The ETL processes
includes: (a) operators, (b) activities and (c) work-flow. By the
means of the WfMC1, we propose a metamodel to handle the
variety of internal and external sources (Figure 1). An ETL workflow
is the global collection of ETL activities and transitions between
1http://www.wfmc.org/
them. A transition determines the execution sequence of
activities to generate a data flow from sources to the target DW.
The ETL activities are defined at the conceptual level in order to
manage U OD. ETL operators manage instances that are stored
in the sources according to a defined format (relational, semantic,
etc). The graph model of LOD is used as a pivot model, and
internal sources are mapped to this format. A mapping efort from
internal sources within pivot model is needed. ETL operators
are redefined using the graph format of LOD, that we consider
as the elected pivot model. ETL process allow extracting the
instances, transform them and load them (ETLOperator Class). We
used the ten generic operators proposed in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] that we classify
into three groups: Source operators, Transform operators and Store
operators. In the next Section, we extend these operators to deal
with our scenarios. The set of operators of each group is defined
as enumerations in the model. An example of redefinition of
operators on LOD graph format is given:
Extract (G, Nj , CS): extracts, from G, the node Nj satisfying
constraint CS;
Context (G, Gc , Ctx ): extracts from G a sub-graph Gc that satisfy
the context defined in Ctx .
3.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>Second V: Value</title>
      <p>The construction of value-based DW is formalized as follows:
given: (i) a set of internal sources SI = {Si1, Si2, ...Sim }, (ii) a set
of external LOD sources : SE = {Se1, Se2, ...Sem }. Each internal
and external source has its own format FormatSj and its
conceptual model CMj describing its universe of speech. (iii) a set of
requirements G to be satisfied. (vi) A DW (to be defined or
operational) with its conceptual model CMDW describing its U OD
and one or more formats FormatDW ={ f1, f2, ... fk }. (v) The
valuerequirement fixed by the company. The added value of the target
warehouse (V alue) regarding a given value-requirement can be
calculated as follows:</p>
      <p>V alue =</p>
      <p>Õ
Si inSI ∪SE</p>
      <sec id="sec-5-1">
        <title>W eiдht (Si ) ∗ V alue(Si )</title>
        <p>(1)
where weiдht (Si ) describes the weight of each source and it can
be estimated for a given organizational sector.</p>
        <p>In our work, the value-requirement concerns the DW
requirement satisfaction which is strongly related to multidimensional
concepts and instances provided by sources. Therefore, we
propose three value metrics associated to each source Si : V alueReq,
V alueMD and V alueInst that are defined as follows, note that
these equations measure the percentage of value added from
external sources in terms of MD concepts and requirements to
be met (rate needs of table 1). :</p>
        <p>V alueReq(Si ) =</p>
      </sec>
      <sec id="sec-5-2">
        <title>Number ReponsesReq(Si )</title>
      </sec>
      <sec id="sec-5-3">
        <title>Number ReponsesReq(DW )</title>
        <p>where Number ReponsesReq(Si ) indicates the number of results
of the queries expressing the initial requirements on the source Si
and Number ReponsesReq(DW ) represents the number of results
of the queries expressing the requirements on the target DW.</p>
        <p>V alueMD(Si ) =</p>
      </sec>
      <sec id="sec-5-4">
        <title>Number _Concepts(Si )</title>
      </sec>
      <sec id="sec-5-5">
        <title>T otal Number _Concepts(DW )</title>
        <p>where Number _Concepts(Si ) is the number of multidimensional
concepts of DW schema by integrating the source i and
(2)
(3)</p>
        <p>T otal Number _Concepts(DW )) describes the total number of
multidimensional concepts of DW.</p>
        <p>V alueInst (Si ) =</p>
      </sec>
      <sec id="sec-5-6">
        <title>Number InstancesInt (Si )</title>
      </sec>
      <sec id="sec-5-7">
        <title>T otalIns(DW )</title>
        <p>(4)
where Number InstancesInt (Si ) and T otalIns(DW ) represent the
number of instances of DW by integrating the source Si and
the total number of instances of the DW.</p>
        <p>
          These metrics could be enriched by considering both
completeness and consistency. Several metrics to measure quality of
referential integrity exists [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. We could revisit them to measure
Value, combined with other Vs.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>OUR SCENARIOS</title>
      <p>Three main scenarios reflecting the policies of a company for
incorporating LOD are proposed :
4.1</p>
    </sec>
    <sec id="sec-7">
      <title>Serial Design</title>
      <p>It is feasible when a company decides to build its DW by
considering both internal and external sources from scratch. This
scenario follows a conventional DW design. LOD is
considered as a new semantic source to manage in addition to internal
sources. An ETL process is defined considering all the sources.
This scenario is not realistic because it requires redefinition of
the whole ETL process each time new requirements are needed
for extracting value from LOD.
4.2</p>
    </sec>
    <sec id="sec-8">
      <title>Parallel design</title>
      <p>This scenario assumes that the target DW is operational and
keeps integrating data from internal sources and LOD. The ETL
process from LOD is generated and then synchronized with
the initial ETL process from internal sources before the loading
step. This proposal describes the reaction of the parts of the ETL
process afected by a flow change. It requires the consolidation of
two parallel ETL flows (internal and external) which is needed to
keep the target warehouse up to date. We formalize the problem
of consolidation, by introducing the main design operation
synchronize. It corresponds to a synchronization between: (i) Current
lfow : existing ETL flows satisfying the n current information
requirements at t time and (ii) New flow : ETL flow satisfying the
upcoming requirement at (t+1) time. Synchronize corresponds
to a workflow pattern and corresponds to a series of operations
commonly encountered in workflow management: (i) AND-Join:
identify both ETL flows, apply potential deadlocks and perform
the join operation between concepts, (ii) OR-Join: corresponds to
a merger operation of concepts and properties done using Merge
operator and (iii) Clean: performs a data cleaning, checks null
values and deletes duplicate data before loading in the target
DW. Synchronize operator is defined using the graph format as
follows:
- Synchronize(G, Gi , Gj , CS): Synchronize two sub-graphs Gi and
Gj based on some criteria CS (AND-JOIN/OR-JOIN).
4.3</p>
    </sec>
    <sec id="sec-9">
      <title>Query-driven Design</title>
      <p>This scenario corresponds to the on-demand ETL to feed the
target DW. Here, data are incrementally fetched from
existing DW and LOD (in case where it is necessary), then loaded
into the DW only when they are needed to answer some cubes
queries. This scenario requires rewriting the cubes queries on
LOD, extract required fragments of LOD (using Context
operator), apply the transformations required (using the Class
TransformOperator) by mean on an ETL process dedicated to LOD and
then materialize (using Store operator) the resulting graphs in
case they are required later. The results are first integrated in a
data cube reserved for LOD data analysis, and final results of
queries are merged with the results of cubes queries executed on
the internal DW in order to display the query result to the end
user.</p>
      <p>This requires the extension of the ETL workflow meta model
by Query operator class (illustrated by red dots in Fig. 1). The
process of this scenario is conceptually illustrated in ETL meta
model by linking Query operator Class to the classes: Source
Operator (for extracting LOD of queries), Store Operator (for
materializing LOD) and Transform Operator in order to
handle the transformations required by the ETL process dedicated
to LOD (e.g., aggregation and join operations). We also added
the methods Rewrite_Query and MergeResult_Queries (unify the
results obtained) to the Query operator Class to manage the
different querying operations mentioned. Merge operator is defined
using the graph format as follows:
- Merдe(G, Gi , Gj ): merges two sub-graphs Gi and Gj into one
graph G.
5</p>
    </sec>
    <sec id="sec-10">
      <title>EXPERIMENTAL STUDY</title>
      <p>In this section, we carry out a set of experiments to show the
efectiveness of our proposal in terms of managing the variety
and augmenting the value of the final</p>
      <p>Experimental Setup. Let us consider the Film Academy Awards
organizations from four countries considered as internal data
sources to be integrated as follows: French Cesar awards (12,
123 004 triples), Deutscher Filmpreis (8 96 962 triples), India IIFA
Awards (15, 3,9 x 105 triples) and USA Oscar (19, 2,5 x 106 record
sets). Let us assume that these organizations collaborate to
globally analyze the cinematography industry. We considered a set of
(15) analytical requirements (eg. the popularity of an actor/actress
by year). The first three data sources are implemented on Oracle
semantic DBMS using N-Quads format, while the fourth one on
a traditional (Relational) Oracle DBMS. Our external resource
corresponds to a fragment of DBpedia extracted using the
context operator applied to Movies. The obtained fragment contains
around 7,9 x 106 graph Quads. Our evaluations of the ETL
processes according to the diferent scenarios were performed on a
laptop computer (HP Elite-Book 840 G3) with an Intel(R) CoreTM
i7-6500U CPU 2.59 GHZ and 8 GB of RAM and a 1 TB hard disk.
We use Windows10 64bits.</p>
      <p>
        Variety Evaluation. The purpose of this evaluation is to
study the impact of the eforts in conceptualizing variety in the
obtained warehouse. In the first experiment, we compare the
impact of our elected graph model against a pivot metamodel
proposed in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] (called graph property pivot model) in terms of
concepts, attributes, relationships and instances. To conduct this
experiment, we considered our three scenarios. Figures 2a and
2b summarize the obtained results by averaging the number of
elements per scenario. It clearly shows that our elected graph
model captures more elements than the pivot metamodel. This
is because all elements satisfying the requirements of the LOD
fragment are 100% materialized in the warehouse.
      </p>
      <p>Added Value Evaluation. The second experiment was
conducted to measure the value captured by the obtained DW.
We use two criteria representing the rate of multidimensional
concepts and integrated data and the rate of satisfied
requirements during the integration of LOD. Figure 3a illustrates the
obtained results. The consideration of LOD fragment increases
the number of multidimensional concepts for the three scenarios.
By comparing the three scenarios, we figure out that they are
almost equivalent.</p>
      <p>To evaluate the second criterion, we have formulated our user
requirements in the form of cubes queries executed once on the
target DW. The execution of the cubes queries was carried out
in four time stages (t0, t1, t2 and t3) during the integration
process. The time t0 corresponds to the time of considering LOD
in addition to internal sources. Figure 3b describes the obtained
results that demonstrate that before taking into account LOD in
the integration process, the user requirements that are satisfied
by internal data sources represent ∼ 65%. Once LOD integration
process has begun, we remark that this rate increases
considerably until reaching a maximum rate of 96%. We also noticed that
the third scenario (Query driven design) gives the best result and
meets the user needs faster than the other scenarios. This can be
explained by the fact that this scenario focuses on integrating
data that correspond to specific queries reflecting user needs.</p>
      <p>
        Table 1 extends the above results and demonstrates the value
added by considering LOD in the design of DW. A comparison
is given between our proposal considering the three scenarios
and previous work [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], on the basis of some criteria identified
during the experimentation. These results clearly indicate that
the consideration of LOD data ofers a value-added in terms of
the final number of dimensions (Dim) and measures (Meas), the
size of the target DW and rates of satisfied requirements than
classical approaches.
      </p>
      <p>Scenarios
Internal Sources
Serial Design
Parallel Design
Query driven design</p>
      <p>Dim/Meas Rate needs Input Size Response time
6/1 60% 550 x 103 1.1
10/7 80% 7.9 x 106 3.2
11/8 84% 3.1 x 106 2.6
12/8 96% 2.9 x 106 1.7</p>
      <p>The evaluation of scenario presented in the example was made
by users who evaluated 80% of concepts and instances in total.
They judged that 97% of the concepts and 91% of the instances
in the sample are correct and 96% of requirements are satisfied.
Additionally, adding LOD only takes 1% of additional time (1-3
secs), which we may consider as fast (cf. Table 1).
6</p>
    </sec>
    <sec id="sec-11">
      <title>CONCLUSION</title>
      <p>In this paper, we show that the DW technology is still alive and
the advances brought by Big Data and the explosion of LOD
contribute to its renaissance. We consider two V’s related to value
and variety. LOD are viewed as external sources that contribute
to increasing the variety of the whole sources. We formalized the
ETL environment and proposed a pivot meta model ETL
worklfow. Based on modeling and meta-modeling eforts, the variety
is nicely managed. Regarding value of the target DW, we have
defined metrics associated to initial requirement satisfaction. Our
scenarios for integrating LOD into the DW design have been
proposed. They are: (i) LOD and internal sources are physically
materialized in the DW and (ii) both LOD and DW query
results are merged. Our validation showed that adding LOD
increases the value of the target warehouse represented by the
initial user satisfaction. We also realized that the incorporation
of the LOD is inexpensive in terms of development and fast
thanks to our variety management.</p>
      <p>We are currently working on the development of a CASE tool
that deals with internal and external sources. Another issues
concerns the risk of integrating the LOD in a DW and detecting
(a) Conceptual elements integrated using Graph LOD and Graph Property.
(b) Instances integrated using Graph LOD and Graph Property.
(a) Number of multidimensional concepts vs. integrated L O D.
(b) Rate of satisfied requirement during the integration of
L O D.
and repairing inconsistency and incompleteness in external data
sources by considering other V’s: Veracity .</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Abelló Gamazo</surname>
          </string-name>
          , E. Gallinucci,
          <string-name>
            <given-names>M.</given-names>
            <surname>Golfarelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Rizzi</given-names>
            <surname>Bach</surname>
          </string-name>
          , and
          <string-name>
            <given-names>O. Romero</given-names>
            <surname>Moral</surname>
          </string-name>
          .
          <article-title>Towards exploratory olap on linked data</article-title>
          .
          <source>In SEBD</source>
          , pages
          <fpage>86</fpage>
          -
          <lpage>93</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Baldacci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Golfarelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Graziani</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          .
          <article-title>Qetl: An approach to on-demand etl from non-owned data sources</article-title>
          .
          <source>DKE</source>
          ,
          <volume>112</volume>
          :
          <fpage>17</fpage>
          -
          <lpage>37</lpage>
          ,
          <year>Nov 2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Berkani</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Bellatreche</surname>
          </string-name>
          .
          <article-title>A variety-sensitive ETL processes</article-title>
          .
          <source>In DEXA (2)</source>
          , pages
          <fpage>201</fpage>
          -
          <lpage>216</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>N.</given-names>
            <surname>Berkani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bellatreche</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Khouri</surname>
          </string-name>
          .
          <article-title>Towards a conceptualization of ETL and physical storage of semantic data warehouses as a service</article-title>
          .
          <source>Cluster Computing</source>
          ,
          <volume>16</volume>
          (
          <issue>4</issue>
          ):
          <fpage>915</fpage>
          -
          <lpage>931</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Berro</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Megdiche</surname>
          </string-name>
          , and
          <string-name>
            <given-names>O.</given-names>
            <surname>Teste</surname>
          </string-name>
          .
          <article-title>Graph-based ETL processes for warehousing statistical open data</article-title>
          .
          <source>In ICEIS 2015</source>
          , pages
          <fpage>271</fpage>
          -
          <lpage>278</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R. P.</given-names>
            <surname>Deb Nath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hose</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T. B.</given-names>
            <surname>Pedersen</surname>
          </string-name>
          .
          <article-title>Towards a programmable semantic extract-transform-load framework for semantic data warehouses</article-title>
          .
          <source>In DOLAP</source>
          , pages
          <fpage>15</fpage>
          -
          <lpage>24</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>Etcheverry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaisman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Zimányi</surname>
          </string-name>
          .
          <article-title>Modeling and querying data warehouses on the semantic web using qb4olap</article-title>
          .
          <source>In DaWAK</source>
          , pages
          <fpage>45</fpage>
          -
          <lpage>56</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Giorgini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Garzetti</surname>
          </string-name>
          .
          <article-title>Goal-oriented requirement analysis for data warehouse design</article-title>
          .
          <source>In ACM DOLAP</source>
          , pages
          <fpage>47</fpage>
          -
          <lpage>56</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Golfarelli</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          .
          <article-title>A survey on temporal data warehousing</article-title>
          .
          <source>IJDWM</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kämpgen</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Harth</surname>
          </string-name>
          .
          <article-title>Transforming statistical linked data for use in OLAP systems</article-title>
          .
          <source>In I-SEMANTICS</source>
          , pages
          <fpage>33</fpage>
          -
          <lpage>40</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kämpgen</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>O'Riain, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Harth</surname>
          </string-name>
          .
          <article-title>Interacting with statistical linked data via OLAP operations</article-title>
          .
          <source>In ESWC</source>
          , pages
          <fpage>87</fpage>
          -
          <lpage>101</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>N.</given-names>
            <surname>Konstantinou</surname>
          </string-name>
          and et al.
          <article-title>The VADA architecture for cost-efective data wrangling</article-title>
          .
          <source>In SIGMOD</source>
          , pages
          <fpage>1599</fpage>
          -
          <lpage>1602</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Matei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chao</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Godwin</surname>
          </string-name>
          .
          <article-title>OLAP for multidimensional semantic web databases</article-title>
          .
          <source>In BIRTE</source>
          , pages
          <fpage>81</fpage>
          -
          <lpage>96</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ordonez</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>García-García</surname>
          </string-name>
          .
          <article-title>Referential integrity quality metrics</article-title>
          .
          <source>Decision Support Systems</source>
          ,
          <volume>44</volume>
          (
          <issue>2</issue>
          ):
          <fpage>495</fpage>
          -
          <lpage>508</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>F.</given-names>
            <surname>Ravat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Song</surname>
          </string-name>
          , and
          <string-name>
            <given-names>O.</given-names>
            <surname>Teste</surname>
          </string-name>
          .
          <article-title>Designing multidimensional cubes from warehoused data and linked open data</article-title>
          .
          <source>In RCIS</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>R.</given-names>
            <surname>Saad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Teste</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Trojahn</surname>
          </string-name>
          .
          <article-title>Olap manipulations on rdf data following a constellation model</article-title>
          . In Workshop on Semantic Statistics,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>T. P.</given-names>
            <surname>Sales</surname>
          </string-name>
          and et al.
          <article-title>The common ontology of value and risk</article-title>
          .
          <source>In ER</source>
          , pages
          <fpage>121</fpage>
          -
          <lpage>135</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>Skoutas</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Simitsis</surname>
          </string-name>
          .
          <article-title>Ontology-based conceptual design of ETL processes for both structured and semi-structured data</article-title>
          .
          <source>Semantic Web</source>
          ,
          <volume>3</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>