<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>To Each His Own: Accommodating Data Variety by a Multimodel Star Schema</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sandro Bimonte</string-name>
          <email>sandro.bimonte@inrae.fr</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yassine Hifdi</string-name>
          <email>hifdiyassine@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohammed Maliari</string-name>
          <email>mohammedmaliari@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrick Marcel</string-name>
          <email>Patrick.Marcel@univ-tours.fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Rizzi</string-name>
          <email>stefano.rizzi@unibo.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DISI, University of Bologna</institution>
          ,
          <addr-line>Bologna</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ENSA</institution>
          ,
          <addr-line>Tangier, Maroc</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LIFAT Laboratory, University Tours</institution>
          ,
          <addr-line>Blois</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University Clermont</institution>
          ,
          <addr-line>TSCF, INRAE, Aubiere</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Recent approaches adopt multimodel databases (MMDBs) to natively handle the variety issues arising from the increasing amounts of heterogeneous data (structured, semi-structured, graphbased, etc.) made available. However, when it comes to analyzing these data, traditional data warehouses (DWs) and OLAP systems fall short because they rely on relational Database Management Systems (DBMSs) for storage and querying, thus constraining data variety into the rigidity of a structured schema. This paper provides a preliminary investigation of the performance of an MMDB when used to store multidimensional data for OLAP analysis. A multimodel DW would store each of its elements according to its native model; among the benefits we envision for this solution, that of bridging the architectural gap between data lakes and DWs, that of reducing the cost for ETL data transformations, and that of ensuring better flexibility, extensibility, and evolvability thanks to the use of schemaless models. To support our investigation we present an implementation, based on the UniBench benchmark dataset, that extends a star schema with JSON, XML, spatial, and key-value data; we also define a sample OLAP workload and use it to test the performance of our solution and compare it with that of a classical star schema. As expected, the full-relational implementation performs better, but we believe that this gap could be balanced by the benefits of multimodel in dealing with variety. Finally, we give our perspective view of the research on this topic.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Big Data is notoriously characterized by (at least) the 3 V’s:
volume, velocity, and variety. To handle velocity and volume, some
distributed file system-based storage (such as Hadoop) and new
Database Management Systems (DBMSs) have been proposed.
In particular, four main categories of NoSQL databases have
been proposed [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]: key-value, extensible record, graph-based,
and document-based.
      </p>
      <p>
        Although NoSQL DBMSs have successfully proved to support
the volume and velocity features, variety is still a challenge [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
Indeed, several practical applications (e.g. retail, agriculture, etc.)
ask for collecting and analyzing data of diferent types: structured
(e.g., relational tables), semi-structured (e.g., XML and JSON), and
unstructured (such as text, images, etc.). Using the right DBMS
for the right data type is essential to grant good storage and
analysis performance. Traditionally, each DBMS has been conceived
for handling a specific data type; for example, relational DBMSs
for structured data, document-based DBMSs for semi-structured
data, etc. Therefore, when an application requires diferent data
types, two solutions are actually possible: (i) integrating all data
into a single DBMS, or (ii) using two or more DBMSs together.
The former solution presents serious drawbacks: first of all, some
types of data cannot be stored and analyzed (e.g., the pure
relational model does not support the storage of images, XML,
arrays, etc. [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]); besides, even when data can be converted and
stored in the target DBMS, querying performances could be
unsatisfactory. The latter approach (known as polyglot persistence
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]) presents important challenges as well, namely, technically
managing more DBMSs, complex query languages, inadequate
performance optimization, etc. Therefore, Multimodel databases
(MMDBs) have recently been proposed to overcome these issues.
A MMDB is a DBMS that natively supports diferent data types
under a single query language to grant performance, scalability,
and fault tolerance [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Remarkably, using a single platform for
multimodel data promises to deliver several benefits to users
besides that of providing a unified query interface; namely, it will
simplify query operations, reduce development and maintenance
issues, speed up development, and eliminate migration problems
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Examples of MMDBs are PostgreSQL and ArangoDB.
PostgreSQL supports the row-oriented, column-oriented, key-value,
and document-oriented data models, ofering XML, HSTORE,
JSON/JSONB data types for storage. ArangoDB supports the
graph-based, key-value, and document-oriented data models.
      </p>
      <p>
        Handling variety while granting at the same time volume and
velocity is even more complex in Data Warehouses (DWs) and
OLAP systems. Indeed, warehoused data result from the
integration of huge volumes of heterogeneous data, and OLAP requires
very good performances for data-intensive analytical queries [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
Traditional DW architectures rely on a single, relational DBMS
for storage and querying1. To ofer better support to volume
while maintaining velocity, some recent works propose the usage
of NoSQL DBMSs; for example, [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] relies on a document-based
DBMS, and [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] on a column-based DBMS. NoSQL proposals for
DWs are based on a single data model, and all data are
transformed to fit with that model (document, graph, etc.). Overall,
although these approaches ofer interesting results in terms of
volume and velocity, they have been mainly conceived and tested
for structured data, without taking into account variety.
1More precisely, this is true for so-called ROLAP architectures. In MOLAP
architectures, data are stored in multidimensional arrays. Finally, in HOLAP architectures,
a MOLAP and a ROLAP systems are coupled.
      </p>
      <p>
        Furthermore, to facilitate OLAP querying, DWs are normally
based on the multidimensional model, which introduces the
concepts of facts, dimensions, and measures to analyze data, so source
data must be forcibly transformed to fit a multidimensional
logical schema following a so-called schema-on-write approach. Since
this is not always painless because of the schemaless nature of
some source data, some recent work (such as [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]) propose to
directly rewrite OLAP queries over document stores that are not
organized according to the multidimensional model, following a
schema-on-read approach (i.e., the multidimensional schema is
not decided at design time and forced in a DW, but decided by
each single user at querying time). However, even this approach
relies on a single DBMS.
      </p>
      <p>An interesting direction towards a solution for efectively
handling the 3 V’s in DW and OLAP systems is represented by
MMDBs. A multimodel data warehouse (MMDW) can store data
according to the multidimensional model and, at the same time,
let each of its elements be natively represented through the most
appropriate model. Among the benefits we envision for MMDWs,
that of bridging the architectural gap between data lakes and
DWs, that of reducing the cost for ETL data transformations, and
that of ensuring better flexibility, extensibility, and evolvability
thanks to the use of schemaless models.</p>
      <p>In this paper we conduct a preliminary investigation of the
performance of MMDWs to store multidimensional data. To this
end we introduce a logical schema for MMDWs and its
implementation on PostgreSQL, which gives native multimodel support.
Our schema extends the classical star schema introducing
semistructured (JSON, XML, and key-value) data in all the
multidimensional elements; thus, it goes in the direction of coupling the
pros of schema-on-write approaches (mainly, good performances
and simple query formulation with no need for query rewriting)
with those of schema-on-read approaches (higher flexibility in
ad-hoc querying).</p>
      <p>
        Due to the lack of a benchmark for multimodel data warehouse,
in this paper we propose our own OLAP workload to evaluate
the performance of our proposal, which we also test against a
full-relational implementation on PostgreSQL. To the best of our
knowledge, no benchmark dataset for DW (either relational or
NoSQL) supports variety; thus, for the experiments we use the
schema and data provided by UniBench [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], a benchmark for
MMDBs that well represents variety.
      </p>
      <p>The paper outline is as follows. After discussing the related
literature in Section 2, in Section 3 we present the UniBench case
study. Sections 4 and 5 introduce our logical schema for MMDWs
and the related OLAP workload, respectively. Section 6 shows
the results of the experiments we made, while Section 7 presents
our vision of future MMDW research. Finally, in Section 8 we
draw the conclusions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        Some recent work concerns warehousing and OLAP using NoSQL
DBMSs of diferent kinds. In [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], three diferent logical
models are proposed, using 1 or N document collections to store
data in document-based DBMSs and highlighting the utility of
nested document and array types [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The same authors also
investigate how to handle complex hierarchies and
summarizability issues with document-based DWs [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The introduction
of spatial data in document-based DWs has been discussed in
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which proposes a new spatial multidimensional model to
avoid redundancy of spatial data and improve performances. A
Key-Value
Ranking and
feedback
      </p>
      <p>Graph
Social
networks</p>
      <p>Relational
Customers
RegUsers</p>
      <p>Vendors</p>
      <p>Orders
JSON</p>
      <p>
        Invoices
Products
logical model for column-based DWs has been proposed by [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
and [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to address volume scalability. In [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], transformation
rules for DW implementation in graph-based DBMSs have been
proposed for better handling social network data. To the best of
our knowledge, only [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] presents a benchmark for comparing
NoSQL DW proposals; specifically, this benchmark is applied to
MongoDB and Hbase. Some works also study the usage of XML
DBMSs for warehousing XML data [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Although XML DWs
represent a first efort towards native storage of semi-structured
data, their querying performances do not scale well with size,
and compression techniques must be adopted [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>Among all these proposals, it is hard to champion one logical
and physical implementation for NoSQL and XML DWs, since
no approach clearly outperforms the other on the 3 V’s.
Moreover, these single-model proposals do not address other issues
related to warehousing big data, such as reducing the cost of ETL,
evolution and improving flexibility.</p>
      <p>
        Recently, some approaches to execute OLAP queries directly
against NoSQL data sources were proposed. In [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], a
schemaon-read approach to automatically extract facts and hierarchies
from document data stores and trigger OLAP queries is proposed.
A similar approach is presented in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]; there, schema variety is
explicitly taken into account by choosing not to design a single
crisp schema where source fields are either included or absent,
but rather to enable an OLAP experience on some sort of “soft”
schema where each source field is present to some extent. In
the same direction, [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] proposes a MapReduce-based algorithm
to compute OLAP cubes on column stores, while [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] aims at
delivering the OLAP experience over a graph-based database.
      </p>
      <p>
        The approaches mentioned above rely on a single-model
database. Conversely, [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] proposes a pay-as-you-go approach which
enables OLAP queries against a polystore supporting relational,
document, and column data models by hiding heterogeneity
behind a dataspace layer. Data integration is carried out on-the-fly
using a set of mappings. Even this approach can be classified as
schema-on-read; the focus is on query rewriting against
heterogeneous databases and not on the performances of the approach.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>CASE STUDY: UNIBENCH</title>
      <p>
        UniBench is a benchmark for multimodel databases proposed
in [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. It includes a retail dataset composed of relational, XML,
JSON, key-value, and graph data as shown in Figure 1, which
makes it a good representative for variety. However, UniBench
was not conceived for OLAP queries. Since our goal is to
handle variety with specific reference to DWs, we had to derive a
multidimensional schema from UniBench. This adaptation
required some modifications, including the addition of descriptive
attributes (e.g., LastName), which allows to better test the
effectiveness of the proposed approach; as a consequence, some
• A Time dimension with levels Day, Month, and Year.
• A Product dimension with one hierarchy including level
Store and some descriptive attributes (e.g., Vendor).
Interestingly, stores are described by a spatial level, Location.
The cloud symbol in the schema denotes that a product can
have some additional descriptive attributes not specified
at design time.
• A Costumer dimension in which two hierarchies are rooted:
one with level Gender, one with UsedBrowser. The
customer also has some descriptive attributes, e.g., LastName.
Attribute Rating is cross-dimensional, i.e., its value is jointly
determined by Product and Customer (a customer can rate several
products). The fact has one measure, TotalPrice.
      </p>
      <p>Finally, since an order is associated to many products, a
manyto-many relationship is set between the fact and the product
dimension (non-strict hierarchy).2
4</p>
    </sec>
    <sec id="sec-4">
      <title>A MULTIMODEL STAR SCHEMA FOR</title>
    </sec>
    <sec id="sec-5">
      <title>UNIBENCH</title>
      <p>In this section we present a MultiModel, MultiDimensional (in
short, M3D) logical schema for the Order fact introduced above.
Essentially, we use a classical star schema with a fact and
dimension tables, extended with semi-structured data in JSON and
XML form, and with spatial data. Starting from a star schema
has several clear advantages: (i) the star schema is supported
by all OLAP servers and already in use in a huge number of
enterprise DWs; (ii) the best practices for designing a star schema
from a conceptual schema are well understood and commonly
adopted by practitioners; (iii) fact-dimension relationships are
ruled by foreign keys so their consistence is natively checked
by the DBMS; (iv) performance optimization of star schema has
been long studied and practiced at both the logical (e.g., via view
materialization) and the physical (e.g., via indexing) level.</p>
      <p>
        Clearly, several possible alternatives arise for modeling the
Order fact with an extended star schema. Defining a set of best
practices for designing an M3D schema that achieves the best
trade-of between the five advantages listed in Section 1 is out of
the scope of this paper; so, we opted for designing the schema
based on a simple guideline: preserve as much as possible the
source data variety, i.e., minimizing the transformations to be
applied to UniBench source data. Figure 3 shows the M3D schema
2We have not considered the graph data of UniBench, since the PostgreSQL DBMS
used for implementation does not support them natively.
that results from applying this guideline to the conceptual schema
in Figure 2. It can be described as follows:
• The fact table, Fact_Order, has one tuple for each order
and references the order customer and date via foreign
keys. Each tuple includes a JSON document that stores
the totalPrice measure and an array of orderlines, each
specifying a product.
• The customer dimension table, Dim_Customer, specifies
each customer’s data in the form of XML documents.
• The temporal dimension table, Dim_Date, stores in each
tuple a JSON document with the order date; to enable
useful aggregations, it also stores the corresponding month
and year.
• The product dimension table, Dim_Product, for each
product stores its location (as a spatial attribute), vendor, and
store, as well as a JSON document with the product name
(title), price, and image. Each product also has a Feedback
attribute that stores all its ratings in key-value form, with
the customer code as a key.
• As shown in Figure 2, each order refers to several
products. To model this non-strict hierarchy, rather than
opting for the classical relational solution (a many-to-many
bridge table [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]), we established a connection between
the InfoOrder document stored in the fact table and the
Dim_Product dimension table via the asin attribute.
An example of instances of the fact table and of the product
dimension table are shown in Figure 4.
      </p>
      <p>The cloud symbol in Figure 2 denotes that the product
dimension can include some additional attributes not specified
at design time (hence, not included in the JSON schema). For
instance, some InfoPrdt documents will have an EU attribute
precising the category of product according to the EU
classification (see Figure 5), while some InfoOrder documents will have a
brand attribute.
5</p>
    </sec>
    <sec id="sec-6">
      <title>AN OLAP WORKLOAD FOR UNIBENCH</title>
      <p>
        The workload we introduce to test our M3D schema is inspired by
that of the classical SSB benchmark [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], itself loosely based on
the TPC-H benchmark. The SSB workload is meant to
functionally cover the diferent types of star schema queries while varying
fact table selectivity. SSB queries are organized in 4 flights, where
each flight is a list of 3 to 4 queries. Query flight 1 has restrictions
on only 1 dimension, flight 2 has restrictions on 2 dimensions,
lfight 3 on 3, and flight 4 represents a what-if sequence of the
OLAP type. We adopt the same approach, while at the same time
classifying queries according to the usage of relational (R)/non
relational (NR) measures, relational/non relational group-by
levels, and relational/non relational selection levels. We also add
a parameter representing the type of join: relational means a
join using two relational attributes, while JSON means a join
between a JSON attribute and a relational one. Q10 and Q12 use
selection attributes that are not part of the JSON schema (brand
and EU, respectively). In Table 1, which presents the workload,
“by” introduces a group-by and “for” a selection.
6
      </p>
    </sec>
    <sec id="sec-7">
      <title>MULTIMODEL VS. FULL-RELATIONAL</title>
      <p>In this section we give a preliminary assessment of the
efectiveness and eficiency of MMDWs as compared to those of a
classical relational implementation.
"type": "object",
"properties":
{ "orderDate":
{ "type": "string" }}}</p>
      <sec id="sec-7-1">
        <title>Dim_Product &lt;PK&gt; Asin : String</title>
      </sec>
      <sec id="sec-7-2">
        <title>Feedback : Hstore</title>
      </sec>
      <sec id="sec-7-3">
        <title>InfoPrdt : JSON</title>
        <sec id="sec-7-3-1">
          <title>Location : Geo</title>
        </sec>
      </sec>
      <sec id="sec-7-4">
        <title>Store : String</title>
        <sec id="sec-7-4-1">
          <title>Vendor : String</title>
          <p>InfoPrdt
"type": "object",
"properties":
{ "title": { "type": "string" },
"price": { "type": "number" },
"imhUrl": { "type": "string" }}}}</p>
        </sec>
      </sec>
      <sec id="sec-7-5">
        <title>Dim_Date</title>
      </sec>
      <sec id="sec-7-6">
        <title>Date</title>
      </sec>
      <sec id="sec-7-7">
        <title>Year</title>
        <p>: JSON
: Date</p>
      </sec>
      <sec id="sec-7-8">
        <title>Month : Date</title>
      </sec>
      <sec id="sec-7-9">
        <title>Fact_Order</title>
        <p>&lt;PK&gt; Id_Order : String
&lt;FK&gt; Id_Cust
: String
&lt;FK&gt; Id_Date : Int</p>
        <p>InfoOrder : JSON
"type": "object",
"properties":
{ "totalPrice":
{ "type": "number" },
"orderLine":
{ "type": "array",
"items":
{ "id": "0",
"type": "object",
"properties":
{ "productId": { "type": "string" },
"asin": { "type": "string" },
"title": { "type": "string" }}}}}}</p>
      </sec>
      <sec id="sec-7-10">
        <title>Dim_Customer</title>
        <p>&lt;PK&gt; Id_Cust : String</p>
        <p>InfoCust : XML
&lt;!ELEMENT row
(firstName,lastName,gender,birthday,creationDate,
browserUsed)&gt;
&lt;!ATTLIST row gender (male | female) &gt;
&lt;!ELEMENT firstName (#PCDATA)&gt;
&lt;!ELEMENT lastName (#PCDATA)&gt;
&lt;!ELEMENT gender (#PCDATA)&gt;
&lt;!ELEMENT birthday (#PCDATA)&gt;
&lt;!ELEMENT creationDate (#PCDATA)&gt;
&lt;!ELEMENT browserUsed (#PCDATA)&gt;
Dim prod WITH EU
and 640,000 orders (|Fact_Order|).</p>
        <p>Query Numb. of selections NR types
Number of orders by months forQg2iv:en months and years 2 —
Number of orders by months forsegilveecnt sctooruensta(nddisyteinarcst o.Id_Order) as Num2berOfOrders, d.MonJSthON
Number of orders by months forfrgoivmen(yesaerlseacntdIdr_atOinrgder, products-&gt;&gt;'asin2’ as Asin JSON,key-value
Number of orders by months, gender forfgriovmenFyaecarts_Order o, 1 XML
Total price by year for given stores jsonb_array_elements(o.InfoOrd11er-&gt;'orderLine') asJSpOrJoSNOd,XNuMctLs
Number of orders by months, gender for given products
Total price by year for given genders ) as op, Fact_Order o, Dim_Date d,1 JSON,XML
Total price by year, gender for given s(torseeslaencdt ysekaeryss(p.Feedback) as Id_Cu2st, Asin, svals(p.FeeJdSbOaNck) as Rating
Total price by year of birth for given browfrsoemrs aDnidmg_ePnrdoedrsuct p 2 XML
Total price by date, customer for given)maosncthps, ratings, stores 3 JSON,key-value
Total price by date for given mownthhse,rgeendd.Yeersa,rb=r'a2n0d2s0' and Rating&gt;4 and o3.Id_Date=d.Id_DatJeSON,XML
Total price by date, customer foragnidveon.mIdo_nCtuhss,tg=ecnpd.eIdrs_,Crautsintgasnd o.Id_Orde3r=op.Id_OrdeJrSaOnNd,XcMp.LA,kseiyn-=voalpu.eAsin
Total price by costumer for givengrEoUupvabluyesd.Month 1 JSON,XML
Q2:
select count(distinct o.Id_Order) as NumberOfOrders, d.Month
from ( select Id_Order, products-&gt;&gt;'asin’ as Asin
from Fact_Order o,
jsonb_array_elements(o.InfoOrder-&gt;'orderLine') as products
) as op, Fact_Order o, Dim_Date d,
( select skeys(p.Feedback) as Id_Cust, Asin, svals(p.Feedback) as Rating
from Dim_Product p
) as cp
where d.Year='2020' and Rating&gt;4 and o.Id_Date=d.Id_Date
and o.Id_Cust=cp.Id_Cust and o.Id_Order=op.Id_Order and cp.Asin=op.Asin
group by d.Month
FQi2g:ure 6: SQL formulation of query Q2 in PostgreSQL over
selecMtco3unt(distinct Id_Order) as NumberOfOrders, d.Month
the D schema
from Fact_Order o, Dim_Date d, Bridge_Cust_Prod cp, Bridge_Ord_Prod op
where d.Year='2020' and Rating&gt;4 and o.Id_Date=d.Id_Date
and o.Id_Cust=cp.Id_Cust and o.Id_Order=op.Id_Order and cp.Asin=op.Asin
group by d.Month</p>
        <p>All the OLAP queries proposed in Section 5 have been
successfully formulated and executed over the M3D schema, which
confirms the feasibility of using PostgreSQL as a platform for
storing and querying MMDWs. Figure 6 shows the SQL formulation
of a sample query in PostgreSQL; note that attributes Feedback
of type key-value and InfoOrder of type JSON are retrieved as
table views to be used for a join or a selection.</p>
        <p>In the following, we present some experiments aimed at
quantitatively comparing the querying performances of the M3D
schema and those of a full-relational star schema (from now
on, FR). For the FR schema we used two bridge tables as shown
in Figure 8. The first one, Bridge_Ord_Prod, stores the
manyto-many relationship between an order and its products. The
second one, Bridge_Cust_Prod, is necessary to store the Rating
cross-dimensional attribute. Noticeably, attributes EU and brand
are not included here since, as explained in Section 4, they were
not known at design time. Clearly, unless some (costly) evolution
of the schema is carried out, these attributes cannot be loaded
and they cannot be used for querying. The FR schema is also
implemented in PostgreSQL; Figure 7 shows the SQL formulation
of query Q2 over the FR schema. A comparison between Figures
6 and 7 suggests that the formulation over the M3D schema is
more complex; however, we wish to emphasize that there is no
real dificulty in formulating queries on an MMDW in
comparison to a traditional star schema, except that some knowledge
of the DBMS-specific operators to manipulate key-value, JSON,
and XML types is required.</p>
        <p>For both implementations, B+trees have been used to index
relational attributes. For the M3D schema, some tests were done
in order to find the best optimization plan for the workload
queries. The results we report below use the following: (i) a Gist
index is used on the Feedback hstore attribute; (ii) B+trees and
Gin indexes are used on JSON attributes. All tests have been run</p>
        <p>Dim_Date
&lt;PK&gt; Id_Date : Int</p>
        <p>Date : Date
Month : Date
Year : Date</p>
        <p>Fact_Order
&lt;PK&gt; Id_Order : String
&lt;FK&gt; Id_Cust : String
&lt;FK&gt; Id_Date : Int</p>
        <p>TotalPrice : Double</p>
        <p>Bridge_Cust_Prod
&lt;PK,FK&gt; Id_Cust : String
&lt;PK,FK&gt; Asin : String</p>
        <p>Rating : Int
Query
Q0
Q1
Q2
Q3
Q4
Q5
Q6
Q7
Q8
Q9
Q10
Q11
Q12
on a Core i5 with 4 CPUs @2.3GHz laptop with 16 GB RAM and
SSD running MacOS Mojave.</p>
        <p>Table 2 shows the query execution in milliseconds against
both implementations. Note that Q10 and Q12 cannot be
executed on the FR schema because they use attributes (brand and
EU, respectively) that were not known at design time so they
are not part of that schema. Not surprisingly, the full-relational
implementation outperforms the multimodel implementation
over most queries. This can partly be explained by recalling that
Q2:
select count(distinct Id_Order) as NumberOfOrders, d.Month
from Fact_Order o, Dim_Date d, Bridge_Cust_Prod cp, Bridge_Ord_Prod op
where d.Year='2020' and Rating&gt;4 and o.Id_Date=d.Id_Date
and o.Id_Cust=cp.Id_Cust and o.Id_Order=op.Id_Order and cp.Asin=op.Asin
group by d.Month
PostgreSQL was originally born as a relational DBMS, so
semistructured and complex data querying is not fully optimized yet.</p>
        <p>In particular, PostgreSQL lacks specific optimization structures
adapted to XML data, thus, the InfoCust attribute cannot be
properly indexed; this impacts queries Q3, Q4, Q6, Q8, and Q10. M3D
is also penalized by the necessity to have a JSON attribute in the
fact table to be joined with a dimension table (namely, InfoOrder).</p>
        <p>Additionally, the fact table in M3D is quite larger than the one in
the FR schema, which results in slower star joins (even using the
JSONB type instead of JSON, the improvement is very small). The
only case where the multimodel implementation significantly
outperforms the full-relational one is Q9; this is due to the use of
a bridge table in the relational implementation and specifically
to the fact that, despite the presence of indexes, the optimizer
uses sequential scan to access the bridge table.</p>
        <p>Table 3 shows the storage size of both implementations.
Unsurprisingly, the relational implementation is more sober than
the multimodel one.</p>
        <p>Though devising complete guidelines and best practices for
multimodel design is out of the scope of this paper, we observe
that:
(1) the relational model is still more eficient, so it should be
used, during logical design, for the data sources that can
be smoothly transformed into relational form (i.e., those
whose transformation does not entail loss of information
content and can be accommodated within the time frame
of ETL);
(2) conversely, the data sources that hardly fit into the fixed
structure of a relational schema, e.g., because their schema
is not completely known in advance, should be left in their
native form.
6.2</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Efectiveness</title>
      <p>
        In this section we provide a qualitative comparison of the two
solutions in terms of efectiveness from three points of view:
• Transformation. The full-relational implementation required
all the UniBench data to be translated in relational form
according to the star schema in Figure 8. While in the M3D
schema the dimension and fact tables are fed with JSON
data with simple INSERT queries, in the FR schema more
steps are required. For instance, just to feed the bridge
table using the ETL Talend tool we need (i) a job for
reading the JSON collection (tFileInputJSOn); (ii) a loop JSON
query to read the array of products of each InfoOrder
document; (iii) a job for reading the Dim_Product dimension
table; and finally (iv) a join operation. This means that
transformations may require a significant time and can be
error-prone, so they may be unsuitable in specific settings
such as those of real-time DWs.
• Flexibility. Diferently from the FR schema, the M 3D one
preserves the data variety existing in the data sources. This
is particularly relevant for instance in self-service
business intelligence scenarios, where data scientist will write
ad-hoc queries to satisfy situational analysis needs [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Besides, mixing diferent models allows, in an MMDW, to
achieve higher flexibility in the modeling solutions taken,
for instance when dealing with many-to-many
relationships.
• Evolution. While the multimodel implementation is
partially schemaless, so it inherently supports evolution, the
situation with the full-relational implementation is quite
diferent. In fact, even adding a couple of simple levels
(as EU and brand in our case study) requires, at the very
least, changing the relational schema of one or more tables,
editing the ETL procedures, and migrating the data from
the old schema to the new one. A more complex
evolution, e.g., one involving a new many-to-many relationship,
would have even more impact because it would require
creating new tables. In case users ask for a full versioning
of the schemata, the efort would be greater still. An M 3D
schema represents a good trade-of here because most
evolutions can be handled seamlessly with no impact on
tables and ETL; clearly, a more invasive evolution (such as
adding a new dimension or measure) would still require
a change to the relational part of the schema and to the
ETL.
7
      </p>
    </sec>
    <sec id="sec-9">
      <title>A PERSPECTIVE ON MMDW RESEARCH</title>
      <p>
        The experiments we conduct in this work are encouraging enough
to set a short- and mid-term perspectives of the research on
MMDWs. The advantages we envision for MMDWs can be
summarized as follows:
(1) An MMDW will natively and eficiently support OLAP
querying over large volumes of multimodel and
multidimensional data, thus ensuring support to both volume,
velocity, and variety.
(2) Storing data in their native model means reducing the
data transformations required; hence, the efort for writing
(time-consuming and error-prone) ETL procedures will be
reduced in MMDWs, and the freshness of data in the DWs
will be increased.
(3) MMDWs will bridge the architectural gap between data
lakes and DWs. A data lake ingests
heterogeneouslystructured raw data from various sources and stores them
in their native format, enabling their processing according
to changing requirements [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Diferently from DWs, data
lakes support storage of any kind of data with low-cost
design, provide increasing analysis capabilities, and ofer
an improvement in data ingestion; however, analysis tasks
are more complex and time-consuming since a
schema-onread approach must be followed. We believe MMDWs will
ofer an efective architectural trade-of by enabling both
OLAP multidimensional analyses and ad-hoc analytics on
the same repository.
(4) Schema evolution is a crucial issue in traditional DW
architectures, since modifying relational schemata to
accommodate new user requirements is a complex and expensive
task. MMDWs can store schemaless data, so they will
ensure a more efective support to schema evolution [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
(5) Again thanks to their support of schemaless data, higher
lfexibility and extensibility will be granted, which will
enhance analysis capabilities thus generating added value
for users [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
(6) More specifically, key-value stores on the one hand,
and the array constructs supported by document-based
databases on the other, provide an alternative solution
to model many-to-many relationships appearing in some
multidimensional schemata.
      </p>
      <p>In our short-term research agenda on MMDWs we mainly
plan to verify and quantify these benefits via an extensive set
of experiments based on a more comprehensive case study. This
will require, for instance, to measure the efort for writing ETL
procedures to transform all data according to a single model;
to assess the increase in querying expressiveness achieved by
MMDWs in function of the amount of data variety; to simulate
dynamic settings so as to evaluate the saving in dealing with
schema evolution. In order to overcome performance limitations
described in the previous section, we think also that new
experiments are mandatory on another multimodel DBMS such as
Oracle, which provides other types of implementation for non
relational data, and also distributed storage and computation.</p>
      <p>
        In the mid-term, the preliminary work we presented in this
paper opens several research issues:
• Multidimensional design from MMDBs. The existing
datadriven approaches to multidimensional design are based
on detecting functional dependencies in single-model data
sources, namely, relational, XML, linked-open data, JSON
[
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. Using a multimodel data source for design requires
integrating diferent techniques into a synergic
methodology.
• Conceptual models. Existing conceptual models for DWs
are mostly aimed at designing multidimensional schemata
with fixed structure. To take full advantage of the flexibility
ensured by MMDWs, new models capable of coping with
schemaless data (as naively done with the cloud symbol
in Figure 2) are needed.
• Best practices for logical design. In presence of variety,
several alternatives emerge for the logical representation of
dimensions and facts [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Indeed, some combinations of
models may be better than others when coupled with star
schemata. A specific set of guidelines for logical design of
MMDWs is thus needed to find the best trade-of between
performances, fidelity to source schemata, extensibility,
and evolvability; this should also include the issues related
to view materialization.
• OLAP benchmark. Efectively benchmarking MMDBs [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]
and non relational DBMSs [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] is still a challenge.
Providing a benchmark for MMDWs is a further challenge,
since it requires defining a dataset representative of DW
volume and multimodel variety, as well as a full range of
representative OLAP queries over this dataset.
• Indexing. PostgreSQL ofers diferent types of indexes over
MMDBs, e.g., B-trees, hash, GiST (for geo data), GIN (for
document and hstore data), etc. Ad hoc indexing strategies
will have to be devised, in presence of variety, to cope with
the specific features of multidimensional data and OLAP
queries.
• OLAP tools. Last but not least, more sophisticated OLAP
tools are required to let users benefit from the additional
lfexibility introduced by MMDWs while ensuring good
performances. Specifically, there is a need for devising
techniques to automatically generate eficient SQL queries
over MMDWs from the (MDX-like or graphical) language
used by the front-end.
      </p>
    </sec>
    <sec id="sec-10">
      <title>8 CONCLUSION</title>
      <p>Handling big data variety, volume, and velocity is an important
challenge for decision-making information systems. On the one
hand, data lakes have been proposed to ensure flexible storage of
raw data, but at the price of making analyses more complex. On
the other hand, classical DW architectures provide an eficient
framework for analyzing transformed and integrated data, but
they fall short in natively handling data variety. Motivated by the
emerging trend of MMDBs, in this work we have investigated
the feasibility of a multimodel approach to DW based on an
extension of the well-known star schema with schemaless data
as dimensions and facts. Our experiments are encouraging as
they show that all queries of our multimodel tailored OLAP
workload can run over the proposed multimodel star schema in
acceptable time compared to a full-relational implementation.
Based on these first results, we have presented many short- and
mid-term research perspectives on MMDW.</p>
    </sec>
    <sec id="sec-11">
      <title>ACKNOWLEDGEMENT</title>
      <p>This work was partially supported by the French National
Research Agency as part of the “Investissements d’Avenir” through
the IDEX-ISITE initiative CAP 20-25 (ANR-16-IDEX-0001), and
the project VGI4bio (ANR-17-CE04-0012).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Alberto</given-names>
            <surname>Abelló</surname>
          </string-name>
          , Jérôme Darmont, Lorena Etcheverry, Matteo Golfarelli, JoseNorberto Mazón, Felix Naumann, Torben Bach Pedersen, Stefano Rizzi, Juan Trujillo, Panos Vassiliadis, and
          <string-name>
            <given-names>Gottfried</given-names>
            <surname>Vossen</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Fusion Cubes: Towards Self-Service Business Intelligence</article-title>
          .
          <source>IJDWM 9</source>
          ,
          <issue>2</issue>
          (
          <year>2013</year>
          ),
          <fpage>66</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Atzeni</surname>
          </string-name>
          , Francesca Bugiotti, and
          <string-name>
            <given-names>Luca</given-names>
            <surname>Rossi</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Uniform access to NoSQL systems</article-title>
          .
          <source>Inf. Syst</source>
          .
          <volume>43</volume>
          (
          <year>2014</year>
          ),
          <fpage>117</fpage>
          -
          <lpage>133</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Nabila</given-names>
            <surname>Berkani</surname>
          </string-name>
          , Ladjel Bellatreche, Selma Khouri, and
          <string-name>
            <given-names>Carlos</given-names>
            <surname>Ordonez</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Value-driven Approach for Designing Extended Data Warehouses</article-title>
          .
          <source>In Proc. DOLAP@EDBT/ICDT</source>
          . Lisbon, Portugal.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Doulkifli</given-names>
            <surname>Boukraâ</surname>
          </string-name>
          , Mohammed Amin Bouchoukh, and
          <string-name>
            <given-names>Omar</given-names>
            <surname>Boussaïd</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Eficient Compression and Storage of XML OLAP Cubes</article-title>
          .
          <source>IJDWM 11</source>
          ,
          <issue>3</issue>
          (
          <year>2015</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Mohamed</given-names>
            <surname>Boussahoua</surname>
          </string-name>
          , Omar Boussaid, and
          <string-name>
            <given-names>Fadila</given-names>
            <surname>Bentayeb</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Logical Schema for Data Warehouse on Column-Oriented NoSQL Databases</article-title>
          .
          <source>In Proc. DEXA</source>
          . Lyon, France,
          <fpage>247</fpage>
          -
          <lpage>256</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Arnaud</given-names>
            <surname>Castelltort</surname>
          </string-name>
          and
          <string-name>
            <given-names>Anne</given-names>
            <surname>Laurent</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>NoSQL Graph-based OLAP Analysis</article-title>
          .
          <source>In Proc. KDIR</source>
          . Rome, Italy,
          <fpage>217</fpage>
          -
          <lpage>224</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Max</given-names>
            <surname>Chevalier</surname>
          </string-name>
          , Mohammed El Malki, Arlind Kopliku, Olivier Teste, and
          <string-name>
            <given-names>Ronan</given-names>
            <surname>Tournier</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Implementation of Multidimensional Databases in ColumnOriented NoSQL Systems</article-title>
          .
          <source>In Proc. ADBIS</source>
          . Poitiers, France,
          <fpage>79</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Max</given-names>
            <surname>Chevalier</surname>
          </string-name>
          , Mohammed El Malki, Arlind Kopliku, Olivier Teste, and
          <string-name>
            <given-names>Ronan</given-names>
            <surname>Tournier</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Implementation of Multidimensional Databases with Document-Oriented NoSQL</article-title>
          .
          <source>In Proc. DaWaK</source>
          . Valencia, Spain,
          <fpage>379</fpage>
          -
          <lpage>390</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Max</given-names>
            <surname>Chevalier</surname>
          </string-name>
          , Mohammed El Malki, Arlind Kopliku, Olivier Teste, and
          <string-name>
            <given-names>Ronan</given-names>
            <surname>Tournier</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Document-Oriented Data Warehouses: Complex Hierarchies and Summarizability</article-title>
          .
          <source>In Proc. UNet</source>
          . Casablanca, Morocco,
          <fpage>671</fpage>
          -
          <lpage>683</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Max</surname>
            <given-names>Chevalier</given-names>
          </string-name>
          , Mohammed El Malki, Arlind Kopliku, Olivier Teste, and
          <string-name>
            <given-names>Ronan</given-names>
            <surname>Tournier</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Document-oriented data warehouses: Models and extended cuboids, extended cuboids in oriented document</article-title>
          .
          <source>In Proc. RCIS</source>
          . Grenoble, France,
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Max</surname>
            <given-names>Chevalier</given-names>
          </string-name>
          , Mohammed El Malki, Arlind Kopliku, Olivier Teste, and
          <string-name>
            <given-names>Ronan</given-names>
            <surname>Tournier</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Document-oriented Models for Data Warehouses - NoSQL Document-oriented for Data Warehouses</article-title>
          .
          <source>In Proc. ICEIS</source>
          . Rome, Italy,
          <fpage>142</fpage>
          -
          <lpage>149</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Mohamed</given-names>
            <surname>Lamine</surname>
          </string-name>
          <string-name>
            <surname>Chouder</surname>
          </string-name>
          , Stefano Rizzi, and
          <string-name>
            <given-names>Rachid</given-names>
            <surname>Chalal</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>EXODuS: Exploratory OLAP over Document Stores</article-title>
          .
          <source>Inf. Syst</source>
          .
          <volume>79</volume>
          (
          <year>2019</year>
          ),
          <fpage>44</fpage>
          -
          <lpage>57</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Khaled</given-names>
            <surname>Dehdouh</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Building OLAP Cubes from Columnar NoSQL Data Warehouses</article-title>
          .
          <source>In Proc. MEDI</source>
          . Almería, Spain.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Ibtisam</surname>
            <given-names>Ferrahi</given-names>
          </string-name>
          , Sandro Bimonte, and
          <string-name>
            <given-names>Kamel</given-names>
            <surname>Boukhalfa</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Model &amp; DBMS Independent Benchmark for Data Warehouses</article-title>
          .
          <source>In Proc. EDA</source>
          . Lyon, France,
          <fpage>101</fpage>
          -
          <lpage>110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Ibtisam</surname>
            <given-names>Ferrahi</given-names>
          </string-name>
          , Sandro Bimonte,
          <string-name>
            <surname>Myoung-Ah Kang</surname>
            , and
            <given-names>Kamel</given-names>
          </string-name>
          <string-name>
            <surname>Boukhalfa</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Design and Implementation of Falling Star - A Non-Redundant SpatioMultidimensional Logical Model for Document Stores</article-title>
          .
          <source>In Proc. ICEIS</source>
          . Porto, Portugal,
          <fpage>343</fpage>
          -
          <lpage>350</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Vijay</surname>
            <given-names>Gadepally</given-names>
          </string-name>
          , Peinan Chen, Jennie Duggan,
          <string-name>
            <given-names>Aaron J.</given-names>
            <surname>Elmore</surname>
          </string-name>
          , Brandon Haynes, Jeremy Kepner, Samuel Madden, Tim Mattson, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Stonebraker</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>The BigDAWG polystore system and architecture</article-title>
          .
          <source>In Proc. HPEC</source>
          . Waltham, MA, USA,
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Enrico</surname>
            <given-names>Gallinucci</given-names>
          </string-name>
          , Matteo Golfarelli, and
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Rizzi</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Approximate OLAP of document-oriented databases: A variety-aware approach</article-title>
          .
          <source>Inf. Syst</source>
          .
          <volume>85</volume>
          (
          <year>2019</year>
          ),
          <fpage>114</fpage>
          -
          <lpage>130</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Golfarelli</surname>
          </string-name>
          and
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Rizzi</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Data Warehouse Design: Modern Principles and Methodologies</article-title>
          .
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          , Inc., New York, NY, USA.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Hamdi</given-names>
            <surname>Ben</surname>
          </string-name>
          <string-name>
            <surname>Hamadou</surname>
          </string-name>
          , Enrico Gallinucci, and
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Golfarelli</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Answering GPSJ Queries in a Polystore: a Dataspace-Based Approach</article-title>
          .
          <source>In Proc. ER</source>
          . Salvador de Bahia, Brazil.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Ralph</given-names>
            <surname>Kimball</surname>
          </string-name>
          and
          <string-name>
            <given-names>Margy</given-names>
            <surname>Ross</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>The data warehouse toolkit: the complete guide to dimensional modeling, 2nd Edition</article-title>
          . Wiley.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Jiaheng</given-names>
            <surname>Lu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Irena</given-names>
            <surname>Holubová</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Multi-model Databases: A New Journey to Handle the Variety of Data</article-title>
          .
          <source>ACM Comput. Surv</source>
          .
          <volume>52</volume>
          ,
          <issue>3</issue>
          (
          <year>2019</year>
          ),
          <volume>55</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>55</lpage>
          :
          <fpage>38</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Mohammed</given-names>
            <surname>El</surname>
          </string-name>
          <string-name>
            <surname>Malki</surname>
          </string-name>
          , Arlind Kopliku, Essaid Sabir, and
          <string-name>
            <given-names>Olivier</given-names>
            <surname>Teste</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Benchmarking Big Data OLAP NoSQL Databases</article-title>
          .
          <source>In Proc. UNet</source>
          . Hammamet, Tunisia,
          <fpage>82</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Patrick E. O'Neil</surname>
          </string-name>
          ,
          <string-name>
            <surname>Elizabeth J. O'Neil</surname>
            ,
            <given-names>Xuedong</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            , and
            <given-names>Stephen</given-names>
          </string-name>
          <string-name>
            <surname>Revilak</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>The Star Schema Benchmark and Augmented Fact Table Indexing</article-title>
          .
          <source>In Proc. TPCTC</source>
          . Lyon, France,
          <fpage>237</fpage>
          -
          <lpage>252</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Zoubir</surname>
            <given-names>Ouaret</given-names>
          </string-name>
          , Rachid Chalal, and
          <string-name>
            <given-names>Omar</given-names>
            <surname>Boussaid</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>An overview of XML warehouse design approaches and techniques</article-title>
          .
          <source>IJICoT 2</source>
          ,
          <issue>2</issue>
          /3 (
          <year>2013</year>
          ),
          <fpage>140</fpage>
          -
          <lpage>170</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Franck</given-names>
            <surname>Ravat</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yan</given-names>
            <surname>Zhao</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Data Lakes: Trends and Perspectives</article-title>
          .
          <source>In Proc. DEXA</source>
          . Linz, Austria,
          <fpage>304</fpage>
          -
          <lpage>313</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Oscar</given-names>
            <surname>Romero</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alberto</given-names>
            <surname>Abelló</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>A Survey of Multidimensional Modeling Methodologies</article-title>
          .
          <source>IJDWM 5</source>
          ,
          <issue>2</issue>
          (
          <year>2009</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Stefanie</surname>
            <given-names>Scherzinger</given-names>
          </string-name>
          , Meike Klettke, and
          <string-name>
            <given-names>Uta</given-names>
            <surname>Störl</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Managing Schema Evolution in NoSQL Data Stores</article-title>
          .
          <source>In Proc. DBPL. Riva del Garda</source>
          , Italy.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Amal</surname>
            <given-names>Sellami</given-names>
          </string-name>
          , Ahlem Nabli, and
          <string-name>
            <given-names>Faïez</given-names>
            <surname>Gargouri</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Transformation of Data Warehouse Schema to NoSQL Graph Data Base</article-title>
          .
          <source>In Proc. ISDA</source>
          . Vellore, India,
          <fpage>410</fpage>
          -
          <lpage>420</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Takeyuki</surname>
            <given-names>Shimura</given-names>
          </string-name>
          , Masatoshi Yoshikawa, and
          <string-name>
            <given-names>Shunsuke</given-names>
            <surname>Uemura</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Storage and Retrieval of XML Documents Using Object-Relational Databases</article-title>
          .
          <source>In Proc. DEXA</source>
          . Florence, Italy,
          <fpage>206</fpage>
          -
          <lpage>217</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Chao</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Jiaheng Lu, Pengfei Xu,
          <string-name>
            <given-names>and Yuxing</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>UniBench: A Benchmark for Multi-model Database Management Systems</article-title>
          .
          <source>In Proc. TPCTC</source>
          . Rio de Janeiro, Brazil,
          <fpage>7</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>