<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>SEBD</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Using Regression to Explain Cube Measures</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Matteo Francia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Rizzi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrick Marcel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DISI - University of Bologna</institution>
          ,
          <addr-line>Bologna</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LIFAT - University of Tours</institution>
          ,
          <addr-line>Blois</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>31</volume>
      <fpage>02</fpage>
      <lpage>05</lpage>
      <abstract>
        <p>The Intentional Analytics Model (IAM) has been devised to couple OLAP and analytics by (i) letting users express their analysis intentions on multidimensional data cubes and (ii) returning enhanced cubes, i.e., multidimensional data annotated with knowledge insights in the form of models (e.g., correlations). Five intention operators were proposed to this end; of these, describe and assess have been investigated in previous papers. In this work we enrich the IAM picture by focusing on the explain operator, whose goal is to provide an answer to the user asking “why does measure  show these values?”. Specifically, we propose a syntax for the operator and discuss how enhanced cubes are built by (i) finding the polynomials that best approximate the relationship between  and the other cube measures, and (ii) highlighting the most interesting one. Finally, we test the operator implementation in terms of eficiency.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Data cube</kwd>
        <kwd>OLAP</kwd>
        <kwd>Analytics</kwd>
        <kwd>Correlation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Despite the huge success of the OLAP paradigm, it is now clear that this paradigm, alone, does
no longer meet the sophisticated requirements of new-generation decision makers. Among the
directions taken by research to enhance OLAP, the Intentional Analytics Model (IAM) suggests to
couple it with analytics [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The IAM approach relies on two main ideas: (i) users explore the data
space by expressing their analysis intentions and (ii) in return they receive both multidimensional
data and knowledge insights in the form of models. To achieve (i) five intention operators
were proposed, namely, describe (describes one or more cube measures at some aggregation
level, possibly focused on some level members), assess (judges one or more cube measures with
reference to some benchmark), explain (reveals the reason behind the values of a measure, for
instance by correlating it with other measures), predict (shows data not in the original cubes,
derived for instance with regression), and suggest (shows data similar to those the current user,
or similar users, have been interested in). As to (ii), first-class citizens of the IAM are enhanced
cubes, defined as multidimensional cubes coupled with highlights, i.e., interesting components
of models automatically extracted from cubes. An overview of the approach is shown in Figure
1. Noticeably, having diferent models automatically computed and evaluated in terms of their
interest relieves the user from the time-wasting efort of trying diferent possibilities.
describe
assess
explain
predict
suggest
cube
data
type
Batteries
Beer
CannedVegetables
Cheese
Chips
ChocolateCandy
Coffee
Cookies
DriedFruit
Eggs
      </p>
      <p>
        Among the five intention operators, describe and assess have been investigated in previous
papers [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. In this paper we enrich the IAM picture by focusing on the explain operator. An
explanation is essentially a description of causation for an observed phenomenon; in practice, it
answers the why? question for that phenomenon by providing a causal model for it [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In our
context, we concentrate on providing explanation models for a measure the user is observing;
thus, the goal of the explain operator will be to provide an answer to the user asking “why does
measure  show these values?”.
      </p>
      <p>
        As envisioned in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], several types of models can be used to this end. To give a
proof-ofconcept for explain, in this paper we restrict to the simplest model type, the one that establishes
a polynomial relationship between  and ′.
      </p>
      <sec id="sec-1-1">
        <title>Example 1. Let a SALES cube be given, and let the user’s intention be</title>
        <p>with SALES explain revenue by type for year=’2022’
First, the subset of facts for 2022 are selected from the SALES cube and aggregated by product type
(in OLAP terms, a slice-and-dice and a roll-up operator are applied). Then, regression analysis is
used to compare the revenue measure with each other cube measure and find the polynomials that
best approximates their relationship. Finally, a measure of interest is computed for the components
(i.e., for the polynomials) obtained, and the most interesting one is shown to the user (in Figure 1,
the one showing that revenue is roughly proportional to quantity).</p>
        <p>The paper outline is as follows. After introducing models and enhanced cubes in Section 2,
in Section 3 we give the syntax of explain and illustrate how models are built. Then, in Section
4 we explain how enhanced cubes are visualized. Finally, in Section 5 we discuss the related
literature and in Section 6 we test the operator implementation and draw the conclusions.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Enhanced cubes</title>
      <p>Models are concise, information-rich knowledge artifacts that represent relationships hiding
in the cube facts. The possible models range from simple functions and measure correlations
to more elaborate techniques such as decision trees, clusterings, etc. A model is bound to (i.e.,
is computed over the levels/measures of) one cube, and is made of a set of components, each
component being a specific relationship among cube facts.</p>
      <p>Definition 1 (Hierarchy and Cube Schema). A hierarchy is a pair ℎ = (ℎ, ⪰ ℎ) where ℎ
is a set of categorical levels, each coupled with a domain including a set of members, and ⪰ ℎ is
a roll-up total order of ℎ. The top level of ⪰ ℎ is called dimension. A cube schema is a pair
 = (,  ) where  is a set of hierarchies and  is a set of numerical measures, each coupled
with one aggregation operator.</p>
      <p>Example 2. For our working example we will use the SALES cube, which includes three
hierarchies and three measures. Formally, SALES = (,  ) with
 = {ℎDate, ℎProduct, ℎStore};  = {quantity, revenue, cost};
date ⪰</p>
      <p>month ⪰ year; product ⪰ type ⪰ category; store ⪰ city ⪰ country</p>
      <p>Aggregation is the basic mechanism to query cubes, and it is captured by the following
definition of group-by set.</p>
      <p>Definition 2 (Group-by Set and Coordinate). Given cube schema  = (,  ), a group-by
set of  is a set of levels, at most one from each hierarchy of . The partial order induced on the
set of all group-by sets of  by the roll-up orders of the hierarchies in , is denoted with ⪰  . A
coordinate of group-by set  is a tuple of members, one for each level of .</p>
      <p>Example 3. Two group-by sets of SALES are 1 = {date, type, country} and 2 =
{month, category}, where 1 ⪰  2. 1 aggregates sales by date, product type, and store
country, 2 by month and category. Example of coordinates of the two group-by sets are,
respectively,  1 = ⟨2022-04-15, Fresh Fruit, Italy⟩ and  2 = ⟨2022-04, Fruit⟩.</p>
      <p>The instances of a cube schema are called cubes and are defined as follows.</p>
      <p>Definition 3 (Cube). A cube over  is a triple  = ( ,  ,  ) where  is a group-by set
of ,  ⊆  , and  is a partial function that maps the coordinates of  to a numerical
value for each measure  ∈  .</p>
      <p>Each coordinate  that participates in  , with its associated measure values, is called a fact
of . With a slight abuse of notation, we will write  ∈  to state that  is a fact of . The
value taken by measure  in the fact corresponding to  is denoted as . . A cube whose
group-by set  includes all and only the dimensions of the hierarchies in  and such that
 =  , is called a base cube, the others are called derived cubes. In OLAP terms, a derived
cube is the result of either a roll-up, a slice-and-dice, or a projection made over a base cube; this
is formalized as follows.</p>
      <p>Definition 4 (Cube Query). A query over cube schema  is a triple  = (, , ) where
 is a group-by set of ,  is a (possibly empty) set of selection predicates each expressed over
one level of , and  ⊆  .</p>
      <p>Example 4. The cube query over SALES used in Example 1 is  = (, , ) where  =
{type},  = {year = ’2022’}, and  = {revenue}. A coordinate of the resulting cube is
⟨Batteries⟩ with associated value 3320 for revenue.</p>
      <p>Definition 5 (Model). A model is a tuple ℳ = (, , , , , ) where: (i)  is the model
type; (ii)  is the algorithm used to compute ; (iii)  is the cube to which the model is bound;
(iv)  is the measure in  to be explained; (v)  is the tuple of levels/measures of  and parameter
values supplied to  to compute the model; (vi)  is the set of model components.
In this paper, to give a proof-of-concept of the explain operator, we restrict to consider a single
type of model, namely, the one that establishes a polynomial relationship between two measures
via regression analysis. In this case,  is the set of measures whose relationship with  is
described; besides, each component  ∈  shows the relationship of  with one measure
 ∈ .</p>
      <p>Definition 6 (Component). For a model of type polynomial, a component  is a triple  =
(, , coeff ) where: (i)  is the measure in  whose relationship with  is described; (ii) 
is the degree of the polynomial used to describe the relationship between  and ; (iii) coeff 
is an array of the  + 1 coeficients of the polynomial   () that best approximates  with
reference to the facts in .</p>
      <sec id="sec-2-1">
        <title>Example 5. A possible model over the SALES cube is characterized by  = regression;  = Polyfit ;  = SALES;</title>
        <p>= revenue;  = {quantity, cost};  = {1, 2};
1 = (quantity, 1, [0.98, 4909.52]); 2 = (cost, 2, [1.1, 22.78, 1409.33])
According to this model, the relationships of revenue with quantity and cost are described,
respectively, as
revenue =  1(quantity) = 0.98 · quantity + 4909.52
revenue =  2(cost) = 1.1 · cost2 − 22.78 · cost + 1409.33</p>
        <p>As the last step in the IAM approach, cube  is enhanced by associating it with a set of
models bound to  and with a highlight, i.e., with the most interesting model component:
Definition 7 (Enhanced cube). An enhanced cube  is a triple of a cube , a set of models
{ℳ1, . . . , ℳ} bound to , and a highlight  = {∈⋃︀=1 }(()).</p>
        <p>
          In our scenario only polynomial models are considered, so an enhanced cube includes a single
model with one component for each measure in . Let  be the component associated to
; we evaluate the interest of , (), as the coeficient of determination R2 [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], which
measures how well the values of  are replicated by the model in  via the variation in the
dependent variable  that is predictable from the independent variable . The better the
model, the closer the value of R2 to 1.
        </p>
        <p>Example 6. With reference to Example 5, it is (1) = 0.61 and (2) = 0.99.
Thus, the highlight is 2.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. The explain operator</title>
      <p>The explain operator provides an answer to the user asking “why is this happening?” “why
does measure  show these values?” by describing the relationship between  and the other
cube measures, possibly focused on one or more level members, at some given granularity. The
cube is enhanced by showing the polynomials that best approximate these relationships, with a
highlight on the most interesting one.</p>
      <sec id="sec-3-1">
        <title>3.1. Syntax</title>
        <p>Let 0 be a base cube over cube schema  = (,  ). The syntax for explain is (optional parts
are in brackets):</p>
        <p>with 0 explain  [ for  ] by 1, . . . ,  [ against 1 [ degree 1 ], . . . ,  [ degree  ] ]
where  ∈  is a measure of ;  is a set of selection predicates, each over one level of ;
{1, . . . , } is a group-by set of ; 1, . . . ,  are measures of  (diferent from ); the ’s
( &gt; 0) are integers denoting, for each , the degree of the polynomial to be computed.
Example 7. Examples of explain intentions on the SALES cube are, besides the one in Example
1,
with SALES explain cost by date, product against quantity degree 1
with SALES explain revenue by year against cost</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Semantics</title>
        <p>
          The execution plan corresponding to a fully-specified intention, i.e., one where all optional
clauses have been specified, is as follows:
1. Execute query  = (, , ), where  = {1, . . . , },  =  , and  =
{, 1, . . . , }. Let  = (0) be the cube resulting from the execution of  over 0.
2. Compute model ℳ = (polynomial, Polyfit , , , {1, . . . , }, {1, . . . , }), where  =
(, , coeff ). The best approximating polynomial of degree  is determined via ordinary
least squares [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], which finds —with a complexity of (2 ||), where || is the number of
facts in — the polynomial coeficients that minimize the sum of squared errors between
each  (independent variables) and  (dependent variable).
3. For each  compute ().
4. Find the highlight  = {1≤ ≤ }(()).
5. Return the enhanced cube  consisting of , {ℳ}, and .
        </p>
        <p>
          Partially-specified intentions are interpreted as follows:
• If the for clause is not specified, we consider  =   .
• If the against clause is not specified, a component is created for each measure in 
(except ).
• If the degree clause is not specified for one or more measures, the value of  is determined
automatically by polynomial fitting [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>Example 8. The first intention in Example 7 is executed by first computing the derived cube 
that aggregates SALES by {date, product} and projects on measures cost and quantity. Then,
a model ℳ including a single component  (a linear polynomial approximating cost in function
of quantity) is determined. Finally, the enhanced cube including , ℳ, and the highlight  is
returned.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Visualizing enhanced cubes</title>
      <p>As previously done for the describe and assess IAM operators, to give an efective visualization
of the enhanced cubes built for explain intentions we couple a text-based representation (a
pivot table and a ranked component list) with a graphical one (a chart) and with an ad-hoc
interaction paradigm. Specifically, the visualization of enhanced cube  = (, ℳ, ) relies on
three distinct but inter-related areas: a table area that shows the facts of  using a pivot table; a
component area that shows a list of model components (i.e., approximating polynomials) sorted
by their interest, with  at the top; a chart area that uses a scatter chart to display, for each
component  of ℳ, the relationship between  and  as well as the function plotting the
approximating polynomial.</p>
      <p>The interaction paradigm we adopt is component-driven: clicking on one component  in
the component area leads to show the corresponding approximating polynomial in the chart
area. The highlight is selected by default.</p>
      <p>Example 9. Figure 2 shows the visualization obtained when the intention in Example 1 is
formulated. On the left, the table area; on the right, the chart area; in the middle, the component area.
The highlight is a quadratic polynomial that approximates revenue in function of cost, so the
chart area shows the relationship between these two measures and the approximating parabola.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Related work</title>
      <p>
        The idea of coupling data and analytical models was born in the 90’s with inductive databases,
where data were coupled with patterns meant as generalizations of the data [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Later on,
datato-model unification was addressed in MauveDB [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which provides a language for specifying
model-based views of data using common statistical models. More recently, Northstar [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
has been proposed as a system to support interactive data science by enabling users to switch
between data exploration and model building. Finally, the coupling of data and models is at the
core of the IAM vision [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], on which this paper relies.
      </p>
      <p>
        The coupling of the OLAP paradigm and data mining to create an approach where concise
patterns are extracted from multidimensional data for user’s evaluation, was the goal of some
approaches commonly labeled as OLAM [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In this context, k-means clustering is used in
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] to dynamically create semantically-rich aggregates of facts other than those statically
provided by dimension hierarchies. Other operators that enrich data with knowledge extraction
results are DIFF [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and RELAX [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Finally, in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] the OLAP paradigm is reused to explore
prediction cubes, i.e., cubes where each fact summarizes a predictive model trained on the data
corresponding to that fact.
      </p>
      <p>
        In an attempt to develop tools for helping users understand data, there have been several
eforts in the research community to devise techniques to model explanations for observations
made on data; see [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] for a comprehensive analysis of the literature and of the trends in
explanation. A common way to give an explanation is to identify the actual cause of the
observed outcome. Given the result of a database query, which database tuple(s) caused that
output to the query? One way to answer this question is to quantify the contribution that
each tuple has to the result and identify the tuples with the highest contributions [
        <xref ref-type="bibr" rid="ref17 ref18">17, 18</xref>
        ]; the
intuition is that tuples with high contribution tend to be interesting explanations to query
answers. Similarly, in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] causality is defined in terms of intervention: an input is a cause to an
output if we can afect the output by changing the value of that input. Causality poses additional
challenges when the query contains aggregates [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], as in our scenario. The DIFF operator [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
tells users why a given aggregated quantity is lower or higher in one cube fact than in another by
returning the set of rows that best explains the observed increase or decrease at the aggregated
level. In Scorpion [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], outliers are explained in terms of properties of the tuples used to compute
these outliers, while [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] explains outliers in aggregation queries through counter-balancing.
LensXPlain [22] explains why some measure value is high or low by identifying subsets of facts
that contributed the most toward such observation. A diferent approach to query explanation
is taken in [23]. The authors focus on multidimensional data where a binary dimension is
present, and explain query results by building explanation tables which provide an interpretable
and informative summary of the factors afecting the binary dimension. Finally, regression
is used to explain query results in the XAXA approach [24]. The authors focus on aggregate
queries with a center-radius selection operator, and give explanations using a set of parametric
piecewise-linear functions acquired through a statistical learning model.
      </p>
      <p>The approach we propose is not competing with the ones mentioned above, but should rather
be seen as a modular framework where any approach to explanation of aggregate data could
be plugged. The added value lies in the IAM paradigm, i.e., in giving users the possibility
of explicitly expressing intentions, in letting the system select the most interesting/suitable
explanations, and showing these explanations together with data.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <p>In this paper we have given a proof-of-concept for explain intentions formulated inside the IAM
framework. The explain syntax is flexible enough to suit users who wish to verify a specific
hypothesis they made about an inter-measure relationship, as well as users who have no clue
so they will let the system find the most interesting relationship.</p>
      <p>The prototype we developed to test our approach relies on the MySQL DBMS to execute
queries on a star schema based on multidimensional metadata. The algorithms used for
regression analysis are imported from the Scikit-Learn Python library. Finally, the web-based
visualization is implemented in JavaScript and exploits the D3 library for chart visualization.</p>
      <p>To verify the feasibility of our approach from the computational point of view, we made
some scalability tests. Two main factors afect performances: the cardinality of the cube to
which a model is bound, || (which determines the time required to compute a single model
component), and the number of cube measures, | | (which determines the number of model
components to be computed).</p>
      <p>To evaluate scalability with reference to cube cardinality, we populated the SALES cube
using the FoodMart data (https://github.com/julianhyde/foodmart-data-mysql) and
considered 10 intentions with increasing cardinalities; in each intention we explained the revenue
measure against both quantity and cost. The tests were run on an Intel(R) Core(TM)i7-6700
CPU@3.40GHz CPU with 8GB RAM; each intention was executed 10 times and the average
results are considered. Remarkably, it turns out that less than one second is necessary to explain
a cube of almost 87000 facts. Additionally, we measured the complexity (as the number of
characters) of writing explain intentions vs. the underlying cube query. It turns out that our
approach saves 85% of complexity with respect to writing cube queries (and without considering
the complexity of extracting regression models, which would make our approach even more
convenient).</p>
      <p>To evaluate scalability with reference to the number of measures, we created a cube with
|| = 106 facts and | | = 10 measures. As expected, our approach scales linearly in the
number of measures, and given 9 measures and 106 facts, the computation of an explanation
takes less than 7 seconds, thus fulfilling the requirement of near-real-time response typical of
analytical workloads. The detailed results of the tests can be found in [25].</p>
      <p>We close the paper by mentioning that the main direction for future research we wish to
pursue is to generalize the definition of model to cope with additional model types.
[22] Z. Miao, A. Lee, S. Roy, LensXPlain: Visualizing and explaining contributing subsets for
aggregate query answers, Proceedings of VLDB Endow. 12 (2019) 1898–1901.
[23] K. E. Gebaly, P. Agrawal, L. Golab, F. Korn, D. Srivastava, Interpretable and informative
explanations of outcomes, Proceedings of VLDB Endow. 8 (2014) 61–72.
[24] F. Savva, C. Anagnostopoulos, P. Triantafillou, Explaining aggregates for exploratory
analytics, in: Proceedings of BigData, Seattle, WA, USA, 2018, pp. 478–487.
[25] M. Francia, S. Rizzi, P. Marcel, The whys and wherefores of cubes, in: Proceedings of
DOLAP, Ioannina, Greece, 2023, pp. 43–50.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Vassiliadis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Marcel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          ,
          <article-title>Beyond roll-up's and drill-down's: An intentional analytics model to reinvent OLAP, Information Systems 85 (</article-title>
          <year>2019</year>
          )
          <fpage>68</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Francia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Golfarelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Marcel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vassiliadis</surname>
          </string-name>
          ,
          <article-title>Assess queries for interactive analysis of data cubes</article-title>
          ,
          <source>in: Proceedings of EDBT</source>
          , Nicosia, Cyprus,
          <year>2021</year>
          , pp.
          <fpage>121</fpage>
          -
          <lpage>132</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Francia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Marcel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Peralta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          ,
          <article-title>Enhancing cubes with models to describe multidimensional data</article-title>
          ,
          <source>Inf. Syst. Frontiers</source>
          <volume>24</volume>
          (
          <year>2022</year>
          )
          <fpage>31</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G. R.</given-names>
            <surname>Mayes</surname>
          </string-name>
          , Theories of Explanation, Internet Encyclopedia of Philosophy (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. G. D.</given-names>
            <surname>Steel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Torrie</surname>
          </string-name>
          ,
          <article-title>Principles and procedures of statistics, with special reference to the biological sciences, McGraw-</article-title>
          <string-name>
            <surname>Hill</surname>
          </string-name>
          , New York,
          <year>1960</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.-N.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Steinbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          , Introduction to data mining,
          <source>Pearson Education India</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E.</given-names>
            <surname>Ostertagová</surname>
          </string-name>
          ,
          <article-title>Modelling using polynomial regression</article-title>
          ,
          <source>Procedia Engineering</source>
          <volume>48</volume>
          (
          <year>2012</year>
          )
          <fpage>500</fpage>
          -
          <lpage>506</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L. D.</given-names>
            <surname>Raedt</surname>
          </string-name>
          ,
          <article-title>A perspective on inductive databases</article-title>
          ,
          <source>SIGKDD Explorations 4</source>
          (
          <year>2002</year>
          )
          <fpage>69</fpage>
          -
          <lpage>77</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Deshpande</surname>
          </string-name>
          , S. Madden,
          <article-title>MauveDB: supporting model-based user views in database systems</article-title>
          ,
          <source>in: Proceedings of SIGMOD</source>
          , Chicago, IL, USA,
          <year>2006</year>
          , pp.
          <fpage>73</fpage>
          -
          <lpage>84</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kraska</surname>
          </string-name>
          ,
          <string-name>
            <surname>Northstar:</surname>
          </string-name>
          <article-title>An interactive data science system</article-title>
          ,
          <source>Proceedings of VLDB Endow</source>
          .
          <volume>11</volume>
          (
          <year>2018</year>
          )
          <fpage>2150</fpage>
          -
          <lpage>2164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <article-title>OLAP mining: Integration of OLAP with data mining</article-title>
          ,
          <source>in: Proceedings of Working Conf. on Database Semantics</source>
          , Leysin, Switzerland,
          <year>1997</year>
          , pp.
          <fpage>3</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>F.</given-names>
            <surname>Bentayeb</surname>
          </string-name>
          , C. Favre,
          <article-title>RoK: Roll-up with the k-means clustering method for recommending OLAP queries</article-title>
          ,
          <source>in: Proceedings of DEXA</source>
          , Linz, Austria,
          <year>2009</year>
          , pp.
          <fpage>501</fpage>
          -
          <lpage>515</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sarawagi</surname>
          </string-name>
          ,
          <article-title>Explaining diferences in multidimensional aggregates</article-title>
          ,
          <source>in: Proceedings of VLDB</source>
          , Edinburgh, Scotland,
          <year>1999</year>
          , pp.
          <fpage>42</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Sathe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sarawagi</surname>
          </string-name>
          ,
          <article-title>Intelligent rollups in multidimensional OLAP data</article-title>
          ,
          <source>in: Proceedings of VLDB</source>
          , Rome, Italy,
          <year>2001</year>
          , pp.
          <fpage>531</fpage>
          -
          <lpage>540</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>B.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ramakrishnan</surname>
          </string-name>
          ,
          <article-title>Prediction cubes</article-title>
          ,
          <source>in: Proceedings of VLDB</source>
          , Trondheim, Norway,
          <year>2005</year>
          , pp.
          <fpage>982</fpage>
          -
          <lpage>993</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>B.</given-names>
            <surname>Glavic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Meliou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <article-title>Trends in explanations: Understanding and debugging data-driven systems</article-title>
          ,
          <source>Found. Trends Databases</source>
          <volume>11</volume>
          (
          <year>2021</year>
          )
          <fpage>226</fpage>
          -
          <lpage>318</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Meliou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Gatterbauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Halpern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Koch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. F.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Suciu</surname>
          </string-name>
          , Causality in databases,
          <source>IEEE Data Eng. Bull</source>
          .
          <volume>33</volume>
          (
          <year>2010</year>
          )
          <fpage>59</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Meliou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Gatterbauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. F.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Suciu</surname>
          </string-name>
          ,
          <article-title>The complexity of causality and responsibility for query answers and non-answers</article-title>
          ,
          <source>Proceedings of VLDB Endow</source>
          .
          <volume>4</volume>
          (
          <year>2010</year>
          )
          <fpage>34</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Suciu</surname>
          </string-name>
          ,
          <article-title>A formal approach to finding explanations for database queries</article-title>
          ,
          <source>in: Proceedings of SIGMOD</source>
          , Snowbird,
          <string-name>
            <surname>UT</surname>
          </string-name>
          , USA,
          <year>2014</year>
          , pp.
          <fpage>1579</fpage>
          -
          <lpage>1590</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>E.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Madden</surname>
          </string-name>
          , Scorpion:
          <article-title>Explaining away outliers in aggregate queries</article-title>
          ,
          <source>Proceedings of VLDB Endow</source>
          .
          <volume>6</volume>
          (
          <year>2013</year>
          )
          <fpage>553</fpage>
          -
          <lpage>564</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Miao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Glavic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <article-title>Going beyond provenance: Explaining query answers with pattern-based counterbalances</article-title>
          ,
          <source>in: Proceedings of SIGMOD</source>
          , Amsterdam, The Netherlands,
          <year>2019</year>
          , pp.
          <fpage>485</fpage>
          -
          <lpage>502</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>