<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Imprecise Data and Knowledge Based OLAP</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ermir Rogova</string-name>
          <email>rogovae@wmin.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Panagiotis Chountas</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Krassimir Atanassov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CLBME - Bulgarian Academy of Sciences</institution>
          ,
          <addr-line>B1. 105, Sofia-1113</addr-line>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Harrow School of Computer Science, Data &amp; Knowledge Management Group University of Westminster</institution>
          <addr-line>Watford Road, Northwick Park, HA1 3TP London</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2007</year>
      </pub-date>
      <fpage>50</fpage>
      <lpage>55</lpage>
      <abstract>
        <p>In this paper we present our approach for extending the OLAP model to include treatment of value uncertainty as part of a multidimensional model inhabited by flexible data and non-rigid hierarchical structures of organisation. A new multidimensional-cubic model named as the IF-Cube is introduced which is able to operate over data with imprecision either in the facts or in the dimensional hierarchies.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In this paper we introduce the semantics of the Intuitionistic Fuzzy cubic
representation in contrast to the basic multidimensional-cubic structures. The basic
cubic operators are extended and enhanced with the aid of [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] Intuitionistic Fuzzy
Logic.
      </p>
      <p>
        Since the emergence of the OLAP technology [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] different proposals have been
made to give support to different types of data and application purposes. One of this
is to extend the relational model (ROLAP) to support the structures and operations
typical of OLAP. Further approaches [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] are based on extended relational systems
to represent data-cubes and operate over them. The other approach is to develop new
models using a multidimensional view of the data [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Nowadays, information and knowledge-based systems need to manage imprecision
in the data and more flexible structures are needed to represent the analysis domain.
New models have appeared to manage incomplete datacube [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], imprecision in the
facts and the definition of fact using different levels in the dimensions [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Semantics of the IF-Cube in contrast to Crisp Cube</title>
      <p>
        Each element of an Intuitionistic fuzzy [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] set has degrees of membership or
truth (μ) and non-membership or falsity (ν), which don’t sum up to 1.0 thus leaving a
degree of hesitation margin (π).
      </p>
      <p>
        As opposed to the classical definition of a fuzzy set given by A′ = {&lt; x, µA′(x) &gt; |x
∈ X} where µA(x) ∈ [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] is the membership function of the fuzzy set A′, an
intuitionistic fuzzy set A is given by:
      </p>
      <p>
        A = {&lt; x, µA(x),vA(x) &gt; |x ∈ X}
where: µA : X → [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] and vA : X → [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] such that 0≤ µA(x) + vA(x)≤1 and µA(x)
vA(x) ∈ [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] denote a degree of membership and a degree of non-membership of x
∈ A, respectively.
      </p>
      <p>Obviously, each fuzzy set may be represented by the following Intuitionistic fuzzy
set A={&lt;x, µA′ (x), (x), 1− µA′ (x)&gt;|x ∈ X}</p>
      <p>For each intuitionistic fuzzy set in X, we will call πA (x) = 1 − µA(x) − vA(x) an
intuitionistic fuzzy index (or a hesitation margin) of x ∈ A which expresses a lack of
knowledge of whether x belongs to A or not. For each x ∈ A 0&lt;πA (x)&lt;1.</p>
      <p>The IF-Cube is an abstract structure that serves as the foundation for the
multidimensional data cube model. Cube C is defined as a five-tuple (D, l, F, O, H)
where:
•
•
•</p>
      <p>D is a set of dimensions
l is a set of levels l1,…, ln,
A dimension di = (l ≤ O, l┴, l┬) dom(di) where l = li i=1...n.
li is a set of values and li ∩ lj = {},
≤ O is a partial order between the elements of l.</p>
      <p>To identify the level l of a dimension, as part of a hierarchy we use dl.</p>
      <p>
        l┴: base level l┬: top level
for each pair of levels li and lj we have the relation
μij : li × lj [
        <xref ref-type="bibr" rid="ref1">0,1</xref>
        ] νij : li × lj [
        <xref ref-type="bibr" rid="ref1">0,1</xref>
        ] 0 &lt; μij + νij &lt; 1
F is a set of fact instances with schema F = {&lt;x, μF(x) , νF(x)&gt;| x∈ X },
where x=&lt;att1, …,attn&gt; is an ordered tuple belonging to a given universe X,
μF(x) and νF(x) are the degree of membership and non-membership of x in
the fact table F respectively.
      </p>
      <p>H is an object type history that corresponds to a cubic structure( l, F, O, H′ )
which allows us to trace back the evolution of a cubic structure after
performing a set of operators i.e. aggregation.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Cubic operators</title>
      <p>Selection (Σ): The selection operator selects a set of fact-instances from a cubic
structure that satisfy a predicate (θ). A predicate (θ) involves a set of atomic
predicates (θ1, …, θn ) associated with the aid of logical operators p ( i.e. ∧, ∨, etc.) .
The set of possible facts (cubic instances) that satisfy the θ should carry a degree of
membership μ and non-membership ν expressed as</p>
      <p>F = {&lt;x, min(μF(x), μ(θ (x))), max(νF(x), ν(θ (x))))&gt; | x∈ X }
Input: Ci = (D, l, F, O, H) and the predicate θ
Output: Co= (D, l, Fo , O, H) where Fo ⊆ F and Fo={f | (f ∈ F) ∧ (f satisfies θ)
Mathematical notation:
∑θ (Ci ) = Co</p>
      <p>Cubic Product ( ⊗ ): This is a binary operator Ci1 ⊗ Ci2 . It is used to relate two
cubes Ci1 and Ci2 assuming that D1 ⊆ D2 and O1 , O2 are reconcilable partial orders.
Thus, l1, l2 could lead to lo being a ragged hierarchy.</p>
      <p>Input: Ci1 = (D1, l1, F1, O1, H1) and Ci2 = (D2,l2, F2, O2, H2)
Output: Co= (Do, lo, Fo, Oo, Ho) where</p>
      <p>Do= D1 ∪ D2 , lo= l1 ∪ l2, Oo= O1 ∪ O2 Ho= H1 ∪ H2, Fo= F1 X F2</p>
      <p>Fo ={&lt;&lt;x, y&gt;, min(μf1(x), μf2(y)), max(νf1(x), νf2(y),)&gt;|&lt;x, y&gt;∈ X×Y}
Mathematical notation: Ci1 ⊗ Ci2 = Co</p>
      <p>Join (Θ): It can be expressed using Cubic Product operator. Ci1 = (D1, l1, F1, O1,
H1)) and Ci2 = (D2 ,l2, F2, O2, H2) are candidates to join if D1 ∩ D2 ≠ 0,
Input: Ci1 = (D1, l1, F1, O1, H1) and Ci2 = (D2,l2, F2, O2, H2)
Output: Co= (Do, lo, Fo, Oo, Ho)
Mathematical notation: Ci1 Θ Ci2 = σp(Ci1 ⊗ Ci2)</p>
      <p>Union (∪): The union operator is a binary operator that finds the union of two
cubes. Ci1 and Ci2 have to be union compatible. The operator also coalesces the
valueequivalent facts using the minimum membership and maximum non-membership.
Input: Ci1 = (D1, l1, F1, O1, H1) and Ci2 = (D2,l2, F2, O2, H2)
Output: Co= (Do, lo, Fo, Oo, Ho) where Do=D1=D2, lo=l1=l2, Oo=O1=O2, Ho=H1=H2,
Fo= F1 ∪ F2 = { &lt; x, max(μF1 (x), μF2(x)), min(ν F1(x),ν F2(x)) &gt; | x ∈ X }
Mathematical notation: Ci1 ∪ Ci2 = Co</p>
      <p>Difference (-):. The difference operator removes the portion of the cube Ci1 that is
common to both cubes. Ci1 and Ci2 have to be union compatible
Input: Ci1 = (D1, l1, F1, O1, H1) and Ci2 = (D2, l2, F2, O2, H2)
Output: Co= (Do, lo, Fo, Oo, Ho) where Do=D1=D2, lo=l1=l2, Oo=O1=O2, Ho=H1=H2,</p>
      <p>Fo= F1 ∩ F2 = { &lt; x, min(μF1(x),μF2(x)), max(νF1(x),νF2(x)) &gt; | x ∈ X }
Mathematical notation: Ci1 – Ci2 = Co</p>
      <p>Aggregation (A): An aggregation operator A is a function A(G) where G = {&lt;x,
μF(x) , νF(x)&gt;| x∈ X } where x=&lt;att1, …,attn&gt; is an ordered tuple belonging to a
given universe X, {att1, …, attn} is the set of attributes of the elements of X, μF(x) and
νF(x) are the degree of membership and non-membership of x. The result is a bag of
the type {&lt;x′, μF(x′ ) , νF(x′)&gt;| x′∈ X }. To this extent, the bag is a group of elements
that can be duplicated and each one has a degree of μ and ν.</p>
      <p>Input: Ci = (D, l, F, O, H) and the function A(G)
Output: Co = (D, lo, Fo , Oo , Ho)</p>
      <p>The definition of the extended group operators allows us to define the extended
group operators Roll up (Δ), and Roll Down (Ω).</p>
      <p>Roll up (Δ): The result of applying Roll up over dimension di at level dlr using the
aggregation operator A over a datacube Ci = (Di ,li ,Fi , O , Hi ) is another datacube
Co = (Do ,lo ,Fo , O , Ho ).</p>
      <p>Input: Ci = (Di ,li ,Fi , O , Hi )
Output: Co = (Do ,lo ,Fo , O , Ho )
ω is the initial state of the cube
An object of type history is a recursive structure H = (l, D, A, H’) is the state of the
cube after performing an
operation on the cube</p>
      <p>The structured history of the datacube allows us to keep all the information when
applying Roll up and get it all back when Roll Down is performed. To be able to apply
the operation of Roll Up we need to make use of the IFSUM aggregation operator.</p>
      <p>Roll Down (Ω): This operator performs the opposite function of the Roll Up
operator. It is used to roll down from the higher levels of the hierarchy with a greater
degree of generalization, to the leaves with the greater degree of precision. The result
of applying Roll Down over a datacube Ci = (D, l, F, O, H) having H=( l’, D’, A’, H’ )
is another datacube Co= (D’, l’, F’, O, H’).</p>
      <p>Input: Ci=(D, l, F, O, H)
Output: Co=(D’,l’, F’, O, H’) where F’ set of fact instances defined by operator A.</p>
      <p>To this extent, the Roll Down operative makes use of the recursive history structure
previously created after performing the Roll Up operator.</p>
      <p>
        The definition of aggregation operator points to the need of defining the IF
extensions for traditional group operators [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], [10], [11] such as SUM, AVG, MI? and
MAX. Based on the standard group operators, we provide their IF extensions and
meaning.
      </p>
      <p>IFSUM : The IFsum aggregate, like its standard counterpart, is only defined for
numeric domains. Given a fact F defined on the schema X (att1, …,attn), let attn-1
defined on the domain U={u1 , …, un ). The fact F consists of fact instances Fi with 1
≤ i ≤ m. The fact instances Fi are assumed to take Intuitionistic Fuzzy values for the
attribute attn-1 for i = 1 to m we have Fi[attn-1] = {&lt;μi(uki), νi(uki)&gt;/ uki | 1 ≤ ki ≤ n } .
The IFsum of the attribute attn-1 of the fact table F is defined by:</p>
      <p>IFSUM((attn-1)(F)) =
{&lt;u&gt;/ y | (( u= min im=1 (μi(uki), νi(uki)) ∧ (y = ∑ki=k1uki ) (∀ k1, …km : 1 ≤ k1, …km ≤ n))}
km</p>
      <p>IFAVG : The IFAVG aggregate, like its standard counterpart, is only defined for
numeric domains. This aggregate makes use of the IFSUM that was discussed
previously and the standard COU?T. The IFAVG can be defined as:</p>
      <p>IFAVG((attn-1)(F) = IFSUM((attn-1)(F)) / COU?T((attn-1)(F))</p>
      <p>IFMAX : The IFMAX aggregate, like its standard counterpart, is only defined for
numeric domains. The IFsum of the attribute attn-1 of the fact table F is defined by:</p>
      <p>IFMAX((attn-1)(F)) =
{&lt;u&gt;y|((u= min im=1 (μi(uki),νi(uki))∧ (y= maxim=1 (μi(uki),νi(uki)))(∀k1,…km :1≤k1,…km≤ n))}</p>
      <p>IFMI) : The IFMI? aggregate, like its standard counterpart, is only defined for
numeric domains. Given a fact F defined on the schema X (att1, …,attn), let attn-1
defined on the domain U={u1 , …, un ). The fact F consists of fact instances fi with 1 ≤
i ≤ m. The fact instances fi are assumed to take Intuitionistic Fuzzy values for the
attribute attn-1 for i = 1 to m we have fi[attn-1] = {&lt;μi(uki), νi(uki)&gt;/ uki | 1 ≤ ki ≤ n } .
The IFsum of the attribute attn-1 of the fact table F is defined by:</p>
      <p>IFMI?((attn-1)(F)) =
{&lt;u&gt;/ y|(( u= minim=1 (μi(uki),νi(uki))∧(y= minim=1 (μi(uki),νi(uki)))(∀k1,…km :1≤ k1,…km≤ n))}</p>
      <p>We can observe that the IFMI? is extended in the same manner as IFMAX aggregate
except for replacing the symbol max in the IFMAX definition with min.</p>
      <p>Once we have defined our Intuitionistic Fuzzy multidimensional model and have
defined the IF cubic-algebra, the concept of knowledge based OLAP is introduced.</p>
    </sec>
    <sec id="sec-4">
      <title>4. The Case for Knowledge Based OLAP-K,OLAP</title>
      <p>Let us consider the Intuitionistic fuzzy set M defined as: {Milk&lt;0.8,0.1&gt;,
WholeMilk&lt;0.7,0.1&gt;, Condensed-Milk&lt;0.4,0.3&gt;}} which is presented in “Fig.1”. Then the
next step is to calculate the &lt;μ, ν&gt; values for “Pasteurized milk”, “Whole Pasteurized milk
” and “Condensed whole milk.”
• If the hierarchical IF structure expresses preferences in a query, the choice of the
maximum values for μ and minimum value ν from the pairs of values &lt;μ, ν&gt;
from the parent elements to the sub elements allows us not to exclude any
possible answer (high possibility necessity degrees). In real cases, the lack of
answers to a query generally makes this choice preferable, because it consists of
widening the query answer rather than restricting it.
• If the hierarchical IF represents an ill-known concept, the choice of the maximum
value for μ and minimum value ν allows us to preserve all the possible values,
but it also makes the answer less specific. In a way, it also participates in
enlarging the query, as a less specific datum may share more common values
with the query (the possibility degree of matching can thus be higher, although
the necessity degree can decrease).</p>
      <p>Pasteurizedmilk</p>
      <p>Milk
&lt;0.8,0.1&gt;
Wholemilk
&lt;0.7,0.1&gt;</p>
      <p>Condensedmilk
&lt;0.4,0.3&gt;</p>
      <p>Pasteurizedmilk
&lt;0.8,0.1&gt;</p>
      <p>Condensedmilk
&lt;0.4,0.3&gt;
Milk
&lt;0.8,0.1&gt;
Wholemilk
&lt;0.7,0.1&gt;
Wholepasteurizedmilk</p>
      <p>Condensedwholemilk</p>
      <p>Wholepasteurizedmilk
&lt;0.8,0.1&gt;</p>
      <p>Condensedwholemilk
&lt;0.7,0.1&gt;
“Fig.2” is a fully weighted Hierarchy after applying the maximum values for μ and
minimum value ν from the pairs of values &lt;μ, ν&gt; from the parent elements to the sub
elements, i.e. from (whole-milk, condensed-milk) to (condensed-whole-milk), from (milk)
to (pasteurized milk), and from (whole-milk, pasteurized milk) to
(pasteurized-wholemilk).</p>
      <p>The complete study of the hierarchical IF requires the formal definition of the IF
hierarchical closure. We will further need to formally define the containment of an IF
hierarchical set to another.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>In this paper we have presented a new multidimensional-cubic model named as
the IF-Cube. The main contribution of this new model is that is able to operate over
data with imprecision in the facts and the summarisation hierarchies. Classical
models imposed a rigid structure that made the models present difficulties when
merging information from different but still reconcilable sources. We introduce the
automatic recommendation of analysis according to the context of users’ explorations
in order to guide the decision making with the aid of Intuitionistic fuzzy set over a
universe that has a hierarchical structure and the corresponding hierarchies.</p>
      <p>There is finally a need to formally define the closure of Intuitionistic fuzzy set over
a universe that has a hierarchical structure as well the containment between different
versions of these sets.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Atanassov</surname>
          </string-name>
          (
          <year>1999</year>
          ).
          <source>Intuitionistic Fuzzy Sets</source>
          , Springer-Verlag, Heidelberg
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>K.</surname>
          </string-name>
          ,
          <source>Atanassov Intuitionistic Fuzzy Sets, Fuzzy Sets and Systems</source>
          ,
          <volume>20</volume>
          ,
          <fpage>87</fpage>
          -
          <lpage>96</lpage>
          , (
          <year>1986</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kimball</surname>
          </string-name>
          ,
          <article-title>The Data Warehouse Toolkit</article-title>
          . New York: John Wiley &amp; Sons,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            , Chaudhuri U. Dayal V.,
            <surname>Ganti</surname>
          </string-name>
          .
          <article-title>Database Technology for Decision Support Systems</article-title>
          . In: Computer, Vol.
          <volume>34</volume>
          , p.
          <fpage>48</fpage>
          -
          <lpage>55</lpage>
          ,
          <year>2001</year>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Jarke</surname>
          </string-name>
          et al.
          <source>Fundamentals of data warehouses</source>
          . Springer, London,
          <year>2002</year>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Thomas</surname>
          </string-name>
          &amp; A.,
          <string-name>
            <surname>Datta</surname>
            ,
            <given-names>A Conceptual</given-names>
          </string-name>
          <string-name>
            <surname>Model</surname>
          </string-name>
          and
          <article-title>Algebra for On-Line Analytical Processing in Decision Support Databases</article-title>
          .
          <source>Information Systems Research</source>
          <volume>12</volume>
          :
          <fpage>83</fpage>
          -
          <lpage>102</lpage>
          ,
          <year>2001</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Dyreson</surname>
          </string-name>
          ,
          <article-title>Information retrieval from an incomplete data cube</article-title>
          , VLDB, Morgan Kaufman Publishers, pp.
          <fpage>532</fpage>
          -
          <lpage>543</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Pedersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jensen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Dyreson</surname>
          </string-name>
          ,
          <article-title>A foundation for capturing and querying complex multidimensional data</article-title>
          ,
          <source>Information Systems</source>
          , vol.
          <volume>26</volume>
          , pp.
          <fpage>383</fpage>
          -
          <lpage>483</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Dubois</surname>
          </string-name>
          et al (
          <year>1988</year>
          ).
          <article-title>Handling Incomplete or Uncertain Data and Vague Queries in Database Applications</article-title>
          , Plenum Press,
          <year>1988</year>
          [10]
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Prade</surname>
          </string-name>
          (
          <year>1993</year>
          ).
          <article-title>Annotated bibliography on fuzzy information processing</article-title>
          .
          <source>Readings on Fuzzy Sets in Intelligent Systems</source>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Prade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dubois</surname>
          </string-name>
          , and R. Yager, Eds. Morgan Kaufmann Publishers Inc.,
          <year>1993</year>
          [11]
          <string-name>
            <given-names>E. Rundensteiner L.</given-names>
            <surname>Bic</surname>
          </string-name>
          . Aggregates in posibilistic databases,
          <source>VLDB'89</source>
          , pp.
          <fpage>287</fpage>
          -
          <lpage>295</lpage>
          ,
          <year>1989</year>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>