<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On Open Problem - Semantics of the Clone Items</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Juraj Macko</string-name>
          <email>fjuraj.mackog@upol.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. Computer Science Palacky University</institution>
          ,
          <addr-line>Olomouc 17. listopadu 12, CZ-77146 Olomouc</addr-line>
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <fpage>130</fpage>
      <lpage>145</lpage>
      <abstract>
        <p>There was presented a list of open problems in the Formal Concept Analysis area at the conference ICFCA 2006. The problem number seven deals with the semantics of the clone items. Namely, for whom can clone items make sense and for whom can make sense the item, which can cause, that clones disappear in the collection of itemsets. In this paper we propose the semantics behind clone items with the couple of examples. De nition of the clone items is very strict and theirs use could be very limited in the real datasets. We introduce method, how to deal with items, which properties are very near to the clones. We also have a look on the items, which causes the disappearing of the clones, or decrease (increase) the degree of property "to be clone". In the experiment part we analyze some known datasets from the clone items point of view. The results bring a couple of new questions for the future research.</p>
      </abstract>
      <kwd-group>
        <kwd>formal concept analysis</kwd>
        <kwd>clone items</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        This paper is structured as follows: The rst part, which is actually cited from
the source, where the problem were de ned [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] describes and de nes the whole
problem - the semantics of the clone items. In the second part is proposed the
semantics of the clone items by putting the problem into the other point of
view. There is also a discussion here, about another possible de nitions of the
clones as presented in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In this part three comprehensive examples can be
found. The third part tries to set a quite new approach to the clone items.
The attributes, which are not clones, but they have properties very close to
clones are considered. A nearly clones are de ned. In this part some results from
the introductory experiments about the clones and nearly clones are presented.
Finally, the conclusion is divided in two parts - conclusion of de ned problem
and conclusion of other proposed issues.
      </p>
    </sec>
    <sec id="sec-2">
      <title>The Problem Setting</title>
      <p>
        The proposed problem of the semantics of the clone items were proposed and
de ned in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] as follows: Let J be a set of items x1; :::; xjJjJ, ldeet Fnebdebay cfoolllleocwtiinong
of subsets of J and let 'a;b be the mapping 'a;b : 2J ! 2
formula:
      </p>
      <p>8 (Xnfag) [ fbg if b 2= X and a 2 X
X ! 'a;b(X) = &lt; (Xnfbg) [ fag if a 2= X and b 2 X</p>
      <p>: X elsewhere
It means swapping items a and b, which are called clone items in F i for any
F 2 F , we have 'a;b(F ) 2 F . A Clone-free collection is, if it does not contain
any clone items.</p>
      <p>Let (X; Y; I) be a formal context such that attributes a 2 Y and b 2 Y are
not clones. Consider the formal sub-context (X; Z; I), where Z Y , such that a
and b are clone in (X; Z; I). Let c 2 Y nZ such that a and b are no longer clone
in (X; Z [ fcg; I). Attributes a and b has symmetrical behaviour in (X; Z; I),
but this behaviour is lost when we add the attribute c to the formal context.
The following question are asked:
1. Does such symmetrical behaviour of a and b make sense for someone?
2. Does it make the sense, that such symmetrical behaviour disappears, when
the attribute c is added?
3. What is semantics behind the attributes a, b, and c?</p>
    </sec>
    <sec id="sec-3">
      <title>Semantics behind Clones 3</title>
      <p>3.1</p>
      <sec id="sec-3-1">
        <title>Semantics behind Clones - Auxiliary Formal De nitions</title>
        <p>The collection of itemsets will be de ned as a formal context (X; Y; I), where X
is a set of objects and Y is a set of attributes. Objects and attributes are related
by I X Y , which means, that the object x 2 X has the attribute y 2 Y . For
A X, B Y and formal context (X; Y; I) we de ne operators
A"I = fy 2 Y j for each x 2 X : hx; yi 2 Ig</p>
        <p>B#I = fx 2 X j for each y 2 Y : hx; yi 2 Ig
The two given attributes a; b 2 Y will be investigated, whether are clones or not.
For this purpose the pivot table will be de ned as the relation R P N ,
where P = fa; bg Y and N is a set of all Nj, where j 2 [1; jN j]. Nj 2 N
represents the set of attributes Nj = fxg"I \ (Y nP ) for each x 2 X such that
fa; bg \ fxg"I 6= ; and fa; bg * fxg"I . The investigated attributes a; b 2 P Y
will be named the pivot attributes and all other considered attributes, hence
n 2 SjjN=1j Nj, we denote as the non-pivot attributes. Nj is a set generated
by pivot attributes (or shortly the generated set). The pivot table has two
rows. The "cross" in pivot table will represent the fact, that in the formal
context there exists at least one row, where the investigated attribute a (or b
respectively) appears together with the attributes in the particular Nj . Formally,
ha; Nj i 2 R i in context (X; Y; I) exists x 2 X such that x"I = fag [ Nj ;
hb; Nj i 2 R i in context (X; Y; I) exists x 2 X such that x"I = fbg [ Nj :
Based on pivot attributes, non-pivot attributes and formal context (X; Y; I)
consider pivot table which is as new formal context (P; N ; R) with operators
for C P and D N de ned as follows</p>
        <p>C"R = fNi 2 N j for each p 2 P : hp; Nj i 2 Rg;</p>
        <p>D#R = fp 2 P j for each Ni 2 N : hp; Nj i 2 Rg
In the pivot table (P; N ; R) we are trying to nd whether fag"R = fbg"R . In
a
b</p>
        <p>g
cg g n32g ;;cn32
;2 3g 3g ;c3 ;
;n1 ;n2 ;n1 ;n1 ;n1 ;n1
nf nf nf nf nf nf
= = = = = =
1 2 3 4 5 6
N N N N N N
other words, we want to know, whether the attribute a appears in given formal
context with the same combination of other attributes, as b appears (in the same
formal context). If yes, the pivot attributes a; b in context (X; Y; I) generates the
same generated sets. In other words, a; b are not unique with respect to the
nonpivot attributes. Such attributes we call clones. When fag"R 6= fbg"R , attributes
a; b are unique with respect to the non-pivot attributes, because generates at least
one di erent generated set. The attribute c 2 Y , which makes a; b unique with
respect to generated sets is called the originality factor of a; b. In Figure 1 we
show examples of the contexts and pivot tables with clones or with the originality
factor respectively. By introducing the pivot table, the whole problem have been
put to the other point of view. The proposed semantics will be explained based
on the previous de nitions.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Discussion and Remarks</title>
        <p>Before the comprehensive examples will be proposed, it is necessary to discuss
previous auxiliary de nition of the clones using the pivot table. There are couple
of problems mainly dealing with ambiguity of the pivot table de nition with
respect to the various de nitions of the clones used by the several authors in
the other works. In the pivot table de nition the set Nj 2 N is de ned as
Nj = fxg"I \ (Y nP ) for each x 2 X such that
1. fa; bg \ fxg"I 6= ; and
2. fa; bg * fxg"I .</p>
        <p>The rst condition tells, that we ignore the itemsets (rows), where neither a
nor b is present. Such items are not interesting when we investigate whethe a and
b are clones, so we will ignore them when the pivot table is de ned. The second
condition excludes itemsets, where we have the both pivot attributes a and b
and the question is: Why we exclude such itemsets from pivot table, when we
can see it in original de nition of the clone items? Recall the original de nition
of the clones:</p>
        <p>X ! 'a;b(X) =
8 (Xnfag) [ fbg if b 2= X and a 2 X
&lt;</p>
        <p>(Xnfbg) [ fag if a 2= X and b 2 X
: X elsewhere</p>
        <p>Items a and b, which are called clone items in F i for any F 2 F , we have
'a;b(F ) 2 F . So we need to have the original itemset and swapped itemset as
well in the whole collection of itemsets. In de nition of ' are interesting the
rows 1 and 2. The row 3 is only technical condition. It means, that ful llment
of swapping condition of itemsets, which does not contain any of a or b or
conversely, when it contains both, is trivial. So we could add them in the pivot
table by skipping the condition fa; bg * fxg"I , but we consider such
information redundant and hence useless. However, the semantics of the clones remains
unchanged. But on the other side, it can in uence the value of the degree of</p>
        <p>I
clones d(a;b) (which will be de ned later). In such case we need to investigate,
which de nition would be more precise for the user. The basic idea of our
semantics of clones (and nearly clones de ned later as well) is, of how original are
items a and b in the whole collection of itemsets. The itemsets which does not
include either a or b will not tell us anything about originality of such items, the
itemsets which include both as well.</p>
        <p>
          The other point for the discussion comes from the problem number six
(presented in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]), which deals with the size of a clone-free Guigues-Duquenne basis.
Namely, whether the clone items are responsible for the combinatorial
explosion of some Guigues-Duquennes basis. The Guigues-Duquennes basis is
nonredundant. All other attribute implications, which holds in given context, can
be derived from this base. In the paper [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] there are presented some partial
results, which includes de nitions and propositions dealing with the clones. The
clones are de ned with respect to pseudo-closed sets in the collection of the
closed itemsets. The one of the basic results is, that in order to detect clone
items, one has to consider meet-irreducible itemsets only (for details see [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]).
The de nition of the clone items given in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] is de ned in more general manner.
It is based not only on the pseudo-closed itemset collection, but it is de ned for
arbitrary collection of itemsets. This fact can cause, that two items may not
appear as clone according the de nition in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], but the are still clones in de nition
according to [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. In the rest of the paper there will be considered the de nition
used in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] only. However, the proposed semantics would be slightly modi ed,
when we would need to use it in the meaning of [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>The other important part is to compare proposed solution with other
attempts or solutions, but the author has no information either about such
attempts or about some real solutions. Hence, according to the author's best
knowledge, the author's proposed solution seems to be novel.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Semantics behind Clones - Examples</title>
        <p>In this part we would like to show on couple of examples, how the clone items
and the originality factor can be used. The originality factor can be desired under
some conditions, but undesired under the other conditions. Inall examples the
same formal context and the pivot tables will be used, but always with the
different meaning of the objects and attributes. The Table 1 represents the original
formal context (X; Z; I) with the clones a and b and it also represents the formal
context (X; Z [ fcg), where the originality factor c is added. The corresponding
pivot tables (P; N ; R) and (P; Nc; Rc) can be seen in the Table 2. A labeling of
the objects and the pivot attributes is done according to the particular sets X
and Y de ned in each example below.</p>
        <p>The sales analysis Let X = fCustomer1; : : : ; Customer8g be a set of
customers and the set of attributes is de ned as Y = fM an; W oman; n1; n2; n3; cg.
The attributes M an and W oman represents the sex of customer and the other
attributes represents the products bought by each customer. The following formal
contexts represents a marketing research of the sales company (the customers and
theirs attributes). In the formal context (X; Z; I) attributes M an and W oman
are clones. In the pivot table (P; N ; R) attributes M an and W oman are pivot
attributes and nj is product bought by customer. On the other hand, in the
formal context (X; Z [ fcg) and the corresponding pivot table (P; Nc; Rc) the
attributes M an and W oman are no longer clones and the attribute c (Product
c) is the originality factor in this case. Namely, for the itemset fM an; n1; n3; cg
there is no corresponding itemset fW oman; n1; n3; cg.</p>
        <p>How can this information be used for the marketing department? Imagine,
that the sales company wants to create packages based on the marketing
research. These packages should consist of the particular products nj . In the rst
2
e
n
e1 eG
en /
/G icra
e e
op m
r A
/uE /an
n m
a o
M Wn1 n2 n3 c
Customer 1 / Animal 1 / Organism 1
Customer 2 / Animal 2 / Organism 2
Customer 3 / Animal 3 / Organism 3
Customer 4 / Animal 4 / Organism 4
Customer 5 / Animal 5 / Organism 5
Customer 6 / Animal 6 / Organism 6
Customer 7 / Animal 7 / Organism 7
Customer 8 / Animal 8 / Organism 8
Table 1. Formal contexts (X; Z; I) and (X; Z [ fcg; Ic)</p>
        <p>N</p>
        <p>3g
2g 3g 3g ;n2</p>
        <p>Nc</p>
        <p>g
cg g n32g ;;cn32
2 3g 3g ;c3 ;
;
;n1 ;n2 ;n1 ;n1</p>
        <p>;n1 ;n2 ;n1 ;n1 ;n1 ;n1
nf nf nf nf nf nf nf nf nf nf
= = = =
1 2 3 4
N N N N
= = = = = =
1 2 3 4 5 6
N N N N N N
case of the formal context (X; Z; I) the company can create the same packages
for man and for woman, because male and female customers buy the same
combinations of products nj. The same packages for two di erent groups can reduce
the total cost of production, because we need to produce only four types of the
packages, namely the packages N1 = fn1; n2g, N2 = fn2; n3g, N3 = fn1; n3g and
N4 = fn1; n2; n3g. With the attribute c added to the formal context, we need six
di erent packages, because only the packages N1 = fn1; n2g and N2 = fn2; n3g
can be produced for men and women at the same time. Other packages are
different for the male and female customers. From this point of view, the originality
factor is undesired and the clones are desired.</p>
        <p>But we can use this information in the other way. Suppose, that the cost
di erence of producing four or six package types is not signi cant, but signi cant
can be a targeted marketing on the male and female customers. The formal
context (X; Z; I), where we have the clone attributes M an and W oman, does
not provide di erentiated information about the male and female customers. On
the other hand, the formal context (X; Z [ f g
c ) does. The attribute c provides
desired information, that the Product c in uences the di erent combination
of the products bought by the male and female customer. It means, that we
can make targeted marketing (namely, the di erent type of packages for the
di erent type of customers) based on the originality factor Product c and its
combinations with the other products. Some combinations of the products with
the originality factor can be used as a topic for advertising to highlight the
di erence between man and woman preferences. However, the clone analysis can
provide the marketing department with the useful information in both cases.
Analysis of the animals Let X = fAnimal1; : : : ; Animal8g be a set of
animals and a set of attributes is de ned as Y = fEurope; America; n1; n2; n3; cg.
The formal context (X; Z; I) in the Table 1, shows the attributes Europe and
America as clones. This fact can be interpreted as follows: In Europe and in
America they live the same types of animals, when we consider the attributes of
the animals n1, n2 and n3 only. The same information can be seen in the pivot
table Table 2. When we add the attribute c, we can see the di erent types of
animals (with the di erent generated sets) in Europe and in America as well
(see Table 2). The information, that exists the originality factor c for attributes
Europe and America can be interpreted as follows: It shows, that Europe and
America are somehow speci c. In Europe are some di erent combinations of
animal's attributes than in America and vice versa and at the same time we see,
that this di erence somehow deals with the attribute c. Biologist can
investigate in more details, what is speci c in Europe and in America, which speci c
attribute of Europe leads to the di erent attributes of the animals in Europe
(and vice versa). Other use of such information is following: From a background
knowledge we know, that there is no reason for di erentiating the animals in
Europe and America just on attribute c. In our dataset we do not have in America
the animal with attributes n1, n3 and c, but with respect to the attribute c we
expect to have the same types of animals in Europe and in America. Thus, we
need to look for such animal in America as well. Our hypothesis is, that in
America lives such animal, because it lives in Europe and based on our background
knowledge there is no reason for c to be the originality factor. From the formal
point of view, we do not have the complete dataset (formal context). Some rows
are missing, and we need to nd such objects in the reality (in this case we are
looking for the animal).</p>
      </sec>
      <sec id="sec-3-4">
        <title>Analysis of genes and the morphological attributes of organisms The</title>
        <p>last example use set X = fOrganism1; : : : ; Organism8g and set of attributes
Y = fGene1; Gene2; n1; n2; n3; cg The attributes Gene1 and Gene2 are clones
in formal context (X; Z; I) in Table 1 and the other attributes represents the
morphological property of the organism. The interpretation can be following:
Organisms with Gene1 and Gene2 has the same combination of morphological
properties Nj, when we consider the morphological properties of organisms n1,
n2 and n3. The same information can be seen in the pivot table Table 2. When we
add the morphological attribute c, we get the formal context (X; Z [ fcg), which
means, that based on attribute c there are some di erent types of the
morphological attributes of organisms with the Gene1 and Gene2 (see the Table 1 and
Table 2). It shows, that the Gene1 and Gene2 probably does not in uence the
sets of the morphological attributes containing only n1; n2; n3, but this Gene1
and Gene2 in uence the sets of the morphological attributes containing c. Thus,
c as the originality factor makes the di erence between these two genes. This
information could be useful for a hypothesis creation in genetics.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Nearly Clones</title>
      <sec id="sec-4-1">
        <title>Degree of Clones and Degree of Originality</title>
        <p>The de nition of the clone items is very strict. Recall, that condition 'a;b(F ) 2 F
needs to be true for any F 2 F . We can see, that adding only one "cross" into
the huge formal context can cause, that two clones disappear. We expect, that
in real dataset such condition can be true very rarely. When we want to use
the clone items meaningfully, we need to have a weaker de nition. For practical
purposes it su ces, that condition 'a;b(F ) 2 F can be true in some reasonable
amount of F 2 F . We de ne degree of clone as
d(Ia;b) = jfag"R \ fbg"R j ;</p>
        <p>jfag"R [ fbg"R j
which can be read as follows: The attributes a and b with respect to the formal
context I are clones in the degree d. For a priori given threshold we de ne a
and b as nearly clones i d(a;b) . Note, that for d(a;b) = 1 the attributes a
and b are clones and for d(a;b) = 0 we say, that they are original attributes.
Consider now the formal context (X; Z; IZ ) and the corresponding pivot table
(P; NZ ; RZ ) (see Figure 2). We can see, that a and b are clones with the degree
d(IaZ;b) = 1. Adding either attribute c1 or attribute c2 to the formal context leads
to decreasing of clone degree for a and b. Namely, d(Iac1;b) = 0 and d(Iac2;b) = 0; 6. In
both cases degree has decreased, but the resulted clone degree is di erent. In the
rst case attributes are original, in the second case attributes are nearly clones
for arbitrary 0; 6. Such situation can be formalized, and de ne the degree
of originality for given ci and context (X; Z; I) as
g(Iac;b) = d(a;b)</p>
        <p>I
d(Iac;b)</p>
        <p>The degree of originality shows, how the attribute, added to the context,
does in uence the degree of clone for given attributes a; b 2 Y and the formal
context (X; Z; I).</p>
        <p>
          3g
2g 3g 3g ;n2
;n1 ;n2 ;n1 ;n1
nf nf nf nf
= = = =
1 2 3 4
N N N N
For the purpose of this paper we arranged two introductory experiments with the
nearly clones, in which we use datasets Mushroom [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], Adults [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and Anonymous
[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] from well known UC Irvine Machine Learning Repository (for the details
see Table 3). In the experiments we used a naive algorithm (the brute-force
search, but with polynomial complexity) for nding the degrees of clones as
de ned above. Looking for more e cient algorithm is out of scope of this paper.
The algorithm was implemented in C, and all experiments have been run on
the computer with an Intel Core i5 CPU, 2.54 Ghz, 6 GB RAM, 64bit W7
Professional.
        </p>
        <p>In the rst experiment we were focused on nding all nearly clone pairs, with
d(a;b) &gt; 0, especially we investigated, if there are some clones (where d(a;b) = 1)
in the real datasets. The results of the rst experiment are shown in Table 3.
(i) nearly clone pairs for d(a;b) &gt; = 0
x-axis - number of nearly clone pairs</p>
        <p>y-axis - degree of clone d(a;b)
z-axis - number of objects processed
(ii) nearly clone pairs for d(a;b) = 0:5
x-axis - number of nearly clone pairs</p>
        <p>y-axis - degree of clone d(a;b)
z-axis - number of objects processed</p>
        <p>In case of the dataset Mushroom, we present also the distribution of the clone
degrees and some other details as well. Figure 3 shows the volume of all pairs
a and b and clone degree d(a;b) &gt; 0, for each scale pattern (from 1000 to 8124
by 1000). In (i) are displayed all pairs with d(a;b) &gt; 0 and part (ii) is more
focused on the amount of pairs where d(a;b) 0; 5 for each investigated scaled
pattern. Note, that the results from numbers of the processed objects in the
dataset M ushroom (namely from 1000 to 7000 depicted in z-axis in the Figure
3) depends on an order of the processing objects. This fact were not investigated
more deeply. However, when we have processed all 8124 objects, the order will
not in uence the result. Figure 4 shows some interesting details. In (i) there are
presented the pairs a, b with d(a;b) = 1, in (ii) the same for 1 &gt; d(a;b) 0; 5 . We
have found 4 clone pairs, and one clone triple. In the clone triple (103; 104; 105)
we can see the transitivity (i.e. when (a; b) are clones and b; c are clones, also a
and c are clones). Such transitivity is not surprising and is direct consequence
of the clone de nition.</p>
        <p>What does such results show and does it appear reasonable? The Figure 4
part (i) shows the clones a and b. The original dataset M ushroom consists of 22
attributes with non-binary values. For the purposes of clone investigation, this
dataset were nominally scaled to the formal context, which is binary indeed. It
is interesting to see, that all clone items represents the value of the same original
attribute. E.g. clones 019 and 021 represents the original attribute Cap Color,
thus its values P urple or W hite respectively. Another example is clone triple 103,
104 and 105 which represents the original attribute Spore P rint Color with the
corresponding values Orange, P urple and W hite. It can be interpreted as
follows: Purple and white color generates the same sets of the non-pivot attributes.
In other words, to each mushroom with the purple cap (the pivot attribute),
there exists corresponding mushroom with the white cap (the pivot attribute),
but all other properties remains he same (non-pivot attributes). Similarly to each
mushroom with the purple spore print color, there exists corresponding
mushroom with the white spore print color and the corresponding mushroom with the
orange one. When we look on the nearly clones in the Figure 4 part (ii), the
attributes 69 and 70 represents the same original attribute stalk color above ring
with the values cinnamon and gray (the details are not shown in the table).
These attributes are not the clones, but the nearly clones with the clone degree
d(a;b) = 0; 96. It can be interpreted similarly as by the clones. Only the di erence
will be in a quanti er. By clones the quanti er was "for each", by the nearly
clones we will have fuzzy quanti er, in this case "for the most". Hence the
interpretation is: For the most mushroom with cinnamon stalk above the ring exists
corresponding mushroom with the corresponding gray stalk above the ring (and
vice versa). The clone degree is very high in this case (d(a;b) = 0; 96) it means
there are only couple of mushrooms with cinnamon stalk above the ring color,
which do not have corresponding mushroom with the gray stalk above the ring
color. However, the for the deeper understanding of such examples, it is required
to ask an expert in mycology.
4.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Experiment 2 - Structure of Nearly Clones in Datasets</title>
        <p>In the second experiment we investigated the structure of nearly clones. Namely,
we have de ned a fuzzy relation T : Y Y ! L, where L = [0; 1] is de ned as
T (a; b) = d(a;b) 2 L. In other words, the fuzzy relation express the degree of the
clone for each pair a; b 2 Y . For the better visualization we display such relation
in so called "bubble chart". The bubble chart displays three dimensional data
in two dimensional chart. The position of the bubble is given by two dimensions
(x and y axis) and the size of the bubble shows the third dimension. The results
from the rst experiments are displayed in bubble chart, where the pairs of
attributes a and b represents two dimensions and the degree of clone d(a;b) is
represented by the size of the bubble. Note, that the fuzzy relation T is indeed
symmetric (i.e. d(a;b) = d(b;a)), but we show only part of the relation, where
a &lt; b. Figures 5, 6 and 7 show the structure of the nearly clones for the datasets
Mushroom, Anonymous and Adults. We can observe very di erent structure of
the nearly clones in each dataset. In Mushroom we can see, that the structure
of the nearly clones is approximately linear. All nearly clones are clustered near
to the line de ned as (y; y). In the case of Adults dataset we can see more
spread, but still approximately linear structure, except of one cluster near point
(0; jY j). The nearly clones of the data set Anonymous forms the di erent, but
kind of regular structure as well. This results leads to the question, what kind of
properties has fuzzy relation of the nearly clones and if properties of such fuzzy
relation correlates with the properties of the formal context, or with properties
of the concept lattice. Until now we know, that such relation is transitive for
clone items d(a;b) = 1 and symmetric for the arbitrary nearly clone items, but
this two properties are trivial. I would be also interesting to nd the semantics
of such fuzzy relation de ned on the nearly clones. All this will be part of the
future investigation.</p>
        <p>original attributes a b
03. cap-color 019 purple=u 021 white=w
05. odor 025 almond=a 026 anise=l
05. odor 030 musty=m 031 none=n
17. veil-color 086 brown=n 087 orange=o
20. spore-print-color 103 orange=o 104 purple=u
20. spore-print-color 103 orange=o 105 white=w
20. spore-print-color 104 purple=u 105 white=w</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Perspectives</title>
      <p>
        The paper was motivated by open problem proposed at ICFCA 2006 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We
hope, that this small open problem is solved now and the reason is presented
in the rst part of the conclusion. This part is structured as a direct answers
on proposed questions. The second part of the conclusion describes ideas, which
overlaps the original open problem and come with some new questions.
Nearly clone pairs for d(a;b) &gt; = 0 , where x,y-axis - a and b pairs
size of bubble = d(a;b)
Nearly clone pairs for d(a;b) &gt; = 0 , where x,y-axis - a and b pairs
size of bubble = d(a;b)
Question 1: Does the symmetrical behaviour of a and b make sense for someone?
Answer 1: Yes, such symmetrical behaviour can identify the same combination
of the non-pivot attributes with respect to pivot attributes and can make sense:
{ for the marketing department to reduce cost of packages - the clone items
enable the same packages for the di erent types of customers (e.g. man and
woman)
{ for biologists to complete the the dataset - the clone items are expected,
because the originality factor c, has no sense based on the background
knowledge. Hence, we some miss rows in the dataset (e.g. we need to nd the new
animals)
{ for genetics - it bring an information that the two genes has no in uence on
a combination of the morphological properties of organisms
{ generally for everyone, who needs an information about the same
combination of non-pivot attributes with respect to the pivot attributes
Question 2: Does it make sense, that such symmetrical behaviour disappear,
when c is added?
Answer 2: Yes, such attribute is called the originality factor for the items a
and b and can be useful:
{ for the marketing department to make a targeted marketing for the di erent
types of customers (e.g. man and woman) using unique combination of the
non-pivot attributes
{ for biologist to nd the di erence between two pivot attributes (e.g. Europe
and America) with respect to other non-pivot properties. The originality
factor c reveals, that the pivot attributes are original and this originality
needs to be investigated deeper.
{ for genetics - it brings an information, that two genes has an in uence on a
combination of the morphological properties of organisms
{ generally for everyone, who needs an information about the reason, why the
non-pivot attributes has the di erent combinations with respect to the pivot
attributes.
      </p>
      <p>Question 3: What is semantics behind a, b, and c?
Answer 3: The attributes a and b are the pivot attributes, all other attributes
are the non-pivot attributes and c is moreover the originality factor for the
attributes a and b. The pivot attributes generates a combination of the
nonpivot attributes in the given context. The attribute c make the attributes a and
b unique, which can be "good" or "bad". It depends on a goal of the analysis.
The second part of conclusion shows, that the clones are very strictly de ned.
Therefore the nearly clones were introduced. The nearly clones operates with
the degree, in which two attributes are clones. Such formalization asks itself for
study of the nearly clones under fuzzy setting (e.g. we have already mentioned,
that structure of nearly clones can be seen as fuzzy relation indeed). The
introductory experiments shows, that the nearly clones in dataset have an interesting
structure, which needs to be investigated more deeply. This paper was
introductory for the nearly clones. As a future work we plan to describe more e cient
algorithm to compute the nearly clones for the given threshold , and algorithm
for identifying the originality factors for another given threshold !. Finally we
hope, that this paper, even it does not come with a great mathematical or
experimental results, brings some interesting ideas to FCA community.
Acknowledgements. The author is very grateful to the reviewers for their
helpful comments and suggestions. Partly supported by IGA (Internal Grant
Agency) of the Palacky University, Olomouc is acknowledged.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Gely</surname>
            <given-names>A</given-names>
          </string-name>
          ,,
          <string-name>
            <surname>Medina</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nourine</surname>
            <given-names>L.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Renaud</surname>
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Uncovering and Reducing Hidden Combinatorics in Guigues-Duquenne Bases</article-title>
          . Springer, Lecture Notes in Computer Science,
          <year>2005</year>
          , Volume
          <volume>3403</volume>
          /
          <year>2005</year>
          ,
          <fpage>235</fpage>
          -
          <lpage>248</lpage>
          , Heidelberg 2005
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>2. more authors: Some open problems in Formal Concept Analysis</article-title>
          .
          <source>ICFCA</source>
          <year>2006</year>
          , Dresden, http://www.upriss.org.uk/fca/fcaopenproblems.html
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Schlimmer</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          : Concept Acquisition Through Representational,
          <source>Adjustment (Technical Report 87-19)</source>
          . (
          <year>1987</year>
          ).
          <article-title>Doctoral disseration</article-title>
          , Department of Information and Computer Science, University of California, Irvine.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kohavi</surname>
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Scaling Up the Accuracy of Naive-Bayes Classi ers: a Decision-Tree Hybrid</article-title>
          <source>Proceedings of the Second International Conference on Knowledge Discovery and Data Mining</source>
          ,
          <year>1996</year>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Breese</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heckerman</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kadie</surname>
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Empirical Analysis of Predictive Algorithms for Collaborative Filtering</article-title>
          .
          <source>Proceedings of the Fourteenth Conference on Uncertainty in Arti cial Intelligence</source>
          , Madison,
          <string-name>
            <surname>WI</surname>
          </string-name>
          ,
          <year>July</year>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>