<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Clustering Algorithms for Economic and Psychological Analysis of Human Behavior</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Boyko[</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>rnytsk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>Lviv79013</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article proposes an approach to analyze the behavior of groups of people and shows how to predict person location for the next month. The clustering algorithms were used in this research. Also was inspected the problem of finding associative rules. We used scalable Apriori algorithms to find the best rules. For analysis, we used the standard mlxtend library to aggregate data by cluster, user login, and time. In this article we explored apriori and k-means clustering algorithms to get user behavior analysis template. In the process, we looked at the problem of finding associative rules that were able to find and describe patterns in large datasets. We used scalable Apriori algorithms to find the best rules. For analysis, we used the standard mlxtend library to aggregate data by cluster, user login, and time. While working, we were faced with the problem of inaccuracy and inconsistency of data with real conditions, and were forced to reduce the minimum support for associative rules.</p>
      </abstract>
      <kwd-group>
        <kwd>Preprocessing</kwd>
        <kwd>clustering</kwd>
        <kwd>associative rules</kwd>
        <kwd>Apriori algorithm</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Nowadays, information technologies offer lots of opportunities for communication
and social networking. Such partial isolation is a huge problem for establishing
sociocommunicative relationships between people. One way to partially solve this problem
is to use the meet city. This application uses geolocation and artificial intelligence
methods to predict and recommend meetings. For example, if you desire to talk to a
similar-interest person during your lunch, this application will help [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1-3</xref>
        ].
      </p>
      <p>
        To solve the problems described above we should use machine learning
technologies. We propose to use clustering and associative rules. While working, we have
formed a pattern of finding algorithms and patterns of behavior of people or groups of
people united by common interests [
        <xref ref-type="bibr" rid="ref8">8, 13</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>Statement of the problem</title>
      <p>To begin the analysis, we select the following data (Fig.1):</p>
      <p>
        First of all, we need to vectorize (binarize) the text data of the PlaceName field in
order to be able to apply clustering using the LabelBinarizer class from the standard
sklearn library [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. As a result, we obtain text format data (list) that returns the
compatibility matrix of given element.
      </p>
      <p>from sklearn.preprocessing import LabelBinarizer
encoder= LabelBinarizer()
encoder.fit(df['PlaceName'])
data=encoder.transform(df['PlaceName'])
df1 = pd.DataFrame(data)</p>
      <p>In Fig.2 columns from 0 to 8 are vectors in multidimensional space. A value of 1
in these columns belongs to a particular value of the previous Placename field for a
concrete type.</p>
      <p>
        Since the k-means algorithm is unstable to anomalies, the data must first be
normalized. To do this, we should apply the MinMax Scaling method which moves all
points of space to the limit [
        <xref ref-type="bibr" rid="ref1">0,1</xref>
        ].
      </p>
      <p>X norm </p>
      <p>X  X min
X max  X min
(1)</p>
      <p>To do this, we will use the Normalizer class from the sklearn library from
sklearn.preprocessing import Normalizer
scaler= Normalizer()
scaler.fit(df1)
scaled_df=scaler.transform(df1)
df1= pd.DataFrame(scaled_df)</p>
      <p>As a result of using of the Normalizer class, we obtain the output data of the
preprocessing function (Fig.3)</p>
      <p>In Fig. 3. we received data as a result of normalization and grouping by the value
of login (Fig. 2). Columns 0-10 are measurements of multidimensional space, and
data is a reflection of vectors in them. All data should be normalized using the
MinMax Scaling method [12-15].</p>
      <p>
        In order for the system to be dynamic and allow users to change their behavior
and, as a consequence, to move to other clusters, their number should be determined
automatically. For this purpose it is necessary to cluster on n different quantities of
clusters, and then to find the point of inflection of the function of dispersion, which is
"elbow". This number will be the optimal number of clusters at the moment [
        <xref ref-type="bibr" rid="ref4 ref7">4, 7</xref>
        ].
      </p>
      <p>
        An alternative to the reproduction of the “elbow” dispersion function is the use of
gap statistics [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6-8</xref>
        ], which are generated on the basis of resampling and Monte Carlo
simulation procedures.
      </p>
      <p>Let E∗n{log(W∗k)} mean an estimate of the average dispersion W∗k obtained by
the bootstrap method when k clusters are generated by random sets of objects from an
input sample of size n. Then the statistic Gapn(k)=E{n{log(W∗k)}-log(Wk) determines
the deviation of the observed dispersion Wk from its expected value, provided that the
null hypothesis holds that the input data form only one cluster.</p>
      <p>To do this, we cluster the data into different number of clusters, find the dispersion
and find the point of inflection by the “elbow” method.</p>
      <p>def clasternumber(df1):
# find number of clusters
cluster_range = range(1, 10)
cluster_errors = []
for num_clusters in cluster_range:
clusters = KMeans(num_clusters)
clusters.fit(df1)
cluster_errors.append(clusters.inertia_)</p>
      <p>Find the inflection point of the function. It corresponds to the number of clusters
4. Therefore, the given dataset must be divided into 4 clusters for best results.</p>
      <p>
        Clustering - grouping objects based on the similarity of their properties, so that
each cluster consists of similar objects, and the objects of different clusters differ
significantly [12-18]. Clustering helps us understand natural grouping or data
structure. The purpose of clustering is to determine the internal grouping of multiple
unlabeled data segments. With the help of clustering we can solve the following problems
of data analysis: search and discovery of knowledge, grouping and recognition of
objects: search for representatives of homogeneous groups (reducing the
dimensionality of data), search for natural clusters and description of their unknown properties,
search for useful and appropriate grouping, search for unusual objects. (detection of
emissions) [
        <xref ref-type="bibr" rid="ref2">2, 19</xref>
        ].
      </p>
      <p>
        Unlike hierarchical methods, when clusters are built up step by step, split
clustering methods examine all segments at once. In doing so, they either attempt to identify
clusters by iteratively moving points between subsets, or to identify clusters as areas
that are densely filled with objects. The first kind of algorithms belong to partitioning
methods with movement (partitioning relocation clustering). They, in turn, are divided
into probabilistic, k-means and k-medoids methods, and concentrate on adjusting
points to corresponding clusters, with a tendency to construct spherical segments. The
second type separation algorithms belong to the group of density-based partitioning
methods. They try to identify dense cohesive data components that are flexible in
terms of their shape. This group is less sensitive to emissions and may find irregular
clusters [
        <xref ref-type="bibr" rid="ref10 ref6 ref7 ref8 ref9">6-10</xref>
        ].
      </p>
      <p>
        Separating moving clustering methods have several advantages: linear complexity
of the algorithm; relative scalability and simplicity; good fit to data with compact,
well-separated spherical clusters. The disadvantages include: a significant decrease in
performance on high-dimensional data; the need to indicate the number of clusters in
advance; sensitivity to initialization, noise and emissions; possibility to get to local
optima’s; inability to cope with clusters of different shapes and densities [
        <xref ref-type="bibr" rid="ref11">11, 14, 16</xref>
        ].
      </p>
      <p>
        To divide users into clusters, you need to use the k-means algorithm, and then
apply the resulting data to the Apriori algorithm to obtain user behavior patterns [
        <xref ref-type="bibr" rid="ref1">1, 15</xref>
        ].
      </p>
      <p>The principle of the k-means algorithm is to find the following cluster centers and
sets of elements of each cluster in the presence of some function F(°), which
expresses the quality of the current division of the set into k clusters, when the total quadratic
deviation of the cluster elements from the centers of these clusters is smallest:
(2)
k
V  </p>
      <p> ( x j   i ) 2
i1 x jSi</p>
      <p>Where k – cluster amount, Si – obtained clusters, i = 1, 2, … , k, µ - vectors` center
of mass, xi ϵ Si.</p>
      <p>{\displaystyle V=\sum _{i=1}^{k}\sum _{x_{j}\in S_{i}}(x_{j}-\mu
_{i})^{2}}</p>
      <p>At first step of the k-means algorithm, we select the cluster centers arbitrarily.
Then, for each element of the set, we calculate the distance from the centers and
attach each element to the cluster. For each of the obtained cluster, we calculate new
value of the center, trying to minimize the function F (°). After that the procedure of
redistributing the elements between the clusters is repeated.</p>
      <p>Algorithm of Clustering using the k-means scheme:
 select k information points as cluster centers until the process of changing
cluster centers is completed;
compose each information point with a cluster whose distance to the center is
minimal;
ensure that each cluster contains at least one point. To do this, each empty
cluster must be supplemented by an arbitrary point located "far" from the
center of the cluster;
replace the center of each cluster with the mean value of the cluster elements;




</p>
      <p>For clustering, we pass to the function numerical normalized data, which are the
coordinates of user vectors in multidimensional space based on their activity:</p>
      <p>The main advantages of the k-means method are its simplicity and speed of
execution. The k-means method is more convenient for clustering large numbers of
observations than the method of hierarchical cluster analysis (in which the dendrograms
become overloaded and lose clarity).</p>
      <p>One of the disadvantages of the simple method is the violation of the connectivity
of elements of one cluster, so different modifications of the method, as well as its
fuzzy k-means methods, are developed, in which, at the first stage of the algorithm,
the membership of one element of a set to several clusters is allowed (with varying
degrees of affiliation).</p>
      <p>Despite the obvious advantages of the method, it also has significant
disadvantages:
 The result of the classification strongly depends on the random starting
positions of the cluster centers
The algorithm is sensitive to emissions, which can distort the mean value
The researcher should determine the number of clusters in advance.
2. At the output, each element is assigned to a specific cluster:</p>
      <p>As a result, we obtained data that is divided into 4 clusters. The size of each
section corresponds to the number of values in the cluster. Next, we need to use this
information to find the rules of conduct. Therefore, after clustering, we should start
looking for associative rules.</p>
      <p>Affinity analysis is one of the common methods of Data Mining. The purpose of
this method is to investigate the relationship between events that occur together. Its
purpose is to identify associations between different events, that means to find rules
for quantifying the relationship between two or more events. These rules are called
association rules.</p>
      <p>The basic concepts in associative rule theory are subject set and transaction. A
subject set is some non-empty set of elements that transactions can include:</p>
      <p>I={i1,i2,…,ik,…,in},
where ik - elements included in the subject sets, k=1..n, n is the number of
elements of the set I.</p>
      <p>Transaction is a set that has some elements of set I that occur together. The
transaction also has a unique TID (Transaction ID).</p>
      <p>There is a certain set of transactions in the database:</p>
      <p>T={t1, t2,…, ti,…, tm},
where ti – some transaction, m – amount of transactions.</p>
      <p>Between transaction elements we can set some regularities in the form of
associative rules: X  ik ik  I, called a condition and Y  jk jk  I , called a
consequence. We should also say that the same set should not be included in the
antecedent and consequent at the same time: ik  jk .</p>
      <p>The associative rule describes the relationship between sets of subjects which
respond to consequence rule and are written down as X  Y . The sets X and Y must
not intersect: X∩Y=Ø. The main indicators of the importance of an associative rule
are support and confidence.</p>
      <p>We should distinguish between support for recruitment and support for associative
rule. In two cases, it is defined as the ratio of the number of transactions having the
specified amount of items to the total number of transactions. The only difference is
that the number of transactions that have the corresponding set is taken to calculate
support for the set, and the number of transactions that have both a condition and a
consequence at the same time to calculate the support of the associative rule.</p>
      <p>Lets assume that we have some set X and associative rule X → Y. Then support of
set X is:</p>
      <p>Supp( X ) </p>
      <p>X (t)</p>
      <p>T
,
Where X (t)  t  Т Х  t.
(3)
(4)
(5)
Support for the associative rule will be equal to:</p>
      <p>Supp( X  Y ) </p>
      <p>X (t)  Y (t)</p>
      <p>T
,
where X (t)  t  Т Х  t, Y (t)  t  Т Y  t.
(6)
(7)
(7)</p>
      <p>Because modern database sizes can reach large enough volumes (up to gigabytes
and terabytes), finding associstive rules requires efficient algorithms that are scalable
and allow you to find a solution to a given problem in an acceptable time.</p>
      <p>Apriori algorithm was designed for relational databases and allows you to
generate frequent datasets from transaction tables.</p>
      <p>Apriori algorithm uses iterative approach. At first step it finds single-element
frequent datasets denoted by the set L1. At the next step dataset L1 is used to find
frequent sets that has two elements, from which we form dataset L2 which is used to find
dataset L3 and so on. To increase the productivity of frequent datasets generation the
anti-monotony property is used. This is based on the following observation: if some
dataset is not frequent: supI  min sup , when if we add some specific object
i we will get new dataset which also would not be frequent:
supI  i  min
sup .</p>
      <p>Using the specified property, frequent k-element sets of Lk data can be obtained by
combining frequent (k-1) element sets. Moreover, in order for some k-element set Lk
to be included in frequent sets Lk, all of its (k -1) element subsets must also be
frequent. If at least one of them is not a frequent set, Lk must be excluded from the set of
frequent subject sets.</p>
      <p>This observation contributes to the creation of a plurality of candidates of
kelemental Ck, sets, which will be a subset of Lk. This subset is obtained by removing
from the Ck infrequent datasets, which is the result of checking the support values of
each of the candidates ck, (ck  Ck,). Based on the antimonotonic property, the set Ck
is generated in two steps. In the first step, the candidate is generated by joining the
members of the set of frequent sets Ck-1, where two members can be joined if they
have k-2 common elements, ie:</p>
      <p>Lk 1  Lk 1  A  B A, B  Lk 1 A  B  k  2</p>
      <p>The next step is to remove from the set of Сk members that include (k-1)-
elemental data sets that are not frequent.</p>
      <p>3</p>
    </sec>
    <sec id="sec-3">
      <title>Numerical experiments</title>
      <p>We used the standard mlxtend library to aggregate data by cluster, user login, and
time.
from mlxtend.frequent_patterns import apriori
rules=apriori(df, min_support=0.01, use_colnames=True)
4</p>
    </sec>
    <sec id="sec-4">
      <title>Analysis of the results</title>
      <p>Algorithm result is shown below:</p>
      <p>In Fig. 8 we can see the work of the associative rule algorithm, where the diameter
of a circle means the support (frequency) of a certain rule, and the arrows to and from
the circle, respectively, indicate the sequence of elements in the rule. As we can see,
the Driving-Shopping rule (with the largest and brightest circle) is most likely. It has
0.046 support and 1.818 color saturation. the least likely is the Amusements-Finance
rule with 0.008 support and a circle saturation of 0.369.</p>
      <p>In Fig. 9 is a diagram showing the 5 most popular places. They are in descending
order. Therefore, it can be assumed that, due to the use of meetcity and its geolocation
capabilities to predict and recommend meetings, Driving is considered the most
visited place.</p>
      <p>5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this article we explored apriori and k-means clustering algorithms to get user
behavior analysis template. In the process, we looked at the problem of finding
associative rules that were able to find and describe patterns in large datasets. We used
scalable Apriori algorithms to find the best rules. For analysis, we used the standard
mlxtend library to aggregate data by cluster, user login, and time.</p>
      <p>While working, we were faced with the problem of inaccuracy and inconsistency
of data with real conditions, and were forced to reduce the minimum support for
associative rules.
12. Ankerst, M., Ester, M., Kriegel, H.-P.: Towards an effective cooperation of the user and
the computer for classification. In: Proc. of the 6th ACM SIGKDD Intern. conf. on
knowledge discovery and data mining, pp. 179-188, Boston, Massachusetts, USA (2000)
13. Peuquet, D.J.: Representations of space and time, N. Y., Guilford Press (2002)
14. Guo, D., Peuquet, D.J., Gahegan, M.: ICEAGE: Interactive clustering and exploration of
large and high-dimensional geodata. Geoinfor-matica, vol. 3, N. 7, pp. 229-253 (2003)
15. Boyko, N.: A look trough methods of intellectual data analysis and their applying in
informational systems. In: Scientific and Technical Conference “Computer Sciences and
Information Technologies (CSIT), 2016 XIth International, pp. 183-185, IEEE (2016).
16. Shakhovska, N., Boyko, N., Zasoba, Y., Benova, E.: Big data processing technologies in
distributed information systems. Procedia Computer Science, 10th International
conference on emerging ubiquitous systems and pervasive networks (EUSPN-2019), 9th
International conference on current and future trends of information and communication
technologies in healthcare (ICTH-2019), Vol. 160, 2019, pp. 561–566, Lviv, Ukraine (2019)
17. Boyko, N., Shakhovska, Kh., Mochurad, L., Campos, J.: Information System of Catering
Selection by Using Clustering Analysis, Proceedings of the 1st International Workshop on
Digital Content &amp; Smart Multimedia (DCSMart 2019), рр. 94-106, Lviv, Ukraine (2019)
18. Kunanets, N., Vasiuta, O., Boikо, N.: Advanced Technologies of Big Data Research in
Distributed Information Systems, Proceedings of the 14th International conference
"Computer sciences and Information technologies" (CSIT 2019), pp. 71-76, Lviv, Ukraine,
(2019)
19. Boyko, N., Basystiuk, O.: Comparison Of Machine Learning Libraries Performance Used
For Machine Translation Based On Recurrent Neural Networks, 2018 IEEE Ukraine
Student, Young Professional and Women in Engineering Congress (UKRSYW), pp.78-82,
Kyiv, Ukraine (2018)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Estivill-Castro</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Amoeba: Hierarchical clustering based on spatial proximity using Delaunay diagram</article-title>
          .
          <source>In: 9th Intern. Symp. on spatial data handling</source>
          , pp.
          <fpage>26</fpage>
          -
          <lpage>41</lpage>
          , Beijing, China (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kang</surname>
          </string-name>
          , H.-Y.,
          <string-name>
            <surname>Lim</surname>
            ,
            <given-names>B.-J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li K</surname>
          </string-name>
          .
          <article-title>-</article-title>
          J.:
          <source>P2P Spatial query processing by Delaunay triangulation , Lecture notes in computer science</source>
          , vol.
          <volume>3428</volume>
          , pp.
          <fpage>136</fpage>
          -
          <lpage>150</lpage>
          , Springer, Heidelberg (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Boehm</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kailing</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kriegel</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kroeger</surname>
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Density connected clus-tering with local subspace preferences</article-title>
          .
          <source>In: Proc. of the 4th IEEE Intern. conf. on data mining</source>
          , pp.
          <fpage>27</fpage>
          -
          <lpage>34</lpage>
          , Los Alamitos: IEEE Computer Society (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Heterogeneous spatial data mining based on grid</article-title>
          ,
          <source>Lecture notes in computer science</source>
          , vol.
          <volume>4683</volume>
          , pp.
          <fpage>503</fpage>
          -
          <lpage>510</lpage>
          ,Springer/Heidelberg (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Harel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koren</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Clustering spatial data using random walks</article-title>
          .
          <source>In: Proc. of the 7th ACM SIGKDD Intern. conf. on knowledge discovery and data mining</source>
          , pp.
          <fpage>281</fpage>
          -
          <lpage>286</lpage>
          , San Francisco, California (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Turton</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Openshaw</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brunsdon</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          et al.:
          <article-title>Testing spacetime and more complex hyperspace geographical analysis tools</article-title>
          .
          <source>In: Innovations in GIS 7</source>
          , pp.
          <fpage>87</fpage>
          -
          <lpage>100</lpage>
          ,
          <string-name>
            <given-names>L.</given-names>
            : Taylor &amp;
            <surname>Francis</surname>
          </string-name>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Tung</surname>
            ,
            <given-names>A.K. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hou</surname>
            ,
            <given-names>J</given-names>
            ., Han, J
          </string-name>
          .:
          <article-title>Spatial clustering in the presence of obstacles</article-title>
          .
          <source>In: The 17th Intern. conf. on data engineering (ICDE'01)</source>
          , pp.
          <fpage>359</fpage>
          -
          <lpage>367</lpage>
          , Heidelberg (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gehrke</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gunopulos</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Automatic sub-space clustering of high dimensional data</article-title>
          .
          <source>In: Data mining knowledge discovery</source>
          , vol.
          <volume>11</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>5</fpage>
          -
          <lpage>33</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Guimei</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jinyan</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sim</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Limsoon</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Distance based subspace clustering with flexible dimension partitioning</article-title>
          .
          <source>In: Proc. of the IEEE 23rd Intern. conf. on digital object identifier</source>
          , vol.
          <volume>15</volume>
          , pp.
          <fpage>1250</fpage>
          -
          <lpage>1254</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Aggarwal</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Finding generalized projected clusters in high dimensional spaces</article-title>
          .
          <source>In: ACM SIGMOD Intern. conf. on management of data</source>
          , pp.
          <fpage>70</fpage>
          -
          <lpage>81</lpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Procopiuc</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>P.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murali</surname>
            ,
            <given-names>T.M. :</given-names>
          </string-name>
          <article-title>A Monte Carlo algorithm for fast projective clustering</article-title>
          .
          <source>In: ACM SIGMOD Intern. conf. on management of data</source>
          , pp.
          <fpage>418</fpage>
          -
          <lpage>427</lpage>
          , Madison, Wisconsin, USA (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>