<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Enabling big data exploration in dynamic contexts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ada Bagozi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Devis Bianchini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valeria De Antonellis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Marini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Ragazzi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Information Engineering University of Brescia Via Branze</institution>
          ,
          <addr-line>38 - 25123 Brescia</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>24</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>According to the Industry 4.0 vision, big data management is among the new challenges for the factory of the future. While many approaches have been developed to investigate data analysis, data visualisation, data collection and management, the impact of big data exploration is still under-estimated. In this paper, we propose an approach for big data exploration in a dynamic context of interconnected systems, such as the Industry 4.0 domain. The approach relies on three main pillars: (i) a multi-dimensional model, that is suited for supporting the iterative and multi-step exploration of big data; (ii) novel data summarisation techniques, based on clustering; (iii) a model of relevance, aimed to focus the attention on relevant data only. Keyword: big data exploration, data summarisation, data relevance, Industry 4.0, smart manufacturing</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Big data management challenges raised from the abundance of real time data
in Cyber-Physical Systems (CPS), enabled by the widespread di usion of IoT
technologies [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Data is emerging as a new industrial asset, to implement
advanced functions like state detection, health assessment, as well as
manufacturing servitization [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In this context, many approaches have addressed issues
related to data collection and storage, analysis and visualisation. Nevertheless,
big data exploration issues are still under-estimated. These issues are very
relevant with reference to the principle of \human-in-the-loop" for Industry 4.0
applications [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], where operators are in charge of taking decisions in unknown
situations, based on their long-term experience. Operators must be supported in
managing the high volume of collected data in order to identify relevant insights
on which these decisions will be based.
      </p>
      <p>In this paper, we discuss the ingredients to enable exploration of real time
data in a dynamic context of interconnected systems, where large amounts of
data must be incrementally collected, organized and analysed on-the- y. Such</p>
      <p>Spindle #1
Tool</p>
      <p>Working stations
(raw material is
positioned here)
Rotating
drum</p>
      <p>
        Engine that
moves
spindle on X,
Y andZ axes
ingredients have been already described in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], where additional details about
the proposed model and techniques, as well as an extended comparison with the
state of the art, can be found. The approach relies on three main pillars: (i)
a multi-dimensional model, to support the iterative and multi-step exploration
of big data; (ii) novel data summarisation techniques, based on clustering; (iii)
a model of relevance, aimed to focus the attention on relevant data only. The
novel contribution of our approach relies on the clustering-based data relevance
evaluation and its use in combination with the multi-dimensional model, in order
to foster big data exploration. In fact, multi-dimensional data modeling, where
information is organised according to dimensions, either at or hierarchically
organized, may ease data exploration [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Furthermore, data summarisation
techniques enable aggregated views over high volume of data. Finally, data relevance
evaluation helps operators to identify data of interest, also when exploration
requirements are not well speci ed or must be iteratively re ned based on collected
data.
      </p>
      <p>The paper is organised as follows: in Section 2 we will motivate the approach
with the help of a smart factory case study and we highlight di erences with
respect to related work; in Section 3 we will provide a general overview of the
approach; the architecture of a framework built on top of the proposed techniques
and models will be presented in Section 4 and preliminary experiments in the
case study will be described in Section 5; nally, Section 6 closes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>Motivations and challenges</title>
      <sec id="sec-2-1">
        <title>Smart factory case study</title>
        <p>Let's consider an Original Equipment Manufacturer (OEM) producing
multispindle machines. As shown in Figure 1, spindles work independently each other
on the raw material, that is positioned on a rotating drum.</p>
        <p>Each spindle is mounted on a unit moved by an electrical engine to perform X,
Y and Z movements. The spindle rotation is impressed by an electrical engine
and its rotation speed is controlled by the machine control. Spindles use di erent
tools (that are selected according to the instructions speci ed within the Part
Program) in order to complete di erent steps in the manufacturing cycle. For
each spindle, the velocity of the three axes (X, Y and Z), the electrical current
absorbed by each engine, the value of rpm for the spindle and the percentage of
power absorbed by the spindle engine (charge coe cient) are measured.
Hereafter, with term features we will refer to the measured quantities. The aim is to
monitor the axle hardening of each spindle and the tool wear. Axle hardening
is monitored by observing changes in the values of energy consumption (spindle
engine charge coe cient) for similar rpm. By detecting di erences in energy
consumption while using di erent tools, spindle hardening can be identi ed as the
possible anomaly during the manufacturing operations. If the increase in energy
consumption is related only to the usage of a speci c tool, this has been
recognised as a symptom of tool wear. We collected real data from three machines,
each one equipped with three spindles and di erent tools. On each spindle, we
monitored the eight features listed above.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Related work</title>
        <p>
          In the considered motivating scenario, experience of human operators still plays
a fundamental role. Data exploration must be performed in near real time and
during incremental data collection. In [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] Exploratory Computing is de ned as
a multi-step process, going beyond traditional exploratory data analysis and
Data Mining techniques. In our approach, we aim at performing a step
forward by describing a combined use of multi-dimensional model, data
summarisation and relevance evaluation techniques to foster exploration while dealing
with disruptive characteristics of big data (e.g., mainly volume and velocity).
The innovation we introduced, compared to On Line Analytical Processing [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]
and faceted search [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], relies on the combined use of data summarisation and
relevance evaluation techniques within a multi-dimensional model speci cally
designed for incremental big data organisation. In [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] OLAP-based exploration
for multi-dimensional data is discussed, but no data relevance techniques are
proposed. Similarly, authors in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] propose the application of query
approximation techniques for data that are incrementally collected, without providing a
way of guiding target users towards data of interest. In [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] an approach to issue
range queries over structured data is described. Semantic windows group data
according to speci c criteria (e.g., all data in a given time interval, all data in
the same geographical area), while sampling techniques are applied to support
users in query resolution. No incremental collection and organisation of data are
performed. The three pillars we will introduce in the following (clustering, data
relevance and multi-dimensional modeling) have been combined to foster
exploration of huge amounts of data incrementally collected from dynamic systems.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Approach overview</title>
      <p>
        Data summarisation. The characteristics of big data, namely, volume,
velocity and variety, pose non trivial issues for data collection and organisation. High
volume calls for techniques and tools to provide a compact view over the large
amount of collected data. Furthermore, we deal with data streams, in which data
is collected in a fast and incremental way. We address these issues by applying
incremental clustering techniques. Clusters o er a two-fold advantage: (a) they
give an overall view over a set of measures, using a reduced amount of
information; (b) they allow to depict the behaviour of the system better than single
records, that might be a ected by noise and false outliers, in order to observe
a given physical phenomenon. Clustering is performed in two steps: (i) in the
rst one, a variant of Clustream algorithm [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is applied, that incrementally
processes incoming data to obtain a set of syntheses, that provide a lossless
representation of measures; (ii) in the second step, syntheses are clustered, in order
to minimise the distance between their centroids within the same cluster and
to maximise the distance between centroids across di erent clusters. Clusters
give a balanced view of the observed physical phenomenon, grouping together
syntheses corresponding to the same working status. The two-steps clustering
algorithm enables an incremental procedure speci cally developed to face data
streams. Considering t as the time interval in which measures are grouped in
syntheses, that in turn are clustered, every t seconds a new clusters set SC is
generated, built on top of the previous iterations.
      </p>
      <p>
        Clustering-based data relevance evaluation. Relevance-based techniques
are used to focus exploration on relevant clusters only. In literature, data
relevance is de ned as the distance from an expected status. In our case, the
expected status corresponds to the normal working conditions of monitored
Cyber Physical Systems. The expected status can be tagged by human operators
while observing the monitored system. Let's denote with S^C = fC^1; C^2; : : : ; C^ng
the clusters set identi ed during normal working conditions, and with SC =
fC1; C2; : : : ; Cmg the current clusters set, where n and m do not necessarily
coincide. Relevant data are recognised when SC di ers from S^C. Therefore, the
proposed relevance techniques are based on clusters set distance between SC and
S^C, denoted with (SC; S^C), and enable to detect clusters movements, clusters
contraction/expansion, changes in the number of clusters. We refer to [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] for
details on (SC; S^C) computation.
      </p>
      <p>Multi-dimensional model for data exploration. The multi-dimensional
model is shown in Figure 2, using the hypercube representation, where axes
represent exploration dimensions and nodes represent clusters sets computed on
measures. Multi-dimensional modeling may help to organise clustered data
according to di erent perspectives, compliant with operators' requirements. Among
the dimensions, we always consider feature spaces, time and operational
parameters. With feature spaces we refer to monitored phenomena (e.g., axle hardening,
tool wear) observed by measuring a set of features. Time is of paramount
importance, since the data summarisation procedure and data relevance techniques are
strictly related to the time dimension. Operational parameters represent
domainspeci c settings or variables: for example, the working mode (G0, fast movement
clusters set of measures
asset hierarchy
multi-spindle machine
spindle1
spindle2</p>
      <p>spindle3
% absorbed
power
time
% absorbed
power
relevant data
(e.g., different
behaviour)
rpm
spindle2
(b) exploration across spindles
% absorbed
power
ti</p>
      <p>rpm
(a) exploration over time
(wco)rekxipnglormatoiodne over
% absorbed
power
ti+tj
relevant data
(e.g., anomalous
rpm behaviour)
rpm
spindle1
of the spindle to catch the tool, or G1, slow movement of the spindle during the
manufacturing), the tool used during manufacturing, the monitored system, the
manufactured product. Dimensions may present hierarchies: for example, tools
can be aggregated into tool types, monitored system follows the asset hierarchy
(enterprise, shop oor, machine, components).</p>
      <p>Exploration scenarios. Data relevance is used to identify dimensions over
which summarised data changed, thus denoting an unexpected behaviour of the
monitored system. Nevertheless, data relevance evaluation performed over all
available dimensions may bring to high computational requirements. Therefore,
exploration scenarios are proposed to constrain relevance evaluation over a subset
of available dimensions, to prevent useless comparisons. In Figure 2, some
examples of exploration scenarios are highlighted. Exploration for anomaly detection
(a) detects anomalies by observing if collected data of a speci c spindle changes
and gets closer/overtakes physical limits of breakage. This exploration scenario
constrains exploration over the time dimension and within the same monitored
system. Exploration for performance comparison (b) focuses comparison between
spindles working in the same conditions. Other exploration scenarios are de ned
to avoid erroneous comparisons. For example, let's consider the exploration over
working mode (c) as shown in Figure 2: a comparison of system behaviour across
G0 and G1 modes does not make sense. Modeling exploration scenarios, in
order to abstract from their implementation, will be further investigated in future
work.
4</p>
    </sec>
    <sec id="sec-4">
      <title>IDEAaS Architecture</title>
      <p>The approach described in this paper has been implemented within the IDEAaS
(Interactive Data Exploration As a Service) framework, whose architecture is
Sensor and</p>
      <p>Control Data
Data Acquisition Service</p>
      <p>CPS Notification Service Anomaly Detection Service</p>
      <p>Monitoring Service
Data Control</p>
      <p>Module</p>
      <p>IDEAaS Core Modules
Data Summarisation</p>
      <p>Module</p>
      <p>Data Relevance
Evaluation Module</p>
      <p>IDEAaS front-end Services</p>
      <p>IDEAaS Cloud-based Services
ModelMetadata</p>
      <p>Col ected Summarised</p>
      <p>Data Data
shown in Figure 3. In the IDEAaS modular architecture, modules are
distinguished in IDEAaS Core Modules (that include Data Control, Data
Summarisation and Data Relevance Evaluation) and IDEAaS front-end Services, exposed
both as web services and as standalone modules. Among these services, Data
Acquisition Service is in charge of storing data collected from CPS. As shown in
Figure 3, data coming from the physical system, properly collected through
sensors and IoT technologies, is sent to the Data Acquisition Service to be processed.
This service operates in order to minimise time spent for data acquisition.
Specifically, data collected is rst saved as JSON documents in a MongoDB NoSQL
database (Collected Data), minimising any other operation, namely data control
and cleaning and data clustering that are performed in parallel. Dimensions of
the multi-dimensional model are stored as metadata into the Model-MetaData
relational database. On collected data, summarisation and relevance evaluation
techniques are applied. Resulting clusters sets are stored as JSON documents
within Summarised Data database (also using MongoDB technology). The CPS
Noti cation Service is in charge of notifying to target operators data of interest
according to the relevance evaluation techniques. Data relevance can be used to
serve di erent purposes, ranging from anomaly detection to performance
monitoring of CPS, for which speci c services will be designed in the future. On top
of CPS Noti cation Service, a Data Exploration GUI will be developed as well,
to allow operators to interact with the system and explore collected data. The
IDEAaS architecture is implemented in Java, on top of a Glass sh Server Open
Source Edition 4.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Preliminary experiments</title>
      <p>We tested the e ciency and e ectiveness of our approach in providing
summarised data for exploration purposes, given the acquisition rate of records in</p>
      <p>Spindle 1 Spindle 2 Spindle 3
dercprmeas coinnescnureemragpsyteion dercprmeas coinnescnureemragpsyteion dercprmeas coinnescnureemragpsyteion
/ / / / / /
/ / / / 15% 16%
20% 20% / / 29% 29%
30% 30% / / / /
40% 40% 20% 20% / /
/ / 30% 30% / /
/ / 20% 20% / /
(a)</p>
      <p>Cluster distances </p>
      <p>Time response clustering 
400000 
the considered case study. We collected 140 millions of records from the three
machines. All records present a timestamp and have been collected every 200ms
(5 records per second). We run experiments on an Intel Core i7-6700HQ, CPU
2.60 GHz, 4 cores, 8 logical cores, RAM 16GB. Collected records of measures
have been saved within MongoDB as JSON documents grouped into collections.
Documents present a simple structure, with at most one level of depth, and
collections have been organised considering the time as main dimension, in
order to speed up both data storage and data extraction for clustering, that is
applied on slots of records on a time interval t. This enabled to storage all
140 millions of records in 1 hour and 14 minutes, with an acquisition rate of
31,531 records per second. We tested clustering and hypercube generation on
real data considering average values on 2 and 3 features (Figure 4a). The worse
response time corresponds to the case where we performed clustering and
relevance evaluation when no previous syntheses have been generated. Also in that
case, IDEAaS framework was able to process 15,600 records in 11.5 seconds
(processing rate of 1,356 records per second). Through the tasks of syntheses
generation and clustering, the processed set of records is reduced to 7,2% on
average. As shown in Figure 4b, we observed a variation in distance between
clusters sets at the cost of decreasing the processing time to 255 records per
second, that is still acceptable. To test e ectiveness of the IDEAaS framework in
detecting working anomalies, we introduced unexpected working states to
simulate spindle hardering by increasing energy consumption and decreasing rpm on
a subset of collected data, as shown in Figure 5a. Experimental results depicted
in Figure 5b show the clustering distance between clusters sets calculated at the
time t and those calculated in case of normal working conditions. The gure
shows how the techniques of the IDEAaS framework allow to timely identify the
unexpected situations induced in the system under observation.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Concluding remarks</title>
      <p>In this paper, we discussed the ingredients to enable exploration of real time
data incrementally collected, organized and analysed on-the- y. Our approach
combines: (i) a multi-dimensional model, that is suited for supporting the
iterative and multi-step nature of data exploration; (ii) e cient data summarisation
techniques, based on clustering; (iii) a model of relevance, to focus the attention
on relevant data only. Future e orts will be devoted to the parallelisation of
clustering and relevance evaluation by relying on data organisation in the
multidimensional model, to the development of data visualisation techniques and to
the de nition of di erent data exploration scenarios.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>L.</given-names>
            <surname>Monostori</surname>
          </string-name>
          ,
          <article-title>Cyber-physical production systems: Roots, expectations</article-title>
          and R&amp;D challenges, in
          <source>: Proc. of the 47th Conf. on Manufacturing Systems</source>
          ,
          <year>2014</year>
          , pp.
          <volume>9</volume>
          {
          <fpage>13</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bagheri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kao</surname>
          </string-name>
          ,
          <article-title>A Cyber-Physical Systems architecture for Industry 4.0-based manufacturing systems</article-title>
          ,
          <source>Manufacturing Letters</source>
          <volume>3</volume>
          (
          <year>2015</year>
          )
          <volume>18</volume>
          {
          <fpage>23</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>F.</given-names>
            <surname>Longo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Nicoletti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Padovano</surname>
          </string-name>
          ,
          <article-title>Smart operators in industry 4.0: A humancentered approach to enhance operators capabilities and competencies within the new start factory context</article-title>
          ,
          <source>Computers &amp; Industrial Engineering</source>
          <volume>113</volume>
          (
          <year>2017</year>
          )
          <volume>144</volume>
          {
          <fpage>159</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>N.</given-names>
            <surname>Kamat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Jayachandran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Tunga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nandi</surname>
          </string-name>
          ,
          <article-title>Distributed and Interactive Cube Exploration</article-title>
          ,
          <source>in: Proc. of 30th Int. Conf. on Data Engineering (ICDE</source>
          <year>2014</year>
          ),
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>M.</given-names>
            <surname>Buoncristiano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Mecca</surname>
          </string-name>
          , E. Quintarelli, Roveri,
          <string-name>
            <given-names>D.</given-names>
            <surname>Santoro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tanca</surname>
          </string-name>
          ,
          <article-title>Database Challenges for Exploratory Computing</article-title>
          ,
          <source>SIGMOD Record 44 (2)</source>
          (
          <year>2015</year>
          )
          <volume>17</volume>
          {
          <fpage>22</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>M.</given-names>
            <surname>Golfarelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          , Data Warehouse Design:
          <article-title>Modern Principles and Methodologies</article-title>
          ,
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>D.</given-names>
            <surname>Tunkelang</surname>
          </string-name>
          ,
          <source>Faceted Search (Synthesis Lectures on Information Concepts</source>
          ,
          <source>Retrieval and Services)</source>
          , Morgan and Claypool Publishers,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>A.</given-names>
            <surname>Wasay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Athanassoulis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Idreos</surname>
          </string-name>
          , Queriosity:
          <source>Automated Data Exploration, in: Proc. of the IEEE International Congress on Big Data</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>A.</given-names>
            <surname>Kalinin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Cetintemel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zdonik</surname>
          </string-name>
          ,
          <article-title>Interactive data exploration using semantic windows</article-title>
          ,
          <source>in: Proc. of the ACM SIGMOD</source>
          <year>2014</year>
          , pp.
          <volume>505</volume>
          {
          <fpage>516</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>A.</given-names>
            <surname>Bagozi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bianchini</surname>
          </string-name>
          , V. De Antonellis,
          <string-name>
            <given-names>A.</given-names>
            <surname>Marini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ragazzi</surname>
          </string-name>
          ,
          <article-title>Summarisation and Relevance Evaluation Techniques for Big Data Exploration: the Smart Factory case study</article-title>
          ,
          <source>in: Proc. of 29th Int. Conference on Advanced Information Systems Engineering (CAISE</source>
          <year>2017</year>
          ),
          <year>2017</year>
          , pp.
          <volume>264</volume>
          {
          <fpage>279</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>C. Aggarwal</surname>
            ,
            <given-names>J</given-names>
            . Han, J
          </string-name>
          .
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>A framework for clustering evolving data streams</article-title>
          ,
          <source>in: Proc. of 29th Int. Conf. on Very Large Data Bases</source>
          ,
          <year>2003</year>
          , pp.
          <volume>81</volume>
          {
          <fpage>92</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>