<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Characterizing Problem Gamblers in New Zealand: A Novel Expression of Process Cubes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Suriadi Suriadi</string-name>
          <email>s.suriadi@massey.ac.nz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Teo Susnjak</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Agate M. Ponder-Sutton</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul A. Watters</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoph Schumacher</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Massey University</institution>
          ,
          <addr-line>Albany, Auckland 0632</addr-line>
          ,
          <country country="NZ">New Zealand</country>
        </aff>
      </contrib-group>
      <fpage>185</fpage>
      <lpage>192</lpage>
      <abstract>
        <p>This paper reports on the challenges and lessons learned from our case study which uniquely integrates a mixture of process mining, data mining, and con rmatory statistical techniques to explore and characterize the variations in gambling behaviours exhibited by gamblers in New Zealand. We demonstrate how we weaved techniques from these three disciplines to understand the variety of behaviours exhibited by gamblers, and to provide assurances of the correctness of our results. This case study also demonstrates how such a combination of techniques provides a rich set of tools to undertake an exploratory data analysis project that is guided by the process cube concept.</p>
      </abstract>
      <kwd-group>
        <kwd>data mining</kwd>
        <kwd>process mining</kwd>
        <kwd>statistics</kwd>
        <kwd>problem gamblers</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>This paper reports on the challenges and lessons learned from our case study
where a mixture of process mining, data mining, and con rmatory statistical
techniques are applied to analyse a data set containing all bets recorded by a
gambling service provider in New Zealand (NZ). This study attempts to identify
and characterize various groups of gamblers (with a focus on problem gamblers)
directly from the data. The nature of this case study is exploratory: we do not
`label' our data with various classes of gamblers; rather, we attempt to learn
how many groups of gamblers can be discerned from the data.</p>
      <p>
        The exploratory nature of this study calls for the use of unsupervised learning
techniques, such as k-means clustering [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] as the starting point of analysis.
However, clustering analysis alone is not su cient as it does not o er any insights into
the behaviour of gamblers seen in each cluster. Our case study demonstrates how
both process mining analysis, with its ability to extract detailed insights from
ne-grained and chronologically-arranged data [
        <xref ref-type="bibr" rid="ref11 ref8">8, 11</xref>
        ], and con rmatory
statistics, can be strategically weaved together with clustering analysis to not only
understand the variety of behaviours exhibited by gamblers, but also to evaluate
our results. Furthermore, we highlight how the combination of these techniques
can be used as an expression of the process cube concept [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>The main contribution of this paper is on the reporting of challenges and
lessons learned in the application of these three classes of analysis techniques,
Copyright c by the paper's authors. Copying permitted only for private and academic
purposes.
highlighting practical challenges that arise in analysing a relatively large size of
data with limited computational resources. Section 2 summarizes the approach
taken in this case study. Section 3 details the analyses performed at each stage
of the case study, along with the challenges and lessons learned. A discussion
about the related work is provided in Section 5, followed by the conclusion.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Approach</title>
      <p>
        While we applied a range of techniques from multiple disciplines, our case study
employed the PM2 [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] approach because: (1) the starting point of our analysis
was an event log (the log used in our study consisted of betting events per
account holder), of which process mining is designed for, and (2) the PM2
methodology is rather detailed in its guidance on each stage, and exible enough to allow
the inclusion of other types of classical data mining techniques (e.g. clustering).
      </p>
      <p>
        We also applied the concept of process cubes - a concept born from the
Online Analytical Processing (OLAP) domain. Using this approach, event logs
are organised into various cells based on multiple dimensions [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and these cells
can then be merged or further split using typical OLAP operations, such as slice,
dice, roll-up, and drill-down. However, as suggested by van der Aalst [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], the
nature of event logs are such that special considerations are needed to perform
those operations. This paper shows some of the techniques that can be applied
to achieve those operations.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Case Study</title>
      <p>This case study focuses on the data processing, data mining, clustering and
analysis, and evaluation stages of the PM2 methodology. The questions that the
stakeholder in our case study (an economist) wanted to address were: (RQ1)
how many groups of gamblers can be discerned from the data set, and (RQ2)
which users were problem gamblers from the data?</p>
      <p>
        The data set was extracted from a NZ gambling provider. It contained
information about all the bets placed in New Zealand from August 2006 to May
2014. The data set extracted was in event log format where each line of data
represents a betting event. Given the size of the raw data (80 GB in compressed
text format) and the limited computational resources at our disposal (three
816GB/i5 Core workstations and one 32GB/i5 Core virtual machine), we focused
our analysis on gambling data from a 9-month window (Aug 2013 to May 2014).
Data Processing. The dataset consisted of 11,311,892 betting events executed
by 91,405 account-holders. The data was pre-processed because (1) our
experiences showed that existing implementations of many process mining techniques
do not scale well; many process mining case studies used event logs that are
signi cantly smaller (between a few thousand events to over 1 million events,
e.g. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]); (2) the raw data is not in an event log format consumable by process
mining analyses.
      </p>
      <p>
        Applying the concept of process cube [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] calls for slicing and dicing event
logs from multiple dimensions. As our case study is exploratory (without any
clear ltering dimensions), we applied an unsupervised k-means clustering [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
technique to nd the `natural separation' of gamblers in the data set. Three
problem gambling features from the data were used as clustering criteria: bet
frequency, the ratio (in dollar) of winnnings over amount lost (win/loss dollar
ratio, and ratio of the number (count) of winning over lost bets (win/loss
ratio). Next, an event log was derived for each cluster using a gambler's account
identi er as the case ID in the event log.For the activity eld, the data was
preprocessed to approximate a gambler's clawback behaviour: an activity name
was created as the concatenation of (1) the outcome of the previous bet (win
or loss) and (2) the amount of money placed in the next immediate bet as a
proportion of the amount of money placed in the previous bet: less than or equal
to the previous bet amount (&lt;= 1), up to double the previous amount (1to2), or
more than double (&gt;2). For example, loss &lt;= 1 would mean that a user having
lost the previous bet, had subsequently placed another bet where the amount
bet was less than or equal to the previous bet. Through data pre-processing, we
obtained 7 manageable-size event logs for each cluster for deeper analysis. Each
cluster is colour-coded for referencing purposes.
      </p>
      <p>
        Challenges. The large dataset was problematic given the lack of scalability of
the software tools used. Attempts to generate an XES le (i.e. the log format
expected by most process mining tools) presented challenges: we were limited
by the number of events that could be imported in Disco (www.fluxicon.com)
tool; we could have used the ProM Tool (www.prcessmining.org) to convert
the original CSV-formatted log into XES; however, the generated XES le would
have been substantially larger than the original CSV data, thus would not have
scaled well.Trace clustering algorithms, e.g. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], could not have been applied
as they would have required the entire event log to be analysed (not feasible
with our resources). We addressed this by using k-means clustering [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] based on
aggregated features at case-level granularity. This required deciding the optimal
number of clusters (k) to be used. We combined the advice from the stakeholder
with the Within Sum of Squares (WSS) analysis which showed that the
cohesiveness of the clusters converged optimally in the range of k=7 and k=8. We
therefore decided to use 7 clusters.
      </p>
      <p>
        Mining and Analysis. Our analysis focused on extracting key markers for
problem gamblers across clusters using process mining techniques. Clawback
behaviour is a key problem gambling psychological feature. We studied the
clawback behaviours by generating process models for each cluster using the Disco
tool. Then, through visual comparisons (similar to Partington et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and
Suriadi et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]), we extracted 9 distinct ow patterns in each cluster.
      </p>
      <p>To assert statistical signi cance w.r.t the di erences in the distribution (in
terms of frequency) of these 9 patterns across all clusters, Kruskal-Wallis tests
in conjunction with Dunn's tests were performed. These tests showed that the
distribution of these patterns across the seven clusters were indeed signi cantly
di erent at the p-value of 0.05. In addition, from box-and-whisker plots, we also
observed that the distribution of these 9 patterns were quite distinct in the blue
and black clusters (see Figure 2, left diagram, for example).</p>
      <p>Pattern 4</p>
      <p>Pattern 3</p>
      <p>Pattern 5</p>
      <p>
        The time elapsed between bets (bet interval) was another marker of
problem gamblers. The event interval analysis plug-in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] was used with the ProM
Tool to extract the bet interval for each bet placed by every gambler in each
cluster. Additionally, the distribution of the median bet interval per gambler
was extracted. Next, Kruskal-Wallis and Dunn's tests were applied to verify
if the distributions of the bet intervals across the 7 clusters were di erent at
p-value 0.05. The box-and-whisker graph (Fig. 2 - right) shows that the blue
and black clusters had signi cantly lower bet interval values compared to other
clusters. The exceptions being cyan and brown clusters, though they were
discounted since their actual bet numbers were very low (e.g. mostly only bet once).
We merged (`roll-up') the blue and black clusters as this was supported by the
pair-wise Dunn's tests (which resulted in a p-value higher than 0.05, indicating
insigni cant di erences).
      </p>
      <p>These analyses allowed us to group and rank the severity of gambling patterns
across the 7 clusters, resulting in 4 clusters that were statistically di erent (thus
addressing RQ1). Our ranking exercise suggested that the blue and black clusters
were more likely to contain problem gamblers, thus warranting further
`drillingdown' analysis into these two clusters. We performed a second-level clustering
of these two clusters into another 7 clusters. For these second-level clusters,
our focus was to establish statistical di erences in terms of gambling clawback
behaviour. We used the frequency distribution of those activities signifying
lossrelated clawback behaviours: loss &lt;=1, loss 1to2, and loss &gt;=2.</p>
      <p>
        Similar to our earlier approach, we used the Kruskal-Wallis and Dunn's tests
results on our second-level clustering to provide a more re ned severity ranking
of problem gamblers within the second-level clusters. Through this analysis,
we addressed RQ2: we narrowed down the most likely problem gamblers to two
clusters (from the second-level clusters) with a combined population size of 6,389.
Challenges Existing comparative analysis of processes [
        <xref ref-type="bibr" rid="ref12 ref8 ref9">8, 9, 12</xref>
        ] seemed to go
no further than visual analysis of process models, or through some forms of
multi-perspective visualisation techniques (e.g. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]). However, such approaches
lack theoretical vigour. As explained above, we addressed this issue by applying
con rmatory statistical testing.Practically, we faced computational limitations
during use of the event interval analysis [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and BPMN Miner [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] plug-ins.
We handled this by further dividing the event log of the largest cluster into
three sub-logs. However, division of large XML-formatted event log was also a
challenge as it contained more than 105 million XML elements/lines. We used
basic, powerful, Linux-based text processing utility, such as less, and sed to
cope with large XML les.
      </p>
      <p>Evaluation. This case study relied heavily on clustering to identify various
classes of gamblers.The main challenge therefore was on how to be sure that
the clusters we obtained were indeed appropriate and meaningful. To do so, we
had to assert that (1) the population within each cluster shared some unique
characteristics and/or behaviours, and (2) the existence of those clusters was
within reasonable expectation of the stakeholders.</p>
      <p>For the former, we applied classi cation analysis on the combined rst- and
second-level clusters (12 clusters in total: 5 from the initial clusters and 7 clusters
from the second-level clusters). We used features that had not previously been
used to generate the k-means clustering, including the median bet interval values,
the frequency of the 9 gambling patterns, the frequency of clawback behaviour
activities, the total sum of money won, and the number of times a gambler won
a bet. These features when used with the random forest algorithm (with 300
trees and 5-fold cross validation) generated a classi cation accuracy of 84.9%
with a standard deviation of =/- 0.8%. These classi cation results, having used
independent features from the clustering process, provided support that the
clusters that were produced were relevant and meaningful.</p>
      <p>The stakeholder also found the results to be reasonable. Some insights will
require further study, e.g. the dominance of Pattern 3 and Pattern 4 (which
suggested that gamblers had bet equal or less than their previous amount after
they had won or lost the previous bet) which contradicted currently-understood
risk-taking behaviour of problem gamblers where the expectation was for them
to increase the amount of money bet following a loss. While further analysis
may be required, there was an explanation for this: the dominance of Pattern
3 and Pattern 4 was evident when we took the all 12 clusters into account
(which could mean that such behaviour is discriminatory for deciding if a gambler
is a problem gambler or not). In the second level clustering, however, it was
the frequency of high-risk clawback behaviour, i.e. the loss 1to2 activity that
is statistically di erent across all second-level clusters (thus in line with the
expected behaviour).
4</p>
    </sec>
    <sec id="sec-4">
      <title>Lessons Learned</title>
      <p>During data processing stage, we found that it is e ective to combine
unsupervised k-means clustering and aggregated case-level attributes to slice a large data
set was highlighted in this analysis. Doing so, we `de ated' the data size from
over 11 million events to just over 94 thousand lines of data (each line represents
one gambler) that was further separated through with k-means clustering.</p>
      <p>To estimate an optimal number of clusters that is conservative enough to
reduce the error probability of not nding enough clusters and ne-grained
enough to account for the complexity of the data, our case study shows that
it can be achieved by doubling the number of initially-suspected clusters and
cross-referencing it with WSS analysis.</p>
      <p>We found the non-parametric assumptions of Kruskal-Wallis tests, paired
with Dunn's tests, made them exible and useful as, based on experience, data
used for process mining analyses is rarely normally distributed in practice.</p>
      <p>
        Our case study also shows that using hierarchical clustering in combination
with con rmatory statistics is a suitable approach to apply the process cube
concept [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Hierarchical clustering based on aggregated case-level attributes
allowed us to `slice and dice' event logs in an unsupervised manner, while
conrmatory statistics informed us as to which groups of processes that should be
`rolled-up' (that is, those two groups whose p-value from the Dunn test indicates
statistical similarities) and which to `drill down' into. This is an alternative
approach to the currently-prevailing approach to slicing processes that is based on
control- ow perspective or simple attribute-based ltering [
        <xref ref-type="bibr" rid="ref3 ref8">3, 8</xref>
        ].
      </p>
      <p>
        Finally, our case study demonstrates how one could apply classi cation
analysis using features that were not used as input to the k-means clustering to
provide some meanings to the clusters that we have managed to extract from
the data. Recall that a detractor of k-means clustering is that it will produce
as many clusters as parametrised. This is an important insight as previous work
which used k-means clustering in a process mining case study [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] also applied
classi cation analysis to provide meaning to each cluster. However, it was done
using the same features as those used for clustering, which resulted in a very
high accuracy rate but adds nothing more to our understanding of each cluster.
      </p>
      <p>
        From a practical perspective, the Inductive Miner [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] implementationed in
the ProM Tool is scalable to even the largest event log in our clusters. However,
the process model was less so: it produced models that exhibited ' ower model'
characteristics. By far, the Disco tool was still the most scalable tool in terms
of extracting process models even on an event log of over 4GB in size.
      </p>
      <p>We question the need for XML-formatted data in process mining. Raw data
often come in a table-like format with elds whose meaning can be
understood through discussions with stakeholders, thus lessening the need for the
self-de ning property of XML. XML les tend to be large and their processing
is resource intensive. Our original CSV-formatted data (&lt; 1.5 GB) was `blown
up' to over 5.8 GB after converting it to the XES format.</p>
    </sec>
    <sec id="sec-5">
      <title>5 Related Work</title>
      <p>
        Early process mining case studies [
        <xref ref-type="bibr" rid="ref10 ref11 ref15">10,11,15</xref>
        ] focused mainly on the application of
standard process mining techniques, such as process discovery and performance
analysis. Later process mining case studies [6{8] applied a combination of process
and data mining techniques. Our case study ts closer to the latter: we applied
a balanced amalgamation of process and data mining techniques. However, we
ventured further by also using a multi faceted approach involving process mining
and well-established con rmatory statistics.
      </p>
      <p>
        This case study con rms some of the observations and experiences reported
in other process mining studies. For example, the use of clustering techniques
to split original event logs into smaller pieces was applied in our case study.
However, instead of using trace/sequence clustering as in [
        <xref ref-type="bibr" rid="ref10 ref2">2,10</xref>
        ], we used k-means
clustering [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] based on aggregated case-level features.
      </p>
      <p>
        While an earlier case study [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] found that it was already feasible to conduct
process mining analysis using the ProM Tool, we found that the quality and the
robustness of the plug-ins in the ProM Tool to be highly variable with some
being robust and scalable enough to handle large data sets (e.g. the Inductive
Miner plug-in), while others tended to perform rather poorly (e.g. [
        <xref ref-type="bibr" rid="ref1 ref13">1, 13</xref>
        ]).
      </p>
      <p>
        While our approach to process comparison was similar to other case studies,
e.g. [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ], ours went further by studying the distribution of observed di
erences, and asserted their statistical signi cance using well-established con
rmatory statistics. Most importantly, this paper reported new lessons learned that
may be novel and helpful to other process mining practitioners.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>We demonstrated how to apply a diverse and complementary set of techniques
from the domains of process mining, data mining, and con rmatory statistics,
as a unique expression of the process cube concept, to characterize and assert
di erences in gambling behaviours exhibited by more than 94 thousand
gamblers in New Zealand. Most importantly, this case study reported a number of
challenges and lessons learned that are novel and would likely be bene cial to
other process mining practitioners.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>R.</given-names>
            <surname>Conforti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Garca-Bauelos</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. La</given-names>
            <surname>Rosa</surname>
          </string-name>
          .
          <article-title>Beyond tasks and gateways: Discovering bpmn models with subprocesses, boundary events and activity markers</article-title>
          .
          <source>In BPM</source>
          , volume
          <volume>8659</volume>
          <source>of LNCS</source>
          , pages
          <volume>101</volume>
          {
          <fpage>117</fpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>J. De Weerdt</surname>
            , S vanden Brouckev,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Vanthienen</surname>
            , and
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Baesens</surname>
          </string-name>
          .
          <article-title>Leveraging process discovery with trace clustering and text mining for intelligent analysis of incident management processes</article-title>
          .
          <source>In IEEE CEC</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Ekanayake</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Garcia-Banuelos</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. La</given-names>
            <surname>Rosa</surname>
          </string-name>
          .
          <article-title>Slice, mine and dice: Complexity-aware automated discovery of business process models</article-title>
          .
          <source>In BPM</source>
          , volume
          <volume>8094</volume>
          <source>of LNCS</source>
          , pages
          <volume>49</volume>
          {
          <fpage>64</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Hartigan</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Wong</surname>
          </string-name>
          .
          <article-title>Algorithm as 136: A k-means clustering algorithm</article-title>
          .
          <source>Journal of the Royal Statistical Society</source>
          . Series C (Applied Statistics),
          <volume>28</volume>
          (
          <issue>1</issue>
          ),
          <year>1979</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>S.J.J.</given-names>
            <surname>Leemans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fahland</surname>
          </string-name>
          , and
          <string-name>
            <surname>W.M.P. van der Aalst.</surname>
          </string-name>
          <article-title>Discovering blockstructured process models from incomplete event logs</article-title>
          .
          <source>In Petri Nets</source>
          , volume
          <volume>8489</volume>
          <source>of LNCS</source>
          , pages
          <volume>91</volume>
          {
          <fpage>110</fpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>J.</given-names>
            <surname>Nakatumba</surname>
          </string-name>
          .
          <article-title>Resource-aware business process management: analysis and support</article-title>
          .
          <source>PhD thesis</source>
          , Eindhoven University of Technology,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>H</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Dumas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M La</given-names>
            <surname>Rosa</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <article-title>M Maggi,</article-title>
          and
          <string-name>
            <given-names>S</given-names>
            <surname>Suriadi</surname>
          </string-name>
          .
          <article-title>Mining business process deviance: A quest for accuracy</article-title>
          .
          <source>In OTM</source>
          <year>2014</year>
          , volume
          <volume>8841</volume>
          <source>of LNCS</source>
          , pages
          <volume>436</volume>
          {
          <fpage>445</fpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>A.</given-names>
            <surname>Partington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.T.</given-names>
            <surname>Wynn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suriadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ouyang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Karnon</surname>
          </string-name>
          .
          <article-title>Process mining for clinical processes: a comparative analysis of four australian hospitals</article-title>
          .
          <source>ACM Trans. on Management Information Systems</source>
          ,
          <volume>5</volume>
          (
          <issue>4</issue>
          ),
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>A.</given-names>
            <surname>Pini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brown</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.T.</given-names>
            <surname>Wynn</surname>
          </string-name>
          .
          <article-title>Process visualization techniques for multiperspective process comparisons</article-title>
          .
          <source>In AP-BPM</source>
          , volume
          <volume>219</volume>
          <source>of LNBIP</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>A.</given-names>
            <surname>Rebuge</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.R.</given-names>
            <surname>Ferreira</surname>
          </string-name>
          .
          <article-title>Business process analysis in healthcare environments: A methodology based on process mining</article-title>
          .
          <source>Inf. Syst.</source>
          ,
          <volume>37</volume>
          (
          <issue>2</issue>
          ):
          <volume>99</volume>
          {
          <fpage>116</fpage>
          ,
          <string-name>
            <surname>April</surname>
          </string-name>
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>A.</given-names>
            <surname>Rozinat</surname>
          </string-name>
          , I.S.M de Jong, C.W. Gunther, and
          <string-name>
            <surname>W.M.P. van der Aalst.</surname>
          </string-name>
          <article-title>Process mining applied to the test process of wafer scanners in ASML</article-title>
          .
          <source>IEEE Trans. on System., Man, and Cybernetics</source>
          , Part C,
          <volume>39</volume>
          (
          <issue>4</issue>
          ):
          <volume>474</volume>
          {
          <fpage>479</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>S</given-names>
            <surname>Suriadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Mans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Wynn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A</given-names>
            <surname>Partington</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J</given-names>
            <surname>Karnon</surname>
          </string-name>
          .
          <article-title>Measuring patient ow variations: A cross-organisational process mining approach</article-title>
          .
          <source>In APBPM</source>
          , volume
          <volume>181</volume>
          <source>of LNBIP</source>
          , pages
          <volume>43</volume>
          {
          <fpage>58</fpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>S.</given-names>
            <surname>Suriadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ouyang</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.M.P. van der Aalst</surname>
          </string-name>
          , and
          <string-name>
            <surname>A.H.M. ter Hofstede</surname>
          </string-name>
          .
          <article-title>Event interval analysis: Why do processes take time? Decision Support Systems</article-title>
          ,
          <volume>79</volume>
          :
          <fpage>77</fpage>
          {
          <fpage>98</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>W.M.P. van der Aalst</surname>
          </string-name>
          .
          <article-title>Process cubes: Slicing, dicing, rolling up and drilling down event data for process mining</article-title>
          .
          <source>In Asia Paci c Conference on Business Process Management</source>
          , volume
          <volume>159</volume>
          <source>of LNBIP</source>
          , pages
          <volume>1</volume>
          {
          <fpage>22</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>W.M.P. van der Aalst</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          <string-name>
            <surname>Reijers</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.J. M. M. Weijters</surname>
            ,
            <given-names>B. F. van Dongen</given-names>
          </string-name>
          ,
          <string-name>
            <surname>A.K.A. Medeiros</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Song</surname>
            , and
            <given-names>H.M.W.</given-names>
          </string-name>
          <string-name>
            <surname>Verbeek</surname>
          </string-name>
          .
          <article-title>Business process mining: An industrial application</article-title>
          .
          <source>Information Systems</source>
          ,
          <volume>32</volume>
          :
          <fpage>713</fpage>
          {
          <fpage>732</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>M. L. van Eck</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>S.J.J.</given-names>
          </string-name>
          <string-name>
            <surname>Leemans</surname>
            , and
            <given-names>W.M.P. van der Aalst.</given-names>
          </string-name>
          <article-title>PM2: A process mining project methodologya</article-title>
          .
          <source>In CAiSE</source>
          , volume
          <volume>9097</volume>
          <source>of LNCS</source>
          , pages
          <volume>297</volume>
          {
          <fpage>313</fpage>
          . Springer International Publishing,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>J. De Weerdt</surname>
            , S. vanden Broucke,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Vanthienen</surname>
            , and
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Baesens</surname>
          </string-name>
          .
          <article-title>Active trace clustering for improved process discovery</article-title>
          .
          <source>Knowledge and Data Engineering</source>
          ,
          <volume>25</volume>
          (
          <issue>12</issue>
          ),
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>