<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Focusing Business Process Lead Time Improvements Using In uence Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Teemu Lehto</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Markku Hinkka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jaakko Hollmen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aalto University, School of Science, Department of Computer Science</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>QPR Software Plc</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <fpage>54</fpage>
      <lpage>67</lpage>
      <abstract>
        <p>Shortening lead times in a business process is important for meetings service level agreements, decreasing inventories and working capital, keeping customers satis ed and in short: staying in business. Process mining methods make it possible to generate a large amount of transaction event data and case attributes that are useful for analysing lead times. However, nding root causes for long lead times is not so straightforward with current process mining methods. In this paper we extend our prevously presented in uence analysis methodology by providing alternative treatment for continuous target variables like lead times and making it possible to give weights for each process case. We extend our contribution measure by presenting the de nitions for binary/continuous as well as weighted/non-weighted needs. Using a publicly available reallife case study from Rabobank's service desk process we demonstrate the e ect of using either continuous or binary approach combined with possible weighting.</p>
      </abstract>
      <kwd-group>
        <kwd>process analysis</kwd>
        <kwd>process improvement</kwd>
        <kwd>process mining</kwd>
        <kwd>lead times</kwd>
        <kwd>root cause analysis</kwd>
        <kwd>data mining</kwd>
        <kwd>in uence analysis</kwd>
        <kwd>contribution</kwd>
        <kwd>working capital</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Every process owner and business leader in the world would be happy to hear
that their own process or at least some part of it can be made faster. Reduced
operational costs, better customer satisfaction and more sales are all potential
bene ts from reducing the process lead times. Since real-life business processes
are often very complex and produce a lot of data, we need to consider many
potential root causes for lead time related problems, including for example
customer speci c requirements, available resources, required competences for
process workers, di erent business models, delivery options and products.</p>
      <p>
        Our previously published in uence analysis methodology shows how the root
causes can be identi ed for generic process related problems [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The limitation of
the already presented generic method is that it only supports binary classi cation
where each case must be considered either success or failure. Analysing lead
times is possible using binary classi cation by de ning for example that every
case taking more than 7 days is failure. However, in many business situations it
is important to take into account the actual duration so that the longer the lead
time is the bigger the problem. Another limitation of our previously published
method is that it did not support case speci c weights. In practice for some use
cases like quality and internal auditing purposes it is acceptable to have equal
weights for all cases since all cases should comply with regulations. For some
other cases like sales order it might be much more important to deliver the large
customer orders in time compared to delivering the small orders.
      </p>
      <p>In this paper we will present a methodology to systematically analyse and
provide actionable root causes for lead times issues in current business processes.
We identify the root causes why some cases have very long lead times and others
are very short. Our method analyses each case attribute and value separately.</p>
      <p>The rest of this paper is organized as follows: Section 2 introduces relevant
background in process mining and data analysis. Section 3 presents our extension
to the in uence analysis methodology introducing the contribution measures for
continuous variables and case-speci c weights. Section 4 shows a real-life example
followed by a section for Discussions and Summary.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        This paper extends the in uence analysis methodology that we have published
earlier [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. We consider in uence analysis as a practical actionable analysis which
utilizes extensively the experiences and ideas from process mining [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], including
speci cally enriching and transforming process-based logs for the purpose of root
cause analysis [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and correlating business process characteristics [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Speci cally
the generic framework presented in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] bene ts from using the formulas and
methods presented in this paper. For example considering the four additional use
cases presented Table 5 in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] the limitation of generic framework illustrations
is that the presented decision tree analysis only tries to show the positive root
causes for given process problem. Our methods as presented in this paper gives
the results in a form of comparative benchmarking thus showing both the most
in uencial root causes for the "bad behavior" as well as most in uencial root
causes for avoiding the "bad behavior". Ability to show simultaneously the root
causes for bad and good behavior makes it possible to quickly see whether the
problem cases have a clear root cause or maybe the good behavior cases have
a common root cause for their good behaviour. There has also been more work
in detection of di erences between groups [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and nding contrast sets [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Our
methodology is based on deviations management [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Even though business process performance has been studied a lot most of
the studies only cover the usage of binary conditions or decision tree approach
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Wetzstein et al. have presented a framework for monitoring and analyzing
in uential factors of business process performance [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. However their method
requires the usage of binary contribution measure and in this paper we will present
the option of using a continuous contribution formula. Grger et al. demonstrate
very relevant data mining approaches for manufacturing process optimization
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] using binary and decision tree approach.
      </p>
      <p>
        Basic idea of in uence analysis [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is to nd root causes for deviations in
the business process. In uence analysis has been used successfully for improving
incoming invoice handling process [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Examples of root causes for long lead
times found with in uence analysis in these four case study companies include
'Contract number blank', 'Currency GBP', 'Business Unit X', 'Invoice Type
EV' and 'Invoice status cancelled' [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Without data analysis tools it would be
very di cult to nd this kind of root causes. Our previously published in uence
analysis methodology consists of following steps:
1. Identify the relevant business process and de ne the case
2. Collect event and case attribute information
3. Create new categorization dimensions
4. Form a binary classi cation of cases such that each case is either problematic
or successful
5. Select a corresponding interestingness measure based on the desired level of
business process improvement e ect
6. Find the best categorization rules and attributes
7. Present the results to business people
      </p>
      <p>In this paper we extend the previous step 4. so that classi cation can be
either binary as previously or we can use a continuous variable for representing
the goodness or badness of a case. Regarding step 5. we only use the as-is
average as the Change Type in this paper as that measure has proven to be
most useful. However, it is also possible to use ideal and other average Change
Types. Regarding step 6. we add new calculation formulas to cover also weighted
versions of both binary and continuous contribution.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Analysis Types for In uence Analysis</title>
      <p>In this section we present four di erent formulas that are to be used as
contribution measures for in uence analysis methodology. These measures are listed
in Table 1. Depending on the performance indicator the contribution formula
can be either binary or continuous. Depending on relative importance of cases
the contribution formula can be weighted or not weighted. In typical business
process analysis situations an actual business problem can often be formulated
with any of these four formulas. Since the formulas give potentially di erent
results it is important to understand that seemingly small di erences in
formulating the problem may lead to large di erences in the analysis results. Thus it
is often bene cial to use multiple contribution formulas for double-checking that
suggested business process improvement areas are correct.</p>
      <p>
        Our previous paper [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] presented the binary performance indicator with equal
weights, corresponding to the contribution formula Binary Contribution (BiCo).
The contribution of this paper is to present three other formulas:
Continuous Contribution (CoCo), Weighted Binary Contribution (wBiCo) and Weighted
Continuous Contribution (wCoCo).
      </p>
      <p>
        Our method and calculations start from understanding the initial size of the
business process problem. As presented in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] one should focus development
resources to improving issues where the size of the problem is large and the size
of required investment is small. Problem size and an example lead time process
for each Contribution Formula is shown in Table 2. When considering business
process lead times we typically want to make the process generally faster
(continuous variable) or then we want to ensure that the lead time of each instance is
shorter than a given target (binary variable). Continuous is used when faster
performance is always better and there is no lower bound. Binary approach is used
for example when each process instance is categorized as successful if it meets
a Service Level Agreement (SLA) and unsuccessful if it exceeds SLA. Following
the power-law distributions in empirical data [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] principle we can use binary
approach by selecting about 20% of worst performing cases to nd explanations
for bad performance.
Amount of problematic cases In service desk process a lead time longer
than 7 days could be considered a problem
case.
      </p>
      <p>Sum of value of problematic Free-of-charge pizza if delivery takes more
cases than 45 minutes. Problem size is equal to the
monetary value of late pizza deliveries.</p>
      <p>Sum of positive overtime Lead time from the customer calling a
compared to average lead helpdesk to the moment the call is answered.
time The shorter the lead time the better it is.
Sum of overtime for each Lead time from the sending an invoice to the
case compared to the moment the payment arrives. When this lead
weighted average lead time time is multiplied by the value of the invoice
multiplied by the weight we get the working capital, ie. using value of
separately for each case invoice as the weight for each case.</p>
      <p>In this paper we consider an actionable business process improvement in
area X, which is a subset of the whole population, as an improvement that
will change the performance of future cases in area X to be improved so that
the performance of area X reaches the the current as-is average performance
of whole population. For example a US company may have delivery challenges
in Dallas region and the improvement project would then be to improve the
performance in Dallas to the same level as other regions. For each measure we
show the formula that we call Contribution% which gives result as a percentage
gure between -100% and +100%. Positive Contribution% indicates how large
part of the current lead time problem can be improved by making the selected
business area to perform on the same level as the initial average for the whole
business. Negative Contribution% tells how much bigger the lead time problem
will become if performance in selected business area is weakened to the current
average level.</p>
      <p>Weighting Weighted contributions can be used in contribution analysis. Simply
we need Weight attribute for each case and we need to replace the 'amount of
cases' values with 'sum of Weights of cases'. Now as an example we could have a
total amount of 13 million EUR orders in the analysis and 2 million EUR orders
are being delivered after the requested delivery date. So we will then run the
contribution analysis to nd the case attributes and values that have the biggest
contribution in terms of EUR to the 2 million that is being delivered late. If there
is one single order of 1.99 million EUR that was delivered late, then obviously
the characteristics of that single order will overrule all other possible ndings,
even though if 100 other orders were delivered late. But that is de nitely just
the wanted nding because in real life if situation is like that then the one order
(almost) fully explains the orders being late and there may be no point in trying
to nd more root causes for late deliveries.
3.1</p>
      <p>Common De nitions
Here we present the common de nitions used in all contribution formulas.
De nition 1. Let C = fc1; : : : ; cN g be a set of cases in the process analysis.
Each case represents a single business process execution instance.
De nition 2. Let Cp = fcp1 ; : : : ; cpN g be a set of problematic cases. Cp
C.</p>
      <p>De nition 3. Let Ca = fca1 ; : : : ; caN g be a set of cases belonging to business
process improvement segment A. Ca C.</p>
      <p>De nition 4. Let dcj be the duration of the case cj .</p>
      <p>De nition 5. Let wcj be the weight of the case cj . We consider linear weights
so that double weight always means double importance. If wcj = 0 then case cj
will have no e ect in the analysis when calculating weighted results.
De nition 6. Let pr be the size of the problem in the original situation
before any business process improvement: BiCo: amount of problem cases, wBiCo:
sum of weights of problem cases, CoCo: sum of overtime compared to average
duration, wCoCo: sum of overtime per case multiplied with weight of the case
compared to the weighted average duration.
3.2</p>
      <p>
        BiCo - Binary Contribution
For binary contribution the problem size is the amount of problematic cases.
Every case needs to be classi ed as problematic or successful as shown in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], ie.
in order to analyse the process lead time one needs to specify a limit such that
exceeding the limit classi es the case as problematic and otherwise it should be
successful.. De nitions for BiCo have already been presented in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. However we
have adopted a new syntax for the de nitions in order to make it easier for the
reader of this paper to compare binary/continuous and weighted/non-weighted
to each other.
      </p>
      <p>Total problem size for BiCo is the amount of problematic cases prBiCo =
Cp = P 1 as shown in equation 1 in Table 8 in Appendix A. Average
cj2Cp
P 1
function for BiCo is the average problem density rho = jCpj = cj2Cp
jCj P 1 as shown
cj2C
in equation 2. Similarly the average problem density for BiCo of subset Ca is</p>
      <p>P 1
a = jCjpC\aCj aj = cj2(CPp\C1a) as shown in equation 3. Finally theContribution%
cj2Ca
for BiCo of subset Ca is conBiCo =</p>
      <p>P 1
cj2Ca</p>
      <p>P 1 as shown in equation 4
cj2C
( a ) P 1</p>
      <p>cj2Ca
prBiCo</p>
      <p>P
= jCp\Caj jCaj = cj2(Cp\Ca)
jCpj jCj cjP2Cp 1
1
3.3</p>
      <p>wBiCo - Weighted Binary Contribution
Weighted Binary Contribution extends the previous sigma-based formulas by
replacing the static equal weight with case speci c weights wcj . Problem size
as de ned in equation 5 in Table 8 in Appendix A is the sum of weights of all
problem cases. Average problem density as de ned in equation 6 in Table 8 is
the sum of weights of all problem cases divided by the sum of weights of all
cases, and in a similar way the average problem density in equation 6 in Table
8 is the sum of weights of all problem cases in subset Ca divided by the sum of
weights of all cases in subset Ca.
3.4</p>
      <p>CoCo - Continuous Contribution
Continuous Contribution allows analysing the lead time variables as continuous
without the need for a xed separation of cases into long and short cases, ie.
without the need of having a binary value for each case. For continuous analysis
we consider a case as problematic if the value of continuous target variable is
bigger than the average in the population as shown in equation 9 in Table 8. The
bigger the positive di erence is the worse the behaviour. On the other hand, if
the continuous target value is less than average then the case is a
better-thanaverage. Using this approach the sum of positive deviations is always the same as
the absolute value of the sum of negative deviations, meaning that the problem
size for the whole population C is always zero. Problem size of any subset Ca
may be nonzero meaning that the cases in subset Ca either have higher or smaller
values for the target variable than the whole population. For analysing lead times
the continuous target variable is any lead time variable of the business process
cases, for example the total end-to-end lead time or any partial lead time from
one activity to another. Average function for continuous analysis types is the
average lead time, which is de ned for the whole population with equation 10
and subset using equation 11 in Table 8.</p>
      <p>Contribution measure for each possible subset Ca for CoCo is calculated as
follows: subtract the average lead time of whole population C from the average
lead time of the subset Ca, multiply this by the amount of cases in subset Ca.
This gives an absolute value of how much more or less time is spent on the cases
in subset Ca as a total compared to average of C. Final step is to divide this
gure by the problem size, ie by the total sum of positive (or negative) cases in
the population, giving the de nition for equation 12 in Table 8.
3.5</p>
      <p>wCoCo - Weighted Continuous Contribution
In this subsection we extend the previous de ned continuous contribution
formulas to supports case speci c weights. It is good to note that weighted continuous
contribution corresponds exactly to the working capital need in a business
process. As an example lets consider the process of building houses where each case
is one house. Working capital needed is proportional to the total cost of each
house and the lead time from starting the constructions to selling the house.
Weighted Continuous Contribution gives this measure when the cost of house
is used as case speci c weight and building time is used as the lead time. The
business improvement activity for reducing working capital for this construction
company then corresponds to conducting in uence analysis using weighted
continuous contribution analysis type to nd our the those subsets that should be
the focus for process improvements.</p>
      <p>Average weighted lead time using case speci c weights is calculated with
equation 14 in Table 8. Di erence to the non-weighted formula is that the lead
time of each case is multiplied by the case speci c weight and nally the result
is divided by the total sum of weights. This weighted lead time is then used to
calculate the total problem size according to equation 13 in Table 8 so that the
absolute di erence of lead time for each case is multiplied by the case speci c
weight and then summed up. In business terms this corresponds to calculating
the extra working capital (positive) or unneeded working capital (negative) for
each case and then summing them together. According to our approach if the
lead time for every case is equally long then the problem size is zero and there
is no extra working capital in the process.</p>
      <p>Finally the contribution calculations for weighted continuous analysis are
done similarly than in non-weighted analysis, ie subtract the weighted average
duration of subset Ca from the total weighted average and multiply this by the
sum of weights in subset Ca. When this is divided by the total problem size we
get the amount of working capital that would be freed if the lead times for cases
Ca could be reduced to the average weighted lead time in the whole population
C as shown in equation 16 in Table 8.
3.6</p>
      <p>Strengths and Weaknesses of Analysis Types
It is not trivial to decide which analysis type should be used in a particular
business process analysis situation. Table 3 shows the strengths and weaknesses
for binary and continuous analysis types and Table 4 respectively for weights.
Often it is desirable to select one analysis type as the primary type for a
particular analysis and then use the other analyses for reviewing, double-checking
and con rming results from a perspective.
{ Does not need any separate { Is sensitive for outliers. If one case
cut-of threshold. Continuous takes million times longer than the
variable like lead time is other cases then the whole analysis is
used directly by the algo- likely to suggest improvement in all
rithm and overtime is calcu- the subsets Ca containing that
particlated from the average dura- ular case.
tion.
Type
Equal
weights
Di erent
weights</p>
    </sec>
    <sec id="sec-4">
      <title>Case Study: Rabobank Group ICT</title>
      <p>
        In this section we show a real life example of using the presented analysis types
with a publicly available data from Rabobank Group ICT from BPI Challenge
2014 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The data contained 46 616 cases and a total of 466 737 events. As a lead
time we consider the total duration of each case including all cases found in the
dataset. Typical process mining analysis discovers that the average duration for
cases is 5.07 days and median duration is 18 hours. For the purpose of comparing
the analysis we set the threshold of problematic cases in the binary analyses to
be 7 days which results in a total of 7 400 (15.9%) problematic cases.
      </p>
      <p>As the weighting for cases we use a formula wcj = (6 Impactcj )(6
U rgencycj )(6 P rioritycj ) where Impact, Urgency and Priority all have values
in (1,2,3,4,5) where 1 means highest importance and 6 is the lowest importance.
With this formula the highest possible weight is whigh = (6 1)(6 1)(6 1) = 125
and lowest possible weight is wlow = (6 5)(6 5)(6 5) = 1. Using these weights
the weighted average lead time drops to 3.97 days. This means that on average
the lead time is shorter for more important cases than for less important cases.
Same nding can also be made from the binary results since the average problem
density is 15.9% and weighted average problem density is 11.1%.</p>
      <p>Top-3 positive and negative contributions for each Analysis Type are shown
in Table 5. For BiCo the highest contribution% is 4.2% for case attribute value
WBS000091 and lowest contribution% is -9.3% for case attribute value WBS000091.
When considering the most bene cial focus are for process improvement
reducing the lead time most we see that BiCo results in WBS000091 and all other
Contribution Formulas result in WBS000088. According to the gures the best
performing area regarding lead time is WBS000073 in all Contibution Formulas
except that in Weighted Binary Contribution the best practice area is #N/B.</p>
      <p>Some interesting results include the behaviour of cases whose attribute
ServiceComp WBS(CBy) has the value WBS000091 which contributes to 4.2% (Top
1) of the total problem in BiCo, 3.4% (Top 2) in wBiCo, but only 1.7% (Top
8) in CoCo and only 1.2% (Top 7) for wCoCo. Reason for higher contribution
in BiCo and lower in CoCo is that average lead time for area WBS000091 is
only 6.14 which is only little longer than the average for the whole process 5.07.
This means that there are many WBS000091 -cases that have lead time a little
bit longer than 7 days. On the other hand the behaviour of area WBS000088 is
the opposite, since it only contributes 2.4% (Top 3 value) of the total problem
in BiCo, 3.7% (Top 1) in wBiCo, much more 10.4% (Top 1) in CoCo and even
more 12.7% (Top 7) in wCoCo. Reason for this behaviour is that the average
lead time for cases in area WBS000088 is 39.2 days which is much longer than
the average lead time 5.07</p>
      <p>Very interesting results include the behaviour of area #N/B which is listed
as best practice area with negative contribution -1.7% in BiCo (Top -2), -8.3%
in wBiCo (Top -1) and -5.9% in wCoCo (Top -2). However it is listed as problem
area with positive contribution 2.6% in CoCo (Top 5 Problem Area!). There
are at least two reasons for this result: rst the high weight cases in #N/B
perform much better than the low weight cases, ie. BiCo contribution gets 6.6
percentage points better with weighting than without and CoCo contribution
gets 8.5 percentage points better. Second reason is that area #N/B performs
consistently worse in Continuous analysis compared to binary analysis, which is
caused by the higher than average lead time of 6.2% in CoCo, which again is
caused by certain amount of very long lead time cases.</p>
      <p>Activity occurrences. In this subsection we show the Rabobank root cause
analysis for long lasting cases using activity occurrence data. As a preprosessing
step we add a new case attribute for each di erent activity name and use the
amount of activity occurrences as the value for that case attribute in each case.
For example if activity Status Change occurs twice for a certain case, then the
value of case attribute Status Change will 2 for that particular case.</p>
      <p>Table 6 shows the top-5 positive and negative root causes for binary
analysis types and 7 for the continuous analysis types. From these tables we make
following observations:
{ Lack of reassignments is the most important negative root cause for a case
to exceed 7 day SLA (BiCo analyses) or generally take a long time (CoCo
analysis). In other words, having zero reassignments makes a case very fast.
{ Contribution values for activity occurrence amounts are much higher than
they are for the case attribute ServiceComp WBS(CBy), which means that
these activity amounts correlate more with the total duration than the case
attribute ServiceComp WBS(CBy).
{ Update from customer(1) is most important positive root cause for long case
duration as can be seen in continuous contributions in table 7. However
for binary contributions in table 6 the Status Change(2) is more important
positive root cause which means that having two occurences of Status Change
is a bigger risk for failing SLA than getting an update from customer.
In this paper we have presented a method for focusing business process
improvement to reduce lead times. We have de ned four di erent analysis types and
shown how they can be used with actual data. Summary of our key ndings is:
1. In uence analysis methodology is able to nd root causes for long lead times.
2. Root causes for long lead times may be substantially di erent when using
a prede ned lead time limit for problematic/successful cases (binary)
compared to when using continuous lead time values.
3. Case speci c weighting can be easily used when analysing both binary and
continuous contribution.
4. When weighting is used together with continuous contribution the analysis
can be directly used as working capital analysis solution.</p>
      <p>Acknowledgements. We thank QPR Software Plc for the practical experiences
from a wide variety of customer cases and for funding our research. The
algorithms presented in this paper have been implemented in a commercial process
mining tool QPR ProcessAnalyzer.</p>
    </sec>
    <sec id="sec-5">
      <title>A Appendix - Summary of Contribution Formulas</title>
      <p>(3)
(7)
(15)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Pazzani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>"Detecting group di erences: Mining contrast sets"</article-title>
          .
          <source>Data Mining and Knowledge Discovery</source>
          <volume>5</volume>
          (
          <issue>3</issue>
          ):
          <fpage>213246</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Clauset</surname>
          </string-name>
          , Aaron, Cosma Rohilla Shalizi, and Mark EJ Newman.
          <article-title>"Power-law distributions in empirical data</article-title>
          .
          <source>" SIAM review 51.4</source>
          (
          <year>2009</year>
          ):
          <fpage>661</fpage>
          -
          <lpage>703</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Grger</surname>
            , Christoph,
            <given-names>Florian</given-names>
          </string-name>
          <string-name>
            <surname>Niedermann</surname>
            , and
            <given-names>Bernhard</given-names>
          </string-name>
          <string-name>
            <surname>Mitschang</surname>
          </string-name>
          .
          <article-title>"Data miningdriven manufacturing process optimization." Proceedings of the world congress on engineering</article-title>
          . Vol.
          <volume>3</volume>
          .
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>De</surname>
            <given-names>Leoni</given-names>
          </string-name>
          , Massimiliano, Wil MP van der Aalst, and
          <string-name>
            <given-names>Marcus</given-names>
            <surname>Dees</surname>
          </string-name>
          .
          <article-title>"A general process mining framework for correlating, predicting and clustering dynamic behavior based on event logs." Information Systems (</article-title>
          <year>2015</year>
          ), http://dx.doi.org/10.1016/j.is.
          <year>2015</year>
          .
          <volume>07</volume>
          .003
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Lehto</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinkka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Hollmen</surname>
          </string-name>
          ,
          <source>J. "Focusing Business Improvements Using Process Mining Based In uence Analysis." International Conference on Business Process Management</source>
          . Springer International Publishing,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Piatetsky-Shapiro</surname>
            , Gregory, and
            <given-names>Christopher J.</given-names>
          </string-name>
          <string-name>
            <surname>Matheus</surname>
          </string-name>
          .
          <article-title>"The interestingness of deviations."</article-title>
          <source>Proceedings of the AAAI-94 workshop on Knowledge Discovery in Databases</source>
          . Vol.
          <volume>1</volume>
          .
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>QPR</given-names>
            <surname>Software Plc</surname>
          </string-name>
          .
          <article-title>"Improved invoicing for Caverion through process insight."</article-title>
          <source>IEEE CIS Task Force on Process Mining, Case Studies</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Rozinat</surname>
          </string-name>
          , Anne, and
          <string-name>
            <surname>Wil MP van der Aalst</surname>
          </string-name>
          .
          <article-title>Decision mining in ProM</article-title>
          . Springer Berlin Hei-delberg,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Suriadi</surname>
          </string-name>
          , Suriadi, Ouyang, Chun, van der Aalst,
          <string-name>
            <surname>Wil</surname>
            <given-names>M.P.</given-names>
          </string-name>
          , &amp; ter Hofstede,
          <source>Arthur</source>
          (
          <year>2013</year>
          )
          <article-title>Root cause analysis with enriched process logs</article-title>
          .
          <source>Lecture Notes in Business Information Processing [Business Process Management Workshops: BPM 2012 International Work-shops Revised Papers]</source>
          ,
          <volume>132</volume>
          , pp.
          <fpage>174</fpage>
          -
          <lpage>186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Van Der Aalst</surname>
          </string-name>
          , Wil, et al.
          <article-title>"Process mining manifesto." Business process management workshops</article-title>
          . Springer Berlin Heidelberg,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Van Dongen</surname>
            ,
            <given-names>B.F. BPI Challenge 2014. Rabobank</given-names>
          </string-name>
          <string-name>
            <surname>Nederland</surname>
          </string-name>
          . Dataset. http://dx.doi.org/10.4121/uuid:
          <fpage>c3e5d162</fpage>
          -0cfd
          <string-name>
            <surname>-</surname>
          </string-name>
          4bb0
          <string-name>
            <surname>-</surname>
          </string-name>
          bd82
          <source>-af5268819c35</source>
          ,
          <year>2014</year>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Vanjoki</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          "
          <source>Automated Purchase to Pay Process Value Modeling and Comparative Process Speeds."</source>
          Lappeenranta University of Technology,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Webb</surname>
            , GI and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Butler</surname>
            and
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Newlands</surname>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>"On Detecting Di erences Between Groups"</article-title>
          .
          <source>KDD'03 Proceedings of the Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Wetzstein</surname>
          </string-name>
          ,
          <string-name>
            <surname>Branimir</surname>
          </string-name>
          , et al.
          <article-title>"Monitoring and analyzing in uential factors of business process performance."</article-title>
          <source>Enterprise Distributed Object Computing Conference</source>
          ,
          <year>2009</year>
          . EDOC'09. IEEE International. IEEE,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>