<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Case Study on the Business Bene ts of Automated Process Discovery</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Matej Puchovsky</string-name>
          <email>matej.puchovsky@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudio Di Ciccio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jan Mendling</string-name>
          <email>jan.mendlingg@wu.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Vienna University of Economics and Business</institution>
        </aff>
      </contrib-group>
      <fpage>35</fpage>
      <lpage>49</lpage>
      <abstract>
        <p>Automated process discovery represents the de ning capability of process mining. By exploiting transactional data from information systems, it aims to extract valuable process knowledge. Through process mining, an important link between two disciplines { data mining and business process management { has been established. However, while methods of both data mining and process management are wellestablished in practice, the potential of process mining for evaluation of business operations has only been recently recognised outside academia. Our quantitative analysis of real-life event log data investigates both the performance and social dimensions of a selected core business process of an Austrian IT service company. It shows that organisations can substantially bene t from adopting automated process discovery methods to visualise, understand and evaluate their processes. This is of particular relevance in today's world of data-driven decision making.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        In order to sustain competitive advantage and superior performance in rapidly
changing environments, companies have intensi ed their e orts towards
structuring their processes in a better and smarter way. Business Process Management
(BPM) is considered an e ective way of managing complex corporate activities.
However, a consistent shift from mere process modelling and simulation towards
monitoring of process execution and data exploitation can be observed nowadays
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Event logs, i.e., data mapping traces of process execution in modern
information systems, can be used to de ne new process models and derive optimisation
possibilities. The techniques of automated process discovery aim at extracting
process knowledge from event logs and represent the initial steps of exploring
capabilities of process mining.
      </p>
      <p>
        Studying real-life enterprise data by applying process mining methods can
deliver valuable insights into the actual execution of business operations and
their performance. This is of particular relevance in today's age of industries
automating their processes via work ow systems [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] with growing support of
process execution by various enterprise information systems. By building
process de nitions, models and exploring execution variations, automated process
discovery has the potential to ll the information gap between business process
departments and domain experts in enterprise settings. In order to assess the
business bene ts of process mining, we cooperated with an Austrian IT service
provider to conduct an industrial case study. The company aimed at analysing
one of their core business processes using process execution data from an
enterprise resource planning system (ERP). Since our industrial partner has had no
previous experience with process mining, the case study focuses on a preliminary
assessment of data exploitation through a process discovery initiative.
      </p>
      <p>The remainder of the paper is as follows. In Section 2, we introduce the
research area of process mining with its fundamental terminology and related work
in the eld. Section 3 describes the studied process together with the structure
of the examined event log. Besides, we present the approach for our case study.
The results of the empirical analysis are covered in Section 4, which is divided
into the process and social views as well as a subsection covering the execution
variants. Section 5 elaborates on the business impacts of the research results.
Finally, Section 6 summarises our ndings and gives suggestions towards future
research.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        Van der Aalst [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] de nes process mining as the link between data science and
process science, making it possible to depict and analyse data in dynamic structures.
By analysing event logs, i.e., records documenting the process executions tracked
by an information system, process mining aims at retrieving process knowledge
{ usually in a form of process models. Depending on the setting, the extracted
models may (i) de ne and document completely new processes, (ii) serve as a
starting point for process improvements, or (iii) be used to align the recorded
executions with the prede ned process models. The idea to discover processes
from the data of work ow management systems was introduced in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Many
techniques have been proposed since then: Pure algorithmic ones [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], heuristic
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], fuzzy [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], genetic (e.g., [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]), etc. ProM [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is currently one of the most used
plug-in based software environment incorporating the implementation of many
process mining techniques. Process mining is nowadays a well-established eld
in research and examples of applications in practical scenarios can be found in
various environments such as healthcare services [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], nancial audits [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
or public sector [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The organisational perspective of process mining focuses
on investigating the involvement of human resources in processes. Schonig et
al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] propose a discovery framework postulating background knowledge about
the organisational structure of the company, its roles and units as a vital input
for analysing the organisational perspective of processes.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Case study</title>
      <p>
        The empirical study of the process mining methods was conducted in
cooperation with an Austrian IT service provider with a diversi ed service portfolio. The
main aim of the analysis was to evaluate the applicability of process mining for
process discovery and to provide solid fundaments for future process
improvements. Our industrial partner has a well-developed business process management
department responsible for all phases of the business process management
lifecycle, as de ned by Dumas et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. All process models and supporting documents
are saved in a centralised process management database, which is also used for
storing all quality management reports.
      </p>
      <p>During our investigation, we analysed the process of o ering a new service
(henceforth, o ering process), which belongs to the core business operations in
the value-chain of our industrial partner. The process describes all the necessary
steps that need to be performed in order to create an o ering proposal with
detailed documents on a speci c service, which is a subject of the o er,
including pricing information for the customer. Upon approval, the o er is sent to the
customer, who decides whether it meets their expectations or not. In case of an
acceptance, a new contract is concluded. If the customer rejects the o er, they
can either return it for adjustments, when a new version of the o ering
documents is created, or decline any further modi cations. Up until now, almost all
information regarding the process has been collected during numerous iterative
sessions of interviews with the domain experts. The industrial partner selected
this process due to the complex structures of its process model, which has
already been simpli ed, as well as an extensive support through an IT system
and suspected loops in process execution that are believed to negatively impact
process duration and often cause lack of transparency. The primarily goal is
therefore to understand how the process actually works, based on the available
execution data.
3.1</p>
      <sec id="sec-3-1">
        <title>Approach</title>
        <p>Our case study aims to give answers to the following business questions:
{ How does the o ering process perform based on the available log data? Here,
the focus is set primarily on process duration, acceptance rates of the o ers,
and number of iterations in the recorded cases. Moreover, we investigate what
variables might have and impact on the process duration. We thus aim to
exploit both the \traditional" log attributes such as timestamps or activities
and \non-standard" attributes with additional process information.
{ Are there any iteration loops recorded in the log, and if yes, how many? How
many di erent variations of process execution can be identi ed and do they
vary signi cantly?
{ How many human resources are involved in the process execution and what
roles can be identi ed in the event log records?</p>
        <p>The analysis was conducted primarily by means of the process mining tool
minit1. Throughout the paper, we di erentiate between the process and social
view, the former dealing with the performance aspects of the process, the latter
explaining its social network. The event log data were processed within both
frequency (i.e., absolute numbers of executions) and performance dimensions
(i.e., duration metrics). Through a comparison of the process execution variants,
i.e., activity sequences shared by more process instances, we aim to identify
similarities between cases, but also to uncover execution anomalies recorded in
the event log. The analysis of the connections between human resources and
their activities in the work ows are used to cover the social perspective of the
examined process.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Description of the event log data</title>
        <p>The event log used in this study was created by the IT system owner via a query
in the EPR system our industrial partner. The extracted cases were restricted
to the work ows started in 2015. Due to the fact that no formal end activity
is de ned in the process, the log possibly contains several incomplete process
instances. We therefore have to resort to speci c attributes to determine the
status of the analysed process instances. Our industrial partner provided us
with a single CSV data table with 12,577 entries (activities), corresponding to
1,265 distinct process instances. Due to privacy concerns of the company, no
other platforms were allowed. Therefore, the fact that we were only able to work
with one data le was a major limitation for our research.</p>
        <p>The recorded events are enriched with the following 11 attributes: CaseID,
Version, Customer, Segment, WF-Start, WF-Status, Activity, Start, End,
Duration, Performer. Overall, 258 employees were involved in the process execution.
The activity names were recorded in German, therefore, we have complemented
these with English translations (see Table 1) and also employ the labels in
English throughout the paper.</p>
        <p>The log attribute work ow status (WFS) plays an important role for our
analysis, even though it is not a typical attribute such as activity name or timestamp.
The values recorded under this attribute provide information about the outcome
of the o ering process or the status of the o er itself. Work ow statuses are
automatically recorded at the beginning of every activity and can be changed after
the activity is executed. All previous activities are then marked with the latest
WFS. This attribute can also be used to determine terminal activities, which
possess one of the three work ow statuses describing the outcome of an o ering
process { accepted, aborted or rejected.</p>
        <p>However, a case can be assigned more than one work ow status, signalising
an important relationship with another attribute { version. A good example of
such case is an o er that requires adjustments after being sent to a customer,
who rejected the original version. For this purpose, a new version of the o ering
documents is created, resulting in a new iteration of the work ow { the activities
in the rst iteration are recorded with new version. If the o er is accepted, the
activities executed in the second loop are recorded with o er accepted. We can
conclude that cases with new version contain at least one loop and that the
version attribute. A closer analysis of both attributes is included in the section
dealing with process performance.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>The following section describes the experimental analysis of the event log data.
4.1</p>
      <sec id="sec-4-1">
        <title>Process view</title>
        <p>In order to be able to build a process map from the event log data, the activity
names and timestamps are essential.The process map is visualised as a directed
graph consisting of n nodes and e edges. Each node represents an event or
activity with an individual name and, based on the chosen dimension, the information
about the execution frequency or duration.</p>
        <p>Our initial point of interest is the process map depicting the sequence of
activities under the frequency dimension, i.e. how often was a particular activity
executed or in how many cases can an activity be found. In order to alter the
abstraction level of the process map, we use di erent complexity levels of activities
(nodes) and paths (edges/transitions), which determine the number of elements
depicted in the model { the higher the complexity level, the more nodes/edges
are visualised in the process map.</p>
        <p>We rst postulate a 0% complexity of both the activities and paths, resulting
in a process model with six most frequently executed activities plus the terminal
nodes start and end. In Figure 1a, the ordered activities contain the descriptive
metric of case count, i.e., the frequency of cases in which the particular activity
was performed. For example, it can be observed that the starting activity O er
creation was performed in 1,173 process instances of the o ering work ow and is
therefore accompanied by the strongest highlight around its node. Based on the
recorded data, we can conclude that the model in Figure 1a represents the most
common path/behaviour of the process. Increasing the percentage of activities
and paths visualised in the model leads to growing complexity of the whole
(a) 0% activity, 0% path
complexity level
(b) 50% activity, 0% path
complexity level
graph. This can be observed in Figure 1b, where 50% of all activities available
in the event log are depicted. We can observe that additional activity types from
the pre-sales phase emerge. Their presence in the model signalises preparing
non-standard o ers, in which a pre-sales team has to be involved to perform the
requirement analysis and deliver a solution design for the desired service. The
proportion of the process instances with the pre-sales phase was slightly over
40%.</p>
        <p>The highest possible level of activity complexity uncovers a total of 18 logged
process activities. At 100% activity and path complexity levels, the generated
process map becomes very cluttered (see Figure 2). The graph gives us a good
perception of how cross-connected the activities in the process are. The o ering
process evidently can be performed in numerous variations. Similarly to the
starting events, the number of possible activities ending the work ow increases
at higher levels of model complexity.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Variants</title>
        <p>A variant can be de ned as a set of process instances that share a speci c
sequence of activities. This means that all cases in a certain variant have the
same start and end event and that all activities between the terminal nodes
are identical and performed in the same chronological order. Variants can be
very helpful for recognising di erent performance behaviour or irregular patterns
not conforming with the a-priori process model. In general, metrics from both
frequency and performance dimensions can be assessed.</p>
        <p>The event log of the o ering process depicts a total of 280 variants with
event count per case varying from 1 to 49. Due to the lack of space, we only
concentrate on the top 24 performance variants that cover around 70% of all
available cases.</p>
        <p>The key goal of our variant comparison is to nd similarities between the 24
investigated variants and thus build variant groups explaining the execution of
the o ering work ow. This allows us to identify two major groups of variants
based on whether they depict o ers for standard or non-standard services, let
them be labelled A and B.</p>
        <p>All variants in group A share the initial activity O er creation and end with
either Customer decision or Takeover (see Fig. 3). There are exactly nine variants
contained in group A with total case coverage of almost 40% (444 cases). Since
no pre-sales activities can be found in the cases of cluster A, we can conclude
that all the involved process instances cover standardised o ers. The presence
of the activity Amend attributes signalises that iteration loops were recorded
in the cases involved. The second identi ed variant group, B, features a higher
variety of the recorded activities. It contains 15 variants with event frequency
varying from 2 to 15. The cases found in this group all contain at least one
activity from the pre-sales phase, thus marking process instances dealing with
non-standardised services. Interestingly, no loop performance was identi ed in
the variant group B, as the recorded cases were all missing the activity Amend
attributes, i.e. a trigger for repeated process executions. Fig. 4 shows that, for
both variant groups, the mean case duration tends to increase with growing
number of activities per case. This can be explained both by the presence of
loops or to the resource-intensive pre-sales phase.
(a) Frequency dimension
(b) Performance dimension
For the analysed process, we de ne three main KPIs: (a) duration, (b) number of
iterations loops in the process, and (c) process outcome. The duration metrics
can be assessed by making use of the start and end timestamps, which are
recorded for every activity in the log. By exploiting the data recorded under the
version attribute we analyse the number of iterations in the recorded process
instances. Finally, the process outcome, i.e., the acceptance rate of the o ers
created through the process, is derived from the work ow statuses of the recorded
cases.</p>
        <p>Several activities stand out when it comes to the duration metrics { Customer
decision accounting the mean duration of over three weeks followed by
Requirement analysis with signi cantly lower mean duration of 1.67 weeks and O er
creation (6.03 days). These results align with the available a-priori process
knowledge as they are known to require the most resources and processing time. The
total duration of the analysed cases varies signi cantly. Therefore, we investigated
the impact of three selected variables { number of activities in a case, resources
involved and the number of versions a work ow possesses (since versions can be
seen as iteration loops) { on the total case duration via a multivariate regression
model: case duration = + a activities + r resources + v versions.</p>
        <p>The coe cients of all activities ( a = 28:79), resources ( r = 77:883) and
versions ( v = 276:154) are signi cantly di erent from 0. We can therefore
conclude that an increase in all selected variables results in an increased case
duration with the number of versions having the highest impact. The model
variables exhibit a mild degree of correlation (r = 0:458). However, the coe cient
of determination amounts to 0:21, meaning that only 21% of the case durations
can be explained by the selected variables.</p>
        <p>The version attribute denotes the version of the o ering documents
throughout the whole process. The document numbering is entered manually by human
resources and begins with V1.0 for new cases. The version number is then
automatically assigned to all new activities in the process until the version, and thus,
the document, is changed (1.x for minor changes or drafts, x.0 for major
alterations). Version numbers can therefore be seen from two major points of view:
(i) business, where altering versions corresponds to changes in the o ering
documents, and (ii) process, with versions being equivalent to the number of loops
in a process instance. By inducing the version number as a metric for process
iteration loops, we can conclude that 1,009 (79.76%) out of the recorded 1,265
cases were executed in only one iteration. These are followed by 192 process
instances (15.18%) with two iterations. Only 64 (around 5%) cases consisted of
three or more loops. These ndings indicate a straight-forward execution of the
process in the majority of examined cases.</p>
        <p>We can also observe a relationship between the version attribute and
workow status, which is depicted in Figure 5, where the x-axis represents the
number of distinct versions found in a case and the y-axis showing the number of
work ow statuses per case. Generally speaking, a change in the work ow status
triggers a change in the version number. We observe that most process instances
are recorded in a 1:1 relationship, where the whole work ow holds only one
version number and one work ow status { e.g., an o ering proposal that has been
accepted after the rst iteration with the version number V1.0, or a proposal
that has been rejected after the rst iteration. The second largest group, 2:2,
is common for process instances, where two distinct document versions were
recorded (e.g., V1.0 and V2.0) together with two distinct work ow statuses.
This behaviour is identi ed in o erings that were returned for adjustments by
the customer (i.e., triggering a change in the version number) and were accepted
after the second iteration.</p>
        <p>Figure 6 depicts a bar chart with all 13 work ow statuses found in the event
log with the corresponding case counts. It shows the apparent dominance of
O er accepted with 715 cases (56.52% of all recorded). According to the data,
only 61 o ers were rejected by the customer. The sum of the cases in Figure 6
is 1,610, which is higher than the total case count in the event log (1,265). This
di erence can be explained by process instances that hold multiple work ow
statuses. Yet, some of the recorded statuses do not provide explicit information
about the process outcome. For instance, the o ering process was aborted in
273 cases. Unfortunately, the data do not provide any additional information
concerning the reasons for terminating the process.</p>
        <p>
          Combining version numbers with the work ow status also allows for
interpreting the \performance e ciency". 70.21% (502) of all accepted o ering
proposals were concluded after the rst iteration with the version number V1.0.
However, we also discovered a signi cant proportion of cases that entered a new
loop (242 cases or 20.23%) or were aborted (221 work ows or 28.48%) after the
rst iteration with the version number V1.0.
Analysing the social graph of the process helps us understand the relationships
between di erent resources involved in the process execution. Our investigation
of the social network was primarily focused on examining its frequency aspects.
We resort on the investigation of relationships of the performers involved in the
process execution at di erent complexity levels. The initial complexity levels
were set to 50% for the resources and 0% for the connections, resulting in the
social network consisting of 32 nodes and 48 edges depicted in Figure 7. In [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ],
Van Steen mentions that the key to understanding the structure of a social
network is a deeper analysis of subgroups found within the network. Since the
investigated social network exhibits very high complexity with 258 resources and
396 connections, this approach is also tting for our case study. Based on the log
data, the top 15 performers (5.81%) account for 60% of all the recorded events.
Here, several relationships between speci c performers stick out (see Table 2).
Noticeable connection counts can also be observed within performers who pass
the workload to themselves. The latter type of relationships appears to have
signi cantly higher total count than the connections between resources, despite
the case occurrences staying very similar. We can therefore conclude that larger
sets of subsequent activities tend to be performed by the same resource, whereas
handing over the workload to other employees occurs less often.
        </p>
        <p>The comparison of the event and case involvement rates, i.e., respectively the
share on the total number of recorded events, and the proportion of cases in which
a particular resource was involved (see Table 3), provides some assumptions
towards the roles of the involved employees. It can be assumed that those with a
Connection Cases Total
BNH49 ! UJH83 143 179
CPI61 ! BNH49 109 127
PXC22 ! GYS85 105 131
CPI61 ! CPI61 227 541
ZHC41 ! ZHC41 111 354</p>
        <p>XJQ61 ! XJQ61 109 276
higher share of executed events perform more activities within a process instance.
On the other hand, the resources with lower event involvement rates, though
active in many cases, are assumed to act as reviewers and approvers of the
o ering proposals who are only responsible for a few speci c activities.</p>
        <p>Considering human resources with regards to the performance aspects of the
process, we set an initial hypothesis that an increasing count in activity instances
per case leads to an increase in the number of employees involved in the process
execution. However, in process instances containing loops, the activities within
the new iterations are assumed to be performed by the same resources who have
already been employed in the execution of previous iterations and know the case
speci cations. This is re ected in lower growth rate of distinct human resources
per case.</p>
        <p>The scatter plot in Figure 8 supports our hypothesis. Since the data
exhibit a rather quick initial increase that then levels out, we computed a
logarithmic trendline for the development of \resource population" per case with
growing activity counts. With the available process knowledge, the log-trendline
can be interpreted as follows: The initial quick increase in resource counts per
case is linked to the di erence between cases with standard and non-standard
services, where the resource-consuming pre-sales phase is required. As the
employees of the pre-sales team di er from those who actually create the o
ering documents, the di erences between these two groups are higher.
Subsequently, this results in a quick increase of resource counts. However, further
increase in activity count per case leads to a slower growth rate in the
resource count. These ndings support our hypothesis that with multiple
iterations in the process, the repeated activities are executed by the same
human resources who have already been involved in previous iterations and the
\intake" rate of new employees levels out. The level-log regression equation,
resources = 0:7005 + 2:6545 ln(activities), shows that an increase in the
activity count per case by one percent, we expect the resource count to increase by
(2:6545=100) human resources. The coe cient of determination of 0:7372
indicates a good t of the regression model with a signi cant result (p &lt; 0:001).
Our study evidenced that applying methods of automated process discovery can
yield useful insights into both the performance and the social structure of
business processes. By analysing event log data of an ERP system, we detected a
signi cant number of di ering execution variants among the recorded cases of
the investigated process. We have shown that duration can be in uenced by the
number of activities per case, resources and number of loops. The presence of a
pre-sales phase in o ering work ows covering non-standard services often causes
a longer case duration and higher numbers of employees involved in the
process. Therefore, we propose the portfolio with standard services of our industrial
partner to be extended by new services that have previously successfully met
customer expectations. This could contribute to reducing the number of cases
with the pre-sales phase and the execution time of the process.</p>
        <p>From the o ering proposals recorded in the investigated log, around 56.5%
were accepted, a gure that our industrial partner aims to improve in the future.
On the other hand, over 70% of the contracts were concluded after only one
process iteration, which is a positive result. Considering the presence of loops in
the process, the investigated cases were performed in a relatively straight-forward
way with almost 80% of cases nished without any loops. Therefore, the initial
suspicion that the majority of cases included several iteration loops was not
con rmed. Still, it is proposed to introduce a comprehensive list of all customer
change requests once the o ering documents are returned for adjustments (and
thus triggering new loop). That way, future performance of the process in the
rst iteration can be improved.</p>
        <p>Based on the events analysis, the case counts and the social network diagram,
the roles of the recorded resources were identi ed. Interestingly, the very high
number of employees involved in the process execution was in fact unknown to
our industrial partner. Even though 258 employees were involved in the examined
cases, most of the activities were performed by only few individuals. Therefore,
it is advised to assess whether to distribute the workload more evenly among the
available employees or to reduce the resource counts for this particular process
to only the key personnel.</p>
        <p>Our research was limited by the restriction on tools that we were allowed to
use by the industrial partner and by the lack of information about the speci
cations of the IT system used as the data source. Nonetheless, it can be concluded
that these results represent a solid basis for future decisions with regards to
process optimisation.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper, we demonstrated how process mining can be used to extract
valuable knowledge about business processes from transactional data. Our study
provides evidence that by exploiting all attributes of the event log, it is possible
to obtain extensive knowledge about the dynamics, performance and the resource
structure of the examined process.</p>
      <p>
        Our investigation was hampered by a relatively high number of events for
which the resource was not speci ed. Such events were also associated to missing
end-timestamps. Therefore, we chose not to include those events in the analysis
of both the process variants and the social network, as they would have yielded
biased results otherwise. The lack of reported information often a ects real-world
logs. It is therefore in the future plans to integrate the techniques of Rogge-Solti
et al. [
        <xref ref-type="bibr" rid="ref17 ref18">17, 18</xref>
        ] to replace missing values with estimates, so as not to disregard
events which may otherwise be relevant because of other information they bring.
Considering the social perspective, we plan to acquire further knowledge about
the organisational structure of the enterprises in future case studies as well as the
role allocation of the available individuals involved in the process execution. This
is of particular importance when verifying whether the designated responsibilities
set for the execution of speci c activities by the enterprise are not violated.
      </p>
      <p>
        In the light of the promising results achieved, it is in our plans to
collaborate further with our industrial partner to deepen the investigation by means
of other process mining techniques, in the spectrum both of process discovery
and conformance checking. Furthermore, the high variability of recorded process
instance executions could also suggest that clarifying results might derive from
declarative process discovery [
        <xref ref-type="bibr" rid="ref12 ref13 ref9">13, 9, 12</xref>
        ]. The declarative approach indeed aims
at understanding the behavioural constraints among activities that are never
violated during the process enactment, rather than specifying all variants as
branches of an imperative model. As a consequence, it is naturally suited to
exible work ows [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.: Process Mining: Data Science in Action. Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.M.P.,
          <string-name>
            <surname>van Dongen</surname>
            ,
            <given-names>B.F.</given-names>
          </string-name>
          , Gunther,
          <string-name>
            <given-names>C.W.</given-names>
            ,
            <surname>Rozinat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Verbeek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Weijters</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          :
          <article-title>Prom: The process mining toolkit</article-title>
          .
          <source>In: BPM (Demos)</source>
          .
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <volume>489</volume>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>van der Aalst</surname>
            , W.M.P., van Hee,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Work ow Management: Models, Methods, and Systems</article-title>
          . MIT Press, Cambridge, MA, USA (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pesic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schonenberg</surname>
          </string-name>
          , H.:
          <article-title>Declarative work ows: Balancing between exibility and support</article-title>
          .
          <source>Computer Science - R&amp;D</source>
          <volume>23</volume>
          (
          <issue>2</issue>
          ),
          <volume>99</volume>
          {
          <fpage>113</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reijers</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weijters</surname>
            ,
            <given-names>A.J.M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>van Dongen</surname>
            ,
            <given-names>B.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Medeiros</surname>
            ,
            <given-names>A.K.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verbeek</surname>
            ,
            <given-names>H.M.W.E.</given-names>
          </string-name>
          :
          <article-title>Business process mining: An industrial application</article-title>
          .
          <source>Inf. Syst</source>
          .
          <volume>32</volume>
          (
          <issue>5</issue>
          ),
          <volume>713</volume>
          {
          <fpage>732</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weijters</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maruster</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Work ow mining: Discovering process models from event logs</article-title>
          .
          <source>IEEE Trans. Knowl. Data Eng</source>
          .
          <volume>16</volume>
          (
          <issue>9</issue>
          ),
          <volume>1128</volume>
          {
          <fpage>1142</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gunopulos</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leymann</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Mining process models from workow logs</article-title>
          .
          <source>In: EDBT. Lecture Notes in Computer Science</source>
          , vol.
          <volume>1377</volume>
          , pp.
          <volume>469</volume>
          {
          <fpage>483</fpage>
          . Springer (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Castellanos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Alves de Medeiros,
          <string-name>
            <given-names>A.K.</given-names>
            ,
            <surname>Mendling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Weber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Weijters</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.J.M.M.:</surname>
          </string-name>
          <article-title>Business process intelligence</article-title>
          . In: Cardoso,
          <string-name>
            <surname>J</surname>
          </string-name>
          . (ed.)
          <article-title>Handbook of Research on Business Process Modeling, chap</article-title>
          . XXI, pp.
          <volume>467</volume>
          {
          <fpage>491</fpage>
          .
          <string-name>
            <given-names>IGI</given-names>
            <surname>Global</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Di</given-names>
            <surname>Ciccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Mecella</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>On the discovery of declarative control ows for artful processes</article-title>
          .
          <source>ACM Trans. Management Inf. Syst</source>
          .
          <volume>5</volume>
          (
          <issue>4</issue>
          ),
          <volume>24</volume>
          :1{
          <fpage>24</fpage>
          :
          <fpage>37</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Dumas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosa</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reijers</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          :
          <source>Fundamentals of Business Process Management</source>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. Gunther, C.W., van der Aalst,
          <string-name>
            <surname>W.M.P.</surname>
          </string-name>
          :
          <article-title>Fuzzy mining - adaptive process simpli - cation based on multi-perspective metrics</article-title>
          .
          <source>In: BPM. Lecture Notes in Computer Science</source>
          , vol.
          <volume>4714</volume>
          , pp.
          <volume>328</volume>
          {
          <fpage>343</fpage>
          . Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Kala</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maggi</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Di</given-names>
            <surname>Ciccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Di Francescomarino</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          :
          <article-title>Apriori and sequence analysis for discovering declarative process models</article-title>
          .
          <source>In: EDOC</source>
          . pp.
          <volume>1</volume>
          {
          <issue>9</issue>
          .
          <string-name>
            <given-names>IEEE</given-names>
            <surname>Computer</surname>
          </string-name>
          <article-title>Society (</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Maggi</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mooij</surname>
            , A.J., van der Aalst,
            <given-names>W.M.P.:</given-names>
          </string-name>
          <article-title>User-guided discovery of declarative process models</article-title>
          .
          <source>In: CIDM</source>
          . pp.
          <volume>192</volume>
          {
          <fpage>199</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Mans</surname>
            ,
            <given-names>R.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schonenberg</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
          </string-name>
          , M.,
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bakker</surname>
            ,
            <given-names>P.J.M.:</given-names>
          </string-name>
          <article-title>Application of process mining in healthcare - A case study in a dutch hospital</article-title>
          .
          <source>In: BIOSTEC (Selected Papers)</source>
          .
          <source>Communications in Computer and Information Science</source>
          , vol.
          <volume>25</volume>
          , pp.
          <volume>425</volume>
          {
          <fpage>438</fpage>
          . Springer (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Mans</surname>
          </string-name>
          , R.:
          <article-title>Process mining in healthcare (</article-title>
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>de Medeiros</surname>
            ,
            <given-names>A.K.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weijters</surname>
            ,
            <given-names>A.J.M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.M.P.:
          <article-title>Genetic process mining: an experimental evaluation</article-title>
          .
          <source>Data Min. Knowl. Discov</source>
          .
          <volume>14</volume>
          (
          <issue>2</issue>
          ),
          <volume>245</volume>
          {
          <fpage>304</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Rogge-Solti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mans</surname>
            , R., van der Aalst,
            <given-names>W.M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weske</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Improving documentation by repairing event logs</article-title>
          .
          <source>In: PoEM. Lecture Notes in Business Information Processing</source>
          , vol.
          <volume>165</volume>
          , pp.
          <volume>129</volume>
          {
          <fpage>144</fpage>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Rogge-Solti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weske</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Prediction of business process durations using nonmarkovian stochastic petri nets</article-title>
          .
          <source>Inf. Syst</source>
          .
          <volume>54</volume>
          ,
          <issue>1</issue>
          {
          <fpage>14</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. Schonig,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Cabanillas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Jablonski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Mendling</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.:</surname>
          </string-name>
          <article-title>A framework for e ciently mining the organisational perspective of business processes</article-title>
          .
          <source>Decision Support Systems</source>
          <volume>89</volume>
          ,
          <fpage>87</fpage>
          {
          <fpage>97</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Van Steen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Graph Theory and Complex Networks: An Introduction (</article-title>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Weijters</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.J.M.M.</surname>
          </string-name>
          ,
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          :
          <article-title>Rediscovering work ow models from event-based data using little thumb</article-title>
          .
          <source>Integrated Computer-Aided Engineering</source>
          <volume>10</volume>
          (
          <issue>2</issue>
          ),
          <volume>151</volume>
          {
          <fpage>162</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Werner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gehrke</surname>
          </string-name>
          , N.:
          <article-title>Multilevel process mining for nancial audits</article-title>
          .
          <source>IEEE Transactions on Services Computing</source>
          <volume>8</volume>
          (
          <issue>6</issue>
          ),
          <volume>820</volume>
          {832 (Nov
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>