<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Stochastic Process Mining, Stochastic Process Discovery, Stochastic Conformance Checking</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>RWTH Aachen University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>The University of Melbourne</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Process mining extracts event logs from information systems to derive insights into organizational business processes. Stochastic process mining specifically emphasizes techniques that integrate the frequency and probability of various process behaviors, allowing organizations to comprehend their business processes by diferentiating between routine and exceptional occurrences. The ambition of my Ph.D. project is to improve state-of-the-art stochastic process discovery techniques and provide novel measures applicable to stochastic conformance checking.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Research Background and Problems</title>
      <p>This section serves to motivate and formalize the research questions for my Ph.D. project.</p>
      <sec id="sec-2-1">
        <title>2.1. Research Question One</title>
        <p>
          Two types of stochastic process modeling formalism have been proposed for stochastic process modeling:
action graph-based models and Petri net-based models. The first type is the stochastic action graph
introduced by Alkhammash et al. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], which is similar to the directed graphs adopted by practitioners.
The nodes and arcs in the model are annotated with numbers that reflect the frequencies of the actions
        </p>
        <p>CEUR</p>
        <p>ceur-ws.org</p>
        <p>EvEevnetnt
lolgog</p>
        <p>StSetpep1 1 CoCnotnrtorlo-fll-oflwow
mmodoedlel</p>
        <p>StSetpep22</p>
        <p>StSotcohcahsatsitcic
mmododelel</p>
        <p>EEvveenntt SStetepp11
lologg</p>
        <p>
          SStotocchhaasstitcic
mmooddeel l
and “can follow” dependencies inferred from the event log. The second type consists of Petri nets
with stochastic extensions. For instance, Leemans et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] introduced the formalization of generalized
stochastic labeled Petri nets, including silent transitions and annotated immediate transitions with
weights. The probability of firing an immediate transition is determined by the relative weights of all
enabled immediate transitions.
        </p>
        <p>Although these formalisms have been studied thoroughly, other process modeling languages, such as
Business Process Model and Notation (BPMN) and causal nets, can also be extended with the stochastic
perspective. In practice, BPMN is a set of diagramming conventions used to describe business processes.
Causal nets are a declarative process modeling formalism that relies on a small number of modeling
constructs, but is expressive. Consequently, the first research question (RQ1) is formalized as: What is
the appropriate modeling formalism to model the stochastic perspective of the process?
Proposed solution Business Process Model and Notation (BPMN) models are widely used by
practitioners for decision making. Causal nets (C-nets) are used in multiple process discovery techniques.
Many Markov models, such as Markov chains, Markov decision processes, or semi-Markov decision
processes, have diferent characteristics. Similarly, transition systems have also been enhanced with
probabilistic information and have been used for prediction. We propose a comparison study to justify
the benefits of the selected stochastic process models compared to other stochastic process modeling
languages.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Research Question Two</title>
        <p>Although conventional process discovery methods excel in many areas, they have the limitation of
ignoring the frequencies of the traces recorded in the event logs in the constructed models. Stochastic
process discovery techniques mine models that pair traces with indications on how likely one can expect
to see them in future executions of the process. Most techniques achieve this through an indirect
method, which involves a conventional process discovery technique to construct the process flow and
then add probability weights based on how often each trace appears, as illustrated in Fig. 1.</p>
        <p>
          The technique in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] is the first two-stage discovery framework to discover generalized stochastic
Petri nets with timed transitions, which allow performance analysis. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] introduced several weight
estimators based on statistics computed on log and model. In 2024, two other algorithms were proposed
to perform the discovery of stochastic processes with optimal stochastic quality guarantee [
          <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
          ].
        </p>
        <p>At the start of my Ph.D., the second research question (RQ2) was formalized as: Given an event log and
a process model that describes the control flow of the process observed in the event log, how do we construct
a stochastic process model that maintains the same control flow while being capable of reproducing the
probability of the observed process?
Proposed solution We propose to define two-stage stochastic process discovery as finding a model with
an optimal stochastic conformance checking measure over a given representation bias. Our strategy is
to turn the given control flow model into a stochastic process model that assigns a weight parameter to
every transition. Then, stochastic discovery is posed as an optimization problem, where values for the
weights must be found so that a stochastic conformance measure is maximized.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Research Question Three</title>
        <p>
          The one-stage techniques operate without relying on an initial control-flow model, but calculate
controllfow and stochastic aspects simultaneously, as illustrated in Fig. 2. Toothpaste Miner [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] is the first
stochastic model record
        </p>
        <p>IT system
stochastic model record
discover</p>
        <p>event log
discover</p>
        <p>
          event log
single-stage technique, which applies a set of reduction and abstraction rules to generate a stochastic
labeled Petri net (SLPN). Another technique GASPD is based on grammatical inference [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], which
discovers a family of direct action graphs from an input event log.
        </p>
        <p>However, these two one-stage approaches do not guarantee the stochastic quality of the constructed
models. Thus, our third research question (RQ3) is: Given an event log, how to directly construct a
stochastic process model of manageable size while capable of reproducing the probability of the observed
process in the event log?
Proposed solution We plan to address inherently competing objectives: Reducing the complexity of
the model while maintaining its stochastic quality. Thus, the one-stage discovery is addressed through
a trade-of between model complexity and accuracy.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Research Question Four</title>
        <p>Another core problem studied in process mining is conformance checking, which quantifies how much
a process model agrees with an event log. The technique that ignores the stochastic perspective
of processes can be misleading. For instance, non-stochastic-aware conformance checking cannot
distinguish the discrepancy between event log [⟨, ⟩ 50, ⟨, ⟩ 50] and a stochastic process model that
describes the stochastic language [⟨, ⟩ 0.9, ⟨, ⟩ 0.1]. However, the model emphasizes that step  should
occur after  more frequently than  after  , while the log suggests that both orders are equally likely.</p>
        <p>Beyond regular Log-to-Model (L2M) scenarios, Model-to-Model (M2M) and Log-to-Log (L2L)
conformance checking also benefit from a stochastic perspective. In many real-life scenarios, processes
are influenced by internal or external requirements and change over time. Consequently, a stochastic
process model designed to describe the behavior of the system can become outdated as time progresses.
One can detect and quantify changes in stochastic behavior by comparing the latest discovered model
with the original model using M2M stochastic conformance. Similarly, event logs that cover long
periods or merge data from multiple organizations may contain diferent versions of the process behavior.
Conclusions drawn from such logs may be misleading or biased when addressing specific regional or
temporal issues. We illustrate these scenarios of stochastic conformance checking in Fig. 3.</p>
        <p>In light of this, we establish our fourth research question (RQ4) as: How to quantify stochastic
conformance checking for log-to-log, log-to-model, and model-to-model conformance scenarios?
Proposed solution In essence, an event log can be considered as a finite sample drawn from a probability
distribution over traces, while stochastic process models represent probability distributions over traces.
We propose adapting established statistical distances between probability distributions for stochastic
conformance checking. Specifically, we review interesting features of statistical distances and discuss
their use in the context of L2L, L2M, and M2M conformance checking. To expand the applicability of
the distances, we propose the necessary adaptations.</p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5. Research Question Five</title>
        <p>
          Although existing techniques [
          <xref ref-type="bibr" rid="ref10 ref11 ref12 ref9">9, 10, 11, 12</xref>
          ] quantify stochastic conformance by computing a numerical
value between a stochastic model and an event log, they are not applicable if an aggregated event log is
not available. Moreover, a single trace is not explicitly matched to the model.
        </p>
        <p>Conformance checking techniques, such as alignments, identify a path allowed by the model with as
few deviations as possible from an observed trace. However, when considering a stochastic perspective,
if the selected path is unlikely according to the stochastic process model, it may not be the most likely
explanation of the path through the model. As a consequence, further diagnostics are based on process
behavior that is less relevant to be followed by design. For instance, a path with a probability of 10%
according to the model and an edit distance of 3 to the trace may be a better match than a path with a
probability of 0.1% and an edit distance of 2. The scenario highlights two possibly competing objectives
when matching the trace to the stochastic model of the process: the probability of the selected path
allowed by the model and its edit distance to the trace.</p>
        <p>Therefore, we formalize our fith research question (RQ5) as: Given an observed trace, how to match it
to a stochastic process model by identifying a likely model path with a low edit distance to the trace?
Proposed solution The trade-of between the probability of the selected path allowed by the model
and its edit distance to the trace, which we illustrate in Fig. 4. The Pareto front consists of model paths
and indicates that any reduction in edit distance would necessarily decrease the model path probability,
and any increase in model path probability would necessarily increase its edit distance to the trace. We
propose a stochastic alignment technique that matches a single trace to a stochastic process model and
produces an alignment that balances the importance of the normative behavior (the probability of the
model path) and the alignment cost (the edit distance between the model path and trace). Business
analysts can explicitly weigh the trade-of with a user-defined parameter.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results and Road map</title>
      <sec id="sec-3-1">
        <title>We first outline the progress to date in this Ph.D. project:</title>
        <p>• Literature review to identify research gaps for stochastic process mining.
• Study 1: Define and implement the two-stage discovery problem with an optimality guarantee.</p>
        <p>
          The results of this study were published in CAiSE 2024 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
• Study 2: Identify statistical distances for the quantification of stochastic conformance. The results
of this study are currently under review.
• Study 3: Introduce stochastic alignments that account for alignment cost and the probability of
the model path. The results of the study are to be presented at the BPM 2025.
        </p>
        <p>To complete the Ph.D. project and thesis, my future research plan is as follows:
• Study 4: Identify alternative process modeling formalisms other than SLPNs to model the
stochastic perspective of the process.
• Study 5: Design a one-stage stochastic process discovery technique, and conduct an extensive
evaluation.</p>
        <p>Studies 4 and 5 are based on the existing work achieved in studies 1, 2, and 3. The quality of the
discovered stochastic process models can be evaluated with the stochastic conformance checking
measures discussed in the study.</p>
        <p>However, two challenges may prevent the project from achieving our target. First, the stochastic
discovery that guarantees stochastic optimality requires solving a nonconvex optimization problem.
This poses a computational challenge as there are multiple locally optimal points, and the computed
result may not be globally optimal. Second, the model-to-model stochastic conformance can be hard to
measure because of potentially infinite process behavior. The conformance between two stochastic
models can be computed by sampling, however, this does not guarantee an exact result.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Outlook</title>
      <p>This Ph.D. project will contribute to the business process management (BPM) community by developing
novel techniques for stochastic process mining. We treat the probability of process behavior as a
ifrst-class citizen due to its close link to simulation, prediction, and recommendation.</p>
      <p>We aim to consider the appropriate representation bias for stochastic process modeling. Then, we
plan to explore two types of stochastic discovery algorithms, i.e., a one-stage approach that directly
constructs a stochastic model from the input event log and a two-stage approach that indirectly
constructs a stochastic model using the event log. Furthermore, two types of stochastic conformance
checking are investigated, one is the statistical distance-based techniques that measure numerically;
the other is an alignment technique that returns an explicit artifact to match a trace with the given
stochastic process model.</p>
      <p>Simultaneously, to bridge the theory-practice divide, we will create user stories that tie all our research
questions together and demonstrate the practical application of our findings in real-world scenarios.
The goal is to illustrate how practitioners can leverage our research to address real-life challenges. For
example, we will showcase how practitioners can perform stochastic alignments to explain deviations
in an observed trace using the discovered stochastic process models.</p>
    </sec>
    <sec id="sec-5">
      <title>Declaration on Generative AI</title>
      <sec id="sec-5-1">
        <title>The author(s) have not employed any Generative AI tools.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. J. J.</given-names>
            <surname>Leemans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polyvyanyy</surname>
          </string-name>
          ,
          <article-title>Stochastic-aware precision and recall measures for conformance checking in process mining</article-title>
          ,
          <source>Inf. Syst</source>
          .
          <volume>115</volume>
          (
          <year>2023</year>
          )
          <fpage>102197</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Alkhammash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polyvyanyy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mofat</surname>
          </string-name>
          ,
          <article-title>Stochastic directly-follows process discovery using grammatical inference</article-title>
          ,
          <source>in: CAiSE</source>
          , volume
          <volume>14663</volume>
          <source>of LNCS</source>
          , Springer,
          <year>2024</year>
          , pp.
          <fpage>87</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S. J. J.</given-names>
            <surname>Leemans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Syring</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
          </string-name>
          ,
          <article-title>Earth movers' stochastic conformance checking</article-title>
          ,
          <source>in: BPM Forum</source>
          , volume
          <volume>360</volume>
          <source>of LNBIP</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>127</fpage>
          -
          <lpage>143</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rogge-Solti</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
          </string-name>
          , M. Weske,
          <article-title>Discovering stochastic petri nets with arbitrary delay distributions from event logs</article-title>
          ,
          <source>in: BPM Workshops</source>
          , volume
          <volume>171</volume>
          <source>of LNBIP</source>
          , Springer,
          <year>2013</year>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>27</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Burke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J. J.</given-names>
            <surname>Leemans</surname>
          </string-name>
          , M. T. Wynn,
          <article-title>Stochastic process discovery by weight estimation</article-title>
          ,
          <source>in: ICPM Workshops</source>
          , volume
          <volume>406</volume>
          <source>of LNBIP</source>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>260</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brockhof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Uysal</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
          </string-name>
          ,
          <article-title>Wasserstein weight estimation for stochastic petri nets</article-title>
          , in: ICPM, IEEE,
          <year>2024</year>
          , pp.
          <fpage>81</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Horváth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ballarini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. L.</given-names>
            <surname>Gall</surname>
          </string-name>
          ,
          <article-title>A framework for optimisation based stochastic process discovery</article-title>
          ,
          <source>in: QEST+FORMATS</source>
          , volume
          <volume>14996</volume>
          <source>of LNCS</source>
          , Springer,
          <year>2024</year>
          , pp.
          <fpage>34</fpage>
          -
          <lpage>51</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Burke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J. J.</given-names>
            <surname>Leemans</surname>
          </string-name>
          , M. T. Wynn,
          <article-title>Discovering stochastic process models by reduction and abstraction</article-title>
          , in: Petri Nets, volume
          <volume>12734</volume>
          <source>of LNCS</source>
          , Springer,
          <year>2021</year>
          , pp.
          <fpage>312</fpage>
          -
          <lpage>336</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Alkhammash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polyvyanyy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mofat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>García-Bañuelos</surname>
          </string-name>
          ,
          <article-title>Entropic relevance: A mechanism for measuring stochastic process models discovered from event data</article-title>
          ,
          <source>Inf. Syst</source>
          .
          <volume>107</volume>
          (
          <year>2022</year>
          )
          <fpage>101922</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S. J. J.</given-names>
            <surname>Leemans</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
            , T. Brockhof,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Polyvyanyy</surname>
          </string-name>
          ,
          <article-title>Stochastic process mining: Earth movers' stochastic conformance</article-title>
          ,
          <source>Inf. Syst</source>
          .
          <volume>102</volume>
          (
          <year>2021</year>
          )
          <fpage>101724</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J. J.</given-names>
            <surname>Leemans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polyvyanyy</surname>
          </string-name>
          ,
          <article-title>The jensen-shannon distance for stochastic conformance checking</article-title>
          ,
          <source>in: ICPM Workshops</source>
          , volume
          <volume>533</volume>
          <source>of LNBIP</source>
          , Springer,
          <year>2024</year>
          , pp.
          <fpage>70</fpage>
          -
          <lpage>83</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E. G.</given-names>
            <surname>Rocha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J. J.</given-names>
            <surname>Leemans</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
          </string-name>
          ,
          <article-title>Stochastic conformance checking based on expected subtrace frequency</article-title>
          , in: ICPM, IEEE,
          <year>2024</year>
          , pp.
          <fpage>73</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S. J. J.</given-names>
            <surname>Leemans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polyvyanyy</surname>
          </string-name>
          ,
          <article-title>Stochastic process discovery: Can it be done optimally?</article-title>
          , in: CAiSE, volume
          <volume>14663</volume>
          <source>of LNCS</source>
          , Springer,
          <year>2024</year>
          , pp.
          <fpage>36</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>