<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Generating Arti cial Event Logs with Su cient Discriminatory Power to Compare Process Discovery Techniques</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Toon Jouck</string-name>
          <email>toon.jouck@uhasselt.be</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>t Depaire</string-name>
          <email>benoit.depaire@uhasselt.be</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hasselt University, Faculty of Business Economics Agoralaan Bldg D</institution>
          ,
          <addr-line>3590 Diepenbeek</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Research Foundation - Flanders (FWO) Egmontstraat 5</institution>
          ,
          <addr-line>1000 Brussels</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Past research revealed issues with arti cial event data used for comparative analysis of process mining algorithms. The aim of this research is to design, implement and validate a framework for producing arti cial event logs which should increase discriminatory power of arti cial event logs when evaluating process discovery techniques. { What model characteristics can we identify which in uence the generated data? { What is the impact of model language bias on the generated data? { Which non-model characteristics exist which in uence the generated data? { What is a proper methodology for generating arti cial data for comparative analysis? { Which tools exist for generating arti cial data and to what extent are they su cient?</p>
      </abstract>
      <kwd-group>
        <kwd>Arti cial Event Logs</kwd>
        <kwd>Event Log Simulation</kwd>
        <kwd>Performance Measurement of Business Processes</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Research Question</title>
      <p>Literature on the comparative analysis of process discovery techniques has
revealed some problems with arti cial data. The data lacked discriminatory power.
We argue that such problems arose due to the absence of a proper framework
to generate arti cial data. This leads to our main research question: how can we
generate arti cial event logs with su cient discriminatory power for a
comparative evaluation of process discovery algorithms? To provide an answer to this
question several other questions need to be answered:</p>
    </sec>
    <sec id="sec-2">
      <title>2 Background</title>
      <p>
        This work focusses on arti cial data used for the comparison of di erent process
discovery techniques, more speci cally the comparison of control- ow techniques.
In past research on process mining many researchers used arti cial data for the
development of and the veri cation of new algorithms (e.g. [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]).
      </p>
      <p>
        In a recent study De Weerdt et al. compare several process discovery
techniques on both arti cial and real data [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The arti cial data used in their
experiments was recovered from past research on the development of a process
discovery algorithm [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Remarkably, the performance of the algorithms did not
seem to be signi cantly di erent for the arti cial data, while real data revealed
signi cant performance di erences. These results indicate that the arti cial test
data used in past research have insu cient discriminatory power.
      </p>
      <p>
        A lot of process discovery techniques have been developed in the last decade.
Since the rst algorithms, process discovery has matured remarkably. However,
it's still not clear which algorithm will perform best in a certain situation. This
has led to an increasing importance of the research on comparing di erent
process discovery techniques [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3 Signi cance</title>
      <p>The comparison of process discovery techniques can be based on both arti cial
and real data. Real data, however, are at a disadvantage when performing such
a comparative analysis.</p>
      <p>Two disadvantages stem from the nature of algorithm comparison and
evaluation. Typically research is performed on a sample of event logs, but conclusions
are preferably generalizable to other event logs. To achieve reliable conclusions,
statistics require su cient observations and samples which are representative for
the considered population. Real data, however, have limited availability and are
typically convenience samples, rather than random samples.</p>
      <p>Another disadvantage of real data is concerned with identifying causal
relationships between process or event log characteristics on the one hand and
algorithms performance on the other. This kind of research requires
experimental data and not observational data (real data).</p>
      <p>
        In contrast, these disadvantages are not present when using arti cial data in
comparative analysis of process discovery techniques if a proper methodology is
used to generate the arti cial data. Such a methodology should focus on creating
arti cial data with su cient discriminatory power to overcome the problems
encountered in past research (e.g. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]). The main contribution of this research
will be drawing up and implementing a general methodology for the generation
of arti cial event logs with su cient discriminatory power in order to evaluate
process mining algorithms.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4 Research design and methods</title>
      <p>Firstly, a structured literature review is performed to get insight into generating
arti cial data and algorithm comparison. The primary sources used to perform
this review are: literature in the domain of process mining and literature from
other domains on (generating) synthetic data.</p>
      <p>Secondly, the general methodology is built and implemented in a tool to
support this new methodology.</p>
      <p>Finally the implemented methodology is tested and validated by repeating
experiments done in past research on comparing process discovery techniques.</p>
      <p>One important limitation of this methodology will be its scope which is
limited to generating arti cial data for analysing control- ow discovery techniques.
Also the reader should be aware that such a general methodology for arti cial
data will not replace the need for real (test) data. Real logs continue to be
necessary for making arti cial event logs more realistic and as a nal review for
process discovery techniques.</p>
    </sec>
    <sec id="sec-5">
      <title>5 Research stage</title>
      <sec id="sec-5-1">
        <title>5.1 A Preliminary Framework</title>
        <p>
          The literature review of articles on the evaluation of process discovery techniques
based on arti cial data (i.a. [
          <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
          ]) reveals that there are only some guidelines
or recurring elements for generating arti cial logs. However, a sound and general
methodology is missing, which decreases the relevance of arti cial logs.
        </p>
        <p>To address this issue a preliminary framework is distilled from the literature
review which focusses on the crucial aspect of randomization. The methodology
can be divided into two stages: the generation of an arti cial process model and
the generation of event logs from this model. Both stages allow the researcher to
de ne the characteristics of the population and produce a representative sample
(see table 1).</p>
        <p>The rst step is to de ne a population of process models, from which arti
cial models are sampled randomly and automatically. In past research this
crucial step in generating arti cial event logs was never made explicit in a general
method or guideline. Mostly processes were drawn manually in an ad-hoc
manner without explicitly de ning the population they were drawn from. However,
it is important that the researcher has insight into the process model population
and can in uence the properties of that population. Therefore, ranges for the list
of controllable properties (see step 1 in table 1) must be set to de ne the
population. Next, values within these ranges are selected randomly and automatically
to de ne a single process model.</p>
        <p>The second step concerns the generation of event logs for each process model
de ned in the previous stage. Again, the researcher must set ranges for several
event log properties, from which exact values are sampled randomly to generate
event logs. The parameters which can be set are shown in step 2 in table 1.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2 Tools for Generating Arti cial Event Logs</title>
        <p>Di erent tools already exist which can help to automatically generate arti cial
event logs. We evaluated two tools considered most appropriate to support the
1. Model Generation
2. Log Generation</p>
        <p>Number of activity types
Choice structural patterns
Choice nested structural patterns
Number of generated process instances
Required completeness
Noise</p>
        <p>
          Imbalance of execution properties
preliminary framework: the PLG tool [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and the BeehiveZ tool [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. At rst sight
both tools seem appropriate because both support the two stages of the proposed
methodology. However, a more detailed evaluation revealed that both tools do
not completely support the proposed methodology and several limitations exist.
The results of the evaluation, summarized in table 2, show that the PLG tool
supports the preliminary framework the best.
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3 A First Step Towards Validation</title>
        <p>Although the framework as presented in table 1 is still preliminary, it was used
in a rst case study to assess if it was a step towards arti cial data with more
discriminatory power.</p>
        <p>
          For this case study we repeat part of the experiment of De Weerdt et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]
in which they evaluated process discovery techniques on both arti cial and real
event logs. Remarkably, the performance of the algorithms did not seem to be
signi cantly di erent for the arti cial data, while real event logs revealed
significant performance di erences.
        </p>
        <p>
          We hypothesize that our methodology can produce arti cial data with more
discriminatory power. Therefore we repeat part of the experiment of De Weerdt
et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] on arti cial data generated with the proposed methodology to see if
our results are closer to the results on real data in De Weerdt et al., than their
own results on arti cial data. If that is true, the case study will provide a rst
support to our hypothesis.
        </p>
        <p>
          We applied the preliminary methodology to generate 35 arti cial event logs
out of two random populations using the PLG tool (with all its limitations). Then
four process discovery algorithms were evaluated in two conformance dimensions,
tness and precision, using the method described in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>
          The results in the tness dimension show that the performance of the tested
algorithms re ect better the results for tness on real data in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], and thus
supports the earlier stated hypothesis. However, the performance di erences
with respect to tness in our experiments were of a di erent order of magnitude
than the performance di erences based on real data found by De Weerdt et
al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Moreover, the results from our experiments don't show any signi cant
di erences in terms of precision in contrast to the results in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] based on real
data.
        </p>
        <p>From these ndings can be concluded that the preliminary methodology is
only a rst step in the direction of increasing the discriminatory power of arti cial
event logs.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.,
          <string-name>
            <surname>Weijters</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maruster</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Work ow mining: discovering process models from event logs. Knowledge and Data Engineering</article-title>
          , IEEE Transactions on
          <volume>16</volume>
          (
          <issue>9</issue>
          ) (
          <year>September 2004</year>
          )
          <volume>1128</volume>
          {
          <fpage>1142</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>de Medeiros</surname>
            ,
            <given-names>A.K.A.</given-names>
          </string-name>
          :
          <article-title>Genetic process mining</article-title>
          .
          <source>PhD thesis</source>
          , Technische Universiteit Eindhoven (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>De Weerdt</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Backer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanthienen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baesens</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>A multi-dimensional quality assessment of state-of-the-art process discovery algorithms using real-life event logs</article-title>
          .
          <source>Information Systems</source>
          <volume>37</volume>
          (
          <issue>7</issue>
          ) (
          <year>November 2012</year>
          )
          <volume>654</volume>
          {
          <fpage>676</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Rozinat</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>de Medeiros</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gnther</surname>
            ,
            <given-names>C.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weijters</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.M.:
          <article-title>Towards an evaluation framework for process mining algorithms</article-title>
          .
          <source>In: BPM Center Report</source>
          , BPMcenter.org (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. vanden Broucke,
          <string-name>
            <given-names>S.K.</given-names>
            ,
            <surname>Delvaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Freitas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Rogova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Vanthienen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Baesens</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Uncovering the relationship between event log characteristics and process discovery techniques</article-title>
          .
          <source>In: Business Process Management Workshops</source>
          , Springer (
          <year>2014</year>
          )
          <volume>41</volume>
          {
          <fpage>53</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Burattin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sperduti</surname>
            , Alessandro,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>PLG: a process log generator</article-title>
          .
          <source>Technical report</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wen</surname>
          </string-name>
          , L.:
          <article-title>E ciently querying business process models with BeehiveZ</article-title>
          .
          <source>In: BPM (Demos)</source>
          , Clermont-Ferrand, France (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>