<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mining Behavioural Patterns from Event Data to Enable Context-Aware Root Cause Analysis (Extended Abstract)</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Greg Van Houdt Faculty of Business Economics UHasselt - Hasselt University Hasselt</institution>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>II. RESEARCH GAPS</title>
      <p>In each of the three aforementioned research goals, several
challenges can be identified. The first goal covers event log
abstraction or augmenting the event log to a higher granularity
level. To start, an evaluation of abstraction quality of existing
pattern detection techniques is required. Abstraction utilises
input patterns of events to be elevated into a single high-level
activity or system state on the one hand, but also an algorithm
to replace these patterns in the event log on the other hand.
In this stage, it is also vital to keep track of the quality of
this high-level event log in terms of fitness, precision and
complexity.</p>
      <p>
        The second goal tackles an investigation of context changes
in the masses of data. Event data typically originates from
process-aware information systems [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] storing a great deal
of information. However, capturing all data does not mean
it is all useable. In most cases, data integration is key to
perform the appropriate analysis for a particular problem due
to a large possible number of data sources. Only with a
correct integration of data can we link all business objects and
their corresponding triggered high-level activities. Important
elements to consider are anomaly detection [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and concept
drift [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Finally, goal three shifts focus to root cause analysis.
Despite the importance of feature selection in root cause analysis,
there is a lack of attention being attributed to the specification
of features [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]–[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. For example, trends in certain variables or
events should also be considered to be a potential feature. An
important question to be asked is to which extent this feature
selection could be automated and, of course, to which degree
the domain expert must be kept in the loop. In that regard,
there is a difference between feature selection [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and feature
engineering [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] which remains to be explored in this context.
      </p>
    </sec>
    <sec id="sec-2">
      <title>III. METHODOLOGY</title>
      <p>
        This PhD research project shall follow the principles of
design science research (DSR). DSR is centred around the
development and study of artefacts aiming to solve a problem
taking the problem context into account [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Additionally, we
attempt to draw methodological plans of action from fields in
which event data is also known. Examples are visual analytics
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and complex event processing [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>For goal 1, we start with an analysis of already developed
abstraction techniques on their ability to elevate event logs (in
the process mining context) to a higher level of granularity.</p>
      <p>The next stage would imply moving beyond process mining
and generalise to recurrent sequences. Goal 2 is about
detecting shifts in the context or behaviour of the system. As
Copyright © 2020 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
a starting point, we can, e.g., evaluate existing concept drift
techniques. Alternatively, we can train a generative model on
the low-level data on different time windows and check if these
models differ in structure significantly. Detected changes can
be represented as new events in the data. Finally, in goal 3, we
start again with an evaluation of existing techniques to discover
areas of improvement regarding their performance in finding
root causes. This should be done for both low-level data as
well as the abstracted data to see if there are differences in
performance. But not only the impact of abstraction should be
tested. Attention must be given to the impact of the context
change events as well.</p>
      <p>This research aims to be valuable not only in theory, but
also in practice. To that regard, we strive for a high degree
of applicability in industry. This implies that validation in
practice is key. To that end, we have identified a number of
partners to aid us with validation processes.</p>
    </sec>
    <sec id="sec-3">
      <title>IV. RELATED WORK</title>
      <p>
        At the time of writing, there is already a strong basis of
event abstraction techniques present. The work of van Zelst
et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] provides a taxonomy on these techniques based
on several properties, for example, the supervision strategy.
We have a profound research interest in the unsupervised
techniques. Most interesting to us are the local process
models [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], global trace segmentation [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], the RefMod-miner
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], and combination based behavioral pattern mining [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
One can then utilise the pattern-based abstraction approach
designed by Mannhardt et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] to obtain the high-level
event log.
      </p>
      <p>
        Having achieved a more simplified event log by removing
part of the clutter, it should be easier to link the higher-level
activities with each other in terms of cause-effect
relationships. Two domains are interesting to investigate further in
this regard, namely visual analytics and root cause analysis.
Visual Analytics is a vast domain in the sense that it is a
combination of several research areas, namely visualisation,
data mining and statistics. At the same time, one cannot forget
the importance of the human factor [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. It can be defined
as the science of analytical reasoning facilitated by interactive
visual interfaces [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which makes the way of processing data
and information transparent [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. A root cause is the most
fundamental reason for an undesirable condition or problem
which, if eliminated or corrected, would have prevented it
from existing or occurring. Therefore, the root cause is always
negative and usually defined in terms of specific or systematic
factors. Root cause analysis can be used to identify the
most apparent improvement opportunities by tagging current
obstacles to efficient operations or activities [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
          </string-name>
          , Process Mining: Data Science in Action. Berlin, Heidelberg: Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>I.</given-names>
            <surname>Lee</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          , “
          <article-title>The Internet of Things (IoT): Applications, investments, and challenges for enterprises</article-title>
          ,
          <source>” Business Horizons</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>I. Lee</surname>
          </string-name>
          , “
          <article-title>The Internet of Things for enterprises: An ecosystem, architecture, and IoT service business model</article-title>
          ,
          <source>” Internet of Things</source>
          , vol.
          <volume>7</volume>
          , p.
          <volume>100078</volume>
          , 9
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Teece</surname>
          </string-name>
          and G. Linden, “
          <article-title>Business models, value capture, and the digital enterprise</article-title>
          ,
          <source>” Journal of Organization Design</source>
          , vol.
          <volume>6</volume>
          , no.
          <issue>1</issue>
          , p.
          <fpage>8</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
          </string-name>
          , T. Weijters, and L. Maruster, “
          <article-title>Workflow mining: discovering process models from event logs</article-title>
          ,
          <source>” IEEE Transactions on Knowledge and Data Engineering</source>
          , vol.
          <volume>16</volume>
          , no.
          <issue>9</issue>
          , pp.
          <fpage>1128</fpage>
          -
          <lpage>1142</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumas</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. M.</surname>
          </string-name>
          <article-title>van der Aalst, and</article-title>
          <string-name>
            <surname>A. H. M.</surname>
          </string-name>
          ter Hofstede,
          <source>ProcessAware Information Systems: Bridging People and Software through Process Technology. Wiley Online Library</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V.</given-names>
            <surname>Chandola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          , “
          <article-title>Anomaly detection: A survey,” ACM computing surveys (CSUR)</article-title>
          , vol.
          <volume>41</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>58</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>G.</given-names>
            <surname>Widmer and M. Kubat</surname>
          </string-name>
          , “
          <article-title>Learning in the presence of concept drift and hidden contexts,” Machine learning</article-title>
          , vol.
          <volume>23</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>69</fpage>
          -
          <lpage>101</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Suriadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ouyang</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. M. Van Der Aalst</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Ter Hofstede</surname>
          </string-name>
          , “
          <article-title>Root cause analysis with enriched process logs</article-title>
          ,
          <source>” in Lecture Notes in Business Information Processing</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B. F.</given-names>
            <surname>Hompes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maaradji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. La</given-names>
            <surname>Rosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Buijs</surname>
          </string-name>
          , and
          <string-name>
            <surname>W. M. Van Der Aalst</surname>
          </string-name>
          , “
          <article-title>Discovering causal factors explaining business process performance variation</article-title>
          ,
          <source>” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Anand</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Sureka</surname>
          </string-name>
          , “Pariket:
          <article-title>Mining business process logs for root cause analysis of anomalous incidents</article-title>
          ,
          <source>” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          , K. Cheng,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Morstatter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. P.</given-names>
            <surname>Trevino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          , and H. Liu, “
          <article-title>Feature selection: A data perspective,” ACM Computing Surveys</article-title>
          , vol.
          <volume>50</volume>
          , no.
          <issue>6</issue>
          , p.
          <fpage>94</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Nargesian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Samulowitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Khurana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. B.</given-names>
            <surname>Khalil</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Turaga</surname>
          </string-name>
          , “
          <article-title>Learning feature engineering for classification,”</article-title>
          <source>in IJCAI International Joint Conference on Artificial Intelligence</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>P.</given-names>
            <surname>Johannesson</surname>
          </string-name>
          and
          <string-name>
            <surname>E. Perjons,</surname>
          </string-name>
          <article-title>An introduction to design science</article-title>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Thomas</surname>
          </string-name>
          and
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Cook</surname>
          </string-name>
          , “
          <article-title>Illuminating the path: The research and development agenda for visual analytics</article-title>
          ,
          <source>” IEEE Computer Society</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Luckham</surname>
          </string-name>
          , “
          <article-title>The power of events: An introduction to complex event processing in distributed enterprise systems</article-title>
          ,
          <source>” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>S. J. van Zelst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Mannhardt</surname>
          </string-name>
          , M. de Leoni,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Koschmider</surname>
          </string-name>
          , “
          <article-title>Event abstraction in process mining: literature review and taxonomy</article-title>
          ,” Granular Computing, May
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>N.</given-names>
            <surname>Tax</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sidorova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Haakma</surname>
          </string-name>
          , and W. M. van der Aalst, “
          <article-title>Mining local process models</article-title>
          ,
          <source>” Journal of Innovation in Digital Ecosystems</source>
          , vol.
          <volume>3</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>183</fpage>
          -
          <lpage>196</lpage>
          , Dec.
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>C. W.</given-names>
            <surname>Gu</surname>
          </string-name>
          <article-title>¨nther, A. Rozinat, and</article-title>
          <string-name>
            <surname>W. M. Van Der Aalst</surname>
          </string-name>
          , “
          <article-title>Activity mining by global trace segmentation</article-title>
          ,” in International Conference on Business Process Management. Springer,
          <year>2009</year>
          , pp.
          <fpage>128</fpage>
          -
          <lpage>139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J.-R.</given-names>
            <surname>Rehse</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Fettke</surname>
          </string-name>
          , “
          <article-title>Clustering Business Process Activities for Identifying Reference Model Components,” in Business Process Management Workshops</article-title>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Daniel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. Z.</given-names>
            <surname>Sheng</surname>
          </string-name>
          , and H. Motahari, Eds. Cham: Springer International Publishing,
          <year>2019</year>
          , vol.
          <volume>342</volume>
          , pp.
          <fpage>5</fpage>
          -
          <lpage>17</lpage>
          , series Title: Lecture Notes in Business Information Processing.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>M.</given-names>
            <surname>Acheli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Grigori</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Weidlich</surname>
          </string-name>
          , “
          <article-title>Efficient Discovery of Compact Maximal Behavioral Patterns from Event Logs,” in Advanced Information Systems Engineering</article-title>
          , P. Giorgini and
          <string-name>
            <given-names>B.</given-names>
            <surname>Weber</surname>
          </string-name>
          , Eds. Cham: Springer International Publishing,
          <year>2019</year>
          , vol.
          <volume>11483</volume>
          , pp.
          <fpage>579</fpage>
          -
          <lpage>594</lpage>
          , series Title: Lecture Notes in Computer Science.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>F.</given-names>
            <surname>Mannhardt</surname>
          </string-name>
          , M. de Leoni,
          <string-name>
            <given-names>H. A.</given-names>
            <surname>Reijers</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
            , and
            <given-names>P. J.</given-names>
          </string-name>
          <string-name>
            <surname>Toussaint</surname>
          </string-name>
          , “
          <string-name>
            <surname>From</surname>
          </string-name>
          Low-Level Events to Activities - A
          <string-name>
            <surname>Pattern-Based</surname>
            <given-names>Approach</given-names>
          </string-name>
          ,” in Business Process Management,
          <string-name>
            <given-names>M. La</given-names>
            <surname>Rosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Loos</surname>
          </string-name>
          , and
          <string-name>
            <given-names>O.</given-names>
            <surname>Pastor</surname>
          </string-name>
          , Eds. Cham: Springer International Publishing,
          <year>2016</year>
          , vol.
          <volume>9850</volume>
          , pp.
          <fpage>125</fpage>
          -
          <lpage>141</lpage>
          , series Title: Lecture Notes in Computer Science.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Keim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Mansmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schneidewind</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          , “
          <article-title>Challenges in visual data analysis</article-title>
          ,
          <source>” in Proceedings of the International Conference on Information Visualisation</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>D.</given-names>
            <surname>Keim</surname>
          </string-name>
          , G. Andrienko,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Fekete</surname>
          </string-name>
          , C. Go¨rg, J. Kohlhammer, and G. Melanc¸on, “
          <article-title>Visual analytics: Definition, process</article-title>
          , and challenges,
          <source>” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Wilson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. D.</given-names>
            <surname>Dell</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. F.</given-names>
            <surname>Anderson</surname>
          </string-name>
          , “
          <article-title>Root Cause Analysis: A Tool for Total Quality Management</article-title>
          ,
          <source>” Journal For Healthcare Quality</source>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>