<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Identifying Cohorts that Di er in their Behaviour: Tool Support</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sander J. J. Leemans</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shiva Shabaninejad</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kanika Goel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hassan Khosravi</string-name>
          <email>h.khosravig@uq.edu.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shazia Sadiq</string-name>
          <email>shazia@itee.uq.edu.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Moe T. Wynn</string-name>
          <email>m.wynng@qut.edu.au</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Queensland University of Technology</institution>
          ,
          <addr-line>Brisbane</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Queensland</institution>
          ,
          <addr-line>Brisbane</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>163</fpage>
      <lpage>167</lpage>
      <abstract>
        <p>Process mining is a specialised form of data analytics that aims to provide data-driven improvement recommendations, derived from event logs. These event logs contain information about the execution of real-world processes, which may be complex. Cohort identi cation recommends drill-down lters for process mining, based on di erences in process. In this paper, we describe its integration in three process mining tools: as a stand-alone ProM plug-in, as part of the visual Miner and (planned) as part of Course Insights.</p>
      </abstract>
      <kwd-group>
        <kwd>Process Mining</kwd>
        <kwd>Feature Selection</kwd>
        <kwd>Filter Recommendation</kwd>
        <kwd>Stochastic Comparative Process Mining</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Process mining, a specialised form of data analytics, provides techniques
using which analysts can extract insights from recorded process behaviour in
event logs. The insights are used to provide data-driven recommendations to
improve business operations. Many real-life processes are complex in nature,
and studying their process models is challenging [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Process mining techniques
to deal with this complexity include ltering, slicing and dicing, and process
cubes.
      </p>
      <p>
        Cohort identi cation aims to identify and recommend sub-sets of the traces
in the log (cohort s). These cohorts are de ned by attributes of traces (e.g. \claim
amount", \gender" or \country"), and cohort identi cation recommends the
attributes and values such that the traces that have the attribute and value
(cohort) di er as much as possible from the traces that do not have the attribute
or have a di erent value (anti-cohort) in terms of the process that is being
followed, which includes the order of steps taken for traces, and how often di erent
sequences appeared in the (anti-)cohort. The di erence between the cohort and
the anti-cohort is expressed as a distance measure [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], where 1 means that their
processes are completely di erent (i.e. no activity appears in both), and 0 means
that their processes are no more di erent than a random division of the combined
log.
      </p>
      <p>That is, cohort identi cation nds groups of cases in event logs that follow
a process that is di erent than the process of other groups. For instance,
consider the following log consisting of 1000 cases of customers purchasing online
and registering for an account, in which each trace is annotated with whether
the customer was a Silver or Gold customer, from an East or West branch:
[hregister; purchasei2S0E0, hregister; purchasei1S0W0 , hregister; purchasei5G0E ,
hregister; purchasei1G0W0 , hpurchase; registeri1S0E0, hpurchase; registeri5S0W ,
hpurchase; registeri2G0E0 , hpurchase; registeri1G0W0 ]. In this log, Gold customers
executed the trace variant hregister; purchasei 145500 = 13 times, while for the other
customers this is 1405000 = 0:45 times. Thus, the likelihood that Gold customers
rst register is lower than for other customers. Cohort identi cation would assess
this for all potential combinations of attributes and values, and provide a ranked
list based on a quanti cation of such di erences.</p>
      <p>
        Cohort identi cation has similar goals as other process comparison techniques
such as trace clustering [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], concept drift detection [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and event attribute
clustering [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], however provides better explainable results: the output is a list of
attribute-value pairs that denote sub-logs of interest. The details of our cohort
identi cation technique are described in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>In this paper, we describe the
integration of cohort identi cation into
three existing open source data
intelligence tools: as a plug-in of the ProM
framework (Section 2), as an
extension of the process mining tool visual
Miner (Section 3), and as an extension
of the learning analytics dashboard
Course Insights (Section 4). The rst
tool demonstrates the use of cohort
identi cation as a stand alone
technique, the second tool depicts the use
of cohort identi cation in conjunction
with other process mining techniques,
and the third tool illustrates the
embedding of cohort identi cation in a
learning analytics context. A screen- Fig. 1: Cohort Identi cation in ProM.
cast is available at3.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Stand-alone Plug-in of ProM</title>
      <p>
        The ProM framework [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is a state-of-art open-source process mining
framework aimed at practitioners and academics. Cohort identi cation has been
implemented as a plug-in of ProM enabling practitioners and academics to access the
technique. Upon starting with as input an event log, two parameters can be set:
the attribute that determines the activity being executed, and the maximum
number of trace attributes that is exhaustively considered. Upon completion,
3 https://vimeo.com/442323972
      </p>
      <p>
        rst a diversi ed set of cohorts is shown (with attribute, value, cohort size, and
distance between the cohort and anti-cohort). Figure 1 shows a screenshot.
Implementation. The implementation is exible, as it provides extension points
to (1) elicit attribute value ranges, (2) measure distance between events, traces
and logs, and (3) truncate cohorts based on size or other aspects. Furthermore,
the implementation is multithreaded and for pruning stores a Map entry, an int[]
and an AtomicBoolean for each attribute value range combination.
Maturity &amp; How to Access. Cohort identi cation is open source and is part
of the ProM 6.10 release; see4. The plug-in has been successfully applied in three
case studies [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
3
      </p>
      <p>visual Miner</p>
      <p>
        The visual Miner (vM) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is an existing process mining tool that enables
endusers to combine advanced academic process mining techniques in an
industrycapable and user-friendly package. The input of vM is an event log. First, vM
applies a process discovery technique to the log to obtain a process model.
Second, vM applies a conformance checking technique and visualises the di erences
between log and the discovered model. Third, it computes detailed frequency
and performance information, and visualises this on the model, amongst others
using animation. Fourth, it allows the user to drill down and focus on parts of
the event log that are of interest, by applying one of several lters. Settings to
any of these techniques and lters can be changed at any time, and vM will
update and redo the necessary steps automatically [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. A screenshot is shown in
Figure 2; for a complete overview of vM's features, please refer to5.
Cohort Identi cation (new). While the vM makes it easy to drill down into
parts of the log or process of particular interest, it was up to the user to manually
analyse the visualisation available to discover potentially interesting parts.
      </p>
      <sec id="sec-2-1">
        <title>4 http://promtools.org</title>
      </sec>
      <sec id="sec-2-2">
        <title>5 http://leemans.ch/publications/ivmProM6.10.pdf</title>
        <p>Selected Attributes
Program, Residential Status</p>
        <p>Minimum Coverage:
20%</p>
        <p>Recommendations
2
Program = Tourism
Program = Engineering and Residential Status = Domestic</p>
        <p>Cohort identi cation suggests lters on attributes on the traces in an event
log, such that applying the lter leads to the largest di erences in process.
Figure 2 shows a screenshot of the integration of cohort identi cation in vM: the
cohorts are computed automatically in the background, and the result is shown
to the user. The rst column shows the trace attribute of the cohort, the second
column the values of the attribute that are in the cohort, the third column the
number of traces in the cohort, while the last column shows the distance between
the cohort and the anti-cohort. Using a click (for the cohort) or shift+click (for
the anti-cohort) one can quickly lter down the event log to the corresponding
traces, in order to study the di erences in process in more detail. The identi ed
cohorts can be exported to an Excel document for further analysis. Embedding
cohort identi cation in vM makes it easy for end-users to conduct detailed
process mining analysis using one plug-in.</p>
        <p>
          Maturity &amp; How to Access. vM has been applied to many process mining
projects by industry partners and academics (see 6 for an overview). Cohort
identi cation has been added in April 2020 and has been successfully applied in three
case studies [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Visual Miner is open source, part of the ProM framework [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]
and can be downloaded from6.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Course Insights</title>
      <p>
        Course Insights is an instructor-facing learning analytics dashboard,
developed at the University of Queensland, that empowers course coordinators to
gain insights and act on student data to enhance student learning and
experience across the course life-cycle at scale. It collates student data from a variety
of learning systems and sources and displays it to instructors all in one
simple and easy to use interface. An essential element is its comparative analysis
functionality, which enables course coordinators to use lters to compare and
contrast di erent student groups based on their demographics, enrolment,
engagement and performance data. An observational study that analysed how the
lters were used by 71 sta members found that commonly only a small subset
of the features was used and lters were rarely applied on top of one another [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
To overcome this challenge, we are implementing cohort identi cation in Course
Insights to recommend insightful lters to instructors [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <sec id="sec-3-1">
        <title>6 http://visualminer.org</title>
        <p>Cohort Identi cation (planned). Figure 3 illustrates the proposed
presentation of lter recommendations to instructors, including the attributes and values,
coverage (fraction of students covered by the lter), and distance (insightfullness
of the lter).</p>
        <p>
          Maturity &amp; How to Access. A case study based on data from a course with
875 students, with high demographic and educational diversity has explored the
potential bene ts of applying cohort identi cation to Course Insights [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The
cohort identi cation is planned to be implemented in Course Insights, which
can be accessed via7. A remaining challenge is to make Course Insights fully
process aware: with cohort identi cation, users can drill down into sub-groups
with di erences in their process, however more support to study these di erences
is necessary.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5 Conclusion</title>
      <p>In this paper, we described how cohort identi cation, which recommends
trace-attribute-based lters to maximise the di erences between processes, is
implemented as a stand-alone ProM plug-in, is integrated in the visual Miner,
and is being integrated in Course Insights. The technique lters sub-logs of
traces (cohorts) de ned by trace attribute value ranges (features) to compare
behavioural di erences in cohorts or to drill down into a particular cohort. The
technique can be used to understand the di erences between two cohorts and
answer questions related to a particular cohort. Cohort identi cation can be
applied in reasonable time to event logs. In the future, we intend to focus
implementing process infrastructure in Course Insights and on automated comparison
techniques to compare the identi ed cohorts.</p>
      <sec id="sec-4-1">
        <title>7 https://analytics.itali.uq.edu.au/dev/insights</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          : Process Mining - Data Science in Action (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bolt</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.M.P., de Leoni, M.:
          <article-title>Finding process variants in event logs (short paper)</article-title>
          .
          <source>In: CoopIS</source>
          . vol.
          <volume>10573</volume>
          , pp.
          <volume>45</volume>
          {
          <issue>52</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. van Dongen,
          <string-name>
            <surname>B.F.</surname>
          </string-name>
          , et al.:
          <article-title>The ProM framework: A new era in process mining tool support</article-title>
          .
          <source>In: Petri Nets</source>
          . pp.
          <volume>444</volume>
          {
          <issue>454</issue>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Leemans</surname>
            ,
            <given-names>S.J.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Syring</surname>
            ,
            <given-names>A.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          :
          <article-title>Earth movers' stochastic conformance checking</article-title>
          .
          <source>In: BPM forum</source>
          . pp.
          <volume>127</volume>
          {
          <issue>143</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Leemans</surname>
            ,
            <given-names>S.J.J.</given-names>
          </string-name>
          , et al.:
          <article-title>Process and deviation exploration with Inductive visual Miner. In: BPM demos</article-title>
          . vol.
          <volume>1295</volume>
          , p.
          <fpage>46</fpage>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Leemans</surname>
            ,
            <given-names>S.J.J.</given-names>
          </string-name>
          , et al.:
          <article-title>Identifying cohorts: Recommending drill-downs based on di erences in behaviour for process mining</article-title>
          .
          <source>In: ER</source>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Maaradji</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosa</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ostovar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Detecting sudden and gradual drifts in business processes from execution traces</article-title>
          .
          <source>TKDE</source>
          <volume>29</volume>
          (
          <issue>10</issue>
          ),
          <volume>2140</volume>
          {
          <fpage>2154</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Shabaninejad</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>Automated insightful drill-down recommendations for learning analytics dashboards</article-title>
          .
          <source>In: LAK</source>
          . p.
          <volume>41</volume>
          {
          <issue>46</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Shabaninejad</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>Recommending insightful drill-downs based on learning processes for learning analytics dashboards</article-title>
          .
          <source>In: AIED</source>
          . pp.
          <volume>486</volume>
          {
          <issue>499</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Weerdt</surname>
            , J.D., vanden Broucke,
            <given-names>S.K.L.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanthienen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baesens</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Active trace clustering for improved process discovery</article-title>
          .
          <source>TKDE</source>
          <volume>25</volume>
          (
          <issue>12</issue>
          ),
          <volume>2708</volume>
          {
          <fpage>2720</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>