<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>PM4Py Web Services: Easy Development, Integration and Deployment of Process Mining Features in any Application Stack</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandro Berti</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sebastiaan J. van Zelst</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wil van der Aalst</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fraunhofer Gesellschaft, Institute for Applied Information Technology (FIT)</institution>
          ,
          <addr-line>Sankt Augustin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Process and Data Science Chair, Lehrstuhl fur Informatik 9 52074 Aachen, RWTH Aachen University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In recent years, process mining emerged as a set of techniques to analyze process data, supported by di erent open-source and commercial solutions. Process mining tools aim to discover process models from the data, perform conformance checking, predict the future behavior of the process and/or provide other analyses that enhance the overall process knowledge. Additionally, commercial vendors provide integration with external software solutions, facilitating the external use of their process mining algorithms. This integration is usually established by means of a set of web services that are accessible from an external software stack. In open-source process mining stacks, only a few solutions provide a corresponding web service. However, extensive documentation is often missing and/or tight integration with the front-end of the tool hampers the integration of the services with other software. Therefore, in this paper, a new open-source Python process mining service stack, PM4Py-WS, is presented. The proposed software supports easy integration with any software stack, provides an extensive documentation of the API and a clear separation between the business logic, (graphical) interface and the services. The aim is to increase the integration of process mining techniques in business intelligence tools.</p>
      </abstract>
      <kwd-group>
        <kwd>Process Mining PM4Py Web Services Process Discovery Case Management Seamless Integration</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Process mining [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a growing branch of data science which, starting from data
stored and processed in information systems, aims to infer information about
the underlying processes, i.e. as captured in the data. Several techniques, e.g.,
process discovery (automated discovery of a process model from the event data),
conformance checking (comparison between an event log and a process model),
prediction (given the current state of a process, predict the remaining time or
the value of some attribute in a future event), etc., have been developed.
      </p>
      <p>
        Process mining is supported by several open-source (ProM [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], RapidProM
[
        <xref ref-type="bibr" rid="ref2 ref9">9,2</xref>
        ], Apromore [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], bupaR [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], PM4Py [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], PMLAB [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) and commercial (Disco,
Celonis, ProcessGold, QPR ProcessAnalyzer, etc.) software. Apart from
RapidProM (an extension of the data science framework RapidMiner) and Apromore,
the majority of the open-source projects provide a standalone tool that only
allows to import an event log and to perform process mining analyses on it.
PM4Py and bupaR provide a set of process mining features as a library, and
this provides integration with the corresponding Python and R ecosystem. At
the same time, some commercial tools, e.g., Celonis and ProcessGold, as well as
the Apromore open-source tool, o er a web-based interface supported by web
services. This leads to some advantages:
{ The possibility to access the information related to the process everywhere,
from di erent devices.
{ Identity and access management (typically unsupported by standalone tools).
{ The possibility for multiple users to collaborate in the same workspace.
A systematic introduction of web services in the process mining eld is o ered
by [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Process mining analyses o ered through web services ideally permit an
easy integration with other software solutions. In the case of Apromore, the
business logic is o ered to the web application mainly through servlets-based
web services. This o ers the possibility for external tools to use the algorithms
integrated in Apromore by querying its web services. However, due to the high
customization on the client-side, required to provide a higher number of features
as application, Apromore does not allow easy embedding of the visual elements
in external applications.
      </p>
      <p>
        In this paper, the PM4Py Web Services (PM4Py-WS), that are built on top
of the recent process mining library PM4Py [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], are presented. The high-level
goals of the web services are (1) to provide an easy entrypoint for the
integration of process mining techniques in external (business intelligence) tools, and,
(2) to provide an extensible platform built on top of modern software
engineering principles and guidelines (testing, documentation, clear separation between
stacks, separation from the front-end). A prototypal process mining web
application, supported by the services, is provided along with the services, in order to
demonstrate that the services work as intended, and to provide some evidence
that PM4Py-WS is easily integrated in any external application stack.
      </p>
      <p>While in Python other ways to build non-trivial data visualization web
interfaces are available, for example Dash by Plot.ly3 that could o er an even simpler
prototyping of process mining visuals based on PM4Py, they do not o er the
same possibility of integration with external applications as exposing a set of
web services, since the interface is more tightly coupled with the backend part.</p>
      <p>The remainder of this paper is structured as follows. In Section 2, the
architecture of the web services is explained. Section 3 provides information regarding
the repository hosting the tool, the maturity of the tool and a reference to the
video demo. Section 4 concludes this paper.</p>
    </sec>
    <sec id="sec-2">
      <title>3 https://dash.plot.ly/introduction</title>
      <sec id="sec-2-1">
        <title>Architecture of the Web Services</title>
        <p>PM4Py-WS is written in Python, i.e., a programming language popular among
data scientists and developers for its simplicity and the vast set of high-performing
data science libraries. The web-services are exposed as asynchronous REST
GET/POST services using the Flask framework4, which supports:
{ Possibility to deploy the services using HTTP/HTTPS.
{ Possibility to use an enterprise-grade server (e.g. IIS, UWSGI).
{ Possibility to manage the Cross-Origin Resource Sharing (CORS5).</p>
        <p>PM4Py Web Services are supported by the algorithms available in the PM4Py
process mining library. The exposed services accept a session identi er (used to
identify the user, and verify the general permissions of the user), and a process
ID (used to identify the process, and verify the permissions of the user on the
given log). Moreover, each service has access to a singleton object hosting all the
software components needed for authentication, session management, user-log
visibility and permissions, log management, and exception handling. Each
different component is provided as a factory method, that means several di erent
implementations are possible. Currently, the following components are provided:
{ Session manager: Responsible for user authentication and veri cation of the
validity of a session. Two di erent session managers are available:
Basic session manager : Supported by a relational database with
user/password information and a separate logs database table.</p>
        <p>Keycloak IAM Session manager 6: Users and sessions are veri ed through
the Keycloak identity and user access control management solution, that
is the most popular enterprise solution in the eld. Thanks to Keycloak,
several applications can share the same credentials and sessions. The
conguration of user/password data, and the settings related to the session
duration, are done in Keycloak.
{ Log management: responsible to manage individual event logs, the
visibility/permissions of the users w.r.t. the logs, and the analysis/ ltering
operations on the log. Each process is managed by a handler that controls the
following operations on the log:</p>
        <p>Loading: depending on the log handler, the log is either persistently
loaded in memory, or loaded in memory when required. In the current
version, two in-memory handlers are provided: a XES handler (that loads
an event log in the XES format, and uses the PM4Py EventLog
structure to store events and cases), and a CSV handler (that loads an event
log in the CSV/Parquet7 format, and uses Pandas dataframes8). These
handlers both load event logs stored as les, however when the PM4Py
library will be more mature, the le dependency might be gradually
dropped.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 http://flask.pocoo.org/</title>
      <p>5 https://de.wikipedia.org/wiki/Cross-Origin_Resource_Sharing
6 https://www.keycloak.org/
7 A popular columnar storage format, see https://parquet.apache.org/
8 A popular data processing framework in Python, see https://pandas.pydata.org/</p>
      <p>Filtering: The ltering algorithms implemented in the PM4Py library are
applied to the log when the ltering services (add lter, remove lter)
are called.</p>
      <p>Analysis: the algorithms implemented in the PM4Py library are applied
to the log, in order to get the process schema, get the social network,
perform conformance checking, etc.
{ Exception handler: Triggered when debug warnings/error messages need to
be logged. The default setting is using the logging utility in Python.
3</p>
      <p>Repository of the Tool, Video Demo and Maturity
PM4Py-WS, along with the prototypal interface in AngularJS, is available at
https://github.com/pm4py/pm4py-ws. The web services can be installed
following the instructions contained in the INSTALL.txt le provided in the
repository. A Docker image9 is also made available with the name javert899/pm4pyws
and could be run through the command docker run -d -p 5000:5000 javert899
/pm4pyws. The docker image uses port 5000, i.e., the web services are exposed
at the URL http://localhost:5000 and the prototypal Angular web interface
is made available at the URL http://localhost:5000/index.html.</p>
      <p>The documentation of the web services is available at the site http://pm4py.
pads.rwth-aachen.de/pm4py-web-services/. A demo of the web services is
made available at http://212.237.8.106:5000/ (for example, the
loginService and the getEndActivities GET services can be tested according to the
documentation). A demo of the prototypal Angular web interface (represented
in Figure 1) is made available at http://212.237.8.106:5000/index.html. A
video demo, that shows how the web services can be easily queried through a
browser, and shows the prototypal Angular web interface in action, is available
at http://pm4py.pads.rwth-aachen.de/pm4py-ws-demo-video/.
9 See https://en.wikipedia.org/wiki/Docker_(software) for an introduction to
Docker</p>
      <p>The web services have just been released and no real-life use case is available
yet. The web interface is in a prototypal status.</p>
      <sec id="sec-3-1">
        <title>Conclusion</title>
        <p>In this paper, we presented PM4Py-WS, a stack of web services developed on
top of the existing process mining python framework PM4Py. PM4Py-WS is
open-source and extendible, and allows to integrate process mining technology
in any external software stack.</p>
        <p>When an algorithm is implemented in the PM4Py library, it is not
immediately available in the PM4Py web services: a service exposing the algorithm
should also be implemented. This does not hamper the scalability of the web
services (in terms of number of available features), but requires an extra e ort by
the developer, that should work on both the PM4Py library and the PM4Py-WS
projects.</p>
        <p>In terms of scalability on big amounts of data, the current PM4Py-WS
handlers are limited to a single core, and the performance su ers from the fact that
not all the CPU cores are used. Some future work will consider to implement
distributed handlers for logs, that would be able to manage bigger amount of
data. We aim to actively develop PM4Py-WS, allowing faster integration and
adoption of cutting-edge research into virtually any external tool.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.: Process Mining - Data Science in Action,
          <source>Second Edition</source>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.,
          <string-name>
            <surname>Bolt</surname>
            , A., van Zelst,
            <given-names>S.J.:</given-names>
          </string-name>
          <article-title>RapidProM: Mine Your Processes and Not Just Your Data</article-title>
          .
          <source>CoRR abs/1703</source>
          .03740 (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Berti</surname>
          </string-name>
          , A., van Zelst, S.J., van der Aalst, W.:
          <article-title>Process Mining for Python (PM4Py): Bridging the Gap Between Process-</article-title>
          and Data Science pp.
          <volume>13</volume>
          {
          <issue>16</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Carmona</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sole</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>PMLAB: an scripting environment for process mining</article-title>
          .
          <source>In: Proceedings of the BPM Demo Sessions 2014 Co-located with the 12th International Conference on Business Process Management (BPM</source>
          <year>2014</year>
          ), Eindhoven, The Netherlands,
          <year>September 10</year>
          ,
          <year>2014</year>
          . p.
          <volume>16</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. van Dongen,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>de Medeiros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.K.</given-names>
            ,
            <surname>Verbeek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Weijters</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.:
          <article-title>The ProM Framework: A New Era in Process Mining Tool Support</article-title>
          . In: International conference
          <article-title>on application and theory of petri nets</article-title>
          . pp.
          <volume>444</volume>
          {
          <fpage>454</fpage>
          . Springer (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Janssenswillen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Depaire</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swennen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jans</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanhoof</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>bupaR: Enabling Reproducible Business Process Analysis</article-title>
          .
          <source>Knowl.-Based Syst</source>
          .
          <volume>163</volume>
          ,
          <issue>927</issue>
          {
          <fpage>930</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>La</given-names>
            <surname>Rosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Reijers</surname>
          </string-name>
          , H.A.,
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.,
          <string-name>
            <surname>Dijkman</surname>
            ,
            <given-names>R.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garc</surname>
          </string-name>
          a-Ban~uelos, L.:
          <article-title>APROMORE: An advanced process model repository</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>38</volume>
          (
          <issue>6</issue>
          ),
          <volume>7029</volume>
          {
          <fpage>7040</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lambrechts</surname>
          </string-name>
          , S., van der Aalst, W.,
          <string-name>
            <surname>Weijters</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Scenario-based process mining: Web servicing and automated scenario generation (</article-title>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Mans</surname>
          </string-name>
          , R., van der Aalst, W.,
          <string-name>
            <surname>Verbeek</surname>
          </string-name>
          , H.:
          <article-title>Supporting Process Mining Work ows with RapidProM</article-title>
          .
          <source>In: BPM (Demos)</source>
          . p.
          <volume>56</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>