<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Nirdizati Light: A Modular Framework for Explainable Predictive Process Monitoring</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrei Buliga</string-name>
          <email>abuliga@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Riccardo Graziosi</string-name>
          <email>rgraziosi@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chiara Di Francescomarino</string-name>
          <email>c.difrancescomarino@unitn.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chiara Ghidini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabrizio Maria Maggi</string-name>
          <email>maggi@inf.unibz.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Williams Rizzi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Massimiliano Ronzani</string-name>
          <email>mronzani@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fondazione Bruno Kessler</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Free University of Bozen-Bolzano Bolzano</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Nexoya</institution>
          ,
          <addr-line>Zurich</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Nirdizati Light is an innovative Python package designed for Explainable Predictive Process Monitoring (XPPM). It addresses the need for a modular, flexible tool to compare predictive models, and generate explanations for the predictions made by the predictive models. By integrating consolidated frameworks libraries for process mining, machine learning, and explainable AI, it ofers a comprehensive approach to predictive model construction and explanation generation. This paper discusses the tool's key features, and its significance in the BPM community.</p>
      </abstract>
      <kwd-group>
        <kwd>Monitoring</kwd>
        <kwd>predictive process monitoring</kwd>
        <kwd>machine learning</kwd>
        <kwd>explainable AI</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>2024.
(M. Ronzani)
†These authors contributed equally.</p>
      <p>In response to these constraints, in this demo paper we introduce a modular and extensible
Python-based version of Nirdizati, by ofering more encoding techniques, newer
state-of-theart predictive models, with a particular focus on novel XAI techniques adapted to the PPM
domain. With Nirdizati Light, users can explore a diverse set of trace encodings, predictive tasks,
predictive models, and explanations, enhancing their ability to make data-driven decisions.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Nirdizati Light innovations for (X)PPM</title>
      <p>
        Predictive Process Monitoring (PPM) is crucial for operational optimisation and informed
decision-making. Fig. 1 shows a general pipeline employed for PPM. However, existing PPM
methods often lack transparency and fail to incorporate domain-specific knowledge, limiting
their efectiveness. The adoption of Deep Learning models in Predictive Process Monitoring
(PPM) has synchronously brought upon the adoption of explanatory techniques intending to
provide explanations for diferent prediction tasks. This has lead to the creation of a novel
subfield, named Explainable Predictive Process Monitoring (XPPM) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Nirdizati Light is a modular Python package that supports PPM by providing a comprehensive
suite of functionalities for Explainable Predictive Process Monitoring (XPPM). Designed with
lfexibility at its core, Nirdizati Light 1 allows users to seamlessly import event logs, experiment
with a range of encoding techniques, and train various predictive models. It integrates popular
libraries such as pm4py [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for event log handling, scikit-learn 2 and PyTorch 3 for model
training, and hyperopt 4 for hyperparameter optimisation. This integration facilitates a cohesive
environment where users can conduct all stages of event log analysis within a single platform.
A standout feature of Nirdizati Light is its modularity, enabling users to efortlessly swap
1The tool is available at the following repository link https://github.com/rgraziosi-fbk/nirdizati-light, while the
video demonstration for the tool can be found at https://tinyurl.com/bdhbwwhz
2https://scikit-learn.org/
3https://pytorch.org/
4https://hyperopt.github.io/hyperopt/
components like encodings, models, and explainable AI (XAI) methods. This flexibility supports
a dynamic experimentation process without being confined to a rigid interface. The tool supports
a diverse array of predictive tasks, including outcome prediction, next activity prediction,
remaining time prediction, and trace duration prediction. This breadth of capabilities allows it
to cater to a wide range of use cases and data characteristics, independently on whether the task
involves classification or regression. Fig. 1 also highlights the main functionalities of Nirdizati
light. We present each of the submodules of the framework below.
      </p>
      <p>Event Log labeling. The Prediction task definition module enables the automatic labeling
of logs with various predictions, including categorical outcomes, numeric values, and next
activities. For categorical outcomes, it allows for multiclass labels from categorical attributes
and next activities, as well as binary labels for outcome predictions. For numerical outcomes, it
supports numeric labels derived from numeric attributes and trace duration.
Trace Encoder/Decoder. The Encoding selection module processes labelled event logs
and converts them into a DataFrame suitable for machine learning. This transformation occurs
through three steps: (i) Encoding information extraction: This step extracts critical attributes
from the event log, such as control-flow (activity names), data flow (trace and event attributes),
and resource-flow (resource-related attributes). This mapping identifies the relevant information
for encoding; (ii) Feature encoding: Using the extracted information, this step determines the
feature set that will represent each trace in the DataFrame; (iii) Data encoding: Finally, the
feature set is transformed into a DataFrame. This includes operations like one-hot encoding
of categorical features and normalization of numeric attributes, ensuring the data is ready for
training predictive models. For this we make use of the scikit-learn library.</p>
      <sec id="sec-3-1">
        <title>Predictive Model Selection + Optimisation. The Model(s) selection module allows users</title>
        <p>to specify and instantiate predictive models. It supports both classification and regression
algorithms. The modular design of Nirdizati Light permits the integration and expansion of
additional predictive algorithms, enhancing its adaptability to diferent requirements. For the
predictive models, Nirdizati Light uses popular Machine Learning/Deep Learning libraries such
as scikit-learn and PyTorch to instantiate the predictive models within the framework.
Hyperparameter optimisation. This module enhances model performance by automating
the tuning of hyperparameters using the hyperopt library. This module receives the training
DataFrame and an instantiated predictive model, then explores multiple hyperparameter
conifgurations to maximize a specified quality metric. This process, although computationally
intensive, significantly improves the accuracy and efectiveness of the predictive models.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Predictive Model Comparison. The Model evaluation module provides a comprehensive</title>
        <p>assessment of predictive models based on two primary classes of metrics: (i) Time metrics:
Evaluate the speed at which the predictive model trains, updates, and generates predictions; (ii)
Accuracy metrics: Assess the model’s predictive performance on the test set.</p>
        <p>This module facilitates detailed comparisons between diferent models, ofering insights
into their performance across various configurations and datasets. Nirdizati Light supports a
streamlined workflow from data preprocessing to model evaluation, making it an invaluable
tool for researchers and practitioners in the BPM community.</p>
        <p>
          Explainability. Nirdizati Light also excels in generating actionable insights through
stateof-the-art XAI methods, incorporating advanced tools such as SHAP (SHapley Additive
exPlanations) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], LiME [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], and DiCE (Diverse Counterfactual Explanations) [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] through the
Explanation method selection module. These methods provide deep, interpretative insights
into model predictions, enhancing their transparency and utility. Furthermore, the tool
emphasizes knowledge-aware explainability, leveraging domain-specific knowledge to produce
explanations that are not only accurate but also meaningful and easy to understand.
Furthermore, we also include a selection of state-of-the-art XPPM techniques [
          <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
          ], which leverage
domain-specific knowledge, either through the form of temporal constraints (LTLf and Declare),
or by providing explanations in terms of process patterns 5. These adapted techniques focus on
both providing the reasons for the prediction made by the model (so-called factual explanations)
and showing the required changes to the input to achieve an alternative outcome (also known
as counterfactual explanations). By integrating these advanced features and methodologies,
Nirdizati Light empowers process analysts and data scientists to unlock profound insights from
event logs and make well-informed decisions. Its ability to support flexible experimentation
and deliver interpretive, domain-specific explanations marks a significant advancement in the
XPPM domain, providing a robust and intuitive platform for comprehensive data analysis.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Concluding Remarks</title>
      <p>This paper introduced Nirdizati Light, a significant advancement in the realm of Explainable
Predictive Process Monitoring (XPPM), addressing the limitations of existing tools like Nirdizati
by ofering a modular, flexible, and powerful Python package that facilitates the construction
and comparison of diferent predictive models and trace encodings for a given event log.
Its architecture supports easy integration and comparison of various encoding techniques,
predictive models, and state-of-the-art explainability methods, while its modularity allows users
to experiment with and adopt the latest advancements in predictive process monitoring, tailoring
solutions to specific use cases. This flexibility is crucial in the PPM domain, where researchers
and practitioners need adaptable tools for a wide range of scenarios and data characteristics.</p>
      <p>
        We assess the current Technology Readiness Level of Nirdizati Light to be a 4, reflecting
its well-defined software structure, its versatility and robustness demonstrated through past
applications in various domains [
        <xref ref-type="bibr" rid="ref12 ref8 ref9">12, 13, 8, 9</xref>
        ]. With its flexible framework and feature set, the tool
ofers researchers and practitioners a tool to enhance their understanding of predictive process
monitoring techniques, and easily extend the framework with additional custom methods.
5See [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] for more details on Declare and LTLf, and [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for more details on process patterns.
      </p>
    </sec>
    <sec id="sec-5">
      <title>4. Acknowledgments</title>
      <p>This work was partially supported by the Italian (MUR) under PRIN project PINPOINT Prot.
2020FNEB27, CUP H23C22000280006 and H45E21000210001, the support is greatly appreciated.
of Lecture Notes in Business Information Processing, Springer, 2020, pp. 141–158. URL:
https://doi.org/10.1007/978-3-030-58638-6_9. doi:10.1007/978- 3- 030- 58638- 6\_9.
[13] M. Ronzani, R. Ferrod, C. Di Francescomarino, E. Sulis, R. Aringhieri, G. Boella, E. Brunetti,
L. Di Caro, M. Dragoni, C. Ghidini, R. Marinello, Unstructured data in predictive process
monitoring: Lexicographic and semantic mapping to ICD-9-CM codes for the home
hospitalization service, in: AIxIA 2021, Revised Selected Papers, volume 13196 of Lecture
Notes in Computer Science, Springer, 2021, pp. 700–715.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Di Francescomarino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ghidini</surname>
          </string-name>
          ,
          <article-title>Predictive process monitoring</article-title>
          ,
          <source>in: Process Mining Handbook</source>
          , volume
          <volume>448</volume>
          <source>of LNBIP</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>320</fpage>
          -
          <lpage>346</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>W.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Simonetto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Di Francescomarino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ghidini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kasekamp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Maggi</surname>
          </string-name>
          ,
          <article-title>Nirdizati 2.0: New features and redesigned backend</article-title>
          ,
          <source>in: Proceedings of the Dissertation Award, Doctoral Consortium, and Demonstration Track at BPM</source>
          <year>2019</year>
          , volume
          <volume>2420</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>154</fpage>
          -
          <lpage>158</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Stierle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Brunk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Weinzierl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zilker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matzner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Becker</surname>
          </string-name>
          ,
          <article-title>Bringing light into the darkness-a systematic literature review on explainable predictive business process monitoring techniques</article-title>
          ,
          <source>ECIS 2021 Research-in-Progress Papers</source>
          <volume>8</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Berti</surname>
          </string-name>
          , S. van Zelst,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schuster</surname>
          </string-name>
          ,
          <article-title>Pm4py: A process mining library for python</article-title>
          ,
          <source>Software Impacts</source>
          <volume>17</volume>
          (
          <year>2023</year>
          )
          <fpage>100556</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Lundberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>A unified approach to interpreting model predictions</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems 30: Annual Conf. on Neural Information Processing Systems</source>
          <year>2017</year>
          „ 2017, pp.
          <fpage>4765</fpage>
          -
          <lpage>4774</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Guestrin</surname>
          </string-name>
          , “
          <article-title>Why should i trust you?” Explaining the predictions of any classifier</article-title>
          ,
          <source>in: Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1135</fpage>
          -
          <lpage>1144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Mothilal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. Tan,</surname>
          </string-name>
          <article-title>Explaining machine learning classifiers through diverse counterfactual explanations</article-title>
          ,
          <source>in: Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>607</fpage>
          -
          <lpage>617</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Buliga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Di Francescomarino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ghidini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Maggi</surname>
          </string-name>
          ,
          <article-title>Counterfactuals and ways to build them: Evaluating approaches in predictive process monitoring</article-title>
          ,
          <source>in: Advanced Information Systems</source>
          Engineering - 35th International Conference, CAiSE
          <year>2023</year>
          , Proceedings, volume
          <volume>13901</volume>
          <source>of LNCS</source>
          , Springer,
          <year>2023</year>
          , pp.
          <fpage>558</fpage>
          -
          <lpage>574</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Buliga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Francescomarino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ghidini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Donadello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Maggi</surname>
          </string-name>
          ,
          <article-title>Guiding the generation of counterfactual explanations through temporal background knowledge for predictive process monitoring</article-title>
          ,
          <source>CoRR abs/2403</source>
          .11642 (
          <year>2024</year>
          ). arXiv:
          <volume>2403</volume>
          .
          <fpage>11642</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Di Ciccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montali</surname>
          </string-name>
          ,
          <article-title>Declarative process specifications: Reasoning, discovery, monitoring</article-title>
          , in: Process Mining Handbook, volume
          <volume>448</volume>
          <source>of LNBIP</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>108</fpage>
          -
          <lpage>152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vazifehdoostirani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Genga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Verhoeven</surname>
          </string-name>
          ,
          <string-name>
            <surname>H. van Laarhoven</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Dijkman</surname>
          </string-name>
          ,
          <article-title>Interactive multi-interest process pattern discovery</article-title>
          ,
          <source>in: Business Process Management - 21st Int. Conf., BPM</source>
          <year>2023</year>
          ,
          <article-title>Proceedings</article-title>
          , volume
          <volume>14159</volume>
          <source>of LNCS</source>
          , Springer,
          <year>2023</year>
          , pp.
          <fpage>303</fpage>
          -
          <lpage>319</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>W.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Francescomarino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Maggi</surname>
          </string-name>
          ,
          <article-title>Explainability in predictive process monitoring: When understanding helps improving</article-title>
          ,
          <source>in: Business Process Management Forum - BPM Forum</source>
          <year>2020</year>
          , Seville, Spain,
          <source>September 13-18</source>
          ,
          <year>2020</year>
          , Proceedings, volume
          <volume>392</volume>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>