<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <article-id pub-id-type="doi">10.1016/J.JBI.2016.04.007</article-id>
      <title-group>
        <article-title>Next Activity Prediction and Elapsed Time Prediction on Process Dataset Vincenzo Dentamaro 1, Donato Impedovo 1 , Giuseppe Pirlo 1 and Gianfranco Semeraro 1-2</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>First Block</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elapsed Time Block</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Studies of Bari</institution>
          ,
          <addr-line>Via Edoardo Orabona, Bari, 4</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University School for Advanced Studies IUSS Pavia</institution>
          ,
          <addr-line>Piazza della Vittoria, Pavia, 15</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <volume>61</volume>
      <issue>1</issue>
      <fpage>198</fpage>
      <lpage>203</lpage>
      <abstract>
        <p>Process mining is a field of research that has gained much attention in recent years because of its ability to analyze and improve processes. Indeed, one of the key aspects of process mining is its ability to predict the activities in the future and the time spent on these activities. In this work is proposed the use of Bidirectional LSTM and Multi-Speed Transformer on a recent dataset called BPIC-2020 related to reimbursement process of the University of Technology of Eidenhoven. Results shows that Multi-Speed Transformer is more capable to performs next activity prediction than the Bi-LSTM. Meanwhile, for the elapsed time prediction, the viceversa is true.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Process Mining is a research domain that gains
a lot of interest thanks to the application of Data
Mining and Machine Learning methods to the
processes [1]–[3]. The application is performed
on processes recorded as timeseries data called
“event log”, accordingly to [4]. In this way several
analysis can be performed, starting from an initial
discovery study [4] (building new models of the
processes recorded) to a conformance and
modeling check to analyze the evolving situation
of processes and, thus, align models, and related
software, to business processes [
        <xref ref-type="bibr" rid="ref6">5</xref>
        ].
      </p>
      <p>The analysis of timeseries is a problem broadly
manage using several types of models. Starting
from the using of classical LSTM [2], [3] to the
use of the Convolutional Neural Networks [1],
[6].</p>
      <p>The importance of the Process Mining
techniques is also highlighted by the Public
Administration (PA) that starts to use such
technology in order to build models that helps to
enhance the quality of the PA’s processes [7]. The
application of Process Mining starts from the
process modelling, using several techniques as
design processes based on collected data [8], or
analysing processes in order to improve them [9].</p>
      <p>The PA involved in the use of Process Mining
techniques is not only related to governance[7]–
[9], but also related to health [10], [11] and
educational system [12], [13].</p>
      <p>In this work, it is proposed the use of an neural
network architecture depicted in Figure 1. The
architecture is based on the use of “block” meant
to be the same neural network architecture used
firstly to analysis the timeseries and then using a
block to predict the next activity and the other one
to predict the elapsed time to the next activity.
This model was presented in [2] using a
BiLSTM. The work is structure presented related
works in Section 2. Then the dataset and the used
method are explained in section 3. The
experimental set-ups to perform experiments are
explained in Section 4. Results and their
discussion are presented in Section 5. Finally in
of the University of Eindhoven (TU/e). The data</p>
      <sec id="sec-1-1">
        <title>Section 6 there are the conclusions.</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>The analysis of the event logs can be
performed for several reason. In this work, it was
proposed the use of the architecture presented by
Gunnarsson B et al in [2]. The main architecture
is depicted in Figure 1. It is noticeable that such
architecture is based on the use of three blocks
that represents the
same
underlying
neural
networks. In this way, a first application performs
a sort of features extraction from the timeseries,
meanwhile
the
other
two
application
are
specialized to predict the next activity and the
elapsed time to the next activity. As in the
majority of works using timeseries (called also
event log), in [2] it was proposed the use of a</p>
      <sec id="sec-2-1">
        <title>Bidirectional LSTM</title>
        <p>(Bi-LSTM) in
order to
capture temporal patterns looking forward and
backward.</p>
        <p>Dentamaro V. et al, propose a new architecture
called</p>
      </sec>
      <sec id="sec-2-2">
        <title>Multi-Speed Transformer [6]. Such</title>
        <p>architecture is based on a firstly application of a
multi branch analysis in order to analysis with
different level of details the data. In this way,
analogously to the use of microscopy, it possible
to identify fine patterns and more gross patterns.</p>
        <p>The</p>
      </sec>
      <sec id="sec-2-3">
        <title>Multi-Speed</title>
        <p>Transformer
using
information improved the State of the Art in terms
of prediction metrics.</p>
        <p>Hence, in this work is proposed the use of both
3.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Materials and Methods</title>
      <p>In this section the used methods to perform
experiments about next activity and elapsed time
prediction are explained along with the used
dataset.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Materials</title>
      <p>In this work, it is used a dataset related to the
activity performed by University of Technology
of Eindhoven (TU/e). In particular, the dataset
that comes from the Business Process Intelligence
Challenge 2020 (BPIC-2020) [14].
were collected from the 2017 to the 2018.</p>
      <p>Furthermore, the collected data are organized
in several event log. Specifically, the event logs
are:
• “Domestic Declarations” that is related to
domestic travel (within the same country).
The
event
log
related
to
“Domestic
Declaration” contains 56437 events recorded
and the recorded events are about 10500
cases;
• “Request for Payement” that contains cases
that could be not related to travels. The event
log
related to
“Request for</p>
      <p>Payment”
contains 36796 events for 6886 cases;
• “International Declarations” is related to
travel outside the country. The event log
contains 72151 events and the recorded
events are about 6449 cases;
• “Travel Permit Data” is related to the
permission to travel. This event log is
composed by 86581 events for 7065 cases;
• “Prepaid Travel Costs” contains data related
to the travel costs prepaided. This events log
is composed 18246 events and the recorded
events are related to 2099 cases.</p>
      <p>Hence, each event log is composed with several
events.</p>
      <p>Within each event log is possible to
identify and separate cases in order to obtain, for
each case, a separate event log, called “trace”.
The “trace” is used to preprocess data in order to
apply methods explained in the next section.
A “trace”  is composed by a sequence of events

 with 1 &lt;  &lt;</p>
      <p>where   is the number of
events for the “trace”  . The event  is composed
as a record of information as: “case id”  , “activity
label”  , “timestamp”  , “attributes j-th”   where
1 &lt;  &lt;  and M is the number of attributes. The
attributes can contains information related to the
case (this information are shared from all the
activity and they don’t change long the trace) and
to the activity (this information are specific to the
recorded activity).
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Methods</title>
      <p>In this subsection, the used methods to perform
prediction about the next activity and the elapsed
time are explained.</p>
      <p>The used methods are based on the use of a
new State of Art neural network called
“MultiSpeed Transformer” [6] and a well-known
Bidirection LSTM (referred also Bi-LSTM). This
two neural networks were used as block in a more
complex architecture in order to perform both
next activity and elapsed time prediction.</p>
      <p>The main architecture is presented in Figure 1.</p>
    </sec>
    <sec id="sec-6">
      <title>4. Experimental Set-Up</title>
      <p>The architecture takes as input a preprocessed
trace as in Table 1. Hence, the sequence of
prefixes is feed to the first block. In this way a first
analysis is performed. The output of the first block
is then passed to both “next activity block” and
“elapsed time block”. Such blocks are equals to
the first block, because they use the same
structure. But the “next activity block” gives as
ouput a distribution of probability to predict the
next activity, meanwhile the “elapsed time block”
gives the predicted elapsed time to the next
activity.
4.1</p>
    </sec>
    <sec id="sec-7">
      <title>Bi-LSTM</title>
      <p>Bidirectional LSTM is a neural network
architecture based on the processing of timeseries.
In particular, the core part is the “LSTM” that
analyse data in order to find temporal patterns.</p>
      <p>Commonly, such analysis is performed in
forward setup, i.e. from the past to the present.</p>
      <p>In Bidirectional LSTM, two layer of LSTM are
used in order to analyse data in forward and
backward way. In this way, both forward and
backward temporal patterns are learnt by the
model.
4.2</p>
    </sec>
    <sec id="sec-8">
      <title>Multi-Speed Transformer</title>
      <p>The Multi-Speed Transformer was presented
by Dentamaro V. et al [6]. This model is based on
the concept of fine and gross analysis of data.
Such analysis are done at different levels of detail
as analysis performed by microscopies at different
resolutions. In this way, different kind of
information are extracted and then concatenated
to obtain the prediction.</p>
    </sec>
    <sec id="sec-9">
      <title>Experimental Set-Ups</title>
      <p>In this section the setups used to perform
experiments are explained.
5.1</p>
    </sec>
    <sec id="sec-10">
      <title>Data Preprocessing</title>
      <p>Each event log is also divided in train and test split
following the set-up used in [2]. Specifically, the
cases are ordered by the timestamp of the related
recorded event. Then a 75/25 split is applied.
After such division, from the train set are
eliminated the cases which are not ended when the
In order to use the dataset BPIC2020, each trace,
in each event log, is firstly sorted accordingly with
the timestamps of each event, then it is
preprocessed extracting information as “prefix”,
“suffix” and “elapsed time”. In the following, this
information are defined:
• “Prefix”: Prefix is defined as a function that
given a trace  , a position  and a windows
size  , it return a sequence of events from  −
 to  :</p>
      <p>( ,  ,  ) = &lt;  { − }, … ,  { } &gt;
• “Suffix”: Suffix is also defined as a function
that given a trace  , a position  and a
windows size  , it return a sequence of events
from  + 1 to  +  :</p>
      <p>( ,  ,  ) = &lt;  { +1}, … ,  { + } &gt;
• “Elapsed Time”: Elapsed time is the time
remaining to the next activity. Hence it is
defined as a function that given a trace  , a
position  it returns the difference between the
timestamp of   +1 and   :</p>
      <p>( ,  ) =   +1.  −   . 
The “.” is intended as the operator that recall
the timestamp  values of the event. The
elapsed time is computed in seconds.</p>
      <p>For both “prefix” and “suffix”, if the windows
size excides the number of events to be selected
zero-padding is applied, e.g. for the prefix, if the
given position is lower than the windows size it
means that “k-w” is negative, hence to overcome
this problem zeros are added.</p>
      <p>Finally, in the Table 1 it is shown an example of
the results obtained from the preprocessing. It is
supposed that the given trace is composed by 5
events:  = &lt;  1,  2,  3,  4,  5 &gt;, and the
windows size is  = 3:</p>
      <p>Table 1 Example of preprocessing of a trace.
For the Suffix the windows size is 1 to obtain only
the next activity to predict</p>
      <p>Prefix(T,p,w)</p>
      <p>Suffix(T,p,1)</p>
      <p>Elapse
Time(T,p)
 2 −  1
 3 −  2
 4 −  3
 5 −  4
t w
first case of the test set started. Successively, the
train set is further divided in train and validation
set in 75/25 setup.</p>
      <p>For each event log, the activity label are one hot
encoded.</p>
    </sec>
    <sec id="sec-11">
      <title>6. Results and Discussion</title>
      <p>Once the experiments were performed,
predictions were evaluated accordingly to the
following metrics:
• Categorical Accuracy: The proportion of the
correctly predicted “next activity label” on the
total number of predicted “next activity label”.
• Root Mean Squared Error: The root applied at
the sum of the squared differences between
predicted “elapsed time” and real “elapsed
time”.</p>
      <p>The results of the performed experiments are
reported in Table 2:
From the results showed in Table 2 it is possible
to notice the in terms of accuracy to predict the
next activity, the proposed use of the Multi-Speed
Transformer is better than the proposed use of the
Bidirectional LSTM (Bi-LSTM). Indeed, the best
performance of the Bi-LSTM is 10% lower than
the lower performance using the Multi-Speed
Transformer.</p>
      <p>Related to the elapsed time, it is possible to notice
that the Root Mean Square Error (RMSE) is
lower for the models using the Bi-LSTM. In this
sense, the model using the Bi-LSTM seems to be
more suitable to predict the elapsed time to the
next activity.</p>
      <p>A more deeply analysis of the results highlights
that the model using Bi-LSTM is truly better than
the model using Multi-Speed Transformer to
predict elapsed time to the next-activity, but the
RMSE in all the models and for all the used event
log are high.</p>
    </sec>
    <sec id="sec-12">
      <title>7. Conclusion</title>
      <p>In this work, the use of Bi-LSTM is compared
with the use of Multi-Speed Transformer as block
of a more complex architecture to build systems
capable of both prediction of the next activity and
elapsed time to the next activity.</p>
      <p>The results shows that multi-speed
Transformer is suitable to predict next activity
given a sequence of activity. Meanwhile, the
prediction of the elapsed time seems to be a task
more suitable for the model using the Bi-LSTM.</p>
      <p>Future works could include more information
about the case or the activity in order to improve
the quality of predictions.</p>
    </sec>
    <sec id="sec-13">
      <title>8. Acknowledgements</title>
      <p>This Word template was created by Aleksandr
Ometov, TAU, Finland. The template is made
available under a Creative Commons License
Attribution-ShareAlike 4.0 International (CC
BYSA 4.0).</p>
    </sec>
    <sec id="sec-14">
      <title>9. References</title>
      <p>[1]
[2]
[3]</p>
      <p>E. Obodoekwe, X. Fang, and K. Lu,
“Convolutional Neural Networks in
Process Mining and Data Analytics for
Prediction Accuracy,” Electronics
(Switzerland), vol. 11, no. 14, Jul. 2022,
doi: 10.3390/ELECTRONICS11142128.
B. R. Gunnarsson, S. vanden Broucke, and
J. De Weerdt, “A Direct Data Aware
LSTM Neural Network Architecture for
Complete Remaining Trace and Runtime
Prediction,” IEEE Trans Serv Comput,
2023, doi: 10.1109/TSC.2023.3245726.
N. Tax, I. Verenich, M. La Rosa, and M.
Dumas, “Predictive Business Process
Monitoring with LSTM Neural
Networks,” Lecture Notes in Computer
Science (including subseries Lecture
Notes in Artificial Intelligence and Lecture</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>Notes in Bioinformatics)</source>
          , vol.
          <volume>10253</volume>
          LNCS, pp.
          <fpage>477</fpage>
          -
          <lpage>492</lpage>
          , Dec.
          <year>2016</year>
          , doi: 10.1007/978-3-
          <fpage>319</fpage>
          -59536-8_
          <fpage>30</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>W. Van der Aalst</surname>
          </string-name>
          , “
          <article-title>Process mining: Data science in action,” Process Mining: Data Science in Action</article-title>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>467</lpage>
          , Jan.
          <year>2016</year>
          , doi: 10.1007/978-3-
          <fpage>662</fpage>
          -49851- 4/COVER.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Tortorella</surname>
          </string-name>
          , “
          <article-title>Assessment and impact analysis for aligning business processes and software systems</article-title>
          ,
          <source>” Proceedings of the ACM Symposium on Applied Computing</source>
          , vol.
          <volume>2</volume>
          , pp.
          <fpage>1338</fpage>
          -
          <lpage>1343</lpage>
          ,
          <year>2005</year>
          , doi: 10.1145/1066677.1066978.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Impedovo</surname>
          </string-name>
          , and G. Pirlo,
          <article-title>“Multi-speed transformer network for neurodegenerative disease assessment and activity recognition,” Comput Methods Programs Biomed</article-title>
          , vol.
          <volume>230</volume>
          , p.
          <fpage>107344</fpage>
          ,
          <string-name>
            <surname>Mar</surname>
          </string-name>
          .
          <year>2023</year>
          , doi: 10.1016/J.CMPB.
          <year>2023</year>
          .
          <volume>107344</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Forliano</surname>
          </string-name>
          ,
          <string-name>
            <surname>P. De Bernardi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Bertello</surname>
            , and
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Temperini</surname>
          </string-name>
          , “
          <article-title>Innovating business processes in public administrations: towards a systemic approach,” Business Process Management Journal</article-title>
          , vol.
          <volume>26</volume>
          , no.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          5, pp.
          <fpage>1203</fpage>
          -
          <lpage>1224</lpage>
          , Oct.
          <year>2020</year>
          , doi: 10.1108/BPMJ-12-2019-0498.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Kiss</surname>
          </string-name>
          and G. Klimkó, “
          <article-title>A Reverse Data-Centric Process Design Methodology for Public Administration Processes</article-title>
          ,
          <source>” Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)</source>
          , vol.
          <volume>11709</volume>
          LNCS, pp.
          <fpage>85</fpage>
          -
          <lpage>99</lpage>
          ,
          <year>2019</year>
          , doi: 10.1007/978- 3-
          <fpage>030</fpage>
          -27523-
          <issue>5</issue>
          _
          <fpage>7</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>G.</given-names>
            <surname>Polancica</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Šumaka</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Pušnik</surname>
          </string-name>
          , “
          <article-title>A Case-based Analysis of Process Modeling for Public Administration System Design,”</article-title>
          <source>Frontiers in Artificial Intelligence and Applications</source>
          , vol.
          <volume>321</volume>
          , pp.
          <fpage>92</fpage>
          -
          <lpage>104</lpage>
          , Dec.
          <year>2019</year>
          , doi: 10.3233/FAIA200009.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Munoz-Gama</surname>
          </string-name>
          et al., “
          <article-title>Process mining for healthcare: Characteristics and challenges,” J Biomed Inform</article-title>
          , vol.
          <volume>127</volume>
          , p.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          103994,
          <string-name>
            <surname>Mar</surname>
          </string-name>
          .
          <year>2022</year>
          , doi: 10.1016/J.JBI.
          <year>2022</year>
          .
          <volume>103994</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>