<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Secrecy in Decentralized Process Mining</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Valerio Goretti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Basile</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Barbaro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudio Di Ciccio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Process mining</institution>
          ,
          <addr-line>Decentralized computing, Confidential Computing, Trusted Execution Envi-</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Sapienza Univeristy of Rome</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Utrecht University</institution>
          ,
          <country country="NL">Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <fpage>14</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>In the contemporary business landscape, collaboration across multiple organizations ofers a multitude of opportunities, including reduced operational costs, enhanced performance, and accelerated technological advancement. The application of process mining techniques in an inter-organizational setting, exploiting the recorded process event data, enables the coordination of joint efort and allows for a deeper understanding of the business. Nevertheless, considerable concerns pertaining to data confidentiality emerge, as organizations frequently demonstrate a reluctance to expose sensitive data demanded for process mining, due to concerns related to privacy and security risks. The presence of conflicting interests among the parties involved can impede the practice of open data sharing. To address these challenges, we propose our approach and toolset named CONFINE, which we developed with the intent of enabling process mining on process event data from multiple providers while preserving the confidentiality and integrity of the original records. To ensure that the presented interaction protocol steps are secure and that the processed information is hidden from both involved and external actors, our approach is based on a decentralized architecture and consists of trusted applications running in Trusted Execution Environments (TEE). In this demo paper, we provide an overview of the core components and functionalities as well as the specific details of its application.</p>
      </abstract>
      <kwd-group>
        <kwd>Metadata description</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org
Mining</p>
    </sec>
    <sec id="sec-2">
      <title>Current version</title>
    </sec>
    <sec id="sec-3">
      <title>Legal code license</title>
    </sec>
    <sec id="sec-4">
      <title>Languages, tools and services used</title>
    </sec>
    <sec id="sec-5">
      <title>Supported operating environment</title>
    </sec>
    <sec id="sec-6">
      <title>Download/Demo URL</title>
    </sec>
    <sec id="sec-7">
      <title>Documentation URL</title>
    </sec>
    <sec id="sec-8">
      <title>Source code repository</title>
    </sec>
    <sec id="sec-9">
      <title>Screencast video</title>
    </sec>
    <sec id="sec-10">
      <title>Value</title>
    </sec>
    <sec id="sec-11">
      <title>CONFINE 1.0</title>
    </sec>
    <sec id="sec-12">
      <title>Apache 2.0</title>
    </sec>
    <sec id="sec-13">
      <title>GNU/Linux</title>
      <p>github.com/Process-in-Chains/CONFINE.git
github.com/Process-in-Chains/CONFINE/blob/main/README.md
github.com/Process-in-Chains/CONFINE
youtu.be/Oaoo6gS_4tw?si=hY0ztcbUrGO8PjIr</p>
      <sec id="sec-13-1">
        <title>1. Introduction</title>
        <p>
          Collaboration between multiple organizations is essential in today’s highly competitive
and development-oriented business environment. By aligning around shared goals,
organizations can streamline operations, reduce redundancies, and ultimately improve
eficiency, performance, and cost-efectiveness. In this context, inter-organizational process
mining enables the coordination of joint eforts, improves overall transparency, and allows
for benchmarking [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Nevertheless, despite the numerous benefits, a number of potential
issues may arise, primarily related to data confidentiality. Information is an asset, and
even in a collaborative environment companies are unwilling to share sensitive data
required to execute process mining algorithms with their partners [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Allowing sensitive
operational data to move across organizational boundaries inevitably raises issues related
to data privacy and security, potentially exposing the data to unauthorized access and
misuse [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. For instance, we may consider a hospitalization process that involves three
diferent parties: a hospital, a pharmaceutical company, and a specialized clinic. In
this scenario, two other entities wish to uncover information on the inter-organizational
process for reporting and auditing purposes: the National Institute of Statistics of the
country where the three organizations reside and the University that hosts the hospital [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
The hospital, specialized clinic, and pharmaceutical company have a partial view of the
overall dynamics of the inter-organizational process as they record operations pertaining
to their own parts. As a result, each player stories a separate event log partition. It
would be mutually beneficial for all parties involved to have access to the findings of an
aggregate data analysis, integrating the event log partitions yielded by the collaborating
organizations. Nevertheless, an intrinsic divergence of interests emerges: granting access
to traces to other organizations may reveal information the parties do not want to disclose.
The ability to conduct process mining on data from multiple sources must, therefore,
meet the necessity to safeguard the privacy of the entities involved and to guarantee the
confidentiality of the information, ensuring that the local event log is never given away
completely in-clear.
        </p>
        <p>
          To solve this conundrum, we present CONFINE, our recently developed approach
and toolkit designed to enhance collaborative information system architectures with
secrecy-preserving process mining capabilities. To secure information secrecy during the
exchange and elaboration of data, our solution resorts to Trusted Execution Environments
(TEEs) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. These are hardware-secured contexts, which guarantee code integrity and data
confidentiality before, during, and after their utilization. Owing to these characteristics,
CONFINE lets information be securely transferred beyond an organization’s perimeter.
Therefore, computing nodes other than the information provisioners can aggregate
and elaborate the original, unaltered process data in a secure, externally inaccessible
vault. Also, CONFINE is capable of providing these guarantees while demanding low
computation overhead and providing scalability.
        </p>
        <p>Process
model
Process
diagnostics
Event log
partitions</p>
        <p>Log Servers’
configuration</p>
        <p>System
configuration</p>
        <p>INPUT
Trusted Execution Environment</p>
        <p>Process
documentation
Log Servers</p>
        <p>CONFINE protocol</p>
        <p>Secure Miner</p>
        <p>OUTPUT
Merged
event log
Provisioning</p>
        <p>Mining</p>
      </sec>
      <sec id="sec-13-2">
        <title>2. The CONFINE Framework</title>
        <p>In the following, we present the core concepts of CONFINE by introducing the main
components and discussing how they interact in the CONFINE protocol.
High-level architecture. Figure 1 displays a high-level schematization of the CONFINE
framework. Our architecture involves diferent information systems running on multiple
machines. An organization can take at least one of the following roles, depending
on the tasks they take on: provisioning if it delivers local partitions of event logs
to be collaboratively mined (e.g., the hospital in our example); mining if it applies
process mining algorithms using event logs retrieved from provisioners (e.g., the National
Institute of Statistics). In our solution, every organization hosts one or more nodes,
incorporating components that depend on the roles played. The nodes hosted by
provisioning organizations implement the Log Server component. The data provided
by Log Servers is fed into the Secure Miner components, which are deployed on nodes
hosted by mining organizations. Notice that every Secure Miner retrieves process data
from one to many Log Servers, as each of the latter can ofer diferent partitions of
the event log (for example, the hospital records activities that are not visible to the
pharmaceutical company, and vice-versa). In CONFINE, multiple organizations can
perform mining tasks by hosting one or more Secure Miners, eliminating the need to
depend on a central authority. This reflects the decentralized nature of our approach.</p>
        <p>
          CONFINE allows miners to request data from provisioners that are part of the
interorganizational context to perform mining operations while ensuring the secrecy and
confidentiality of the event logs exchanged. These properties are ensured since the
Secure Miner is a trusted application executed within a trusted execution environment
(TEE). TEEs create isolated contexts separate from the operating system, safeguarding
code and data with hardware-based encryption mechanisms. They use dedicated CPU
instructions to manage encrypted data within a reserved memory portion [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The CPU
encrypts this memory with a random decryption key generated at each power cycle. By
enforcing strict memory access controls, TEEs prevent applications from accessing or
altering each other’s memory space, thus enhancing system security. The CONFINE
protocol securely interconnects Log Servers with the Secure Miners.
The CONFINE protocol. In our protocol, the interaction between components is divided
into four sequential steps, namely (i) initialization, (ii) remote attestation, (iii) data
transmission, and (iv) computation. The Secure Miner starts the protocol execution with
the initialization step, during which it gets informed about the case distribution of the log
partitions in the Log Servers; this data request includes identity evidence of the mining
organization. In this phase, e.g., the National Institute of Statistics requests the
hospital’s Log Server to list aggregate information about the cases they can share in
order to pre-plan the subsequent data requests. Next, the remote attestation phase occurs.
The purpose of remote attestation is to establish trust between the Log Servers and
the Secure Miner. Since the actual process data are going to be transmitted next, a
stronger certification is required. This phase is thus based on the RATS RFC standard [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]
(the basis of TEE attestation schemes) and has three main objectives: (i) to provide the
Log Servers evidence that the data request for a log partition originates from a trusted
application running within a TEE; (ii) to verify that the trusted application is indeed
the authentic Secure Miner software entity; (iii) to identify the owner of the Secure
Miner application. Once the trusted nature of the Secure Miner is verified, it retrieves
from the Log Servers their log partitions to internally merge the cases. Based on the
information retrieved during the initialization phase, it plans the requests accordingly.
Each Log Server retrieves its event log and filters it based on the case id specified by the
miner. Given the typically limited capacity of a TEE working memory, Log Servers are
required to divide the filtered log into separate segments of up to a certain size. In the
computation phase, the process mining technique selected by the organization is applied
to the received data. The mining organizations can use the Secure Miner in either an
incremental or a non-incremental manner. The distinction between these approaches
concerns the timing in which the computation phase is executed. With the incremental
approach, the Secure Miner begins the computation phase concurrently with the data
transmission phase: As soon as all Log Servers have sent their log segments pertaining
to a specific case, the mining algorithm is executed to produce and update a partial
mining result. In contrast, with the non-incremental approach, the Secure Miner waits
until all full partitions are received from the Log Servers, and only then it starts the
mining algorithm. The incremental approach tends to save on TEE memory (the full log
does not need to get loaded before mining begins), but it poses a restriction on the mining
algorithm in use as it must be apt for treating partial inputs and produce updates.
        </p>
      </sec>
      <sec id="sec-13-3">
        <title>3. Maturity</title>
        <p>To showcase the capabilities of CONFINE, we propose a prototype implementation and
run it with artificially generated and real-world data. Our Secure Miner implementation
is an Intel-SGX1 trusted app developed in GO.We adopt the EGo2 framework to deploy
the Secure Miner trusted app into the Intel SGX TEE. As depicted in Fig. 1, the Secure
Miner takes as input a configuration of Log Servers, and a set of parameters to execute
the protocol, and process documentation (a collective term indicating cues for the mining
algorithms, such as a collaborative process specification for conformance checking). Users
can interact with Secure Miner using a terminal interface through which commands are
forwarded to the TEE. Figure 2 shows a screenshot of this interface. We developed the
Log Server in GO. Communication between the Secure Miner and Log Server relies
upon the HTTP protocol secured via TLS.</p>
        <p>
          Our Secure Miner implementation incorporates the Heuristics Miner [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] algorithm
for process discovery and a porting of the PM4Py Conformance Declare3 algorithm for
the conformance checking of declarative process specifications [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The results produced
by the HeuristicsMiner can be examined using workflow net visualization tools like
WoPeD [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. We provide both an incremental variant, which updates partial results
as complete cases are collected, and a non-incremental variant, starting only after the
entire merged event log is collected. We validated our prototypical implementation of
CONFINE on real-world event logs, including BPIC 2013 [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and Sepsis [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], as well
as a synthetic log generated from a healthcare scenario. We evaluated the convergence
of the CONFINE protocol and the scalability of the Secure Miner’s runtime memory
usage. These tests were conducted in both simulation mode and SGX native mode.
        </p>
        <p>We plan to improve the CONFINE toolset in diferent directions. Given the stringency
of the hardware requirement imposed by CONFINE, we envision real-world production
scenarios involving the Secure Miner instances being deployed on remote machines
owned by service providers. To take steps towards this direction, we plan to implement
remote APIs through which mining organizations can remotely submit commands to
the Secure Miners and securely receive the output of the computation. At present, the
Secure Miner prototype ofers a limited choice of process mining algorithms. We aim to
widen the spectrum of applicable algorithms by designing a plug-in integration system
enabling implementers to add computation modules that are auditable by Log Server
via remote attestation. Furthermore, the integration of a graphical user interface would
facilitate the interaction with the Log Server and the Secure Miner trusted app. These
improvements pave the path for future work.
1https://www.intel.com/content/www/us/en/developer/tools/software-guard-extensions/overview.html.
Accessed: September 26, 2024.
2https://www.edgeless.systems/products/ego. Accessed: September 26, 2024.
3https://processintelligence.solutions/static/api/2.7.11/pm4py.algo.conformance.declare.html.
Accessed: September 26, 2024.</p>
      </sec>
      <sec id="sec-13-4">
        <title>4. Availability</title>
        <p>
          A more detailed description of the CONFINE approach is available in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. Our
implementation can be downloaded from github.com/Process-in-Chains/CONFINE. The readme
ifle of the repository guides the installation of the Secure Miner and the Log Server
components. The video demonstration of CONFINE is included in the repository’s
readme and can be directly watched at www.youtube.com/watch?v=Oaoo6gS_4tw.
        </p>
      </sec>
      <sec id="sec-13-5">
        <title>Acknowledgments</title>
        <p>This research work was partly funded by MUR under PRIN grant B87G22000450001
(PINPOINT), the Latium Region under PO FSE+ grant B83C22004050009 (PPMPP),
Sapienza University of Rome under grant RG123188B3F7414A (ASGARD), and by the
EU-NGEU under the NRRP MUR grant PE00000014 (SERICS).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
          </string-name>
          ,
          <article-title>Intra-and inter-organizational process mining: Discovering processes within and between organizations</article-title>
          , in: PoEM,
          <year>2011</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Challenges and opportunities in collaborative business process management: Overview of recent advances and introduction to the special issue</article-title>
          ,
          <source>Inf. Syst. Front</source>
          .
          <volume>11</volume>
          (
          <year>2009</year>
          )
          <fpage>201</fpage>
          -
          <lpage>209</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ostern</surname>
          </string-name>
          ,
          <string-name>
            <surname>Koljada</surname>
          </string-name>
          , et al.,
          <article-title>Trust mining: analyzing trust in collaborative business processes</article-title>
          , IEEE Access (
          <year>2021</year>
          )
          <fpage>65044</fpage>
          -
          <lpage>65065</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Jans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hosseinpour</surname>
          </string-name>
          ,
          <article-title>How active learning and process mining can act as continuous auditing catalyst</article-title>
          ,
          <source>Int. J. Accounting Inf. Systems</source>
          <volume>32</volume>
          (
          <year>2019</year>
          )
          <fpage>44</fpage>
          -
          <lpage>58</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sabt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Achemlal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bouabdallah</surname>
          </string-name>
          ,
          <article-title>Trusted execution environment: What it is, and what it is not</article-title>
          , in: 2015 IEEE TrustCom/BigDataSE/ISPA,
          <year>2015</year>
          , pp.
          <fpage>57</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>V.</given-names>
            <surname>Costan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Devadas</surname>
          </string-name>
          , Intel SGX explained,
          <source>Cryptology ePrint Archive</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Birkholz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Thaler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Richardson</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Remote ATtestation procedureS (RATS) Architecture</surname>
          </string-name>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>A. J. M. M. Weijters</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. K. Alves De Medeiros</surname>
          </string-name>
          ,
          <article-title>Process mining with the HeuristicsMiner algorithm</article-title>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Di Ciccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montali</surname>
          </string-name>
          ,
          <article-title>Declarative process specifications: Reasoning, discovery, monitoring</article-title>
          , in: Process Mining Handbook, Springer,
          <year>2022</year>
          , pp.
          <fpage>108</fpage>
          -
          <lpage>152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Freytag</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Allgaier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Burattin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Danek-Bulius</surname>
          </string-name>
          ,
          <article-title>Woped - A “proof-of-concept” platform for experimental BPM research projects</article-title>
          ,
          <source>in: BPM (Demos)</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>W.</given-names>
            <surname>Steeman</surname>
          </string-name>
          , BPI challenge
          <year>2013</year>
          , incidents,
          <year>2013</year>
          . doi:
          <volume>10</volume>
          .4121/UUID: 500573E6
          <string-name>
            <surname>-ACCC-</surname>
          </string-name>
          4B0C-
          <fpage>9576</fpage>
          -
          <lpage>AA5468B10CEE</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>F.</given-names>
            <surname>Mannhardt</surname>
          </string-name>
          , Sepsis cases - event log,
          <year>2016</year>
          . doi:
          <volume>10</volume>
          .4121/UUID:
          <fpage>915D2BFB</fpage>
          -7E84
          <string-name>
            <surname>-</surname>
          </string-name>
          49AD
          <string-name>
            <surname>-A286-DC35F063A460.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>V.</given-names>
            <surname>Goretti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Barbaro</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Di Ciccio, Trusted execution environment for decentralized process mining</article-title>
          , in: CAiSE, Springer,
          <year>2024</year>
          , pp.
          <fpage>509</fpage>
          -
          <lpage>527</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>