<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Social Process Mining Cockpit: A Collaboration Pat- tern Detection Tool for Enterprise Collaboration Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jonas Blatt</string-name>
          <email>jonasblatt@uni-koblenz.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrick Delfmann</string-name>
          <email>delfmann@uni-koblenz.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Just</string-name>
          <email>martinjust@uni-koblenz.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Petra Schubert</string-name>
          <email>schubert@uni-koblenz.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Koblenz</institution>
          ,
          <addr-line>Universitätsstraße 1, Koblenz</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Social Process Mining (SPM) combines the research fields Process Mining (PM) and Social Collaboration Analytics. The SPM-Cockpit is a prototype that can detect and analyze Collaboration Patterns in event logs of Enterprise Collaboration Systems (ECS) semi-automatically. We use this tool to investigate activity patterns that occur while users work collaboratively in ECS. To do so, we adapt methods from Process Mining, Frequent Subgraph Mining, and Graph Clustering and bundle them in the SPM-Cockpit. The prototype is a first step towards investigating, understanding and categorizing collaboration activity in ECS automatically.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Social Process Mining</kwd>
        <kwd>Collaboration Pattern</kwd>
        <kwd>Enterprise Collaboration Systems</kwd>
        <kwd>Process Mining</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In recent years, many workers have shifted to a remote workplace [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. At least the COVID-19
pandemic led companies to invest in collaborative software to stay competitive in such a remote work
environment [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Employees had to work together in an efficient manner so that they still fulfilled
their workload. Enterprise Collaboration Systems (ECS) provide functionality, such as wikis,
forums, file sharing, or blog components, that improves collaborative work in an enterprise
environment. Such components encourage a flexible digital workplace where collaboration is
essential to the daily routine. In general, in an ECS, users work with social documents (e. g., wikis
articles, blog entries, forums threads, or text files) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Process Mining (PM) is an established research domain that analyzes event logs extracted
from software systems [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Such event logs contain records that describe what activities happened
at what time in the system during a particular process instance [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Additional event log
information may describe the process context, for instance, the user who executed the activities, data
used as in- or outputs for activities, and further environmental information. PM techniques use
such event logs to discover process models, check the conformance of processes with prescribed
behavior, or provide means for process enhancement [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. As ECS also log event data, we can apply
PM in principle. However, due to the inherent malleability of ECS, users use such systems in an
ad-hoc manner, fulfill their tasks in an unstructured order, and are not bound to a governing
process [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For instance, a user may first create a blog post and then update a wiki page, or the user
does this the other way around. Thus, logs can be unstructured and complex, and applying PM
algorithms to native ECS event data may result in so-called spaghetti models, which are hard to
interpret, even for domain experts.
      </p>
      <p>The SPM-Cockpit aims to detect patterns of collaborative activities of users in ECS. SPM is the
acronym for Social Process Mining, which combines the research fields of Process Mining and
Social Collaboration Analytics in ECS. As there are many steps and several methods involved in the
process of SPM, we need a tool that bundles these methods and guides the users in their
application to ease access. Thus, the SPM-Cockpit provides tailored methods for preprocessing log data,
for the discovery of process models, frequent subgraph mining and graph set clustering
algorithms for the detection of collaboration patterns, and for creating collaboration pattern
repositories (explained in the following section). Furthermore, the cockpit guides the users through the
SPM process. The outcomes of the SPM-Cockpit are collections of collaboration patterns (CPs) that
help to describe collaboration. CPs are subsections of process models that can be found frequently
in process models mined from ECS logs. The primary target user for the SPM-Cockpit is an ECS
analyst or a researcher who aims to improve the collaborative work processes of ECS. Thus, s/he
can use the cockpit to mine CPs from ECS process models, which describe the collaborative
interactions of different users with social documents. These outcomes illustrate the typical
collaboration activities in ECS. Eventually, the CPs can be used as building blocks that can be embedded in
Business Process Management environments to improve collaborative business processes and
ECS.</p>
      <p>In the remainder, we describe the SPM-Cockpit and its features in Section 2. Section 3
describes the maturity of the tool. Finally, Section 4 concludes the demo with an outlook to future
work with further feasible improvements.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Tool Description &amp; Features</title>
      <p>
        The SPM-Cockpit is a web application that guides the user through the process of SPM. It is built
with the Django web-framework2 and uses the following complementary technologies: a
relational database (MariaDB®) for storing configurations and user sessions, a Redis® database3 for
caching traces, a Neo4j®4 graph database for storing the collaboration pattern repositories, and
Celery5 for distributing computationally expensive tasks with task queues. In addition, the
application uses the process mining library pm4py [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] for handling event logs and applying the process
discovery algorithms. We deploy the presented main components within docker containers.
Figure 1 shows the basic architecture of the application with its containers and the data flow
between them.
      </p>
      <p>Cache</p>
      <p>Message
Broker</p>
      <p>Celery
 Messages
Cache</p>
      <sec id="sec-2-1">
        <title>Celery runner</title>
        <p>DB access
DB access
application</p>
      </sec>
      <sec id="sec-2-2">
        <title>SPM-Cockpit</title>
        <p>DB access
DB access</p>
      </sec>
      <sec id="sec-2-3">
        <title>Docker container</title>
      </sec>
      <sec id="sec-2-4">
        <title>Data Flow</title>
        <p>The BPMN model (Figure 2) shows the underlying process of the application with the SPM
functionalities as activities and steps. If the results of an intermediate step do not satisfy the
expected outcome, the user can step back and reconfigure the parameters and redo the
corresponding activity (e. g., redo the process mining discovery with another algorithm). The SPM-Cockpit
divides the proposed analysis into six main activities: 1. Importing an Event Log, 2. Preprocessing
&amp; Process Mining Configuration, 3. Exploring Process Models, 4. Frequent Subgraph Mining,
2 https://www.djangoproject.com (last access: 24th of July, 2023)
3 https://redis.io (last access: 24th of July, 2023)
4 https://neo4j.com (last access: 24th of July, 2023)
5 https://docs.celeryq.dev/en/stable (last access: 24th of July, 2023)
it
-ckopC
PSM</p>
        <p>ECSEXvEeSntLog</p>
        <p>Importingan
EventLog</p>
        <p>PPrCreoopcnreofiscgseusrMasitinnioignng&amp; ProEcexspsloMrinogdels</p>
        <p>RedisCache</p>
        <p>ProceosksayM?odels
no</p>
        <p>GraphDatabase
Frequent Subgraphsokay?
Subgraph
Mining
no
GraphSet
Clustering</p>
        <p>Clustersokay?
no</p>
        <p>Managing
Colaboration
Patern
Repository
5. Graph Set Clustering, and 6. Managing Collaboration Pattern Repositories. The entry point of
the SPM-Cockpit is the landing page with an overview of these SPM steps (Figure 3a). The
following paragraphs explain each activity and describe their purpose, functions, and in- and outputs.</p>
        <p>
          The first step is Importing an Event Log. ECS event logs have some special properties that are
relevant for the further SPM process. We assume that the baseline case (the process instance) for
such an event log is based on the so-called workspace (i. e., a particular working environment in
an ECS that was created, for instance, for a project). For example, all events that were fired in one
workspace belong to one trace in the event log. We further require the events to have the
following attributes: the associated social document, the executing user (resource), and the timestamp
when the event was fired. As eXtensible Event Stream (XES) [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] is the standard for event logs in
the PM domain, we require this as input format.
        </p>
        <p>
          After the upload is finished, the tool starts the second activity (preprocessing) with the default
parameters. Thus, the second step, Preprocessing &amp; Process Mining Configuration, triggers the
preprocessing phase and defines how the process models should be mined. To detect which persons
are working together on which social documents, the traces of the workspaces are further split
into sub-traces. The intent is that we split each trace (the workspaces) into smaller traces based
on those social documents that are jointly worked on by several users. Then, for each sub-trace,
we identify ECS users who have worked in the same timeframe (of the trace) on the same
documents and group those traces with direct or indirect overlaps in common users and timeframes.
Thus, we create multiple sub-logs (the groups of sub-traces) based on the collaborative activities
of the ECS users. Then, for each of these sub-logs, we can mine process models that express the
collaboration. The process models are generated based on the PM configuration. For now, the
cockpit user can decide whether to discover Directly Follows Graphs [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] or to discover process
models using the Heuristic Miner [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. However, further algorithms can be easily included in future
versions of the SPM-Cockpit. The cockpit user can view the resulting process models in the next
step.
        </p>
        <p>The intermediate step, Exploring Process Models, lets the cockpit user investigate whether the
generated process models are sufficient for further analysis. Therefore, the SPM-Cockpit provides
visualizations of the generated process models with basic navigation, pan and zoom functionality.
The user can inspect the process models and step back if s/he has concerns regarding the PM
configuration.</p>
        <p>
          The Frequent Subgraph Mining step uses the previously generated process models and
searches for patterns using Frequent Subgraph Mining (FSM) [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The idea is that through the
discovery of frequent subgraphs, we identify patterns, which occur regularly when users work
collaboratively on social documents in ECS. The result is a set of collaboration patterns. The
current implementation of the FSM algorithm is gSpan [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], which searches for frequent
substructures in graphs by building a depth-first search tree while adding forward and backward edges.
This algorithm needs three parameters: The min and max number of vertices determine the min
and max number of nodes (i. e., process activities in our context) and the support parameter
defines how often the respective subgraph should occur to be present in the result set. The user
defines these parameters and starts the FSM algorithm. Afterwards, s/he can inspect the resulting
subgraphs and can step back if the subgraphs are too small, or the number of results is too small
or too high. If this is the case, s/he can re-run the FSM algorithm with reconfigured parameters.
        </p>
        <p>During the Graph Set Clustering step, similar patterns (subgraphs) are clustered into groups.
Such groups then represent similar collaboration patterns rather than identical ones. The reason
is that we assume that a typical collaboration scenario does not always look the same but rather
similar. These groups/clusters can be viewed in the next step in the Collaboration Pattern
Repository. The user can choose a hierarchical clustering algorithm or a partitioning clustering
algorithm. Where the first requires the analyst to define the number of clusters, the latter finds the
number of clusters based on the dissimilarities of the given graph set6. Again, based on the results,
the user can redo the clustering process with a new configuration.</p>
        <p>In the last step, Managing Collaboration Pattern Repository, the user can investigate the
resulting clusters, which we store in a Collaboration Pattern Repository (CPR). Each cluster consists of
a set of patterns (subgraphs) and can be visualized with a representative graph (e. g., the initial
graphs from the clustering step). Additionally, the user can add a name and a description for the
respective cluster as they represent similar collaboration patterns. The analyst can now use these
patterns for further analysis that can provide insights into how the collaboration was conducted
in the ECS. An example of two found patterns is illustrated in Figure 3b. These show the
collaborative activities of users working on a wiki page and a file. As the CPRs are stored in a graph
database (Neo4j), other existing CPRs can be loaded and inspected as well. Furthermore, the
SPMCockpit provides the functionality for down-/uploading the CPRs.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Maturity of the Tool</title>
      <p>We deployed one instance of the SPM-Cockpit7 on a server. Furthermore, we published the
related user documentation of the SPM-Cockpit8 and the technical documentation of the
underlining Python library9. We provide some ECS event logs, which one can select in the first step as
sample data (instead of uploading an event log) directly in the SPM-Cockpit. We successfully
applied the SPM-Cockpit with these event logs, which we demonstrate in the SPM-Cockpit
documentation as a walkthrough video tutorial10.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion and Future Work</title>
      <p>The SPM-Cockpit is a tool for the analysis of ECS event data and provides a bundled toolset for
the discovery and investigation of collaboration patterns. Thus, this tool extends the analysis of
6 Additional clustering parameters can be defined by the user, but we refer here to the documentation.
7 https://w3id.org/spm/cockpit (last access: 24th of July, 2023)
8 https://w3id.org/spm/docs/cockpit (last access: 24th of July, 2023)
9 https://w3id.org/spm/docs/colpadef (last access: 24th of July, 2023)
10 https://w3id.org/spm/docs/cockpit/tutorial &amp; https://youtu.be/4PYqSRKEPDw (last access: 24th of July, 2023)
unstructured processes in (social) document-orientated systems. For this purpose, we apply
algorithms from the PM domain, graph-based algorithms, and clustering algorithms. For future
improvements, we aim to:
• improve the preprocessing by including a parameter for temporal overlaps of the sub-traces to
refine the sub-log generation.
• include further process discovery algorithms.
• extend subgraph discovery by implementing a relaxed Frequent Subgraph Mining algorithm
that also finds subgraphs that are similar to the search pattern rather than being isomorphic.
This way, similar subgraphs that would normally have a too low support in a “classic” FSM
will not be missed.
• implement and evaluate further clustering algorithms to find the best one for SPM.
•
apply existing CPRs to analyze unseen ECS event logs, while calculating metrics based on
the patterns and the unexplored event log.</p>
      <p>While we focus on ECS event data in SPM, further systems logs may be investigated with the
help of the SPM-Cockpit in future research.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work was partly funded by the Deutsche Forschungsgemeinschaft (DFG, German Research
Foundation) – project number 445182359.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Williams</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Grams</surname>
          </string-name>
          , “
          <source>Remote Working Study</source>
          <year>2022</year>
          ,”,
          <year>2022</year>
          , [Online]. Available: https://kola.opus.hbz-nrw.de/frontdoor/index/index/docId/2310
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Gartner</surname>
          </string-name>
          , “
          <source>Gartner Forecasts Worldwide Social Software and Collaboration Market to Grow</source>
          <volume>17</volume>
          % in
          <year>2021</year>
          .”
          <year>2021</year>
          . [Online]. Available: https://www.gartner.com/en/newsroom/pressreleases/2021-03-23
          <article-title>-gartner-forecasts-worldwide-social-software-and-collaboration-market-to-</article-title>
          <string-name>
            <surname>grow-</surname>
          </string-name>
          17
          <string-name>
            <surname>-</surname>
          </string-name>
          percent-in-2021
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mosen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Schubert</surname>
          </string-name>
          , “The Structure of Social Documents,”
          <source>in Proceedings of the Annual HICSS</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>2825</fpage>
          -
          <lpage>2834</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>W. van der Aalst</surname>
          </string-name>
          et al., “Process Mining Manifesto,” in International conference on business process management,
          <year>2012</year>
          , pp.
          <fpage>169</fpage>
          -
          <lpage>194</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>J. De Weerdt and M. T. Wynn</surname>
          </string-name>
          , “Foundations of Process Event Data,” in Process Mining Handbook, Springer Science and Business Media Deutschland GmbH,
          <year>2022</year>
          , pp.
          <fpage>193</fpage>
          -
          <lpage>211</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Berti</surname>
          </string-name>
          and
          <string-name>
            <surname>S. J. van Zelst</surname>
          </string-name>
          , “
          <article-title>Process Mining for Python (PM4Py): Bridging the Gap Between Process-</article-title>
          and
          <string-name>
            <surname>Data Science</surname>
          </string-name>
          ,”
          <source>in Proceedings of the ICPM Demo Track</source>
          <year>2019</year>
          , Aachen, Germany,
          <year>2019</year>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>IEEE</given-names>
            <surname>Computational Intelligence</surname>
          </string-name>
          <string-name>
            <surname>Society</surname>
          </string-name>
          ,
          <article-title>Standard for eXtensible Event Stream (XES) for Achieving Interoperability in Event Logs</article-title>
          and
          <string-name>
            <given-names>Event</given-names>
            <surname>Streams</surname>
          </string-name>
          .
          <year>2016</year>
          . [Online]. Available: https://ieeexplore.ieee.org/servlet/opac?punumber=
          <fpage>7740856</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>W. M. P. van der Aalst</surname>
          </string-name>
          , Ed., “Foundations of Process Discovery,” in Process Mining Handbook,
          <source>in Lecture Notes in Business Information Processing</source>
          , vol.
          <volume>448</volume>
          . Springer International Publishing,
          <year>2022</year>
          , pp.
          <fpage>37</fpage>
          -
          <lpage>75</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Weijters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. M. van Der</given-names>
            <surname>Aalst</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A</given-names>
            .
            <surname>A. De Medeiros</surname>
          </string-name>
          , “
          <article-title>Process mining with the heuristics miner-algorithm,”</article-title>
          <source>Tech. Univ. Eindh. Tech Rep WP</source>
          , vol.
          <volume>166</volume>
          , no.
          <source>July</source>
          <year>2017</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>34</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Coenen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Zito</surname>
          </string-name>
          , “
          <article-title>A survey of frequent subgraph mining algorithms</article-title>
          ,” Knowl. Eng. Rev., vol.
          <volume>28</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>75</fpage>
          -
          <lpage>105</lpage>
          ,
          <year>2013</year>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>X.</given-names>
            <surname>Yan</surname>
          </string-name>
          and J. Han, “
          <article-title>gSpan: graph-based substructure pattern mining</article-title>
          ,” in
          <source>2002 IEEE International Conference on Data Mining</source>
          ,
          <year>2002</year>
          . Proceedings.,
          <source>IEEE Comput. Soc</source>
          , pp.
          <fpage>721</fpage>
          -
          <lpage>724</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>