<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Standardizing Process Data Exploitation by means of a Process Instance Metamodel</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Antonio Cancela</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonia M. Reina Quintero</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alejandro Garc a-Garc a</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mar a Teresa Gomez-Lopez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad de Sevilla</institution>
          ,
          <addr-line>Sevilla</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <fpage>55</fpage>
      <lpage>59</lpage>
      <abstract>
        <p>ion perspective, we can realize that all the points of view share a common ground: the business process model and its instantiation are in the kernel of all of them. In this paper, we propose the use of a Business Process Instance Metamodel, which serves as a common interface to make independent the applications producing business process data from those applications that consume and exploit it. A tool has been implemented as a proof of concept to facilitate the matching between data from di erent data sources and the metamodel.</p>
      </abstract>
      <kwd-group>
        <kwd>Process Instance Metamodel</kwd>
        <kwd>Data Model</kwd>
        <kwd>Model Mapping</kwd>
        <kwd>Domain Knowledge</kwd>
        <kwd>Process Data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Activities undertaken by companies can be choreographed by business processes
to achieve their goals. Many companies use Business Process Management
Systems (BPMSs) to support the automation of business processes, but others use
software applications that are implemented ad hoc, having their own data sources
with a domain-speci c structure. As a consequence, business data cannot be
exploited in a general way and, the analysis carried out using the stored data must
also be developed ad hoc.</p>
      <p>
        Some of the most common ways of exploiting data generated during process
execution are: the creation of execution traces for process discovery algorithms to
obtain the system behavior [
        <xref ref-type="bibr" rid="ref1 ref4">1, 4</xref>
        ]; querying process data to help decision-making
in business process scenarios [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]; using ontology-based reasoning techniques [
        <xref ref-type="bibr" rid="ref3 ref5">5,
3</xref>
        ]. These di erent scenarios consume data in di erent formats, and some
formatting tasks can be tedious and complex, for example, the creation of traces
from business process executions may be a tedious and complex task [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], since
data can be stored in heterogeneous repositories.
      </p>
      <p>
        On the one hand, the need for a data representation model that let
applications that generate business data be independent from those that consume them
was detected in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and a RDF data model was proposed as an
interoperability mechanism. On the other hand, the necessity to de ne a Business Process
Instance Metamodel was introduced in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. However, none of the previous
contributions o ers the possibility of a reusable data model valid for di erent data
sources, di erent BPM techniques and even though di erent domains without
customization.
      </p>
      <p>
        In this paper, we propose the use of a Business Process Instance Metamodel
as an intermediate layer to bring closer the domain speci c data produced by
business processes and how they can be exploited. The de nition of an
intermediate metamodel helps us to abstract domain speci c knowledge from the
analyzed data, being easier to apply di erent process analytics techniques. The
approach is based on the de nition of mappings between the data source and the
expected elements in a Business Process de ned as a Business Process Instance
Metamodel. The bene ts produced by the use of an intermediate metamodel are
the reduction of the analysis time, and the exploitation of data in a more
appropriated way [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In fact, the use of the intermediate metamodel is a bene t itself,
since provides a standard way of accessing business process data and improving
the interoperability among di erent organizations.
      </p>
      <p>The paper is organized as follows: Firstly, Section 2 introduces the main
ideas of our approach and describes brie y the Business Instance Metamodel.
Secondly, Section 3 presents a tool implemented as a proof of concept to de ne
the matching between an Oracle database and the Process Instance metamodel.
And nally, the paper is concluded and some further work is pointed out.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Proposal</title>
      <p>Business process data exploitation depends highly on the technology that
supports business process execution as well as the way data is structured. Some
companies have BPMSs to support the execution, but others have software
applications that have been implemented ad hoc. As a consequence, there is no
standard approach to exploit these data. Thus, for example, the data conversion
needed to generate an event log for a process mining tool is totally di erent if
data are stored in a relational database or the data storage of a SAP system.</p>
      <p>To make data exploitation technology-agnostic, our approach is inspired by
the guidelines provided by Model-Driven Architecture to structure speci cations,
in such a way that we propose a Business Process Instance Metamodel that helps
us to separate the technological details and the structure of the information from
the data itself. In other words, the business process instance metamodel can be
seen as an intermediate artifact that is not dependent on the business domain
nor focused on a subset of tools. Thus, it allows us to make applications that
produce business process data independent from those applications that consume
them. Figure 1 depicts how the process instance metamodel acts as an interface
for both, data producers and consumers.</p>
      <p>
        The Business Process Instance Metamodel depicted in the center of Figure 1
is detailed in Figure 2. The metamodel has been speci ed with EMF [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Note
that it is a very simple model which is mainly centered on the most basic entities
related to business process instances together with their attributes.
      </p>
      <p>The root of the metamodel is the ProcessEngine metaclass and represents
the BPMS or software application that is in charge of processes execution. The
process engine can be in charge of di erent processes. The ProcessDe nition
metaclass represents the formal de nition of the process, that is, what we call
the Business Process Model. One business process can be executed many times
and the ProcessInstance metaclass models these executions or instances. A
business process is composed of di erent activities and the Activity metaclass
models them. Finally, the ActivityInstance metaclass represents the
execution of an activity and it is related to the Activity metaclass (note that an
activity may be executed many times) and to the ProcessInstance metaclass
(an activity may be executed in the context of di erent business processes). This
metamodel allows us to exploit business data in di erent contexts independently
of the storage technology and structure. We only need to de ne mappings from
the concrete technology to the Process Instance Metamodel. Then, the
information stored as a BP instance Metamodel may be used to generate event log
traces (both in XES or MXML format), to be queried for decision-making or to
be semantized and to reason about it.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Proof of Concept Implementation</title>
      <p>We have implemented a tool that assists business experts in de ning the mapping
between the data repository and the Process Instance Metamodel, facilitating
the later data analysis. The tool allows, once the data repository and metamodel
are connected by the mapping de nition, to analyze and use the data from a
Business Process point of view and exploit those data by applying any of the
techniques used in the di erent business analytics contexts. The tool has been
developed as a web application. Figure 3 shows a screen-shot that captures
the mapping de nition process. Further details about the tool can be found in
http://www.idea.us.es/portfolio-item/process-data-matching-tool/.</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and further work</title>
      <p>This paper shows how the use of an intermediate metamodel can help to
standardize the exploitation of business process data by de ning a common
infrastructure that may be used in di erent business process analytics contexts.</p>
      <p>As further work, we consider interesting to enrich the way of de ning the
matching, making more exible the tool and allowing the building of more
complex processes and exploitation of more complex data sources.
This work has been partially funded by the Ministry of Science and Technology of
Spain (TIN2015-63502-C3-2-R and TIN2016-75394-R) and the European Regional
Development Fund (ERDF/FEDER).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          :
          <article-title>Process discovery from event data: Relating models and logs through abstractions</article-title>
          .
          <source>Wiley Interdisc. Rew.: Data Mining and Knowledge Discovery</source>
          <volume>8</volume>
          (
          <issue>3</issue>
          ) (
          <year>2018</year>
          ), https://doi.org/10.1002/widm.1244
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Buijs</surname>
          </string-name>
          , J.:
          <article-title>Mapping data sources to xes in a generic way</article-title>
          .
          <source>Department of Mathematics and Computer Science</source>
          , Eindhoven University of Technology (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cairns</surname>
            ,
            <given-names>A.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ondo</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gueni</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fhima</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwarcfeld</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joubert</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khelifa</surname>
          </string-name>
          , N.:
          <article-title>Using semantic lifting for improving educational process models discovery and analysis</article-title>
          .
          <source>In: SIMPDA</source>
          . pp.
          <volume>150</volume>
          {
          <issue>161</issue>
          (
          <year>2014</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>1293</volume>
          / paper11.pdf
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Calvanese</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montali</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Syamsiyah</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          :
          <article-title>Ontologydriven extraction of event logs from relational databases</article-title>
          .
          <source>In: Business Process Management Workshops</source>
          . pp.
          <volume>140</volume>
          {
          <issue>153</issue>
          (
          <year>2015</year>
          ), http://dx.doi.org/10.1007/ 978-3-
          <fpage>319</fpage>
          -42887-1_
          <fpage>12</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Giordano</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dupre</surname>
          </string-name>
          , D.T.:
          <article-title>Enriched modeling and reasoning on business processes with ontologies and answer set programming</article-title>
          .
          <source>In: Business Process Management Forum - BPM Forum</source>
          <year>2018</year>
          , Sydney,
          <string-name>
            <surname>NSW</surname>
          </string-name>
          , Australia, September 9-
          <issue>14</issue>
          ,
          <year>2018</year>
          , Proceedings. pp.
          <volume>71</volume>
          {
          <issue>88</issue>
          (
          <year>2018</year>
          ), https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -98651-7\_
          <fpage>5</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gomez-Lopez</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reina</surname>
            <given-names>Quintero</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>A.M.</given-names>
            ,
            <surname>Parody</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Perez</surname>
          </string-name>
          <string-name>
            <surname>Alvarez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.M.</given-names>
            ,
            <surname>Reichert</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.:</surname>
          </string-name>
          <article-title>An architecture for querying business process, business process instances, and business data models</article-title>
          . In: Teniente,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Weidlich</surname>
          </string-name>
          , M. (eds.) Business Process Management Workshops. pp.
          <volume>757</volume>
          {
          <fpage>769</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Leida</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majeed</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Colombo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A lightweight rdf data model for business process analysis</article-title>
          . In: Cudre-Mauroux,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Ceravolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Gasevic</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.)
          <article-title>Data-Driven Process Discovery and Analysis</article-title>
          . pp.
          <volume>1</volume>
          {
          <fpage>23</fpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mannhardt</surname>
          </string-name>
          , F.,
          <string-name>
            <surname>de Leoni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reijers</surname>
          </string-name>
          , H.A.,
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toussaint</surname>
            ,
            <given-names>P.J.:</given-names>
          </string-name>
          <article-title>Guided process discovery - A pattern-based approach</article-title>
          .
          <source>Inf. Syst</source>
          .
          <volume>76</volume>
          ,
          <issue>1</issue>
          {
          <fpage>18</fpage>
          (
          <year>2018</year>
          ), https://doi.org/10.1016/j.is.
          <year>2018</year>
          .
          <volume>01</volume>
          .009
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Polyvyanyy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ouyang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barros</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          :
          <article-title>Process querying: Enabling business intelligence through query-based process analytics</article-title>
          .
          <source>Decision Support Systems</source>
          <volume>100</volume>
          ,
          <fpage>41</fpage>
          {
          <fpage>56</fpage>
          (
          <year>2017</year>
          ), https://doi.org/10.1016/j.dss.
          <year>2017</year>
          .
          <volume>04</volume>
          . 011
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Steinberg</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Budinsky</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paternostro</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merks</surname>
          </string-name>
          , E.: EMF:
          <article-title>Eclipse Modeling Framework 2.0</article-title>
          .
          <string-name>
            <surname>Addison-Wesley</surname>
            <given-names>Professional</given-names>
          </string-name>
          ,
          <volume>2nd</volume>
          <fpage>edn</fpage>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>