<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automated Process Model Annotation Support: Building Blocks and Parameters</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michael Fellmann</string-name>
          <email>michael.fellmann@uni-rostock.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felix Oehmgen</string-name>
          <email>felix.oehmgen@uni-rostock.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Computer Science, University of Rostock</institution>
          ,
          <addr-line>Rostock</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In business process modeling, semi-formal models typically rely on natural language used to express the labels of model elements. This can easily lead to ambiguities and misinterpretations. To mitigate this issue, the combination of process models with formal ontologies or predefined vocabularies has often been suggested. A cornerstone of such suggestions is to annotate elements from process models with ontologies or predefined vocabularies. Although annotation is suggested in such works, past and current approaches rarely discuss building blocks, parameters and strategies for automating the tedious and errorprone manual task. In this paper, we hence first describe the nature of the annotation task. We then identify building blocks and parameters for automated systems and describe an implementation of an annotation system we used to conduct first empirical studies on the effect of parameters. The paper at hand in sum presents design options and parameters for (semi-) automatically linking semi-formal process models with more formal knowledge representations. It hence may be a source of inspiration for further explorations and experiments on that topic.</p>
      </abstract>
      <kwd-group>
        <kwd>Business Process</kwd>
        <kwd>Semantic Annotation</kwd>
        <kwd>Automatic Matching</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        In business process modeling, semi-formal modeling languages such as BPMN are used
to specify which activities occur in which order within business processes. Whereas the
order of the activities is specified using constructs of the respective modeling language,
the individual semantics of a model element such as “Check order” is bound to natural
language. However, if models have to be interpreted by machines, e.g. for offering
modeling support, querying on a semantic level [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] or content analysis, a more formal,
machine processable semantics of modeling elements is required [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. More use cases
that would be possible if an automated annotation could be realized are described in
more detail in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. In the past, several approaches tried to formalize the semantics of
individual model elements by annotating elements of ontologies or other predefined
vocabularies that to some degree formally specify the semantics of a model element.
However, such approaches suffer from a major limitation: Annotation is a highly
manual and tedious task. The user has to select suitable elements of an ontology by browsing
the ontology or doing a keyword-based search in the labels of the ontology. Even if the
system is capable of presenting some annotation suggestions, e.g. based on lexical
similarity of labels, the user has to make sure that annotations match the appropriate context
in the process model by inspecting the structure of the ontology that typically is
organized in a hierarchy. For example, if the ontology contains two activities labelled with
“Accept invitation”, it is important whether this activity is part of the hiring process
(where the applicant accepts e.g. a job interview) or the planning process for business
trips (where the employee accepts an invitation of a business partner). In other words,
the semantic context of an element that is to be annotated must be considered. Since
only a very limited number of highly automated context-sensitive approaches for
process model annotation is available so far (see [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] for an overview on current and past
annotation approaches, [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] for an implementation using Markov Logic), this
contribution is meant to facilitate developing, comparing and optimizing such approaches. To
bootstrap systematic research in this direction, we describe building blocks and
parameters (in short: design options) for automated annotation. With this, interest in a very
promising research topic should be raised; both in regard to scientific outcome as well
as practical usefulness (for use cases, see e.g. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]).
      </p>
      <p>The remainder is structured as follows. In Section 2, the annotation task is described
and three major building blocks for semantic annotation are identified. In Section 3,
these building blocks along with their parameters are described in more detail. In
Section 4, first considerations and results for/of an empirical analysis are given. In Section
5, related work is discussed and in Section 6 the article is concluded.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Description of the Annotation Task</title>
      <sec id="sec-2-1">
        <title>2.1 Fundamental Characteristics of the Annotation Task</title>
        <p>
          Semantic annotation as investigated in this paper means linking process model tasks
(e.g. a task such as “Check order”) with elements of an ontology or vocabulary such as
“Order checking”). We denote these elements as “concepts”. In regard to the
characteristics of the ontology or vocabulary used for annotation, we assume that it is structured
in a hierarchical way, that semantics of the hierarchy is “part-of” and that there is a
partial ordering between siblings in the hierarchy. This assumption seems to be justified
when considering major examples of vocabularies or ontologies such as the PCF
(Process Classification Framework), a publicly available collection of approx. thousand
enterprise activities which is also available industry-specific versions [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Another
example is the MIT Process Handbook [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], a large collection of enterprise knowledge
integrated into an ontology where activities are also ordered in a part-of-hierarchy.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Deriving Building Blocks for IT-Support by Observing Human Annotators</title>
        <p>In order to understand which building blocks are required for an automated annotation
approach, it is helpful to observe and interview human annotators about their strategy.
We did so by observing and interviewing students who manually annotated business
process models as a part of a tutorial. Process models were specified in the BPMN
language and annotated with elements of the PCF (Process Classification Framework)
taxonomy. 50 undergraduate students with good knowledge in process modelling
participated in small groups in the exercise in the years 2012-2014 and annotated 23
models in a group effort. Since this empirical work is not in the center of the article at hand,
we only roughly report the insights we gained. A recurring pattern that has been
observed both directly and by interviewing the students has been that annotation roughly
followed a 3-step procedure: First, keyword search was performed to search for
relevant elements of the PCF taxonomy. Second, in case that multiple relevant elements of
the taxonomy were found, the context of these elements was considered and items of
the taxonomy were preferred that better correspond to the overall topic of the process.
For example, if the topic of the process was Human Resources (HR), participants
preferred activities belonging to the category “6. Human Resources” of the PCF taxonomy.
Third, in a last step, the selection of an item for annotation was reviewed considering
the annotation of preceding and following model elements to verify that it is meaningful
and fits the process context. In this step, the partial ordering of the activity taxonomy
was taken into consideration meaning that if activities in the taxonomy appeared to
occur in a meaningful order (e.g. check order, approve order, execute order),
participants strived to not violate that order in the annotations. In this step, also activities that
are on a similar hierarchy level (i.e. that are not more specific or detailed) than those
selected for the surrounding model elements have been preferred, if possible. In sum,
roughly three steps were executed: (1) retrieve annotation candidates by lexical
matching, (2) put annotation candidates into context and select the most meaningful and (3)
optimize annotation in regard to the annotations of surrounding model elements in
terms of order and hierarchy level. These three steps inspire corresponding building
blocks of an automated annotation approach which we refer to as element annotation,
context detection and annotation fitting. They are described in the following along with
adjustment parameters.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 Building Blocks and Parameters</title>
      <sec id="sec-3-1">
        <title>3.1 Element Annotation</title>
        <p>For annotating process model activities, relevant activity concepts in the taxonomy
have to be found. It is thus necessary to match model element labels against activity
concepts from the vocabulary as it is illustrated in Fig. 1. To match process labels
against vocabulary concepts, we basically need a similarity function   () that
returns the similarity between a process activity  ∈  and an activity concept  ∈ 
between 0 and 1.</p>
        <p>
          ( ,  ) ∈ [
          <xref ref-type="bibr" rid="ref1">0,1</xref>
          ]
(1)
        </p>
        <p>
          Using this function, a set of annotation candidates  (for “metadata”) can be
computed containing process elements  that match to vocabulary concepts  with a
matching value  ∈ [
          <xref ref-type="bibr" rid="ref1">0,1</xref>
          ] being above a similarity threshold  ℎ  and that occurs between
a minimum level  ℎ (to exclude root node) and maximum level  ℎ (to prevent
too fine-grained annotations) hierarchical position in the taxonomy. The hierarchical
position for a concept  ∈  is given by the function ℎ( ).
        </p>
        <p>= { ( ,  ,  ) |  ∈  ∧  ∈  ∧  ≥  ℎ  
∧  ℎ
≤ ℎ( ) ≤  ℎ
}</p>
        <p>(2)
Confirm order to
customer
Send notification about</p>
        <p>order acceptance</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Identification of Context</title>
        <p>be assigned to a category):
If the business topic such as e.g. “Human Resources” of a process model is known, then
this knowledge can be leveraged to improve the annotation result. To do so, it could be
used to discriminate between activity concepts with a comparable lexical matching
value that are candidates for annotation. Hence it is required to detect the general topic
of a model which we call category in the following. A category  ∈  (for “domain”)
can be interpreted as activity concepts that are sub-concepts of the taxonomy root, i.e.
 = { | ∈  ⋀( , 
) ∈  } with</p>
        <p>being the set of hierarchy relations of the
concepts in the taxonomy. In the simplest form, a category may be specified for the whole
model by the user. If that is not possible, a category for the whole model may be derived
in an automated way. However, unfortunately there may be models with multiple
categories (i.e. multiple topics in one model such as HR and financial planning) and hence
it is not clear which category is dominating the model. Such an example is illustrated
by Fig. 2. In order to cope with the possibility of multi-category models, the model
needs to be partitioned into fragments  ∈  containing subsets of activities  ⊆ 
referring to the same category (with a default-subset   for parts of the model that cannot

 = ∪ =1   ∪  
where   ∩   = ∅ ,  ≠ 
(3)</p>
        <p>The function  ( ) returns the category  ∈  for a given fragment  and function
 ( ) returns the corresponding fragment  ∈  for a given activity  ∈  . Each
fragment is associated to exactly one category, i.e. ∀ ∈  ∶ ∃ ∈ 
∧ | ( )| = 1.
Likewise, each activity should be contained in exactly one fragment, i.e. ∀ ∈  ∶ ∃ ∈
 ∧ | ( )| = 1.</p>
        <p>About
"Human Resources"
About
"Financial Planning"</p>
        <p>The task of partitioning the model into fragments that are associated to a category is
dependent on two parameters. First, a minimum size   of a fragment has to be
specified in order to prevent partitioning the model in activity fragments containing single
activities and hence losing the usefulness of categories to discriminate between
candidate activity concepts. Second, a lower threshold value for the minimum lexical
matching value   has to be set specifying the minimum average lexical similarity value
between all activities contained in a fragment and sub-concepts in the category  ∈ 
that may be assigned to the fragment. Hence a splitting function  () takes the
process model and these two parameters as input and generates a set  of process fragments
as output.</p>
        <p>In order to use categories associated to fragments to augment element annotation,
the implementation of the similarity function   (a, c) has to be extended. Such an
extended similarity function will give a higher similarity value if the category given by
 ( ( )), i.e. the category associated to the fragment an activity belongs to, matches
the category of  . In order to detect the latter, a function  ′( ) which returns the
category  ∈  for a given activity concept  ∈  is defined. In order to control the
influence of category matches, i.e. if  ( ( )) =  ′( ), a weight   is added to the refined
similarity function   ′ ().</p>
        <p>
          ′ ( ,  ,  
) ∈ [
          <xref ref-type="bibr" rid="ref1">0,1</xref>
          ]
(4)
(5)
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3 Annotation Fitting</title>
        <p>In order to increase the semantic quality of the annotation, annotations may be “fitted”.
This in essence means to choose between activities that receive comparable matching
values in the element annotation step according to notion of betweeness and differences
in the hierarchy level. Fig. 3 illustrates this with a small example of three activities that
are symbolized on the left side. The discrimination problem is considered for the middle
activity. Comparably well matching activity concepts are depicted as grey shaded small
circles on the right side. According to the notion of betweeness, the bottom most
activity concept c3 can be neglected since it is not in the area of preferred annotation
candidates (surrounded by a dotted line). This is due to the fact that it is not between the
already selected best matching activity concepts for the previous and following process
activity which are illustrated as solid black filled circles. Further, according to the
principle of preferring a similar hierarchical level, c2 can also be skipped. Hence amongst
similar lexical matches, c1 is superior to c2 and c3 and hence c1 is selected for annotation.
?
c2 ?</p>
        <p>Area of preferred
annotation candidates
c1 ?
c3 ?</p>
        <p>The procedure introduced so far to select activity concepts for annotation roughly
resemble to interpolation routines in common image manipulation software where the
color of a pixel is calculated according to its neighbors (e.g. Gaussian Filter). We stick
to that analogy and call this procedure of fitting semantic interpolation. Analogously,
the radius of interpolation may be more than just one pixel which in our case is one
preceding and one following activity. The radius may be extended to all preceding and
following elements which in literature is also referred to as the corona of a process
model element. In this way, the size of the corona may range from 1 to n, i.e. from all
preceding and following activities reachable via one step or n steps in the process graph.
Beyond the radius parameter for semantic interpolation, it is important that a fitting
function is able to adjust the influence of the lexical matching in relation to the
influence of hierarchy. Taking this into account, we define a function  () as follows:
,  
) =  
(6)
where  are the annotation candidates (cf. formula 2),  is the radius used in
semantic interpolation,   the weight of the lexical matching result and   is the
weight of the hierarchy match. The latter refers to how the difference of the hierarchy
levels of two activity concepts  ∈  given by |ℎ(  ) − ℎ(  )| with  ≠  affects the
semantic interpolation. The function produces a fitted annotation set   ⊆  with
just one annotation per process activity, i.e. | | = |  |.
3.4</p>
      </sec>
      <sec id="sec-3-4">
        <title>Overview of Possible Configurations</title>
        <p>The building blocks specified in in the previous sections may be combined leading to
different configurations of the overall automatic annotation system. The following
Table 1 describes these configurations. The first column provides a number for each
configuration. The next three columns indicate if the building blocks (B1-B3) are used that
are specified in the previous sections. The next column Configuration Description
provides a short name (in bold) and description of this configuration variant. The next
column Configuration Parameters provides a list of accumulated configuration
parameters resulting from the different building blocks described in the previous sections.
The configurations and parameters described in Table 1 may be used in the
development, comparison and optimization of different implementation strategies and hence
support a systematically evaluation of automated annotation approaches.</p>
        <p>Configuration Parameters
–  ℎ 
–  ℎ
–  ℎ
similarity threshold
min. hierarchy level of activity concepts
max. hierar. level of activity concepts
Parameters of configuration variant 1 plus the
following additional parameters:
minimum size of a fragment
minimum average lexical similarity value
weight of category matches
Parameters of configuration variant 1 plus the
following additional parameters:
To gain first insights regarding the implementation of the building blocks introduced
so far, a project was set up to create a simple algorithm. The goal of this algorithm is to
match labels of process model activities to a similar or equal counterpart in a
standardized framework. The standardized annotations are provided by the Process
Classification Framework (PCF). It consists of twelve main hierarchies which are structured into
sub-hierarchies of four levels. The hierarchies are in order of increased detail: Process
category, process group, process and activity. The first five Process Categories contain
operating process while the other describe management and support processes. The
models were created by students in a non-related task. The labels were manually
matched to the best corresponding concept in PCF thus creating a gold standard. This
standard is the basis on testing the accuracy of the algorithm and offers valuable
information by analyzing the characteristics of correct matches.</p>
      </sec>
      <sec id="sec-3-5">
        <title>4.2 Simple Algorithm Structure</title>
        <p>The algorithm is designed to apply a sequence of techniques that each individually
affect the outcome. The workflow below shows each step the algorithm takes in order to
find the best match for a label. The round-edged rectangles represent techniques. The
sharp-edged rectangles indicate the resource data as well as the resulting outcome.</p>
        <p>In the following, we describe how the implementation reflect the three building
blocks introduced in the previous sections.</p>
        <p>Element annotation. The process starts with the resources holding the information
schemas of in this case a model and the PCF. For convenience the labels of the
individual model elements are called activity labels and the annotations in the PCF are called
PCF elements. Each step aims at deriving information about the activity label on
different levels. The first method compares solely the characters of two strings. The result
is an editing distance showing lexical similarity. In this algorithm, the method used is
called the Sorensen-Dice Coefficient. The result of this comparison between activity
labels and PCF elements are cross products stored in a similarity matrix. This matrix is
the core of the algorithm holding a similarity value for all possible pairs thus enabling
an analysis for the best match. As stated in Section 3.1, the scores range between 0 and
1 with 1 being a 100% match. The following steps aim at modifying this value to single
out the best match. We did not restrict the similarity threshold  ℎ  , the min.
hierarchy level  ℎ was set to 1 and the max hierarchy level  ℎ was set to unbounded.</p>
        <p>Context detection. To identify the context of a model helps to narrow down the list
of possible matches according to their domain. For this analysis the theory in Section
3.2 is taken and applied on the test case, although at the moment we identify the context
of a process model only globally, i.e. not partitioning the model in fragments
(parameters   and   hence are irrelevant). The PCF is already classified into 12 domain
specific hierarchies with a total of 4 layers. The hierarchy level process group (layer 2)
thereby serves as the reference hierarchy. To match an activity label to the correct
process group means in this case that the results are narrowed down to a list of max 60
PCF elements. The method uses information extracted by an external tool to derive
certain words from the activity labels called synonyms as well as from the process
groups called keywords which indicate domain affiliations. Keywords are words that
show representative value for a process group in the PCF list. A word suits a keyword
if it is present more than 3 times in a sub hierarchy (Fig. 9). This application roughly
corresponds to building block 2 from Section 3.4. The matching process compares
words by lexical matching. Semantic matching does not take place yet.</p>
        <p>The same procedure takes place to create the synonyms list. The synonyms are
generated by analyzing the activity labels for representative words. Since a gold standard
is provided, a list of representative words on the activity label side towards process
groups can be extracted (Fig. 10). The further calculation is the same as with the
keywords. The synonyms list is however furthermore in need of constant updating for
application on other models. The list created in this case is solely based on the gold
standard and still requires verification for other models. Both lists are a preliminary result
and are stored as a reference table. Matching activity labels to this list enables a
contextrelated comparison and highlights matches on a more abstract level. Controlling the
influence of the these matches corresponds to parameter   .</p>
        <p>Annotation fitting. The last step implies on a theoretical level (cf. Section 3.2)
analyzing the predecessors and successors (i.e. the semantic context) of an activity label
in the respective model. The content of models can be affiliated to certain process
groups. The aim of this step is to punish and reward similarity scores by manipulating
the overall score of a model in a process group. This behavior encourages scores of
certain process groups to increase more in order to narrow down the highest similarity
scores onto few process groups. For instance this step should increase the scores in
process group 2.1 and 2.2 for model “1” above all other process groups to reward
similarity scores in this sub-matrix (Fig. 8). This pattern of models showing an affinity to
certain process groups was discovered by analyzing the gold standard but the algorithm
further assumes that the results of the previous steps create the same pattern. The
algorithm therefore manipulates the calculated scores. In terms of parameters, radius  is
unbounded and weights   and   are implicitly set to 0.</p>
        <p>After all methods are processed, an analysis is run to extract the best matches from
the similarity matrix with the highest similarity score. Thus two scores are calculated
showing the percentage of correct activity label to PCF element match and the
percentage of correct activity label to process group match (for the sake of brevity, process
group is also simply called hierarchy in Fig. 5-10).</p>
      </sec>
      <sec id="sec-3-6">
        <title>4.3 Test and Results</title>
        <p>The basic function of the algorithm is to create a matrix containing measures of
similarity ranging between 0 and 1 across all possible matches and to manipulate these
measures following each step. Each step can be adjusted on the amount of influence on
the similarity measure. The testing phase was conducted in order to determine the
effectiveness of each step at varying influence. The test set includes 430 annotations
derived of 33 independent models. The base result for comparison consists solely on
lexical matching. The results of the test is displayed in three graphs (Fig. 5, 6, 7) found in
the appendix. In total synonym matching achieved the best results and has a big effect
on matching a label to the correct PCF element. It manages to increase the percentage
of correct matches from 12% to 18% for direct matches and from 25% to 57% for
hierarchy matches, i.e. where at least the process group (level 2) is detected correctly.
The other steps show little to no improvement.</p>
        <p>The analysis and test are overall a first try at matching model labels. A test of the
algorithm with combined influence of all steps is still missing and keyword and
synonym matching step is heavily based on the characteristics of the gold standard.
Furthermore steps like semantic matching that have not been implemented yet. The analysis
could however confirm that semantic patterns such as the overall topic of a process
model can be detected and that the PCF can in fact be leveraged as a valuable standard
framework. The analysis moreover showed promising results concerning the creation
and application of a synonym list.</p>
        <p>
          Comparing our results to the method proposed by Leopold et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], this method
achieves about 76% correctness for annotating all activities with a concept from the
correct main category and 44% of correctness for the process group level. In the light
of these numbers (especially correctness on the process group level for which we
achieve 57%), our approach seems promising. However, it requires a manually created
synonym list. A detailed comparison is left open for future work.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5 Related Work</title>
      <p>
        Most approaches developed so far only suggest manual annotation (cf. an overview
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]). For example, [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] describe a mapping relation of a BPEL4WS process to an
OWLS ontology as well as relations between concepts from the OWL-S profile ontology to
domain ontologies; [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5–7</xref>
        ] develops an approach for adding properties of model
elements or establishing relations to separate annotation models. A model for semantically
annotating business process models is devised in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. There are however some works
considering the annotation support by tools. For example, Bögl et al [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] describe a
semantic linkage of Event-driven Process Chain (EPC) functions and events to ontology
instances supported by a lexicon (WordNet), term extraction and stemming. Similarly,
annotation approaches for BPMN models with ontologies have been developed [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10–
12</xref>
        ] and partly supported using various lexical analysis techniques. Also, the annotation
of process models with other domain specific ontologies such as the SCOR model for
supply chain management has been explored [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] as well as annotations of process
models with goal models [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        However, the only approach that we are aware that considers context information
(e.g. in the form of preceding or following annotations) when calculating an annotation
suggestion is that of Leopold et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. The approach makes use of a Markov
Logicbased formalization and considers automated annotation as an optimization problem.
Further, in the field of execution-level (i.e. runtime) processes, the structure and
lifecycle of involved objects in the process is considered [
        <xref ref-type="bibr" rid="ref15 ref16 ref17">15–17</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>6 Discussion and Conclusion</title>
      <p>
        As of today, annotation of process models is rarely automated. Also, rarely prototypes
are shown. Regarding the semantics of annotation, context information is (apart from
[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]) almost never used [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. This is a surprising research gap that exists even today –
after almost one decade of research on semantic technologies applied to BPM that
started with simple process model annotation proposals like [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Therefore, a research
opportunity lies in developing (semi-)automated annotation approaches in order to first
leverage existing standards such as PCF and second to make use of the wealth of
semantic technologies (e.g. for search and matching of models on the semantic level)
when process models could automatically be annotated. In this paper, we first have
described the nature of the annotation task and how humans perform it. We then
identify building blocks and parameters for automated systems that imitate human
annotation behavior. We then conduct first empirical studies on the effect of parameters. It
turned out that context information such as the topic of a process model is indeed very
important for an automated annotation approach. All in all, this contribution aims to
inspire more research on methods in (semi-)automatic approaches capable of linking
semi-formal process models with more formal knowledge representations. With this,
new use cases are possible as described in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] shifting the automated interpretation of
process models to a new and more semantic level.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Literature</title>
    </sec>
    <sec id="sec-7">
      <title>Appendix</title>
      <p>Selected results from the implementation (Fig. 5-7). Annotation quality (precision,
yaxis) is shown in relation to various parameter values (x-axis). Direct Hit means correct
annotation in regard to the gold standard. Hierarchy means correct annotation at the
process group level (level 2) of the Process Classification Framework.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fellmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Semantic Process Modeling - Design and Implementation of an Ontology-Based Representation of Business Processes</article-title>
          .
          <source>Bus. Inf. Syst. Eng</source>
          .
          <volume>1</volume>
          ,
          <fpage>438</fpage>
          -
          <lpage>451</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. APQC:
          <article-title>Process Classification Framework (PCF</article-title>
          ),
          <source>Version 5.2.0</source>
          . (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Malone</surname>
            ,
            <given-names>T.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crowston</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herman</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          :
          <article-title>Organizing Business Knowledge: The MIT Process Handbook</article-title>
          . The MIT Press (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Aslam</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herrmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Expressing Business Process Models as OWL-S Ontologies</article-title>
          . In: Eder,
          <string-name>
            <given-names>J.</given-names>
            and
            <surname>Dustdar</surname>
          </string-name>
          , S. (eds.) Business Process Management Workshops. pp.
          <fpage>400</fpage>
          -
          <lpage>415</lpage>
          . Springer Berlin Heidelberg (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Fill</surname>
          </string-name>
          , H.-G.:
          <article-title>Using Semantically Annotated Models for Supporting Business Process Benchmarking</article-title>
          . In: Grabis,
          <string-name>
            <given-names>J.</given-names>
            and
            <surname>Kirikova</surname>
          </string-name>
          , M. (eds.) Perspectives in Business Informatics Research. pp.
          <fpage>29</fpage>
          -
          <lpage>43</lpage>
          . Springer Berlin Heidelberg (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Fill</surname>
          </string-name>
          , H.-G.,
          <string-name>
            <surname>Schremser</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karagiannis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>A Generic Approach for the Semantic Annotation of Conceptual Models Using a Service-Oriented Architecture</article-title>
          .
          <source>Int. J. Knowl. Manag. 9</source>
          , (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Fill</surname>
          </string-name>
          , H.-G.:
          <article-title>On the Social Network Based Semantic Annotation of Conceptual Models</article-title>
          . In: Buchmann,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Kifor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.V.</given-names>
            , and
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <surname>J</surname>
          </string-name>
          . (eds.) Knowledge Science, Engineering and Management. pp.
          <fpage>138</fpage>
          -
          <lpage>149</lpage>
          . Springer International Publishing (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mturi</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johannesson</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A context-based process semantic annotation model for a process model repository</article-title>
          .
          <source>Bus. Process Manag. J</source>
          .
          <volume>19</volume>
          ,
          <fpage>404</fpage>
          -
          <lpage>430</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Bögl</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schrefl</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pomberger</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          :
          <article-title>Semantic Annotation of EPC Models in Engineering Domains to Facilitate an Automated Identification of Common Modelling Practices</article-title>
          . In: Filipe,
          <string-name>
            <given-names>J.</given-names>
            and
            <surname>Cordeiro</surname>
          </string-name>
          ,
          <string-name>
            <surname>J</surname>
          </string-name>
          . (eds.)
          <source>Enterprise Information Systems</source>
          . pp.
          <fpage>155</fpage>
          -
          <lpage>171</lpage>
          . Springer Berlin Heidelberg (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Francescomarino</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tonella</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Supporting Ontology-Based Semantic Annotation of Business Processes with Automated Suggestions</article-title>
          . In: Halpin,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Krogstie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Nurcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Proper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Soffer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            , and
            <surname>Ukor</surname>
          </string-name>
          , R. (eds.) Enterprise,
          <article-title>Business-Process and Information Systems Modeling</article-title>
          . pp.
          <fpage>211</fpage>
          -
          <lpage>223</lpage>
          . Springer Berlin Heidelberg (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Di</given-names>
            <surname>Francescomarino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Tonella</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <source>Supporting Ontology-Based Semantic Annotation of Business Processes with Automated Suggestions: Int. J. Inf. Syst. Model. Des</source>
          .
          <volume>1</volume>
          ,
          <fpage>59</fpage>
          -
          <lpage>84</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Rospocher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Francescomarino</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghidini</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serafini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tonella</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Collaborative Specification of Semantically Annotated Business Processes</article-title>
          . In: Rinderle-Ma,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Sadiq</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            , and
            <surname>Leymann</surname>
          </string-name>
          ,
          <string-name>
            <surname>F</surname>
          </string-name>
          . (eds.) Business Process Management Workshops. pp.
          <fpage>305</fpage>
          -
          <lpage>317</lpage>
          . Springer Berlin Heidelberg (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cai</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>An Ontological Approach for Semantic Annotation of Supply Chain Process Models</article-title>
          . In: Meersman,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Dillon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            , and
            <surname>Herrero</surname>
          </string-name>
          , P. (eds.) On the Move to Meaningful
          <source>Internet Systems: OTM 2010</source>
          . pp.
          <fpage>540</fpage>
          -
          <lpage>554</lpage>
          . Springer Berlin Heidelberg (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Semantic Annotation for Process Models: Facilitating Process Knowledge Management via Semantic Interoperability</article-title>
          . Department of Computer and Information Science Norwegian University of Science and Technology, Trondheim, Norway (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Born</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dörr</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>User-Friendly Semantic Annotation in Business Process Modeling</article-title>
          . In: Weske,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Hacid</surname>
          </string-name>
          , M.-S., and
          <string-name>
            <surname>Godart</surname>
          </string-name>
          , C. (eds.) Web Information Systems Engineering - WISE 2007 Workshops. pp.
          <fpage>260</fpage>
          -
          <lpage>271</lpage>
          . Springer Berlin Heidelberg (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Born</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoffmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaczmarek</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kowalkiewicz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markovic</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scicluna</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Semantic Annotation and Composition of Business Processes with Maestro</article-title>
          . In: Bechhofer,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Hauswirth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Hoffmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            , and
            <surname>Koubarakis</surname>
          </string-name>
          , M. (eds.)
          <source>The Semantic Web: Research and Applications</source>
          . pp.
          <fpage>772</fpage>
          -
          <lpage>776</lpage>
          . Springer Berlin Heidelberg (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Born</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoffmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaczmarek</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kowalkiewicz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markovic</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scicluna</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Supporting Execution-Level Business Process Modeling with Semantic Technologies</article-title>
          . In: Zhou,
          <string-name>
            <given-names>X.</given-names>
            ,
            <surname>Yokota</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            , and
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Q</surname>
          </string-name>
          . (eds.)
          <article-title>Database Systems for Advanced Applications</article-title>
          . pp.
          <fpage>759</fpage>
          -
          <lpage>763</lpage>
          . Springer Berlin Heidelberg (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Fellmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Towards Automated Process Model Annotation with Activity Taxonomies: Use Cases and State of the Art</article-title>
          . In: Abramowicz,
          <string-name>
            <surname>W</surname>
          </string-name>
          . (ed.)
          <source>Business Information Systems</source>
          . pp.
          <fpage>74</fpage>
          -
          <lpage>90</lpage>
          . Springer, Cham (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Leopold</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meilicke</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fellmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pittke</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stuckenschmidt</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendling</surname>
          </string-name>
          , J.:
          <source>Towards the Automated Annotation of Process Models</source>
          . In: Zdravkovic,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kirikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            , and
            <surname>Johannesson</surname>
          </string-name>
          , P. (eds.)
          <source>Advanced Information Systems Engineering</source>
          . pp.
          <fpage>401</fpage>
          -
          <lpage>416</lpage>
          . Springer International Publishing (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Fellmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Process Model Verification with SemQuu</article-title>
          . In: Nüttgens,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Thomas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            , and
            <surname>Weber</surname>
          </string-name>
          ,
          <string-name>
            <surname>B</surname>
          </string-name>
          . (eds.)
          <source>Enterprise Modelling and Information Systems Architectures (EMISA</source>
          <year>2011</year>
          ), Hamburg, Germany. pp.
          <fpage>231</fpage>
          -
          <lpage>236</lpage>
          . Köllen,
          <string-name>
            <surname>Bonn</surname>
          </string-name>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>