<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On Business Process Model Reviews</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Hasso Plattner Institute, University of Potsdam</institution>
          ,
          <addr-line>Germany bpt.hpi.uni-potsdam.de</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2010</year>
      </pub-date>
      <fpage>31</fpage>
      <lpage>42</lpage>
      <abstract>
        <p>In process reviews, domain experts validate the model against reality. In general, reviews are conducted in an iterative manner. Better reviews can build consensus faster and save iterations, i.e. time and money. In an exploratory study, student clerks were asked to provide feedback to models from their domain. In this paper, we report on the study and the review performance. We explore typical issues raised in reviews and derive implications for practitioners and further studies. We identi ed education as the most in uential factor on review performance in our sample set.</p>
      </abstract>
      <kwd-group>
        <kwd>Process Modeling</kwd>
        <kwd>Reviews</kwd>
        <kwd>Performance</kwd>
        <kwd>BPMN</kwd>
        <kwd>t</kwd>
        <kwd>BPM</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Visualized process models serve as a communication vehicle in business process
management. Moreover, they become the blueprint for software implementations.
On the path from the initial business process elicitation to software support,
review cycles are required. Models are created once and get iterated several times.
Iterations typically involve feedback cycles with domain experts. They have to
ensure that the domain knowledge is properly represented in the model. Better
review performance promises less iterations, which in turn translates to time
and money saved on projects. But what can you expect from domain expert's
reviews? How can you in uence the performance of the reviewing task?</p>
      <p>
        Empirical research on business process modeling has largely investigated the
roles of models [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ] and modelers [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Condensed ndings from empirical research
even led to modeling guidelines [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Process reviews have not been addressed
comparably.
      </p>
      <p>
        We did a pre-study to assess the experiment setup for t.BPM [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. It is a
tangible toolkit to enable BPMN process modeling on a table. As part of this,
university freshmen were introduced to BPMN and lled in a feedback test about
a given model (adopted from [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]). It contained a graph with 14 tasks and ve block
structured exclusive and parallel sections. Activities were labelled with A,B,C...
The students easily passed a test on understandability (also adopted from [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]). It
turned out that all of them, had a strong formal background. Even though the
freshmen were barely educated in process modeling, they could map the process
semantics to known concepts from mathematics and physics, such as logical
equations and circuit diagrams. This obviously in uenced their performance. We
concluded that, to get meaningful data about process reviews, a more realistic
setup is needed.
      </p>
      <p>In this paper, we rst introduce our study design in Section 2. The data
is evaluated and discussed in Section 3. Additional insights are drawn from
investigating a related study in Section 4. We close the paper in Section 5 with a
discussion of the ndings and implications.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Study Design</title>
      <sec id="sec-2-1">
        <title>2.1 Setup</title>
        <p>
          The sample population, used in research studies, should be representatives of
the population to which the researchers wish to generalize [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Thus, we wanted
domain experts to provide feedback to domain speci c processes. From interviews
with BPM consultants we identi ed a typical scenario in which process consultants
give workshops to elicit the domain knowledge, model the processes, and send
them out as email attachment. Domain experts are asked to provide feedback.
The model gets iterated. Part of the workshops with the consultant would be
reserved to educate the participants about the goal of BPM and the notation
used for process modeling.
        </p>
        <p>
          To emulate best practices in the eld, we designed the following exploratory
study for subjects at the trade school in Potsdam. Students there are learners
to become o ce or industrial clerks. They get practical training on the job and
theoretical background for their profession at the trade school. As clerks, we
consider them to be representatives of the population to be generalized on. We
chose the domain processes Moving to a new at and Getting a new job. The
seventeen students (18-22 years) are considered to be domain experts, meaning
they do know the context and can comment on the processes. Additionally, we
designed a two page introduction into BPM and a one page modeling sample
(topic: Making Pasta). The sample page contained a legend of the BPMN elements
used. On that same page four pragmatical hints for to process modeling were
provided. In particular, we suggested the balanced use of gateways, an eighty
percent rule for relevance to set granularity, verb-object style activity labels as
suggested by Mendling et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and a notational convention for conditions at
gateways.
        </p>
        <p>
          The introduction and sample sheet were designed to condition the subjects.
They replace the guidance provided by the modeling experts in the workshops.
The written form enforced the same type of treatment for all subjects. This
was embedded in a larger experiment design to test the e ect of t.BPM [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] on
subjects. The hypotheses were that t.BPM modeling would yield positive e ects
on individuals, including more feedback in process model reviews. While the
experiment result are to be published, this study explores partial data with focus
on feedback performance. The experiment procedure is depicted in Figure 1.
        </p>
        <sec id="sec-2-1-1">
          <title>Model Reviews</title>
          <p>repeated measurement design (random order)
BPM Intro Sample</p>
          <p>Feedback test</p>
          <p>Interview
Process Modeling
10 industrial clerks
7 office clerks</p>
          <p>Subjects</p>
          <p>Conditioning</p>
          <p>Treatment</p>
          <p>Evaluation</p>
          <p>Each student got the BPM introduction and the sample. In general, time was
not limited but tracked for each stage. Students were then randomly assigned to
do either a structured interview or model their process on the table using BPMN
elements. In that treatment step, they were asked to describe procurement
processes, such as purchasing expensive hardware. Afterwards subjects were
randomly given one of the process models for feedback. The treatment was
repeated for each subject. In the second run, they got the alternative treatment
and use the alternative feedback test.</p>
          <p>In other words, the setup was a repeated measurement design in which all
subjects get the same treatment in di erent orders. Subjects were assigned
randomly. All subjects did interviews and process modeling. And all subjects did
get both feedback tests, again randomly assigned.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Process Models used in Reviews</title>
        <p>The process models used in the study are depicted in Figure 2. The models are
annotated with the issues which we intentionally built into them.</p>
        <p>Issues were chosen to belong to the area of language or domain. For example,
two deadlocks were built into each process. This can be found by formal language
analysis and requires no knowledge about the process domain. Nevertheless, we
built-in these language related problems as indicators for the subjects' semantical
understanding of the modeling language. The focus of this study are issues linked
to the domain. They can only be interpreted if context information is available.
Within the domain we consider three main categories: labeling, information
granularity and logical mismatch. Labeling covers unsuitable naming of process
elements, i.e. activity labeled with states not actions. Two obviously unsuited
labels were build into the model (see Figure 2). Information granularity deals with
too much or missing information in the process model. We left out an obvious
activity and document per process model. Finally, logical mismatch describes
wrong information in the model which contradicts the reality. We misplaced an
activity to generate an issue of this category. An overview of the built-in problems
is given in Table 1 when we report on the review performance.</p>
        <p>One sample issue, a missing control ow connector, was marked up in the
model to indicate how to give feedback. We asked reviewers very broadly to
"provide feedback". We assume that guiding questions and a clear focus, e.g.
communicating the goal of the modeling e ort, would have steered the reviews.
Our goal is to explore. Therefore, neither guiding questions nor a goal were
provided.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3 Variables</title>
        <p>For this investigation feedback is the dependent variable. We quantify feedback
by counting the number items provided in a review. Feedback is distinguished
into intentionally built-in problems and additional comments. Categorization
and quanti cation of feedback items was done by expert reviews. We refer to
the sum of all issues raised as feedback. While the quality of feedback matters
most, we start with quantity for our exploration. The initial assumption was that
variation in the amount of feedback could be explained by the treatment method
(t.BPM vs. interviews). Data analysis revealed no in uence by treatment method
(details in Section 3.4).</p>
        <p>Thus, we decided to explore other available information to explain the variance
in the data set. In Section 3.4 we investigate the time, the participant's education,
and sex as independent variables. Guiding questions and modeling goal were
consciously excluded as variables from this study.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 Data Exploration</title>
      <sec id="sec-3-1">
        <title>3.1 Data Analysis Instruments</title>
        <p>The data in the sample set was tested and is normally distributed1. Signi cance
was tested with a one-tailed t-test, abbreviated here with p. Correlation between
variables was calculated using Pearson's correlation coe cient r. It is a normalized
measure of dependence between two quantities where 0 indicates no correlation,
-1 is a perfect negative correlation and 1 is a perfect positive correlation. In
Section 3.4 we use Multi Regression Analysis to explain variation with signi cantly
in uential factors. The coe cient of determination R2 describes the proportion
of variation in the data set that can be explained with the regression model. As
an example, R2 = :30 means that thirty percent of the data variation can be
explained by a particular regression model.</p>
        <p>Based on the repeated measurements design we treat each test as an
independent sample (n=34). We keep in mind that pairs of samples result from a single
person, but we'll see that splitting them up yields no negative e ect on the data
analysis. In summary,
1 True for Kolmogorov-Smirnov and Shapiro-Wilk test
1
l
c
o
a
k
d
1
g ic
e d
t
t u
i
B a
. n
t e
ib g
e s
r a
h
c w
s ,
e r</p>
        <p>e
b t
s n
e u
s r
s a
e d
z
o e
rp ie</p>
        <p>b
s r
g h
n c
u s
rb d
e n</p>
        <p>u
ew re
B
n m
e m
h u
c N
s
i
s r
s e
la in
k e
n it
e
d m
s e
a h</p>
        <p>c
d i
,ll re
e e
d B
.
t
s
e
d
r
ü
d
r y
ü it</p>
        <p>v l
wti e
nac lab
e
ng nt
i
s e
is m
mcu
o
d</p>
        <p>o e
e M d
s n
ah ien reö
P ud ts</p>
        <p>h
k ts ic
c is
n t
g vi
. . .
data is normally distributed
r = [ 1::1] describes the correlation of two quantities
p is a one-tailed t-test, p &lt; :05 is considered as signi cant
R2 = [0::1] is the explained variance in the regression model
all numbers are based on a sample set of n = 34</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Review Performance</title>
        <p>The review performance of the subjects was quite poor. Most problems were not
found with an overall success rate of less then thirty percent. Table 1 lists the
built-in issues and shows how often it was by a single reviewer in one or both of
the feedback tests.</p>
        <p>While the issue of a wrong activity order was always found, by all reviewers
in all feedback tests, the opposite is the case for the deadlock1 which results from
a loop back. If reviewers found an issue only once, it indicates that they did not
systematically checked for this type of issue.</p>
        <p>Category</p>
        <p>Built-in Issues</p>
        <sec id="sec-3-2-1">
          <title>Execution Semantics deadlock1 (back loop)</title>
          <p>(Language) deadlock2 (bad block)</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Labels activity labeled as state</title>
          <p>(Domain) data object labeled as activity</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>Information Granularity missing document</title>
          <p>(Domain) missing activity</p>
        </sec>
        <sec id="sec-3-2-4">
          <title>Logical Mismatch wrong activity order (Domain)</title>
          <p>Investigating the individual performance, we found that review performance
varies between one and six reported built-in issues. On average, only two were
found per person and indeed, in nineteen of the thirty four cases, only one
built-in issue, the wrong activity order, was found. Reviewers gave 2.2 additional
comments. In summary, over all tests, subjects reported back 4.2 feedback items
on average. It shows that, although only a few problems were found (69 out of
238 in total), the reviewers still had a lot to share about the process with 75
items delivered as additional feedback.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3 Distribution of comments on topics</title>
        <p>The role of comments is to capture additional issues which were not intentionally
built-in. Domain speci c issues with processes are not necessarily modeling
mistakes, they might be a conscious decision to capture a certain aspect, or not.
If an issue was arguable, it was counted as a comment. Almost all comments were
counted, except for two. They were dropped as questions about the notation, not
comments on the process. Comments were aggregated if they centered around
one single issue.</p>
        <p>Simply counting comments was not meaningful, so we categorized them,
see Table 2. Two researchers reviewed each comment, negotiated the type and
category. Despite the potential experimenter bias, the advantage of this qualitative
approach is to discover new issues and categories.</p>
        <p>As shown in Table 2, most comments seek to inject additional information
into the process model (30 of 75), rather than leaving them out (only 11).</p>
        <p>Parallelizing or sequentializing activities was surprisingly popular (11 of 75
comments). As an example, subjects commented for the process "Moving to a
new at " that they would not clean until the are done with painting or that
changing the address with the authorities should be started earlier in the process
(in parallel). In our opinion it indicates, that the reviewers understood the
semantics of the model as well as the domain. Two reviewers found fundamental
optimization potential. As an example, if multiple o ered ats were researched
early on, we do not need to loop back and start over with research all the
time. This observation is acknowledged by introducing a new issue and category.
However, one might argue that this new category relates to process design (to-be
situation) whereas feedback is typically focussed on validation (as-is situation).</p>
        <p>Most interesting to note are the three issues marked with * in Table 2.
Those three categories originate from four reviewers. Two of them criticized that
start and end events were not properly labelled. One argued that the activity
"Moving" should be a word-object style label. One reviewer raised the question for
process scoping. In particular, he commented, that the process Moving to a new
at should be completed after the rental contract was signed. The subsequent
activities are not in the scope of the process. While the authors do not agree with
this opinion, it brings up process scoping as an issue addressed in reviews. Notes
taken during the pre-study interview indicate that all four subjects were involved
in process modeling activities within their company. We conclude, that they
brought in additional process knowledge which was not part of the conditioning
for this study. With nine out of eleven issue types being new, this qualitative
assessment of feedback widened our or repertoire of issues addressed in process
model reviews.
3.4 In uential factors
The initial assumption was that t.BPM modeling in uences the reviewer's
performance, which did not happen. Indeed, subjects performed quite stable in both
feedback tests independent of treatment order or type, see Table 3 for details.</p>
        <p>We even found that the amount of feedback does not signi cantly di er
between the rst and the second feedback test. For that reason we decided to
treat all thirty-four feedbacks as independent samples (n = 34). We also compared
the mean scores for the two di erent feedback models. They do not signi cantly
di er which indicates that both models were equally hard or easy to understand.
We therefore conclude that model type, treatment and order have no in uence on
the reviewer's performance.</p>
        <p>Sex, education, and time taken to conduct the review had a signi cant in uence
on the performance. For education and sex the results are depicted in Table 3.
Education emerges as the most dominant factor with the highest e ect size and
strongest signi cance.</p>
        <p>In uence Factor
(independent variable)</p>
        <sec id="sec-3-3-1">
          <title>Treatment order</title>
        </sec>
        <sec id="sec-3-3-2">
          <title>Treatment type</title>
        </sec>
        <sec id="sec-3-3-3">
          <title>Model type</title>
        </sec>
        <sec id="sec-3-3-4">
          <title>1st/2nd Test</title>
        </sec>
        <sec id="sec-3-3-5">
          <title>Education Sex</title>
          <p>Alternatives E ect Size</p>
          <p>(x feedback)
1st t.BPM / 2nd interview 4.5556 / 4.7778
1st interview / 2nd t.BPM 3.875 / 3.75
t.BPM / interview 4.1765 / 4.3529
moving / job nding 4.0588 / 4.4706</p>
          <p>1st / 2nd 4.2353 / 4.2941
o ce / industrial 2.50 / 5.45
male / female 4.95 / 2.92</p>
          <p>Signi cance
(one sided t-test)
.354
.418
.331
.15
.442
.000019
.001</p>
          <p>To explain the signi cant in uence of education , we had post-study
interviews with the principal of the trade school. We were informed that o ce clerks
undergo a much stricter selection procedure and have better school achievements.
On their job, they switch departments more easily and are often involved in
supply chain optimization. Therefore, education for industrial clerks at the trade
school does also include process notations, although in a very limited scope. Some
students are also involved in process elicitation and modeling at their companies.</p>
          <p>A boxplot in Figure 3a depicts the scattering of feedback for both groups. It
visualizes that industrial clerks have much more to say about the process, up to
eleven items in a single feedback test, while o ce clerks typically give two (at
most ve) items in a feedback test. In numbers, eight of fourteen tests done by
o ce clerks reported one or two issues as feedback.</p>
          <p>This dramatic di erence due to education puts new light on our pre-study
experience with HPI freshmen students. It raises the general question for
transportability of empirical ndings, if there is such a big gap between rather close
professions.</p>
          <p>The in uence of sex is likewise signi cant with a considerable e ect size.
However, there is also a large overlap of sex and education in our sample set. Out
!"#$$
/!$#$$
$
%
3
+"%#$$
+
,**
,&amp;(&amp;#$$
)
#
(
,-&amp;%2'#$$
%
&amp;("#$$
$#$$
_^$';C2?8$($eU6f</p>
          <p>(**'$+,$-+./0
!"#$%&amp;'()
(a) Boxplot comparing education:
o ce clerks only provide 1-5 items with
a median of 2
^UQe QUee ]UQe SeUee S^UQe</p>
          <p>'"&amp;()*+%)*((,-#./)'(0')1&amp;"23
(b) Scatterplot and regression curve:</p>
        </sec>
        <sec id="sec-3-3-6">
          <title>The longer a feedback test takes, the more feedback is gathered</title>
          <p>!"#$%&amp;
of eleven industrial clerks in the study, eight were male and three were female.
Whereas out of six o ce clerks, two were male and four were female.</p>
          <p>We conducted a hierarchical multi regression analysis to determine the actual
in uence of this variable. This multi regression model has a coe cient of
determination of R2 = :508, which means that it can explain 50.8 % of the variance in
the data2. In that model, the contribution of sex boils down to explain 0,1% of
the overall variance (R2 = :001). The st!a"#n$%&amp;dardized multi regression equation is:
F EEDBACK = :404 education+:346 timefeedback :093 timeintro +:043 sex
The in uence of time was determined using Pearson's correlation coe cient r.
The time taken to complete the feedback test correlates signi cantly positive with
the amount of feedback given (p = :00008; r = :6). That means, subjects that
take more time for the feedback test, give more feedback. Figure 3b depicts the
correlation in a scatterplot with a linear regression line. In the hierarchical multi
regression model timefeedback is the second strongest in uence and contributes
10,4 % to the explanation of variance. While o ce clerks take about ve minutes
on average to complete the feedback test, industrial clerks take 8.3 minutes on
average. We assume that subjects with less understanding have less to contribute
and therefore need less time. Alternatively, subjects that investigate the process
more deeply, nd more issues but this of course needs more time. Similarly, people
that need more time to read the BPM introduction perform worse in the feedback
test (p = 0:0375; r = :39). However, this has only a minor contribution of 0,9 %
to the overall explanation.</p>
          <p>Concluding the review of the in uential factors, we can explain 50.8% of the
overall variation using a Hierarchical Multi Regression Analysis which considers
the four variables education, timefeedback, timeintro, sex. The main in uential
factor is education with the highest signi cance and e ect. Education alone can
2 Re2ducation = :393129</p>
          <p>Rt2imefeedback = :104</p>
          <p>Rt2imeintro = :009216
explain 39.3% of the data in the Multi Regression Model. The signi cance and
e ect size found for sex (see Figure 3), diminishes in the Multi Regression Model.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>3.5 Limitations Discussion</title>
        <p>The validity of this explorative study is limited by the decisions taken for
its practical implementation. In particular, one might argue that the domain
processes from a private background might limit the transportability of ndings
to business domains (external validity). And of course, the de nition of an "issue"
as well as its categorization is subjective (internal validity).</p>
        <p>The small sample set, with the in uence factors reported earlier, also limits
the generalizability of ndings. Larger sets with more controlled variables should
be used for hypotheses testing. In this exploratory study, the small sample set
enabled us to look deeply into the reviews (qualitative research). Thereby, we
identi ed new issues that we did not see before.</p>
        <p>Throughout the study and its evaluation we took the following
countermeasures to limited the experimenter bias:</p>
        <p>We standardized conditioning for the subjects using written documents.</p>
        <p>Two researchers coded the feedback and negotiated categories.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 Related Work on Reviews</title>
      <p>
        In 2002, Moody et. al. assessed a quality framework for conceptual models using
process modeling [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. The subjects were 194 third year students in Information
Systems (IS) which had to model a process and then peer review three processes
modeled by others. A set of 20 process models and their reviews was qualitatively
investigated.
      </p>
      <p>Category defects</p>
      <sec id="sec-4-1">
        <title>Syntax missing ows</title>
        <p>(Language) wrongly speci ed decision point</p>
      </sec>
      <sec id="sec-4-2">
        <title>Labels (Domain) poor naming of tasks</title>
      </sec>
      <sec id="sec-4-3">
        <title>Information Granularity missing roles</title>
        <p>(Domain) missing ressources</p>
        <p>missing activity
Logical Mismatch lacking decision point
(Domain) wrong activity order
a ected models
50%
35%
27%
50%
44%
25%
30%
19%</p>
        <p>
          The authors state that "Many of the models were of quite poor quality, and
counting the number of errors did not give interesting results." [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Thus the
"errors" were classi ed and the reviews were assessed. He uses the notion of
defects to summarize the issues. Table 4 shows the defects.
        </p>
        <p>
          Interestingly, subsequent expert reviews found 6.6 defects per model of which
2.4 got reported by the reviewers. In other words, "on average, 64% of the defects
went unreported" [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. These numbers compare well with our seven intentionally
built-in problems of which 2 were found (success rate &lt; 29%) on average.
        </p>
        <p>While this is the nearest known relative to our study, several fundamental
di erences hamper a proper comparison of numbers from both studies. To name
the most important ones,</p>
        <p>
          The review reported in Table 4 relates to modeling defects. In the study,
reviews are evaluated by reporting true/false negatives/positives.
The notion of defect used by [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] is much stronger than our notion of issues.
Quality and defects per model did vary in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], while we had a stable set of
pre-de ned issues per model.
        </p>
        <p>IS students have a very di erent education. They are method experts rather
than domain experts.</p>
        <p>Nevertheless, we learn from this study defect types that can be build into
models for review tests. This further extends our set of feedback issues. Most
important, we learn that proper education does not guarantee good process
reviews. Thus, further research is needed to de-mystify the task of reviewing.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5 Discussion</title>
      <p>Reviews by domain experts are a critical part of model validation and need more
scienti c investigation. Better review performance can avoid additional iterations
needed in process analysis and design. This equals money and time saved on a
project. We conducted an explorative study using qualitative and quantitative
methods.</p>
      <p>Findings from this study are the issue types and the in uence factors. The
identi ed issue types can be used to create better models for review tests with
a larger variety of built-in issues. The distribution of issues raised by reviewers
is also a nding. It can be used as a starting point to guide reviewers in their
task. In other words, issues that are often missed might be worth a hint. Thus,
reviewers can systematically check for them. A guideline for reviewers was out of
scope for this work.</p>
      <p>By statistical evaluation, we found education, sex and time as in uential
factors in the sample set. In particular, education dominated our ndings.
Although we had similar previous experiences with university freshmen, we did
not anticipate education to be as in uential within o ce and industrial clerks.
We conclude that the subject group should be as homogeneous as possible to
exclude those in uences on the data set in future investigations. At the same
time the model should involve a large variety of issues. Thus, it is possible to
create the variance needed for insightful results. Our ndings are limited by the
small sample set and the dominance of education as an in uential factor.
Implications for practitioners are phrased as suggestions to process modelers
that do review cycles with domain experts. We suggest to,
choose your reviewers wisely (huge di erences in review performance)
one reviewer per model is not enough (on avg. &gt; 60% of issues not found)
Further research can built on the ndings from this study to build a proper
controlled experiment. In particular, the in uential factors identi ed here should
be xed to rule them out. When designing models for review experiments, future
research can take advantage of the domain related problems identi ed in this
study. That can help to create models with a larger variety of problems built in.</p>
      <p>In this study, we left out the aspects of a modeling goal and guiding review
questions. We assume that they signi cantly in uence the performance of
reviewers. For example, guiding questions can link to frequently unreported issue
types. In future work, we intend to investigate the in uence of these aspects on
reviewing performance.</p>
      <p>Acknowledgements
We gratefully acknowledge the support of Karin Telschow and Markus Guentert
to setup and conduct the t.BPM experiment series. Special credits to Karin for
her great support during the data exploration phase.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Holschke</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rake</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levina</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Granularity as a Cognitive Factor in the E ectiveness of Business Process Model Reuse</article-title>
          .
          <source>In: Proceedings of the 7th International Conference on Business Process Management</source>
          , Springer (
          <year>2009</year>
          )
          <fpage>260</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Melcher</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reijers</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seese</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>On measuring the understandability of process models (experimental results)</article-title>
          .
          <source>In: 1st Int. Workshop on Empirical Research in Business Process Management (ER-BPM)</source>
          .
          <article-title>(</article-title>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Recker</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dreiling</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Does it matter which process modelling language we teach or use?</article-title>
          <source>In: 18th Australasian Conference on Information Systems</source>
          . (
          <year>2007</year>
          )
          <volume>356</volume>
          {
          <fpage>366</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Mendling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reijers</surname>
          </string-name>
          , H., van der Aalst, W.:
          <article-title>Seven process modeling guidelines</article-title>
          .
          <source>Information and Software Technology (IST)</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Grosskopf</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edelman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weske</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Tangible business process modeling - methodology and experiment design</article-title>
          . In Mutschler,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Wieringa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Recker</surname>
          </string-name>
          , J.,
          <source>eds.: 1st Int. Workshop on Empirical Research in Business Process Management (ER-BPM'09)</source>
          .
          <source>(September</source>
          <year>2009</year>
          )
          <volume>53</volume>
          {
          <fpage>64</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cooper</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schindler</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Business Research Methods.
          <volume>10</volume>
          edn.
          <string-name>
            <surname>McGraw-Hill Higher Education</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Moody, D.,
          <string-name>
            <surname>Sindre</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brasethvik</surname>
          </string-name>
          , T., S lvberg, A.:
          <article-title>Evaluating the quality of process models: empirical analysis of a quality framework</article-title>
          .
          <source>In: 21st Int. Conference on Conceptual Modeling{ER</source>
          . (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Sindre</surname>
          </string-name>
          , G., Moody, D.,
          <string-name>
            <surname>Brasethvik</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solvberg</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Introducing peer review in an IS analysis course</article-title>
          .
          <source>Journal of Information Systems Education</source>
          <volume>14</volume>
          (
          <issue>1</issue>
          ) (
          <year>2003</year>
          )
          <volume>101</volume>
          {
          <fpage>120</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>