<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Scenario-based Resilience Evaluation and Improvement of Microservice Architectures: An Experience Report</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sebastian Frank</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alireza Hakamian</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lion Wagner</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dominik Kesim</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jóakim von Kistowski</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>André van Hoorn</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DATEV eG</institution>
          ,
          <addr-line>Nürnberg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Software Engineering, University of Stuttgart.</institution>
          <addr-line>Stuttgart</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Context. Microservice-based architectures are expected to be resilient. However, various systems still sufer severe quality degradation from changes, e.g., service failures or workload variations. Problem. In practice, the elicitation of resilience requirements and the quantitative evaluation of whether the system meets these requirements is not systematic or not even conducted. Objective. We explore (1) the scenario-based Architecture Trade-Of Analysis Method (ATAM) for resilience requirement elicitation and (2) resilience testing through chaos experiments for architecture assessment and improvement. Method. In an industrial case study, we design a structured ATAM-based workshop, including the system's stakeholders, to elicit resilience requirements. We specify these requirements into the ATAM scenario template. We transform those scenarios into resilience experiments to quantitatively evaluate and improve system resilience. Result. We identified 12 resilience scenarios. We use and extend ChaosToolkit to automate and execute two scenarios. We quantitatively evaluate resilience requirements and suggest resilience improvements in the scope of both scenarios. We share lessons learned from the case study. In particular, our work provides evidence that an ATAM-based workshop is intuitive to stakeholders in an industrial setting. Conclusion. Our approach helps requirement and quality engineering teams in the process of resilience requirements elicitation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>ture Trade-Of Analysis Method (ATAM) [ 8] for (1) the
system’s resilience requirement elicitation and (2)
reAn intrinsic quality property of the microservices archi- silience testing through resilience experiments (aka chaos
tectural style is resilience, i.e., the system meets perfor- experiments) for architecture assessment and
improvemance and other Quality of Service (QoS) requirements ment. We hypothesize that ATAM has already been used
despite diferent failure modes or workload variations [ 1]. in practice to elicit and specify quality requirements other
However, real-world postmortems [2] show that systems than resilience, e.g., availability, performance, and
mainsufer either unacceptable QoS degradation, or recovery tainability, and can be adopted for efectively eliciting
retime. It is necessary to assure system resilience in the silience requirements and evaluating them through chaos
context of microservice-based architectures. experiments. Therefore, our research question is: How</p>
      <p>Practitioners use Chaos Engineering [3], including to leverage ATAM to elicit resilience requirements, which
tools such as CTK [4, 5] or Chaos Monkey1, for resilience can be utilized to evaluate resilience through resilience
extesting. They need to (1) think about hazards [6] as causes periments and suggest architectural improvements
quanof QoS degradation, (2) set up chaos experiments by spec- titatively?
ifying failure mode types and hypotheses of expected We use an ATAM-based workshop to elicit and specify
quality behavior, and (3) execute each experiment to de- resilience requirements by involving system
stakeholdtect deviations from the hypotheses. First, this approach ers. ATAM allows us to describe resilience requirements
lacks the systematic identification of causes of a hazard as scenarios in semi-structured textual language. The
through hazard analysis methods. We contributed to scenario template consists of the following elements:
this problem in our previous work [7], which serves as a (1) source, (2) stimuli, (3) artifact(s), (4) system’s
envifoundation for this paper. Particularly, we now integrate ronment, (5) its response, and (6) response measure.
hazard analysis into a more systematic elicitation process The designed structured workshop aims to identify
and use a more formal description of requirements (sce- hazards and architectural design decisions. During the
narios). Second, the approach lacks a systematic process workshop, we employ a hazard analysis based on the
of eliciting and refining resilience requirements. Fault Tree Analysis (FTA) [6]. The result is a set of 12</p>
      <p>In the context of an industrial case study, our objective resilience scenarios, which we turn into experiments.
is to explore the application and adoption of the Architec- Next, we use CTK to automate these experiments, and
conduct a measurement-based resilience evaluation.
FurECSA 2021© 2C02o1mCoppyaringhitofnor Vthoislpuapmereb,y Vitsäaxutjhöo,rsS.wUseedpeermni,tt1ed3u-n1d7er SCerepattiveember 2021 thermore, we improve system resilience by applying a
CPWrEooUrckReshdoinpgs 1IhStphN:/c1e6u1r3-w-t0s.o7r3gtpCCso:/Em/mUgoiRntshLWuicebon.srcekoAstmthrib/ouNptioenPt4flir.x0o/Inccteehrenaadotiiosnnmagl s(oCCn(CBkYEe4yU.0)R.-WS.org) irmespilrioevnecmeepnatttbeyrnre[-9e,x1e]c,untianmgetlhyerreetrsyp.ecWtievevsacliednaatreioth.e</p>
      <p>To summarize, the paper makes the following contri- scenarios are two activities in requirements
engineerbutions: ing [11] that benefit the elicitation process. According
to Pohl [11], scenario development benefits elicitation
• Leveraging ATAM and FTA in an industrial system by making goals understandable for stakeholders and
to elicit resilience requirements, then evaluating the may refine or identify new goals. Our work uses
scerequirements, and improving system’s resilience. nario development without goal-oriented modeling as all
• Automating scenario execution using CTK for measure- stakeholders know the system’s high-level quality goal.
ment-based evaluation. To our knowledge, this is the first work using ATAM
for eliciting and specifying resilience requirements before
evaluating the resilience through experiments.
• We share lessons learned that benefit both
practitioners and researchers regarding resilience requirement
elicitation, evaluation, and improvement.
• Artifacts — including scenarios, resilience experiments,
and results of the experimentation — are available
online [10].</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
    </sec>
    <sec id="sec-3">
      <title>3. Research Methodology</title>
      <p>Section 3.1 explains the domain context and describes
the high-level architecture of the case study system.
Section 3.2 summarizes our research methodology.</p>
      <sec id="sec-3-1">
        <title>3.1. Domain Context</title>
        <p>A workshop is an efective technique for requirement The case study system’s purpose is to calculate payments.
elicitation [11]. In our case, the workshop’s prepara- An accounting department’s wage clerks use the
paytion and conduction are based on the scenario template ment accounting system to calculate each registered
emof Bass et al. [12]. Our diference to existing works on ployee’s income taxes. The payment accounting system
measurement-based resilience evaluation is that we have has to gather data from health insurance providers and
an explicit step on eliciting resilience requirements. In send its results to the corresponding tax ofice to execute
the next paragraphs, we elaborate on this in more detail. the calculations. This process presumes that a company</p>
        <p>Cámara et al. [13, 14, 15] propose an approach for that wants to use the payment accounting system
proresilience analysis of self-adaptive systems. The core vides its employee and tax information to the health
idea consists of three parts: (1) specification of resilience insurance provider and tax ofices.
properties using Probabilistic Computation Tree Logic, All of the payment accounting system’s tasks are
cur(2) modeling causes of a hazard, e.g., high-load using ex- rently taken care of by a monolithic legacy system. In
perimentation and collecting traces of system behavior, peak times, up to 13 million calculation requests have to
and (3) verification of resilience properties using model be handled in a day or single night. Under normal
circumchecking. In contrast to model checking-based verifica- stances, this number is significantly lower. In order to
tion, we evaluate a resilience scenario’s response measure handle such varying loads more eficiently, stakeholders
by analyzing collected measurements. Furthermore, Cá- desired a better scaling system. Therefore, the old system
mara et al. do not focus on the elicitation of resilience will be replaced by a more scalable microservice-based
requirements using requirement engineering methods. Spring application in the coming years. The investigated</p>
        <p>Chaos Engineering [5, 3] is a technique for evaluating part of the system under study, which is still under
desystem resilience through injecting failures [16]. There velopment, consists of seven services. It is deployed to
are works on both (1) using engineering methods to iden- a Platform as a Service (PaaS), i.e., Cloud Foundry (CF).
tify failure modes [7], i.e., causes of a hazard systemati- Together with the industrial partner, we decided on a
cally before failure injection, and (2) ad-hoc failure injec- scenario-based approach, as our industrial partner
altion with no systematic failure mode identification [ 2]. ready employed ATAM for other quality attributes.
However, they do not explicitly specify resilience
requirements and lack a methodical way for requirement
elicitation. Our work is a step toward closing this gap. 3.2. Research Procedure</p>
        <p>In the context of resilience requirement elicitation, Yin To answer our research question How to leverage ATAM to
et al. [17] propose a goal-oriented technique for represent- elicit resilience requirements, which can be utilized to
evaling resilience requirements. The high-level idea is to rep- uate resilience through resilience experiments and suggest
resent a resilience goal — e.g., all requests are processed architectural improvements quantitatively?, we conduct
correctly — and identify possible causes of hazards that the following steps:
act as obstacles for achieving a resilience goal — e.g., node
failure. However, they do not discuss how to identify haz- 1. We gather relevant system stakeholders, i.e., product
ards and their causes. Goal orientation and developing owners, software architects, and quality engineers,
into an ATAM-based workshop. The objective is to For this reason, we refer to the result of this session
identify resilience scenarios that lead to QoS degrada- as a fault graph. Note that the (directed acyclic) fault
tion and downtime. graph can be transformed into an equivalent fault tree
by creating duplicate sub-trees for nodes having more
2. We derive resilience experiments from the scenarios. than one parent.</p>
        <p>The experiment description comprises the stimuli,
artifact, and response according to the scenario.</p>
        <p>Session 3: Resilience Scenarios for collecting and
prioritizing resilience scenarios based on the previously
identi3. We use CTK to automate the execution of the re- fied hazards. We provided a scenario template based on
silience experiments. We assess the response measure the layout used in ATAM. Then, the stakeholders jointly
by analyzing the QoS metrics measurements. created scenarios by informally analyzing the fault graph
in a sequence driven by the associated severity (in
de4. After executing resilience experiments, we apply suit- scending order) of the hazards.</p>
        <p>able resilience patterns. We re-execute the resilience Session 4: Retrospective to collect feedback about the
experiments to assess the pattern’s efect by compar- workshop from the participants and to inform them about
ing QoS-related behavior with and without the re- the next steps, which comprise (1) refinement resilience
silience pattern. requirements, and (2) execution of resilience experiments.</p>
      </sec>
      <sec id="sec-3-2">
        <title>4.2. Workshop Results</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Elicitation and specification of scenarios</title>
      <p>Elicited Architecture Description: Figure 1 shows the
component diagram of the system as specified in the first
This section elaborates on the planning, execution, and session of the workshop. It describes a snapshot of the
results of the workshop. system as used in the workshop and the subsequent
activities. It represents a typical microservice-based
architec4.1. Elicitation through Structured ture. As such, the system is deployed to a CF and contains
several services. Each service has its own PostgreSQL</p>
      <p>Workshop database. The only exception is the Calculations service,
Before the workshop, we received documentation regard- which employs a Mongo database. The API-Gateway
sering the architecture of the system. This allowed us to vice handles all incoming connections and routes all
comspecify an architecture model of the case study system, munications. A Eureka service is employed to provide
including a component diagram and an explanation of service discovery for all internal components. The
Fronthe implemented components. Using ATAM, we required tend service is the only external component that a user
to know key architectural design decisions. Therefore, can directly access. The Calculations service is the
cenknowing the architecture description in advance allowed tral hub of the system since the calculation of payments
us to focus more on the hazard analysis and developing is the system’s main feature. Once this service receives a
resilience scenarios. calculation request from the gateway, it collects all
nec</p>
      <p>The full-day workshop consisted of four sessions lever- essary data asynchronously from the other services. The
aging diferent methods, as described next. The modera- Companies service is used to handle data the Frontend
tors explained each technique and method at the begin- displays, but is not relevant for the calculation.
ning of each session. The participants were stakeholders Hazard Analysis: Figure 2 shows the fault graph
creof the system and comprised two software architects, one ated in the second workshop session. The stakeholders
product owner, and one quality assurance engineer. agreed on unavailability or long response of settlement</p>
      <p>Session 1: Introduction and Architecture Description for calculations as the main system hazard. Therefore, user’s
achieving a common understanding of the workshop settlement can not be calculated is the top event in the
process and the system’s architecture. (1) We resolved fault graph. We analyzed possible causes from the top
misunderstandings regarding the elicited architecture event until we reached basic events that we could not
description through asking questions, and (2) refined the further decompose. We connected diferent causes by
prepared architectural models. logical operators, i.e., AND and OR. For example, users</p>
      <p>Session 2: Hazard Analysis to identify potential causes can not calculate their settlement if it is not processed
for degradation in QoS. Index cards were used as a means in time. This can occur when the assigned instance stalls
to collect hazards. Afterward, the participants arranged OR responds to slow. We argue that the latter can be
the hazards and their causes in a fault-tree-like fashion. experienced if the system receives a sudden (work)load
To not break the participants’ creative flow, we relaxed peak AND its (auto) scaling does not work correctly. The
the strict construction rules of fault trees, e.g., we allowed hazards at the leaf nodes are potential candidates for
events having multiple parents, which resulted in a graph. fault/failure injection during resilience experiments and
«Service»
Payments
«Service»
Eureka
«Service»
Calculations
«Service»
Taxes
«Service»
Employees
«Service»</p>
      <p>SocialInsureances
«Service»
WorkingHours
«Service»
Companies</p>
      <p>A
I-aPG rve«S
teaw ic»e
y
«Service»</p>
      <p>Frontend
The API Gateway
can communicate
with each Service
directly</p>
      <p>IEnvteernmt ediate stakeholders chose the responses and response measures
EUvnednetveloped based on their SLOs.</p>
      <p>AND Gate The stakeholders elaborated 12 resilience scenarios,
(LParcoragcnTeeni)msosCteelbideenints bDeatCaacpatnunreodt bCeacElacxunelnacotuitotend CIanlccourlraeticotn BOaRsiGcaEtveent vsuamriamtiaornizseodf iannTuanbelexp1e. cStceednlaoraidosp0e1akto,i0n4claurdeindgifelriennetar
and exponentially increasing loads. Scenario 05 and 06
dehRaIsneTsstotipmaoonenscsleoew InsstatalnecdGeSaetrevwicaey dSaoneesrsvwinceoert Sewrvitihceetreraocnhrsicwaelrs CInacloDcwnuaistltahaitteionnt wCiaVthcelruWslairootnionng TIencRchoenrtirrceyaclty rsaecnvrdiob0lev8etaharereofaaubinlouduretgamotfeidawdsalieynwgfaaleirleusefrarevisli.ucLreeaissn.tslSytca,enSnccaeer.niSoacr0ieo9na1ar1nidoan10d07
S(acBaualtdion)g PLoeaadk deSAoxeentcesrsrvrwannicseoaehrtles Mcidrdalsehweasre DaistaddfbIouniaeilnssstgeawnwhcoielrekData Loss DDracbeatoeptaartlciwrceieseacntenttelenyodrts
1aa2sndednetedscc-uhrisnbeiecrsat,hleeislesfamuieleusnrcteasuoosffemtdhubelytCiptFhlepelimantsifdtoadrnmlecw,esda.irfeAerceotnrotdrbseupsgulosc,yhEureka ment artifacts and issues intrinsic to individual services
crashes Middcrleaowsthahereesr mInirsggteraatnstceed oofthCEeFrr-rrToeyrlspaetesd Outage Dabfeatetrareoccouavtnaengrodet caoolflvtshecreednsiyafesrrteieonmst csctaoantmeaspferociftsetahltelhsseeyerssvtetiacmbelsai.schcToehdredsienonugvrtcioreotsnh.emIniedtneotntsa-l,
tified system domain context, e.g., payslip calculation
Figure 2: Cleaned Fault Graph periods or simply services being non-idle independent
of the diferent calculations. The response and response
can be initiated by tools such as CTK. The stakeholders measures were specified by the stakeholders based on
selected and prioritized the set of resilience experiments. their internal SLOs.</p>
      <p>Resilience Scenarios: We gave the participants an empty Retrospection: The brief retrospective at the end of the
table according to the ATAM scenario template with the workshop showed that the participants were satisfied
columns (1) source, (2) stimulus, (3) artifacts, (4) environ- with the agenda, content, and outcomes. However,
comment, (5) response, and (6) response measure. Further, ments were made concerning time management.
we explained the meaning of each table column to the
participants. By using index cards again, the participants
steadily added content to the table. We began by iden- 5. Resilience Evaluation
tifying possible sources. The stimuli and artifacts were
then derived from the previously created fault graph. The This section aims to evaluate the case study system’s
environment represents diferent time periods when the resilience. Therefore, we implemented a subset of the
identified stimuli occur. The responses are the stakehold- previously elicited resilience scenarios into resilience
exers’ assumptions about how the system should respond periments using CTK. We compare the system’s behavior
to the particular stimulus. The response measures are against the expected behavior described in the scenarios’
based on their internal Service Level Objectives (SLOs). response part.</p>
      <p>For example, a workload peak resulting in a system
failure was transposed into multiple scenarios. Users of the 5.1. Experiment Setup
system are the source of the scenario since they cause
the load peak. The respective stimulus is the workload 5.1.1. Examined Software System
peak itself. A service was chosen as the artifact to
represent that a load peak can influence all service instances.</p>
      <p>As the environment, the payslip calculation period was
chosen to imply an existing base workload. At last, the
Due to legal constraints and to maintain anonymity, our
industrial partner provided us with a mocked version
as a proxy for the real payroll accounting system. This
version, shown in Figure 3, is used throughout this paper
n
i
m
1
w
o
l
e
b
s
i
e
c
s
e
s
a
c
e
h
t
f
o
n %
i 9
m 9
5 in</p>
      <p>s
in 1 ifed an
ith n itsa tsn
w ih s i
d it e y
ife r a
it w a w
o s s e
n ev O ta
s i L g
te rr S f
g a d o
r n n e
eop itao trsa itm
lev fiitc tsa nw
e o e o
D N R D
d
e
k
c
i
p y</p>
      <p>a
e
b w</p>
      <p>e
n t
a a
c g
t ,r
co g fied u o
.lcca innn iton ,on fiiteod tebd rrse
,e ru ts it n ro w</p>
      <p>e la e b o
ra ce g u b a h</p>
      <p>s
w n re lc ll s
a a a i i d
run itsn levop trco lrew ssce tenn ttrsa
se 1 e b la ro p ro se
U ≥ D A c P u F r
w w p
D</p>
      <p>D A N
,
e
m
i
t
n
i
&amp;
t
c
e
r
r
n
o
i
t
a
l
u
c
l
a
c
e
g
a
g
n
i
r
u
e
c
n
a
t
s
n
I
y
r
d
d
u
o
l
C
n</p>
      <p>g w ta
u
o
F u le r
e
r r
a o
B d e
d p
i O</p>
      <p>M
n
o
i
t
a
l
u
c
l
a
c
p
i
l
s
y
a
p
g
n
i
r
d
l
o
u d
d o
t i
o re</p>
      <p>N p
o
(c d (c d
a</p>
      <p>a
k o k o
a l a l
e g e g
p n p n</p>
      <p>i i
s d s d s
lu lao re lo re se</p>
      <p>a a a
u g c g c ta
im in in )tr in in )tr in
tS sea lly tsa sea lly tsa rm
r ia r ia e
c t ld c t ld t
in n o in n o e
ra ) en (c ra ) en (c cn
ien ttra xop eak ien ttra xop eak tsan
L s E p L s E p I
s
2
w
o
l
e
b
s
i
e
m
i
t
n e
i s</p>
      <p>n
m o</p>
      <p>p
1 s</p>
      <p>e
w r
lo n
e o
b it
is la</p>
      <p>u
e c</p>
      <p>l
im a
t c
n e
w g
o a
D W
s
t
r
a
t
s
e
r e
R im
s t
e n
c i
i
v d
r n
se a</p>
      <p>t
d c
n e
a r</p>
      <p>r
e o
g c
a
s n
s o
e it
m la
r u
ro lc
r a
p
o
h
s
k
r
o
w
e
g
n
i
e
c
S
:
1
E C h</p>
      <p>t
e
o n
n o
, ,
n n
o o
i i
t ta le d
a
l l b e
u u la t
lca leb lca ia a
c la c av re
eg ia e t c</p>
      <p>g o
a av a</p>
      <p>n s
w e w e o
g c g c i
n n n n r
i a i a
r t r t a
u s u s</p>
      <p>n
D in D in</p>
      <p>,S gn R
ign ice iiv ice le
d v e v b
n r c r a
e e e e</p>
      <p>S S R S T
p
i
l
s
y
a
p
,
s
e
s
a
e c
r
u e</p>
      <p>h
s t
ea fo )s</p>
      <p>e
M e
e %y
s 9 lo
n 9 p
o
sp ,in m</p>
      <p>E
e s 0</p>
      <p>0
R 1 3</p>
      <p>(
≤ s
n 0
o 2
i
t
la ≤
u n
c o
l
a i
c t</p>
      <p>a
e l</p>
      <p>u
ag lc</p>
      <p>a
W c
s
t
s
e
u
q e
e m
r i
l t
l
A in
d
n
a
y
l
t
c
e
r
r
o
c
e d
s e
on ldn
p a
s
e h
R re
a
d
t io
en rep
m n
n o
o it
r a
i l
v u
n lc
E a</p>
      <p>c
t
c
a
f
it e
r ic
A rv
p
i
l
s
y
a
P
e
S
d
l
e
c r
r e
u s
o U
S
e
as the system under test. It implements a similar business
logic but with less computational overhead. The system
uses typical patterns of the microservice architectural
style, i.e., API-Gateway-service as a central gateway that
manages all incoming requests and Eureka [18] to provide
service discovery. The payslip-service utilizes an H2
inmemory database and the third-party API Jollyday. It can
forward requests to the payslip-service2. Requests can
also be sent directly to payslip-service2 using a diferent</p>
      <p>The following six endpoints are used during the
experINTERNAL_DEP. — Calls the payslip-service2 via</p>
      <p>DB_READ — Reads an entry from the database of the
r
u EXTERNAL_DEP. — Calls the third-party API
Jollyd
payslip-service.
payslip-service.
days via payslip-service.
payslip-service.
service responds.</p>
      <p>payslip-service2.</p>
      <p>DB_WRITE — Writes an entry into the database of the
GATEWAY_PING — Checks whether the
API-GatewayUNAFF._SERVICE — Sends a request directly to</p>
      <p>The actual payment accounting system is deployed to
a paid CF. Due to financial constraints and legal issues,
the mock system is deployed to a local CF environment
[19], which has similar properties as a paid CF. As CF is a
constraint given by the stakeholders, we did not consider
other cloud providers.
5.1.2. Experiment Tools
four tools, i.e., CTK, load generator, hypothesis
validator, and dashboard. During an experiment, these tools
interact with the system to monitor the experiments and
provide detailed insights, e.g., response times of calls to
individual endpoints.</p>
      <p>To execute the experiments, we used CTK [20], which
can execute and monitor chaos tests and has drivers for
various PaaS solutions. We leveraged the CF driver to
terminate a service instance at a specific point in time</p>
      <sec id="sec-4-1">
        <title>System Under</title>
      </sec>
      <sec id="sec-4-2">
        <title>Test</title>
        <p>generate
load</p>
      </sec>
      <sec id="sec-4-3">
        <title>Loadgenerator write results</title>
      </sec>
      <sec id="sec-4-4">
        <title>Hypothesis</title>
      </sec>
      <sec id="sec-4-5">
        <title>Validation query results</title>
      </sec>
      <sec id="sec-4-6">
        <title>InfluxDB</title>
        <p>execute experiment
retrieve
results</p>
      </sec>
      <sec id="sec-4-7">
        <title>ChaosToolkit</title>
        <p>Load Profile
and validate the steady-state hypothesis. The load that
the system receives is controlled by an adapted version
of the load generator from the TeaStore microservices
benchmark [21] that monitors response times, number
of successful, dropped, and failed requests. The collected
data is written into an InfluxDB [ 22] for a time series
based evaluation. During the evaluation, a Spring
service collects the necessary data from the InfluxDB and
calculates whether a hypothesis holds. We also created a
dashboard application that provides convenient features,
like synchronized starting of CTK and the load
generator, live monitoring, and automated CTK setup. Since
the dashboard does not add functionalities in executing
experiments, it is not part of Figure 4.</p>
        <sec id="sec-4-7-1">
          <title>5.2. Experiment Execution</title>
          <p>Based on Scenarios 04 and 05, we implemented three
resilience experiments. The first experiment investigates
a load peak with an exponential increase (Scenario 04),
while the remaining two investigate instance
termination due to an internal CF error for random instances
(Scenario 05) and specifically the payslip-service
(Scenario 05’). The selection of experiments is based on the
industrial partner’s preferences. In all experiments, the
efect on all endpoints is examined. In the following, we
will only discuss the results of a subset of endpoints for
Scenario 05’. The residual results can be found in the
supplementary material [10].</p>
          <p>The design of the experiment related to Scenario 05’ is
given in Table 2. The target service of this experiment is
the payslip-service, which holds the core business logic of
the mock system. We use CTK to terminate running CF
application instances to simulate the scenario’s stimulus.
The stimulus refers to an error that occurs in CF, which
leads to a loss of an application instance. We assume that
the blast radius only afects the payslip-service and that
CF registers the loss of the payslip-service instance and
starts a new instance. Our hypothesis is that the response
measure of Scenario 05 still holds.</p>
          <p>During the experiments, the system is exposed to an
almost constant, synthetic load. We generated a load
profile with a target load of 20 requests per second and
Hypothesis
Blast Radius</p>
          <p>payslip-service
Terminate payslip-service
application instance
Response measure of</p>
          <p>Scenario 05 holds
payslip-service
some noise. The requests are evenly distributed over all
six endpoints. To assess whether the system still responds
correctly and in time, we measure response times of the
requests and compute their success rate.</p>
        </sec>
        <sec id="sec-4-7-2">
          <title>5.3. Experiment Results</title>
          <p>Figure 5 shows the steady-state, injection, and recovery
phases of the experiment for endpoints INTERNAL_DEP.,
GATEWAY_PING, and UNAFF._SERVICE. In the
steadystate phase, we assume that the system is working as
expected, i.e., the response times satisfy the SLOs. In the
injection phase, CTK terminates the payslip-service
instance. In the recovery phase, we assume that the system
recovers and returns to a steady state, i.e., the response
times satisfy the SLOs. We omitted the load generator’s
warmup and cooldown phase due to readability and
analysis purposes, which refers to the overall first and last
300 s. Further, a 30 s binning was applied, and extreme
outliers (&gt;100 ms) are not shown.</p>
          <p>The success rates at the endpoints INTERNAL_DEP.
(Figure 5a), DB_READ, EXTERNAL_DEP., and DB_WRITE
drop to 0 % as the payslip-service is terminated after
600 s and rises back to 100 % as it recovers in about
1.5 min. During this downtime, no response times are
recorded since no requests arrive at the payslip-service.
During the steady-state and recovery phase, the response
times are stable at around 20 ms and 15 ms, respectively.
During the injection phase, there is a slight increase as
the payslip-service has restarted. The results for
GATEWAY_PING (Figure 5c) and UNAFF._SERVICE (Figure 5e)
show a similar structure. However, the load generator did
not record any successful or failed requests during the
downtime. Therefore, no success rate could be calculated.</p>
        </sec>
        <sec id="sec-4-7-3">
          <title>5.4. Discussion of Results</title>
          <p>As visible in Figure 5 (left side), the response time and
success rate values are almost identical in the steady
state phase and the recovery phase. Furthermore, the
increase in the success rate indicates that the
payslipservice becomes available after 30 s to 60 s. Thus, the CF
platform can re-instantiate the payslip-service quickly,
leading to a quick recovery of the system.</p>
          <p>Response times are slightly higher while the
payslipservice is re-instantiated, which was expected as normal
cold-start behavior. Endpoints GATEWAY_PING and</p>
          <p>S04 Direct Run 1
ID 0 Response Times and Success Rate of</p>
          <p>HTTP Requests</p>
          <p>S04 Direct Run 2
ID 0 Response Times and Success Rate of</p>
          <p>HTTP Requests
0
0
1
0
0
1
0
0
1
0
0
0</p>
          <p>0
300 600 800 1200</p>
          <p>ESxp0e4rimDeinretDcturRatuionn 1(s)</p>
          <p>ID 4 Response Times and Success Rate of
(a) INTERNAL_DEPH.(TwTitPhoRuterqeutreys)ts
300 600 800 1200</p>
          <p>ESxp0e4rimDeinretDcturRatuionn 2(s)
ID 4 Response Times and Success Rate of</p>
          <p>(b) INTERNALH_DTTEPP. (Rweitqhureestrtys)
0
0
0</p>
          <p>0
300 600 800 1200</p>
          <p>ESxp0e4rimDeinretDcturRatuionn 1(s)</p>
          <p>ID 5 Response Times and Success Rate of
(c) GATEWAY_PINGH(TwTitPhoRuterqeutreys)ts
300 600 800 1200</p>
          <p>ESxp0e4rimDeinretDcturRatuionn 2(s)
ID 5 Response Times and Success Rate of</p>
          <p>(d) GATEWAYH_PTINTPG (Rweitqhureestrtys)
)s ●
(m57 ●
s
ieTm0 ●
e 5
opn ●●● ●●●● ●●● ●●● ●●●●● ●● ●● ●●● ●●●●● ●● ●
s
se 52
R
●
● ●</p>
          <p>●
●● ● ● ●●
cuS57
c
e
s
s0
R5
a
t
e
(
%
5
2
)
0
0
1
cuS57
c
e
s
s0
R5
a
t
e
(
%
5
2
)
0
0
1
cuS75
c
e
s
s0
R5
a
t
e
(
%
5
2
)</p>
          <p>Success Rate</p>
          <p>Response times
Experiment Phases</p>
          <p>Steady state
Injection
Recovery
Success Rate</p>
          <p>Response times
Experiment Phases</p>
          <p>Steady state
Injection
Recovery
Success Rate</p>
          <p>Response times
Experiment Phases</p>
          <p>Steady state
Injection
Recovery
0
0
1
0
0
1
0
0
1
0
0
0</p>
          <p>0
300</p>
          <p>600 800
Experiment Duration (s)
1200
300</p>
          <p>600 800
Experiment Duration (s)
1200
(e) UNAFF._SERVICE (without retry)
(f) UNAFF._SERVICE (with retry)
UNAFF._SERVICE should remain unafected during the quest responses at the endpoints GATEWAY_PING and
injection because the payslip-service is not required to UNAFF._SERVICE during injection, which indicates no
answer the requests. Nevertheless, response times at end- requests exist in the system. Another possibility is that
point GATEWAY_PING are afected, which indicates a requests have been dropped. Looking at the raw data
propagation of the failure efects from the payslip-service tables disproves this argument as there are no dropped
to the API-Gateway-service. requests. Another explanation is that no requests arrived</p>
          <p>After the injection started, the success rate drops to at the system, which leads to a lack of data in the time
0 % at the endpoints INTERNAL_DEP., DB_READ, EX- frame between approximately 600 s and 660 s.
TERNAL_DEP., and DB_WRITE. The CTK terminates We hypothesized that the response measure of
Scethe single payslip-service instance. The load generator nario 05 holds, i.e., requests are answered in time (99 % in
lfags all requests as failed, leading to a success rate of less than 1 s) and correctly. As the response times are far
0 %. The plots show neither successful nor failing re- below 1 s, our hypothesis regarding the response times
is technically fulfilled. However, several requests are not in Section 5, each plot is divided into the steady state
answered at all, which is indicated by the dropped success phase, injection phase, and recovery phase. Table 3 shows
rate. We consider these as incorrect response. Therefore, the associated statistical values.
we assume that the hypothesis regarding correctness is In general, similar behavior can be observed at all the
not fulfilled. endpoints. Comparing the plots at left and right of the
Figure 5, shows that the mean response times in the
steady state phase do not vary significantly when the
6. Resilience Improvement retry pattern is activated. Although, at the beginning of
the injection phase, far more high response times can be
The previous section’s experiments showed that the sys- observed. In addition, the boxplots show a slightly higher
tem does not respond as described in Scenario 05 to a interquartile range in the plot where the retry pattern is
failure of an instance of the payslip-service. While the integrated.
response times are technically below 1 s in 99 % of all The plots also show that the success rate does not
cases, requests are temporarily not answered at all, and drop to zero anymore when the pattern is active. For
thus, not correctly. Therefore, we aim to improve the the endpoints INTERNAL_DEP., the success rate drops
system’s success rate concerning Scenario 05 by applying to approximately 70 %. For the two endpoints
GATEresilience pattern(s). We then determine the eficacy of WAY_PING and UNAFF._SERVICE, requests are arriving
improvements to the system’s resilience by re-executing and the success rate remains at 100 %.
the experiments. The application of the retry pattern can explain the
response time spikes during the injection (see the Figure 5).
6.1. Architectural Modifications Requests sent shortly before the restart of the
payslipThe system under test was fortified with a retry pat- service fail, but are retried by the API-Gateway-service
tern [9], i.e., the API-Gateway-service sends another re- until the payslip-service recovered after approximately
quest to the payslip-service if a request fails or remains 10 s. However, as several retries have been aggregated,
unanswered. The retry pattern seems to be a reasonable the payslip-service will have to handle a high amount of
choice since response times are far below the threshold requests upon recovery, resulting in a visible spike in
of 1 s, as indicated by the previous experiment. Due to its response times.
specific purpose, the system has to accept requests near The endpoints UNAFF._SERVICE and GATEWAY_PING
real-time and always answer correctly. Thus, resilience do not depend on the payslip-service. This explains the
patterns that rely on backup or restricting behavior, like high success rate at these endpoints.
circuit breakers or flow limiters, are unsuited. To avoid In contrast to the experiment without the retry
patbad retry behavior, we configured the Spring-Retry as tern, the success rate does not drop entirely. Therefore,
follows. We set the maximum number of retries of each the retry pattern improves the scenario satisfaction as
payslip-service request to be 4, the initial delay to 10 ms, it increased the percentage of correct responses while
the factor for the exponential increase to 3, and the max- keeping the response times below 1 s.
imum delay to 150 ms — resulting in retries after 10 ms,
30 ms, 90 ms, and 150 ms.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>7. Discussion</title>
      <sec id="sec-5-1">
        <title>6.2. Experiment Results and Discussion</title>
      </sec>
      <sec id="sec-5-2">
        <title>7.1. Key Lessons Learned</title>
        <p>Each plot on the right side of Figure 5 visualizes the
system’s response times and success rates with the retry
pattern for an endpoint. As in the experiment presented
Lesson 1: Elicitation of resilience requirements
involves hazard analysis. It is essential to include
stakeholders with diferent roles and particular expertise in
the business domain to quickly prepare a list of relevant 7.2.1. Workshop
hazards. Other roles, such as software developers and
infrastructure engineers, help to identify causes of hazards
that stem from software and its running environment. Conclusion validity One threat is the reliability of</p>
        <p>Lesson 2: ATAM is a useful method to adopt re- measures, which means repeating the workshop yields
silience elicitation. Stakeholders of the software project the same resilience requirements list. Elicitation of
rewere already familiar with scenario development for qual- silience requirements involves human judgment. Hence,
ity requirements. Therefore, the structure of the scenario it is a subjective measure. Therefore, we can not entirely
template of Bass et al. [12] was intuitive for the stake- rule out this threat.
holders. Internal validity One threat is instrumentation, which</p>
        <p>Lesson 3: Loose adoption of formalisms is already means our tools and techniques were not suitable. We
good enough. Researchers and practitioners have used conducted a one-day structured workshop and used the
fault tree formalism for both qualitative and quantitative scenario template of Bass et al. [12] for eliciting resilience
hazard analysis in safety engineering. To identify the requirements. We refined all the resilience requirements
causes of a hazard, we did not have to comply with fault through several iterations after the workshop and
valitree formalism rigorously. The informal way of construct- dated them against the workshop participants.
ing a fault tree was easy to understand for stakeholders. Construct validity For us, the main threat in this</p>
        <p>Lesson 4: The ATAM workshop requires consid- category is mono-method bias, which means we did not
erable refinement that can be done “ofline”. The use other elicitation methods. Therefore, there is a threat
outcome of the well prepared one-day workshop needed that elicited resilience requirements are biased. We can
further refinement. In particular, it was necessary to re- not entirely rule out this threat as we did not apply other
ifne the stimulus and response measures parts of each methods and cross-check the results.
scenario, e.g., we modeled the workload and tried to ex- External validity The heterogeneity poses a threat,
press the scenarios in temporal logic. This revealed that i.e., diferent roles and expertise of participants.
Workthe initial requirements were partially ambiguous and im- shops with less heterogeneity in the stakeholders could
precise, which was easy to resolve through clarification lead to no resilience requirements. We can not entirely
requests to the stakeholders. Therefore, we hypothesize rule out this threat.
that formalization benefits both validation and
quantitative evaluation of resilience requirements and that an 7.2.2. Experiment design
explicit (ofline) formalization step could complement the We used the mock system for quantitative evaluation
proposed workshop well. of resilience requirements that are based on the actual</p>
        <p>Lesson 5: A tightly planned one-day workshop is system. There is a threat that evaluation results are
insuficient. We managed to collect resilience scenarios in accurate. However, the purpose of the experiments is to
a one-day workshop because it was well prepared (know- exemplary show how elicited requirements and derived
ing the architecture description) and well-conducted (strict experiments can help to improve the system — we do not
time management). Refinement can be done ofline by claim the accuracy of the quantitative results.
Furthermore skilled engineers in formalizing stimuli and re- more, due to legal issues, we used CF Dev [19]. We faced
sponse measures (similar to writing SLOs). However, instability, e.g., resource drainage of Dev nodes, in the
it is important to ask for feedback to check the validity environment during experimentation. There is a threat
of the requirements. of a negative impact on results due to this instability. To</p>
        <p>Lesson 6: The resilience elicitation helps to re- counteract this threat, we re-executed experiments to
ifne “classical” QoS requirements. All response mea- gain insight into approximate measurements, ensuring
sures are based on non-resilience specifications that make reliable data with no unintended node or service crash.
them imprecise. For example, maximum degradation and
time to recovery was not specified. Thus, it is unclear
whether experimentation shows acceptable or unaccept- 8. Conclusion
able degradation in performance or availability quality.</p>
      </sec>
      <sec id="sec-5-3">
        <title>7.2. Threats to Validity</title>
        <p>We discuss the threats to validity for the workshop and
our experiment design.</p>
        <p>The successful development of resilience scenarios
depends on the outcome of the hazard analysis. Our
approach to scenario-based resilience evaluation assumes
a business domain expert to derive an initial list of
hazards. FTA can then be a means to analyze the hazards
and derive resilience scenarios. We plan to (1) extend our
process with an explicit formalization step after the
workshop for refinement of the scenarios, (2) formally verify
response measures of resilience scenarios, and (3) create
processes for continuous hazard analysis when a system
faces changes, e.g., updates and refinement/development
of resilience scenarios.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work has been supported by the Baden-Württemberg
Stiftung (ORCAS — Eficient Resilience Benchmarking
of Microservice Architectures) and the German Federal
Ministry of Education and Research (Software Campus
2.0 — Microproject: DiSpel).</p>
      <p>Our artifacts [10] comprise (i) the resilience scenarios
and (ii) the data and R scripts as a CodeOcean capsule.</p>
      <p>We are working on making parts of the created/modified
experiment tools available as open-source software. For
confidentiality reasons, the system under test cannot be
published.
[10] S. Frank et al., Supplementary material, 2020.</p>
      <p>Artifacts: https://doi.org/10.5281/zenodo.5142006
(Scenarios); https://doi.org/10.24433/CO.0520280.v1
(Code Ocean capsule).
[11] K. Pohl, Requirements Engineering - Fundamentals,</p>
      <p>Principles, and Techniques, Springer, 2010.
[12] L. Bass, P. Clements, R. Kazman, Software
Architecture in Practice, 2 ed., Addison-Wesley Longman</p>
      <p>Publishing Co., Inc., USA, 2003.
[13] J. Cámara, R. de Lemos, Evaluation of resilience
in self-adaptive systems using probabilistic
modelchecking, in: Proc. 7th Int. Symposium on Software
Engineering for Adaptive and Self-Managing
Systems (SEAMS), 2012, pp. 53–62.
[14] J. Cámara, R. de Lemos, M. Vieira, R. Almeida,</p>
      <p>R. Ventura, Architecture-based resilience
evaluation for self-adaptive systems, Computing 95 (2013)
689–722.
[15] J. Cámara, R. de Lemos, N. Laranjeiro, R. Ventura,</p>
      <p>M. Vieira, Robustness-driven resilience
evaluation of self-adaptive software systems, IEEE
Transactions on Dependable and Secure Computing 14
(2017) 50–64.
[16] R. Natella, D. Cotroneo, H. Madeira, Assessing
dependability with software fault injection: A survey,
[1] S. Newman, Building Microservices, O’Reilly, 2015. ACM Computing Surveys (CSUR) 48 (2016) 44:1–
[2] V. Heorhiadi, S. Rajagopalan, H. Jamjoom, M. K. 44:55.</p>
      <p>Reiter, V. Sekar, Gremlin: Systematic resilience [17] K. Yin, Q. Du, W. Wang, J. Qiu, J. Xu, On
testing of microservices, in: Proc. 36th IEEE Int. representing and eliciting resilience requirements
Conf. on Distributed Computing Systems (ICDCS), of microservice architecture systems, CoRR
2016, pp. 57–66. abs/1909.13096 (2020). URL: https://arxiv.org/abs/
[3] A. Basiri, N. Behnam, R. de Rooij, L. Hochstein, 1909.13096v3. arXiv:1909.13096.</p>
      <p>L. Kosewski, J. Reynolds, C. Rosenthal, Chaos engi- [18] Netflix Inc., Eureka, 2020. URL: https://github.com/
neering, IEEE Softw. 33 (2016) 35–41. Netflix/eureka.
[4] Chaos toolkit, 2020. URL: https://github.com/ [19] Cloud Foundry Foundation, Cloud foundry dev
chaostoolkit. documentation, 2020. URL: https://github.com/
[5] R. Miles, Learning Chaos Engineering – Discover- cloudfoundry-incubator/cfdev.
ing and Overcoming System Weaknesses through [20] Chaos Toolkit, Chaos toolkit documentation, 2020.</p>
      <p>Experimentation, O’Reilly Media, Inc., 2019. URL: https://chaostoolkit.org.
[6] N. G. Leveson, Safeware — System Safety and Com- [21] J. von Kistowski, S. Eismann, N. Schmitt, A. Bauer,
puters: A Guide to Preventing Accidents and Losses J. Grohmann, S. Kounev, Teastore: A micro-service
Caused by Technology, Addison-Wesley, 1995. reference application for benchmarking, modeling
[7] D. Kesim, A. van Hoorn, S. Frank, M. Häussler, Iden- and resource management research, in: Proc. IEEE
tifying and prioritizing chaos experiments by using 26th Int. Symp. on Modeling, Analysis, and
Simulaestablished risk analysis techniques, in: Proc. 31st tion of Computer and Telecommunication Systems
Int. Symposium on Software Reliability Engineer- (MASCOTS), 2018, pp. 223–236.</p>
      <p>ing (ISSRE), 2020. [22] InfluxData Inc., InfluxDB website, 2020. URL: https:
[8] R. Kazman, M. Klein, M. Barbacci, T. Longstaf, //www.influxdata.com/.</p>
      <p>H. Lipson, J. Carriere, The architecture tradeof
analysis method, in: Proc. 4th IEEE Int. Conf. on
Engineering of Complex Computer Systems (ICECCS),
1998, pp. 68–78.
[9] M. T. Nygard, Release It!: Design and Deploy</p>
      <p>Production-ready Software, Pragmatic Bookshelf,
2018.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>