<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Survey of Big Data Pipeline Orchestration Tools from the Perspective of the DataCloud Pro ject∗</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mihhail Matskin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shirin Tahmasebi</string-name>
          <email>shirint@kth.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amirhossein Layegh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amir H. Payberah</string-name>
          <email>payberah@kth.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aleena Thomas</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nikolay Nikolov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dumitru Roman</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>KTH Royal Institute of Technology</institution>
          ,
          <addr-line>Stockholm</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SINTEF AS</institution>
          ,
          <addr-line>Oslo</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <fpage>63</fpage>
      <lpage>78</lpage>
      <abstract>
        <p>This paper presents a survey of existing tools for Big Data pipeline orchestration based on a comparative framework developed in the DataCloud project. We propose criteria for evaluating the tools to support reusability, flexible pipeline communication modes, and separation of concerns in Big Data pipeline descriptions. This survey aims to identify research and technological gaps and to recommend approaches for filling them. Further work in the DataCloud project is oriented towards the design, implementation, and practical evaluation of the recommended approaches.</p>
      </abstract>
      <kwd-group>
        <kwd>Big Data pipeline</kwd>
        <kwd>Orchestration tools</kwd>
        <kwd>Reusability</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The availability of massive amounts of data has tremendously changed the data
collection process and analysis over the last few years. The concept of Big Data
and the supporting solutions have allowed dealing with potentially unlimited
heterogeneous data in diferent formats within a practically acceptable time.
However, the growth of data has increased both opportunities and challenges.
In terms of opportunities, data processing is being heavily invested in to
empower the decision-making process of organizations in possession of Big Data.
Furthermore, the data analytics process is becoming complex due to the
characteristics of Big Data, the sophisticated tools and technologies involved, diferent
interests among stakeholders, often changing business needs, and the lack of a
standardized process for the lifecycle of Big Data pipelines [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>
        Because of the complexity of Big Data analysis tasks, the software
supporting such analysis requires a combination of a broad spectrum of trusted software
components. Such a combination involves integrating components into pipelines
that take care of the pipeline execution and data transfer. The design and usage
of Big Data pipelines increase the eficiency of data analysis, while at the same
time require support for designing and managing the pipelines. Many
organizations recognize the significance of Big Data pipelines, however there are still
critical challenges in their implementation, such as the heterogeneity of involved
stakeholders and limited knowledge reuse [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In this paper, we refer to the tools
that support the description and execution of Big Data pipelines as Big Data
pipeline orchestration tools. Although there exist many orchestration tools for
Big Data pipelines, they have not focused on crucial areas such as reusability
and separation of concerns.
      </p>
      <p>
        The problem of combining diferent components into an executable process is
not new. Workflow systems are systems that support the integration of steps of
a semi- or fully automated procedure into a manageable process. Business
worklfows are workflow systems oriented towards automating business processes [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Scientific workflows are other types of workflow systems oriented towards
automation of scientific experiments [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Recently, Big Data workflows are becoming
prevalent and refer to modeling processes containing various Big Data analytic
or processing steps [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. The main characteristics of Big Data workflows are the
dynamics and heterogeneity of data sources and processing components, which
typically require diferent orchestration models. Big Data pipelines are special
cases of Big Data workflows where workflows are more oriented towards the
endusers. However, since there is no clear boundary between Big Data workflows and
Big Data pipelines, in this work we do not make an explicit diference between
them.
      </p>
      <p>The approach in this paper is based on (1) extracting requirements for Big
Data pipelines from the business cases of the DataCloud project as well as
existing scientific literature and software tools around Big Date pipelines, and (2)
analyzing which of the existing solutions (if any) can satisfy the identified
requirements. For the requirements extraction, we defined the following Research
Questions (RQ), which guided our work:
– RQ1. How to bridge the technological gap between diferent experts involved
in the Big Data pipeline design, implementation, and management?
– RQ2. How to support the reuse of previously developed knowledge and
solutions in designing and implementing Big Data pipelines?
– RQ3. How to support debugging of Big Data pipelines?
In this paper, we survey existing Big Data pipeline orchestration tools, identify
research and technological gaps, and suggest approaches to filling the gaps. This
work is done in the context of the HORIZON 2020 project DataCloud3. The
rest of the paper is organized as follows. Section 2 provides a brief overview of
the DataCloud project, including its objectives and expected results. Section 3
identifies the requirements and the classifiers for building a comparison table of
the existing tools relevant to the DataCloud project perspective, and includes
the comparison tables to identify gaps in the current solutions. The final Section
4 summarizes the main findings and refers to ongoing work in the DataCloud
project to solve the identified gaps.
2</p>
      <p>DataCloud Project Perspective
In this section, we provide a brief overview the DataCloud project and present
some requirements for Big Data pipeline orchestration tools identified while
working on the project.
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>The DataCloud project</title>
      <p>The DataCloud project, which runs between 2021 and 2023, aims to develop a
novel paradigm for Big Data processing over heterogeneous resources, including
the Cloud/Edge/Fog Computing Continuum. The core concept of the project is
Big Data pipelines, whose complete lifecycle is supported by several processing
capabilities. The DataCloud project utilizes this paradigm to solve issues for a
broad set of business cases coming from Small and Medium-sized Enterprises
(SMEs) and large organizations with dificulties in capitalizing on Big Data due
to the lack of technical expertise and suitable processing capabilities.</p>
      <p>In the DataCloud project, we develop a set of new languages, methods,
infrastructures, and software prototypes for discovering, simulating, deploying,
and adapting Big Data pipelines on heterogeneous and untrusted resources. The
project underlines the separation of concerns and separates the design from the
run-time aspects of their deployment. This separation allows domain experts
without significant technical/programming knowledge to participate in the
definition and management of Big Data pipelines. Moreover, the DataCloud
solutions allow the incorporation of Big Data pipelines in organizations’ business
processes more eficiently and make them more accessible to a broader set of
stakeholders regardless of the hardware infrastructure. DataCloud assumes a
typical Big Data pipeline lifecycle involving a set of high-level processing steps
executed in a loop (see Figure 1). The project aims to deliver its solutions to
support the Big Data pipelines lifecycle as a set of interoperable tools that form the
DataCloud toolbox. The toolbox includes the following components (see Figure
2):
– DIS-PIPE: Provides integration of process mining techniques and Artificial
Intelligence (AI) algorithms to learn the structure of Big Data pipelines. The
learning is based on extracting, processing, and interpreting huge amounts
of event data collected from heterogeneous data sources.
– DEF-PIPE: Provides support for the visual design and description of Big
Data pipelines based on a Domain Specific Language (DSL). The tool
includes the means to store and load the pipeline definitions that enable the
reuse of previously developed solutions. It also enables data experts and
domain experts to define the content by configuring individual steps and
injecting code or customizing generic predefined step templates.
– SIM-PIPE: Simulates the enactment of Big Data pipelines and provides
pipelines testing functionality, including a sandbox for evaluating
individual pipeline step performance. Furthermore, SIM-PIPE provides a simulator
to analytically predict the performance of the overall Big Data pipelines
across the Computing Continuum resources.
– R-MARKET: Deploys a decentralized backbone resource network based on
a hybrid permissioned and permissionless blockchain. This component
provides a marketplace for resources and enables transparent provisioning of
resources that increase the overall trust.
– DEP-PIPE: Enables elastic and scalable deployment of Big Data pipelines
with real-time event detection and automated decision making.
– ADA-PIPE: Provides a data-aware algorithm for intelligent and adaptive
provisioning of resources and services across the Computing Continuum as
well as intelligent resource reconfiguration.</p>
      <p>We evaluate the DataCloud solutions on five business cases provided by the
DataCloud business partners, which cover a broad spectrum of Big Data pipeline
applications:
– Smart mobile marketing campaigns.
– Automatic live sports content annotation.
– Digital health system.
– Predicting deformations in ceramics.</p>
      <p>– Analytics of manufacturing assets.
2.2</p>
    </sec>
    <sec id="sec-3">
      <title>Identifying Requirements</title>
      <p>The diversity and complexity of modeling data, processing Big Data pipelines,
and the heterogeneity of Computing Continuum platforms require a
multidisciplinary efort using expert knowledge of the domain, data, and technical
knowledge of the computational environment. However, the collaboration among
domain, data, and technical experts requires repeated communication cycles
introducing significant overhead and barriers to success. Therefore, it is crucial
and challenging to provide tools to bridge the technological and knowledge gaps
between all relevant experts and enable them to collaborate while keeping the
separation of concerns. Nevertheless, there are many design solutions for Big
Data pipelines, so reusing them can boost the design and the development of
new pipelines. The recent development of containers (as a technology allowing
eficient and reliable installation of software components on diferent platforms)
provides a good foundation for implementing reusable solutions. This is why
containerization is a promising approach to support reusable solutions.</p>
      <p>In designing Big Data pipelines, we need to consider the computational
resources and decide about the deployment of the pipelines. However, it is not easy
to make such a decision in the design phase in many cases, and running pipelines
in the production computational environment might be expensive. Therefore, to
make such a process more eficient, we need a simulation and debugging
environment for pipelines that operates without deploying the pipeline. This is why
combining pipeline descriptions with simulation tools is an essential requirement
for modern systems.</p>
      <p>By summarizing the above aspects and taking into account the properties of
the DataCloud project described in Section 2.1, we outline the following
requirements for Big Data pipeline orchestration tools:
– Req1: Provide separation of concerns between design and run-time aspects.
– Req2: Provide convenient means for describing pipelines including visual and
textual (DSL-based) interfaces.
– Req3: Support reusability of the previously developed steps and pipelines in
designing new pipelines.
– Req4: Provide flexible data transfer between steps in pipelines.
– Req5: Support containerization for nodes and pipeline descriptions.
– Req6: Provide smooth integration of description and simulation of
components.
3</p>
      <p>Overview of Existing Solutions
In this section, we describe a set of criteria that we consider important for Big
Data pipeline orchestration tools and compare current solutions with respect to
those criteria. Most of the criteria refer to the requirements identified in Section
2.2.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Criteria for Comparison</title>
      <p>The traditional way of representing pipelines is DSL that is incorporated into
an orchestration environment. We performed an analysis of existing tools based
on the following classifiers reflecting the requirements identiefid in the previous
section.</p>
      <p>Type of Workflow/Pipeline. We consider three types of workflows/pipelines:
business workflows , scientific workflows, and Big Data workflows (defined in
Section 1). Each tool falls into one of these three types according to its main
applications.</p>
      <p>
        Workflow/Pipeline Model. These categories are not mutually exclusive.
Different possible workflow models can be categorized as follows:
– Script-based: In these type of workflows the composition of nodes is described
using scripting languages. These workflows are useful for expert users to
design complex applications more flexibly and concisely [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ].
– Event-based: These workflows are characterized by a discrete set of states.
      </p>
      <p>
        A transition from one state to another happens on the occurrence of events
emitted asynchronously or by an external trigger. Hence, in event-based
workflows, users define event rules to declare under what circumstances state
transitions should occur. This provides a responsive orchestration process [
        <xref ref-type="bibr" rid="ref50 ref6">6,
50</xref>
        ].
– Adaptive: This model allows designing adaptive or context-aware workflows
to consider the runtime situations and exceptions, such as a failure in
preprocessing an input file. Such workflows can respond to dynamic
environmental needs efectively [
        <xref ref-type="bibr" rid="ref1 ref49 ref58 ref6">1, 6, 49, 58</xref>
        ].
– Declarative: In a declarative approach, a minimal set of requirements, which
are often expressed by a set of constraints, is defined. Therefore, the
execution of the workflow is allowed until this set of constraints is satisfied. This
is advantageous for increasing flexibility and is especially useful in highly
unpredictable contexts, in which there are a large number of allowed and
possible alternatives [
        <xref ref-type="bibr" rid="ref13 ref14 ref7">7, 13, 14</xref>
        ].
– Procedural: This workflow model explicitly specifies the sequence of steps
and tasks known as control-flow. Thus, during the execution of the model, it
is possible to execute a process only as explicitly specified in the control-flow
[
        <xref ref-type="bibr" rid="ref14 ref44">14, 44</xref>
        ].
Separation of Concerns. This criterion is related to the Req1 requirement.
If a tool has a mechanism for the separation of high-level workflow definition
concerns from step-specicfi implementation and deployment details, it supports
separation of concerns for stakeholders (the value of the classifier is Yes in the
tables in the rest of the paper). Otherwise, it is concluded that separation of
concerns is not a focus [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] (the value of the classifier is No).
      </p>
      <p>Type of Language. This criterion is related to the Req2 requirement. We
consider the following types of languages:
– General-Purpose Language (GPL): GPL is a highly applicable language (e.g.,</p>
      <p>Python, R, Java) across a variety of application domains.
– Domain-Specific Language (DSL): DSL is a language that is specially
designed for a specific problem domain. The main advantage of using a DSL
is that domain experts, who have little knowledge outside of their domain,
can eficiently design relevant parts of the pipeline logic. On the other hand,
its main drawbacks are the limited portability across diferent environments
and low applicability outside the discrete domain.</p>
      <p>Input Supported. This criterion is related to the Req2 requirement. Here,
we consider two possible input types for designing and implementing Big Data
pipelines: text-based input and graphical or visual input. We assume that tools
with visual input types may also support text-based input types.
Ease of Use. This criterion is related to the Req2 requirement. The possible
values for this classifier are Hard, Medium, and Easy, which refer to the level
of expertise a user needs to have to be able to use the tool. For example, if a
tool allows designing a workflow through a clear graphical interface, the tool is
considered to be Easy to use. However, if the tool relies on using a lightweight,
human-readable text format, such as YAML, XML, and JSON, the level of
easeof-use is considered Medium. Otherwise, if programming knowledge and skills are
needed for using the tool, it is concluded that the tool is Hard to use. Sometimes
it is dificult to place a tool into one only category and we allow a combination
of values, for example, Medium/Hard or Easy/Medium.</p>
      <p>
        Focus on Reusability. This criterion is related to the Req3 requirement.
Reusability refers to the characteristic of a designed workflow, whereby it can
be used to create another similar workflow. The reusability is a tool’s focus if
(i) it provides a visual drag-and-drop feature for reusing the previously-designed
workflows, or (ii) it supports a search for previously-designed similar solutions
to be imported as text-based workflows for reusing in designing the current one.
Otherwise, it is concluded that reusability is not a focus [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. For the former
group of tools, the value of the classifier is Yes, and the value of the classifier for
the latter is No.
      </p>
      <p>Reusable elements. This criterion is related to the Req3 requirement. In this
research, the reusability of each tool is evaluated in three aspects:
– Whether or not it is possible to reuse the definition of the entire workflow.
– Whether or not it is possible to reuse the definition of each step.
– Whether or not it is possible to reuse the step implementation.</p>
      <p>The possible values for each of these aspects are Yes, No, and Partial, which
means that although no specicfi way is designed to share assets, sharing can be
done by manually copying scripts.</p>
      <p>Nested Step Definition This criterion is related to the Req3 requirement.
This is a binary classifier, the value of which would be Yes, if the tool has a
specific way for defining and using a step inside and as a part of another more
granular step. Otherwise, its value would be No.</p>
      <p>Configurable Data Transmission Medium Definition. This criterion is
related to the Req4 requirement. This classifier has two possible values; Yes
and No, depending on whether or not the tool allows users to choose between
multiple data transmission mediums (such as shared file system, web services,
RPC, file transfer, HTTP, and FTP) for transferring data between steps during
the pipeline execution.</p>
      <p>Configurable Communication Medium Definition. This criterion is
related to the Req4 requirement. This classifier has two possible values; Yes and
No, depending on whether or not the tool supports choosing the medium of
control flow communication between steps during the pipeline definition,
invocation, and execution. For example, this could be done using RESTful API,
message queues, or RPC.</p>
      <p>
        Containerization. This criterion is related to the Req5 requirement. In
diferent workflow types, containers can be used to automate the deployment process,
enabling better scalability. We consider three possible approaches for
containerization [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]:
– Workflow-level: In this approach, the entire workflow is wrapped inside a
container allowing scalability of the full workflow.
– Step-level: In this approach, each step is encapsulated and wrapped inside a
container allowing scalability of individual steps.
– Encapsulation: Here, a Big Data pipeline framework/tool is available as a
container image for installation and local or distributed usage.
      </p>
      <p>Integrates a Simulation Tool. This criterion is related to the Req6
requirement. Here, we consider if a tool integrates a simulation tool. The value of this
classifier is either Yes or No.</p>
      <p>Monitoring. If the tool provides means for monitoring the execution of the
pipeline. Depending on how the execution can be monitored, it can be further
classified into the following three sub-categories:
– Runtime: If the monitoring is available real-time.
– Logging : If the run logs are available at the end of the execution.
– No: If the tool ofers no monitoring support.
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Tools Comparison</title>
      <p>In Tables 1-4, we compare a representative number of the most popular Big
Data pipeline orchestration tools with respect to the classifiers described in the
previous subsection. In these tables denotes Yes and denotes No.
to rs
wflokro WeritnE in in
g e</p>
      <p>a
p t
wflokro WeritnE i
a
ifc
s
s
es Ufo esaE C r
d
a
.
t
u
p
n</p>
      <p>I
m le
b
7
l
e
d
o
M
w
lfo
k
r
o
W</p>
      <p>, ]4 ] ]6 , ]
lfireohwAcap ,[]303 lfrookwWrog][4 [rcea40ydhm ]42 []59lreepK i[]ftrcaa65T [irsee5apPm l[eo32Fwak [-reo1CTAG [eea28akknm ]25 [,sse12ag43u
A A P tS M C S P</p>
      <p>T T T T
L L L L
S S S S</p>
      <p>L L L L</p>
      <p>S P S S
5
L
S
D
p n
a o
wflokro WeritnE i
a
c
i
f
s
s
gniroti no M t
m
i
n
u
es Ufo esaE C d r
a m
l u</p>
      <p>i d
. e a
m
u le
p
n
I
t
u
y
r</p>
      <p>R
a
2
a
ta ific ific ific ific ific</p>
      <p>t t t t t
D n n n n n
g ie ie ie ie ie
i c c c c c
R R L L R R R R
l l
a a
i i
t t
r r
a a</p>
      <p>P P
t
x
e
P
L</p>
      <p>L L</p>
      <p>L L L</p>
      <p>L
P P S P P S</p>
      <p>l
t t t t t a
x x x x x u
e e e e e s</p>
      <p>i</p>
      <p>T T T T T
T</p>
      <p>T</p>
      <p>T</p>
      <p>V</p>
      <p>V</p>
      <p>D G G D G G D G G D
s
l
o
o
T
ifc
i
t
n
e
i
c
S
]
5
3
[
w
o
l
F
t
x
e
]
9
2
[
2
7
d y
r s
a a</p>
      <p>P m m Pm m
d iu iu iu iu y
r d d d d s
a e e e e a
M H</p>
      <p>H E</p>
      <p>E
H M M M M
noiti fine DpetS n s</p>
      <p>S b
a
u
io e
t R
noitatne mel p mI petS t
y
r l
a a</p>
      <p>i
m r
a
wflokro WeritnE iif ita
a</p>
      <p>l
c
s r
s - a
es Ufo esaE C -sy iud id rd id sy sy rd
a
l</p>
      <p>u u
. a e e a e a a a
m
u le
lfo
k
r
o
W
T
L</p>
      <p>L</p>
      <p>L
R</p>
      <p>R</p>
      <p>R R
3
7
E</p>
      <p>E</p>
      <p>E
M</p>
      <p>M
V</p>
      <p>V</p>
      <p>V
D D D D D</p>
      <p>D
sed lra8iev sed lra sed eiv eitv sed sed lra sed ev itev lra sed ev itev lra sed sed lra sed sed sed ev itev lra
i-trcapbS ceduroP ltrceaaD i-trcapbS ceduroP i-trcapbS tadpA lrceaaD i-trcapbS -teanbvE rceoduP i-trcapbS itadpA lrceaaD rceoduP i-trcapbS itadpA lrceaaD rceoduP i-trcapbS -teabnvE rceoduP i-trcapbS i-trcapbS -teavbnE itadpA lrceaaD rceoduP
y
s
a
l
a
u
s
i
L
S
a
t
a
g
i
,
8
3
[
D</p>
      <p>]
E 9
R 3
e
y
s
a
l
a
u
s
i
L
S
a
t
a
g
i
]
7
2
[
E
M
I
N
K
m
u
i
d
e
t
x
e
T
L
S
s
s
e
D</p>
      <p>D</p>
      <p>B S
B
noitatne mel p mI petS</p>
      <p>m
noiti fine DpetS n su</p>
      <p>S b</p>
      <p>a
io e
t R
a
wflokro WeritnE iif ita
c l
s r
s a</p>
      <p>P
t 4</p>
      <p>t
ex ex
T
t
u
lfo
k
r
o
W
T
a
t
a
g
i
]
1
5
,
8
4
[
r
e
t
t
i
k
S
a
t
a
g
i
]
1
1
[
r
e
t
s
g
a
D
e</p>
      <p>e
4
7
Pm P
M H
P
]
7
4
[
w
lfo
e
R
y
s
a
l
a
u
s
i
V
L
S
s
s
e
o [9
C
L
P</p>
      <p>G
T</p>
      <p>T
a
t
a
g
i
]
6
4
,
5
4
[
t
c
e
f
e
r</p>
      <p>D
]
6
3
[
e
h
c
a
p
d d d l d d e l e l d d e l d d e l
se ev se se ev ra se se ev itv ra iv ra se se ev itv ra se se ev itv ra
i-trcabp itapdA i-trcabp -teabnv itapdA rceoduP i-trcabpS -teanbvE itapdA lrceaa rceodu ltrceaa rceodu i-tabp -teanb itapd lrcaaD ceoduP i-trcabp -teavbn itapdA lrceaa rceodu</p>
      <p>D P D P rcS vE A e r</p>
      <p>S S E S E D P</p>
      <p>Conclusions
In this paper we investigated existing Big Data pipeline orchestration tools. By
analysing them (in Section 3.2) we show that important requirements defined
in the DataCloud project are not (or are only partially) satisfied by the
currently available tools. In particular, only a few tools support a graphical input
language for the description of pipelines. While several tools allow some levels
of reusability, most reusability aspects (such as support for searching available
suitable solutions) are not implemented. Moreover, integration with other tools
(including simulation tools) is not supported by many of them. However, we can
mention that Apache Airflow, Argo Workflow, and Snakemake are the closest
tools to our requirements among the reviewed ones. Nevertheless, although some
aspects of some requirements are considered in some tools, no tool supports all
required aspects.</p>
      <p>
        To address these problems in the DataCloud project, we are developing a
DEF-PIPE component (see Section 2.1) for designing Big Data pipelines. This
component will provide means for the description and manipulation of pipelines
and the environment. It will also support a complete graphical and textual
interface, accumulation and reuse of solutions across diferent applications,
flexible integration with a simulation tool, and separation of concerns between the
description of design-time and run-time aspects. Preliminary results in this
direction are reported in [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basten</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verbeek</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verkoulen</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voorhoeve</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Adaptive workflow: On the interplay between flexibility and support</article-title>
          .
          <source>In: Enterprise Inf. Systems</source>
          , pp.
          <fpage>63</fpage>
          -
          <lpage>70</lpage>
          . Springer (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Accio: Accio - Workflow
          <string-name>
            <surname>Authoring</surname>
          </string-name>
          . https://privamov.github.io/accio/docs/ creating-workflows.html, [Online; accessed 01-April-2021]
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Airflow</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          : Tutorial. https://airflow.apache.org/docs/apache-airflow/ stable/tutorial, [Online; accessed 25-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>4. Argo: Argo Workflow - Docs. https://argoproj.github.io/argo-workflows/, [Online; accessed 21-April-2021]</mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Barga</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , D.:
          <article-title>Workflows for e-Science, chap</article-title>
          .
          <source>Scientific versus Business Workflows</source>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>16</lpage>
          . Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Barika</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garg</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zomaya</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moorsel</surname>
            ,
            <given-names>A.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ranjan</surname>
          </string-name>
          , R.:
          <article-title>Orchestrating big data analysis workflows in the cloud: research challenges, survey, and future directions</article-title>
          .
          <source>ACM Computing Surveys (CSUR) 52(5)</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>41</lpage>
          (
          <year>2019</year>
          )
          <article-title>4There exist visual and graphical tools for working with workflows, but not for creating and defining workflows. 5Despite the fact that Java programming language is used for defining data sources, data sinks, and data processors, since they are all converted to RDF notation, and also since the graphical tool is used for defining the whole control oflw, the language is DSL, not GPL</article-title>
          .
          <article-title>6Data sources, data sinks, and data processors are implemented in Java programming language. However, for creating the whole workflow, there exists a GUI interface. 7The DSL provides domain-specific syntax and is built on top of Python 3. 8Provides explicit support for external automation tools to deploy experiments.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Bernardi</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimitile</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Di</given-names>
            <surname>Lucca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Maggi</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.M.:</surname>
          </string-name>
          <article-title>Using declarative worklfow languages to develop process-centric web applications</article-title>
          .
          <source>In: 2012 IEEE 16th International Enterprise Distributed Object Computing Conference Workshops</source>
          . pp.
          <fpage>56</fpage>
          -
          <lpage>65</lpage>
          . IEEE (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>BioDepot: BioDepot-Workflow-Builder - General Information</surname>
          </string-name>
          . https://bwb. readthedocs.io/en/latest/#general-information, [Online; accessed 01-April2021]
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>BMC: BMC Control-M - Docker</surname>
          </string-name>
          <article-title>Image with Embedded Agent</article-title>
          . https: //docs.bmc.com/docs/display/public/ctmapitutorials/Manage+workload+ in+Docker+Containers, [Online; accessed 26-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>BPipe: BPipe - Pipeline Even</surname>
          </string-name>
          . http://docs.bpipe.org/Guides/ PipelineEvents/, [Online; accessed 02-April-2021]
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>11. Dagster: Dagster - Concepts. https://docs.dagster.io/concepts, [Online; accessed 26-March-2021]</mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Deelman</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vahi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juve</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rynge</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Callaghan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maechling</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mayani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Da</surname>
            <given-names>Silva</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.F.</given-names>
            ,
            <surname>Livny</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          , et al.:
          <article-title>Pegasus, a workflow management system for science automation</article-title>
          .
          <source>Future Generation Computer Systems</source>
          <volume>46</volume>
          ,
          <fpage>17</fpage>
          -
          <lpage>35</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Demeyer</surname>
            , R., Van Assche,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Langevine</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanhoof</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Declarative workflows to eficiently manage flexible and advanced business processes</article-title>
          .
          <source>In: Proceedings of the 12th international ACM SIGPLAN symposium on Principles and practice of declarative programming</source>
          . pp.
          <fpage>209</fpage>
          -
          <lpage>218</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>van Der Aalst</surname>
            ,
            <given-names>W.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pesic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schonenberg</surname>
          </string-name>
          , H.:
          <article-title>Declarative workflows: Balancing between flexibility and support</article-title>
          .
          <source>Computer Science-Research and Development</source>
          <volume>23</volume>
          (
          <issue>2</issue>
          ),
          <fpage>99</fpage>
          -
          <lpage>113</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Dessalk</surname>
            ,
            <given-names>Y.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nikolov</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matskin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soylu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Scalable execution of big data workflows using software containers</article-title>
          .
          <source>In: Proceedings of the 12th International Conference on Management of Digital EcoSystems</source>
          . pp.
          <fpage>76</fpage>
          -
          <lpage>83</lpage>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Developers</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>CGAT-core documentation</article-title>
          . https://cgat-core.readthedocs. io/en/latest/, [Online; accessed 21-April-2021]
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Dray</surname>
          </string-name>
          <article-title>: Dray Overview</article-title>
          . https://github.com/CenturyLinkLabs/dray (
          <year>2015</year>
          ), [Online; accessed 17-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. Flyte: Flyte Documentation. https://docs.flyte.org/en/latest/index.html (
          <year>2020</year>
          ), [Online; accessed 17-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Foundation</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <string-name>
            <surname>YAWL - Yet Another</surname>
          </string-name>
          Workflow Language. http://www. yawlfoundation.org/documents/YAWL_leaflet-final.
          <source>pdf</source>
          (
          <year>2007</year>
          ), [Online; accessed 19-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. Galaxy: Galaxy Tutorials. https://galaxyproject.org/learn/ (
          <year>2005</year>
          ), [Online; accessed 17-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Google: GoogleWorkflow - Error Handling</surname>
          </string-name>
          Syntax. https://cloud.google.com/ workflows/docs, [Online; accessed 31-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <given-names>Henning</given-names>
            <surname>Baars</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.E.</surname>
          </string-name>
          :
          <article-title>From data warehouses to analytical atoms - the internet of things as a centrifugal force in business intelligence and analytics</article-title>
          .
          <source>In: wenty-Fourth European Conference on Information Systems. ECIS</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. Hyperloom: Hyperloom - Basic
          <string-name>
            <surname>Terms</surname>
          </string-name>
          . https://loom-it4i.readthedocs.io/en/ latest/intro.html#basicterms, [Online; accessed 01-April-2021]
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Jimenez</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sevilla</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watkins</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maltzahn</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lofstead</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohror</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arpaci-Dusseau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arpaci-Dusseau</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>The popper convention: Making reproducible systems evaluation practical</article-title>
          .
          <source>In: 2017 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW)</source>
          . pp.
          <fpage>1561</fpage>
          -
          <lpage>1570</lpage>
          (
          <year>2017</year>
          ). https://doi.org/10.1109/IPDPSW.
          <year>2017</year>
          .157
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Kashlev</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Big data workflows: A reference architecture and the dataview system</article-title>
          .
          <source>Services Transactions on Big Data (STBD) 4</source>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>26. Keboola: Keboola - Overview. https://help.keboola.com/overview/, [Online; accessed 02-April-2021]</mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27. Knime: KNIME - Extensions. .https://docs.knime.com/2019-06/analytics_ platform_quickstart_guide/index.html
          <article-title>#extend-knime-analytics-platform</article-title>
          ., [Online; accessed 31-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28. Ko¨ster, J.,
          <string-name>
            <surname>Rahmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Snakemake-a scalable bioinformatics workflow engine</article-title>
          .
          <source>Bioinformatics</source>
          <volume>28</volume>
          (
          <issue>19</issue>
          ),
          <fpage>2520</fpage>
          -
          <lpage>2522</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29. Kubeflow: Kubeflow - Pipelin. https://www.kubeflow.org/docs/components/ pipelines/overview/pipelines-overview/
          <article-title>#what-is-a-pipeline</article-title>
          <string-name>
            <surname>,</surname>
          </string-name>
          [Online; accessed 01-April-2021]
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <given-names>Lal</given-names>
            <surname>Chattaraj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Villamariona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.: Apache</given-names>
            <surname>Airflow Tutorial - DAGs</surname>
          </string-name>
          , Tasks, Operators, Sensors, Hooks &amp; XCom. https://www.qubole.
          <article-title>com/tech-blog/ apache-airflow-tutorial-dags-tasks-operators-sensors-hooks-xcom/</article-title>
          , [Online; accessed 25-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <article-title>MachineFlow: MachineFlow Overview</article-title>
          . https://github.com/sean-mcclure/ machine_flow (
          <year>2018</year>
          ), [Online; accessed 17-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>32. MakeFlow: MakeFlow - Overview. https://cctools.readthedocs.io/en/ latest/makeflow/#overview, [Online; accessed 02-April-2021]</mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Marozzo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Talia</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trunfio</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>Js4cloud: script-based workflow programming for scalable data analysis on cloud platforms</article-title>
          .
          <source>Concurrency and Computation: Practice and Experience</source>
          <volume>27</volume>
          (
          <issue>17</issue>
          ),
          <fpage>5214</fpage>
          -
          <lpage>5237</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>34. Netflix: Conductor - Configuration. https://netflix.github.io/conductor/ configuration, [Online; accessed 25-March-2021]</mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Nextflow: NextFlow - Completion Handler</surname>
          </string-name>
          . https://www.nextflow.io/docs/ latest/metadata.html#completion-handler, [Online; accessed 25-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>NiFi</surname>
          </string-name>
          , A.:
          <string-name>
            <surname>Apache NiFi - Expression Language</surname>
          </string-name>
          Guide. https://nifi.apache.org/ docs/nifi-docs/html/expression-language-guide.html, [Online; accessed 26- March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          37.
          <string-name>
            <surname>Nikolov</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dessalk</surname>
            ,
            <given-names>Y.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>A.Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soylu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matskin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Payberah</surname>
            ,
            <given-names>A.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Conceptualization and scalable execution of big data workflows using domain-specific languages and software containers</article-title>
          .
          <source>Internet of Things</source>
          p.
          <volume>100440</volume>
          (
          <year>2021</year>
          ). https://doi.org/https://doi.org/10.1016/j.iot.
          <year>2021</year>
          .
          <volume>100440</volume>
          , https: //www.sciencedirect.com/science/article/pii/S2542660521000834
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          38.
          <string-name>
            <surname>Node-RED: Node-RED - Flow Control</surname>
          </string-name>
          . https://cookbook.nodered.org/, [Online; accessed 01-April-2021]
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          39.
          <string-name>
            <surname>Node-RED</surname>
          </string-name>
          :
          <article-title>Node-RED - Subflow</article-title>
          . https://nodered.org/docs/creating-nodes/, [Online; accessed 01-April-2021]
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          40.
          <string-name>
            <surname>Novella</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Emami</surname>
            <given-names>Khoonsari</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Herman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Whitenack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Capuccini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Burman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kultima</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Spjuth</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.</surname>
          </string-name>
          :
          <article-title>Container-based bioinformatics with pachyderm</article-title>
          .
          <source>Bioinformatics</source>
          <volume>35</volume>
          (
          <issue>5</issue>
          ),
          <fpage>839</fpage>
          -
          <lpage>846</lpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          41.
          <string-name>
            <surname>Oozie</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Oozie</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Apache Oozie - Coordinator Job</surname>
          </string-name>
          . .https://oozie. apache.
          <source>org/docs/5</source>
          .2.1/CoordinatorFunctionalSpec.html#a1._Coordinator_ Overview, [Online; accessed 26-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>42. Pachyderm: Pachyderm - Introduction. https://docs.pachyderm.com/latest/ how-tos/create-pipeline/, [Online; accessed 25-March-2021]</mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>43. Pegasus: Pegasus - Documentation. https://pegasus.isi.edu/documentation, [Online; accessed 02-April-2021]</mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          44.
          <string-name>
            <surname>Pesic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schonenberg</surname>
          </string-name>
          , H., van der Aalst, W.:
          <article-title>Declarative workflow</article-title>
          .
          <source>In: Modern Business Process Automation</source>
          , pp.
          <fpage>175</fpage>
          -
          <lpage>201</lpage>
          . Springer (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          45.
          <string-name>
            <surname>Prefect: Prefect - Core Concepts</surname>
          </string-name>
          . https://docs.prefect.io/core/concepts/, [Online; accessed 25-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>46. Prefect: Prefect - Latest API. https://docs.prefect.io/api/latest/, [Online; accessed 25-March-2021]</mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>47. Reflow: Reflow - Overview. https://github.com/grailbio/reflow, [Online; accessed 25-March-2021]</mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          48.
          <string-name>
            <surname>Saey</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A dsl for distributed, reactive workflows (</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          49.
          <string-name>
            <surname>Seiger</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huber</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schlegel</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Toward an execution system for self-healing workflows in cyber-physical systems</article-title>
          .
          <source>Software &amp; Systems Modeling</source>
          <volume>17</volume>
          (
          <issue>2</issue>
          ),
          <fpage>551</fpage>
          -
          <lpage>572</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          50.
          <string-name>
            <surname>Semeniuta</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Falkman</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Epypes: a framework for building event-driven data processing pipelines</article-title>
          .
          <source>PeerJ Computer Science</source>
          <volume>5</volume>
          ,
          <issue>e176</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>51. Skitter: Skitter - Documents. https://soft.vub.ac.be/~mathsaey/skitter/ docs/latest/readme.html, [Online; accessed 26-March-2021]</mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>52. Snakemake: Snakemake - Tutorial. https://snakemake.readthedocs.io/en/ stable/tutorial/tutorial.html#snakemake-tutorial, [Online; accessed 26- March-2021]</mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          53. StreamFlow: StreamFlow GitHub. https://github.com/alpha-unito/ streamflow (
          <year>2020</year>
          ), [Online; accessed 17-March-2021]
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          54.
          <string-name>
            <surname>Streampipe</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          : StreamPipes - Tutoria. https://streampipes.apache.org/docs/ docs/dev-guide-tutorialsources/, [Online; accessed 01-April-2021]
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          55.
          <string-name>
            <surname>Toil</surname>
          </string-name>
          <article-title>: Toil - Workflows with Multiple Job</article-title>
          . https://toil. readthedocs.io/en/latest/developingWorkflows/developing.html
          <article-title># workflows-with-multiple-</article-title>
          <string-name>
            <surname>jobs</surname>
          </string-name>
          , [Online; accessed 01-April-2021]
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          56. Trifacta: Trifacta - Function. https://docs.trifacta.com/display/SS/Wrangle+ Language#WrangleLanguage-Functions.
          <volume>1</volume>
          (
          <issue>2013</issue>
          ), [Online; accessed 26-March2021]
        </mixed-citation>
      </ref>
      <ref id="ref57">
        <mixed-citation>
          57.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Script of scripts: A pragmatic workflow system for daily computational research</article-title>
          .
          <source>PLoS Computational Biology</source>
          <volume>15</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref58">
        <mixed-citation>
          58.
          <string-name>
            <surname>Wieland</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwarz</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , Breitenbu¨cher, U.,
          <string-name>
            <surname>Leymann</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Towards situationaware adaptive workflows: Sitopt-a general purpose situation-aware workflow management system</article-title>
          .
          <source>In: 2015 IEEE International Conference on Pervasive Computing and Communication Workshops (PerCom Workshops)</source>
          . pp.
          <fpage>32</fpage>
          -
          <lpage>37</lpage>
          . IEEE (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref59">
        <mixed-citation>
          59. Wikipedia: Kepler - HierarchicalWorkflow. https://en.wikipedia.org/wiki/ Kepler_scientific_
          <article-title>workflow_system#Hierarchical_workflows (</article-title>
          <year>2015</year>
          ), [Online; accessed 01-April-2021]
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>