<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>IWSG</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Parsl: Scalable Parallel Scripting in Python</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Yadu Babuji</institution>
          ,
          <addr-line>Kyle Chard , Ian Foster , Daniel S. Katz</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>13</volume>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>-Computational and data-driven research practices have significantly changed over the past decade to encompass new analysis models such as interactive and online computing. Science gateways are simultaneously evolving to support this transforming landscape with the aim to enable transparent, scalable execution of a variety of analyses. Science gateways often rely on workflow management systems to represent and execute analyses efficiently and reliably. However, integrating workflow systems in science gateways can be challenging, especially as analyses become more interactive and dynamic, requiring sophisticated orchestration and management of applications and data, and customization for specific execution environments. Parsl (Parallel Scripting Library), a Python library for programming and executing data-oriented workflows in parallel, addresses these problems. Developers simply annotate a Python script with Parsl directives wrapping either Python functions or calls to external applications. Parsl manages the execution of the script on clusters, clouds, grids, and other resources; orchestrates required data movement; and manages the execution of Python functions and external applications in parallel. The Parsl library can be easily integrated into Python-based gateways, allowing for simple management and scaling of workflows. Parsl, Parallel scripting, Python, Scientific Workflows-</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        Data-driven research methodologies have had a disruptive
impact on science, enabling new types of exploration and
facilitating new discoveries [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Underlying these
methodologies are new tools and technologies such as Jupyter
notebooks for interactive analysis, scripting languages for
flexible exploration, and a suite of libraries like Pandas and
scikit-learn that facilitate cutting-edge analyses.
      </p>
      <p>
        Science gateways [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] have long supported the varied needs
of users, providing intuitive interfaces for end users to
access both data and computing capabilities. Science
gateway frameworks, such as Apache Airavata [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and
WSPGRADE/gUSE [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], often rely on workflow frameworks to
represent and execute analyses that benefit from extensibility,
scalability, and robustness [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. However, there are two
significant challenges associated with current approaches: 1) many
workflow engines are focused on many task applications rather
than interactive, online, or machine learning analyses; and 2)
workflow engines are not easily integrated into external
services (e.g., gateways) due to issues such as language mismatch
and the need for intermediate workflow representations.
      </p>
      <p>
        Here we present Parsl, a Python parallel scripting library
that supports the development and execution of asynchronous
and implicitly parallel data-oriented workflows. Building on
the model used by the Swift workflow language [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], Parsl
brings parallel workflow capabilities to scripts, applications,
and gateways implemented in Python. Parsl scripts allow
selected Python functions and external applications (called
Apps) to be connected by shared input/output data objects
into flexible parallel workflows. Parsl abstracts the specific
execution environment, allowing the same script to be
executed on arbitrary multicore processors, clusters, clouds, and
supercomputers.
      </p>
      <p>When a Parsl script is executed, the Parsl library causes
annotated functions (Apps) to be intercepted by the Parsl
execution fabric, which captures and serializes their
parameters, analyzes their dependencies, and runs them on selected
resources, referred to as sites. The execution fabric brings
dependency awareness to Apps by introducing data futures as
the inputs and outputs of Apps. Apps that use a data future as
an input can be enqueued but will be blocked until that data
future has been written. This feature allows Apps to execute
in parallel whenever they do not share dependencies or their
data dependencies have been resolved.</p>
      <p>
        Fig. 1 depicts how Parsl interacts with its environment,
including code, data, and resources. Parsl provides several
advantages to science gateways: it allows a single script to
be executed on any computing infrastructures from clouds to
supercomputers; it provides fault tolerance, automated
elasticity, and support for various execution models; it handles data
management by staging local data through its secure message
queue and by managing wide area transfers with Globus [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ];
and it can be trivially integrated via its Python interface.
      </p>
      <p>In this paper we describe Parsl, highlighting how it allows
standard Python scripts and science gateways to be augmented
to execute complex workflows and facilitate parallel execution.
We describe Parsl’s unique capabilities and present several
example workflows that are common in science gateways from
computational chemistry, materials science, and biology, to
highlight the power of the approach.</p>
    </sec>
    <sec id="sec-2">
      <title>II. WORKFLOW MODELS</title>
      <p>Parsl is designed to support not only traditional
manytask workflow models but also new analysis models that are
and will be increasingly supported by science gateways (e.g.,
online and interactive computing). We briefly describe three
such workflow models that can be supported by Parsl.</p>
      <p>
        Workflows have long been applied to a range of
manytask applications, for example protein-ligand docking for drug
screening [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Here, workflows are used to orchestrate a series
of external applications to be applied to a large set of input
data. For example, in drug screening, dozens of proteins are
evaluated against hundreds of thousands of drug candidates to
identify the location and orientation of a ligand that binds to
a protein receptor. The top candidates are then processed with
detailed molecular dynamics simulations to identify the most
likely combinations to be used for further experimentation.
Gateways such as MoSGrid [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and Galaxy [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] support such
workflows.
      </p>
      <p>
        Discovery science represents a new research methodology
based on explorative, interactive analysis. The general model
centers around analysis of large volumes of data with the aim
to find unknown patterns. Notebook environments, such as
Jupyter, provide an ideal interface in which researchers can
discover and explore large data volumes using a variety of
analytics approaches. Such methods are used in a wide range
of studies from computing the stopping power of electrons
through materials to measuring discursive influence across
scholarship [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Gateways such as Cloud Kotta [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and
HubZero [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] expose Jupyter notebook interfaces for
interactive computing.
      </p>
      <p>
        Exploding data acquisition rates from scientific instruments,
such as light sources, microscopes, and telescopes,
necessitate rapid analysis to avoid data loss and enable online
experiment steering. Real-time (or online) computing, such
as that conducted at the Advanced Photon Source, allows for
data streamed from beamline computers to be processed in
real-time on a large cluster, with the aim to make real-time
decisions during experiments [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>III. PARSL MODEL</title>
      <p>The Parsl architecture is shown in Fig. 2. Parsl scripts are
decomposed into a simple dependency graph by the DataFlow
Kernel (DFK). The DFK manages execution of individual
Parsl Apps on a variety of sites. Unlike parallel scripting
languages like Swift, in which every variable and piece of code
is asynchronous, Parsl relies on users to annotate functions that
will be run asynchronously based on data dependencies. The
DFK provides a lightweight data management layer in which
Python objects and files are staged to an execution site via a
dedicated communication channel or Globus.</p>
      <p>Dataflow Kernel: The DFK provides a single lightweight
abstraction on top of different execution resources. This
abstraction is at the heart of Parsl’s ability to transparently
support different execution fabrics.
@python_app
def hello():</p>
      <p>return 'Hello World!'
@bash_app
def hello(inputs=[], outputs=[],</p>
      <p>stdout=None, stderr=None):
return 'echo "Hello World"'</p>
    </sec>
    <sec id="sec-4">
      <title>Listing 1: Two examples of Parsl Apps.</title>
      <p>Parsl launches asynchronous Apps and passes futures to
other Apps in lieu of computing results synchronously. The
DFK is responsible for managing a script’s execution, making
ordinary functions aware of futures and ensuring the
execution of these functions are conditional on the resolution of
all dependent futures. This enables completely asynchronous
management of all launched tasks with the data dependencies
alone determining the order of execution.</p>
      <p>Apps: A Parsl script is comprised of standard Python code
plus a number of Apps—annotated units of Python code
or external applications that specify their input and output
characteristics and that may be run in parallel. An App may
be defined by wrapping an existing function or the execution
of an external command-line application using Bash scripting
with the @App decorator. Listing 1 shows examples of these
two types of Parsl Apps.</p>
      <p>Futures: Parsl Apps are completely asynchronous. When
an App is invoked, there is no guarantee of when the result
will be returned. Instead of directly returning a result, Parsl
returns an AppFuture: a construct that includes the real result
as well as the status and exceptions for that asynchronous
function invocation. Parsl also supplies methods to examine
the future construct, including checking status, blocking on
completion, and retrieving results. Parsl leverages Python’s
concurrent.futures module for this purpose.</p>
      <p>Parsl also introduces a model for managing the
asynchronous output files generated by an App invocation as
DataFutures. DataFutures extend the AppFuture model by
providing support for a range of operations related to files.</p>
      <sec id="sec-4-1">
        <title>A. Execution</title>
        <p>When instantiating the DFK, developers specify the
specific execution providers and executors that will be used for
executing the parallel components of the script. Execution
providers are simple abstractions over computational resources
and executors provide an abstraction layer for executing tasks.</p>
        <p>
          Parsl’s execution interface is called libsubmit [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]—a
simple Python library that provides a common interface to
execution resources. Libsubmit’s interface defines operations such as
submission, status, and job management. It currently supports
a variety of providers including Amazon Web Services,
Microsoft Azure, and Jetstream clouds as well as Cobalt, Slurm,
Torque, GridEngine, and HTCondor Local Resource Managers
(LRM). New execution providers can be easily added by
implementing libsubmit’s execution provider interface.
        </p>
        <p>Depending on the the selected execution provider, there are
a number of ways to submit workload to that resource. For
example, for local execution, threads can be used, while for a
cluster, pilot jobs or specialized launchers can be used. Parsl
supports these different methods via its executor interface.
Parsl currently supports three executors:</p>
        <p>ThreadPoolExecutor for multi-thread execution on local
resources.</p>
        <p>IPyParallelExecutor for both local and remote execution
using a pilot job model. The IPythonParallel controller
is deployed locally and IPythonParallel engines are
deployed on execution nodes. IPythonParallel then manages
the execution of tasks on connected engines.</p>
        <p>
          Swift/TurbineExecutor for extreme-scale execution
using the Swift/T (Turbine) [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] model to enable distributed
task execution across an MPI environment. This executor
is typically used on supercomputers.
        </p>
        <p>It is important to note that Parsl scripts are not tied to a
specific executor or execution provider. Furthermore, a single
Parsl script may leverage multiple executors and execution
providers concurrently—a model we refer to as multi site.
This allows Parsl developers to mix and match resources and
execution models to meet their needs. For example, enabling
a computational simulation to run on specialized HPC nodes,
simple data manipulation tasks to be executed locally using
threads, and visualizations to be rendered on GPU nodes.</p>
      </sec>
      <sec id="sec-4-2">
        <title>B. Uniform execution model</title>
        <p>Providing a uniform representation of heterogeneous
resources is one of the most difficult challenges for parallel
execution. Parsl provides an abstraction based on resource
units called blocks. A block is a single unit of resources
that is obtained from an execution provider. Within a block
are a number of nodes. Parsl can then create TaskBlocks
within and across (e.g., for MPI jobs) nodes. A TaskBlock
is a virtual suballocation in which individual tasks can be
launched. Figure 3 shows three different block configurations.
The first configuration represents the most simple model in
which a block is comprised of a single node with a single
TaskBlock. The second configuration, with several TaskBlocks
in a single node, is well suited for executing many, single
threaded applications on a multicore node. The final
configuration shows a block comprised of several nodes and offering
several TaskBlocks. This configuration is generally used by</p>
        <p>MPI applications that span nodes. It requires specific MPI
launchers supported by the target system such as aprun, srun,
mpirun, and mpiexec.</p>
      </sec>
      <sec id="sec-4-3">
        <title>C. Parallelism and elasticity</title>
        <p>Rather than precompile a static representation of the entire
workflow, Parsl implements a dynamic dependency graph
in which the graph is constructed as tasks are enqueued.
As the Parsl script executes the workflow, new tasks are
added to a queue for execution, tasks are then executed
asynchronously when their dependencies are met. Parsl uses
the selected executor(s) to manage task execution on the
execution provider(s).</p>
        <p>As Parsl manages a dynamic dependency graph it does
not know the full “width” of a particular workflow a priori.
Further, as a workflow executes, the needs of the tasks may
change as too might the capacity available on execution
providers. Thus, Parsl must elastically scale the resources it is
using. To do so, it includes an extensible flow control system
to monitor outstanding tasks and available compute capacity.
This monitor, which can be extended or implemented by users,
determines when to trigger scaling (in or out) events.</p>
        <p>Parsl provides a simple user-managed model for
controlling elasticity. It allows users to prescribe the minimum and
maximum number of blocks to be used on a given execution
provider and a parameter (p) to control the level of parallelism.
Where parallelism is expressed as the ratio of TaskBlocks to
active tasks. Each TaskBlock is capable of executing a single
task at any given time. Therefore, a parallelism value of 1
represents aggressive scaling in which as many resources as
possible will be used; parallelism close to 0 represents the
opposite situation in which few resources (i.e., 1 TaskBlock)
will be used.</p>
      </sec>
      <sec id="sec-4-4">
        <title>D. Data management</title>
        <p>Parsl is designed to enable implementation of dataflow
patterns in which data passed between Apps manages the
flow of execution. Dataflow programming models are popular
as they can cleanly express, via implicit parallelism, the
concurrency needed by many applications in a simple and
intuitive way.</p>
        <p>Parsl aims to abstract not only parallel execution but also
execution location, which in turn requires data location
abstraction. For Python Apps, Parsl uses a direct channel between
(a) A block comprised of a node with one
TaskBlock.</p>
        <p>(b) A block comprised of a node with
several TaskBlocks.</p>
        <p>
          (c) A block comprised of four nodes with
two TaskBlocks.
checkpoint files when the DataFlow Kernel is initialized and
written out to checkpoint files when explicitly requested.
the script and executors using Python object serialization. For
files, Parsl implements a simple abstraction that can be used
to reference data irrespective of its location. At present this
model is limited to local and Globus [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] accessible files.
        </p>
        <p>
          The Parsl file abstraction is used to pass
locationindependent references between Apps. It requires that the
developer initially define a file’s location (e.g., /local/path/file
or globus://endpoint/file). The file may then be passed to each
App and, when executed, Parsl will translate the location
to a locally accessible file path. In the case of Globus, an
explicit staging model is supported in which the developer
must select the execution site to which the file should be
transferred. Parsl uses the Globus SDK and its native App
authentication model [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] to authenticate with the Globus
service and securely move data between endpoints.
        </p>
        <p>We present three workflows implemented using Parsl to
illustrate how it can satisfy the needs of different application
domains. While these workflows have not yet been
implemented in science gateways they represent use cases that
would benefit from gateway models.</p>
        <p>
          SwiftSeq [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] is a bioinformatics workflow that supports
aligning and genotyping gene panels, exomes, and whole
genomes. The Parsl-based workflow is comprised of
approximately 10 applications that communicate by writing and
reading files. While applications must often execute in sequence,
there are also opportunities for parallelism. First, the workflow
is often executed on many samples, each of which can be
E. Caching analyzed in parallel; second, the large genetic sequences can
        </p>
        <p>
          When developing a workflow, developers often execute the be divided up and analyzed in parallel; and finally, some of
same workflow with incremental changes over and over, this the applications themselves can also be executed in parallel.
scenario is especially prevalent in interactive computing work- SwiftSeq benefits not only from Parsl’s ability to specify such
flows. Often large fragments of the workflow have not been parallelism, but also from its ability to express a complex
changed yet are computed again, wasting valuable developer workflow, manage the flow of data between Apps, recover
time and computation resources. Caching of Apps (often called from errors, and execute on many computational resources.
memoization) solves this problem by saving results from Apps Parsl has been used in computational chemistry to
dethat have completed so that they can be re-used. Parsl’s velop molecular dynamics workflows. In one example,
PACKcaching model stores App results in an index alongside the MOL [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] is used to assemble initial starting configurations
App function, input parameters, and hash of the function body. of ionic liquid molecules with a protein (e.g., Trp-cage),
If caching is enabled, by an annotation on the App function before a GPU-accelerated version of Amber [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] is used
or globally at the workflow level, the cache is interrogated to energy minimize, heat, equilibrate, and run production
before each App executes. Caching is supported for Python molecular dynamics simulations. The workflow relies on three
and Bash Apps. Users must explicitly enable caching to avoid separate applications that are executed iteratively to perform
issues with non-deterministic applications. different functions. PACKMOL is used to generate the system
configuration, AmberTools are used to create input coordinate
F. Checkpointing and parameter files for simulations, and Amber is used to run
        </p>
        <p>
          Large scale workflows are prone to errors due to node various simulations. Parsl allows a wide range of different
failures, application or environment errors, and myriad other system configurations to be considered in parallel, and it also
issues. Parsl provides fault tolerance via an incremental check- allows simple error handling logic to be expressed.
pointing model, where each checkpoint call saves all results In materials science, researchers have used Parsl to predict
that have been updated since the last checkpoint was created. the electronic stopping power of materials. Stopping power
When loading checkpoints, if entries with results from multiple is the predominant energy-loss mechanism for charged
parfunctions (with identical hashes) are encountered, only the last ticles and is important for applications related to radiation
entry read will be considered. Checkpoints are loaded from protection. Historically, the stopping power for a material is
computed using analytical models such as the Lindhard model
or using Time-Dependent Density Functional Theory
(TDDFT). However, these methods are not suitable in all cases
and are computationally expensive. The Parsl-based workflow
uses TD-DFT calculations of a proton passing through a
material [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], transforms that data to a representation compatible
with machine learning, and then executes a number of machine
learning algorithms to learn a predictive model. It finally
applies these models from various directions to calculate a
three dimensional model of stopping power for a material.
        </p>
        <p>Parsl was used as it was able to trivially parallelize the existing
Python codebase, support the composition of a sophisticated
machine learning pipeline in a Jupyter notebook, and facilitate
scalable execution of the pipeline from within the notebook
on large-scale computing resources at the Argonne Leadership
Computing Facility.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>V. RELATED WORK</title>
      <p>
        Many workflow systems have been developed to
facilitate the expression and execution of arbitrary, data-oriented
workflows, for example, the Swift parallel scripting language.
Other systems include Pegasus [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] and Galaxy [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. A
weakness of these systems, however, is the need to develop
a workflow representation in a separate representation (e.g.,
a graph) Parsl provides similar capabilities, directly in a
programming language that is broadly adopted by scientific
users and increasingly science gateways.
      </p>
      <p>
        There are a number of Python-based workflow tools that
better match common research environments, for example,
Dask [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], Apache Airflow [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], Luigi [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], and
FireWorks [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
      </p>
      <p>Dask is a parallel computing library designed for parallel
analytics. It allows users to trivially migrate their single-node
analyses to a parallel execution environment. Unlike Parsl,
Dask scripts use Dask-specific functions in place of common
libraries and programming constructs, for example using the
Dask DataFrame in place of the Pandas DataFrame. Like Parsl,
Dask decomposes a script into a dependent task graph that
controls the execution of code blocks. Parsl focuses on a broader
problem, including the ability to execute arbitrary applications
on heterogeneous computing resources and providing support
for managing data dependencies between these executions.</p>
      <p>Apache Airflow is a workflow engine written in Python.
Developers can express directed acyclic graphs of independent
tasks. The Airflow scheduler is then responsible for executing
the tasks on distributed workers according to their
dependencies. Unlike Parsl’s implicit workflow model, Airflow relies
on users expressing their workflows as explicit tasks and with
explicit relationships between those tasks. Thus, the job of
the user is to essentially describe a task dependency graph in
Python.</p>
      <p>Luigi scripts are created by writing Python classes that
extend the Luigi task model: developers implement functions
that manage input and output data, the code that will be run,
as well as the explicit dependencies on other tasks. Unlike
Parsl, Luigi focuses on Python tasks rather than orchestrating
execution of external applications. Further, Luigi offers a
execution model that deploys workers on a single cluster; it is
not designed to support multiple sites, provide elastic resource
management, or handle wide area data staging.</p>
      <p>FireWorks is a Python-based workflow engine designed
for executing high-throughput workflows on supercomputers.
Workflows are described in Python, JSON, or YAML and
as a collection of tasks which are connected together into a
“FireWork” for execution. The centralized server manages the
workflow, using a MongoDB database to provide persistence
and to support reliable execution on distributed resources.
FireWorkers are deployed on compute resources to execute tasks,
they connect to the centralized server to request tasks, execute
them, and return results. Unlike Parsl, FireWorks focuses on
the reliable execution of long running jobs and therefore may
not be suitable for short running jobs or applications that
demand a high submission rate.</p>
    </sec>
    <sec id="sec-6">
      <title>VI. SUMMARY</title>
      <p>Parsl provides an easy-to-use model that can be easily
integrated in science gateways to support the management
and execution of workflows composed of Python functions
and external applications. Science gateways benefit from the
extensibility, scalability, and robustness of the Parsl model to
manage execution of potentially complex workflows on
arbitrary computational resources. Parsl is specifically designed to
address new workflow modalities, such as interactive
computing in Jupyter notebooks, and provides a seamless and
transparent way to scale these analyses from within the notebook.
Parsl abstracts the complexity of interacting with different
resource fabrics and execution models. It instead supports the
development of resource-independent Python scripts. It also
includes a number of advanced capabilities such as automated
elasticity, support for multi-site execution, fault tolerance, and
automated direct and wide area data management.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGMENT</title>
      <p>This work was supported in part by NSF award
ACI1550588 and DOE contract DE-AC02-06CH11357.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A. W.</given-names>
            <surname>Toga</surname>
          </string-name>
          , I. Foster,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kesselman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Madduri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. W.</given-names>
            <surname>Deutsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. D.</given-names>
            <surname>Price</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Glusman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. D.</given-names>
            <surname>Heavner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. D.</given-names>
            <surname>Dinov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ames</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Van Horn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kramer</surname>
          </string-name>
          , and L. Hood, “
          <article-title>Big biomedical data as the key resource for discovery science</article-title>
          ,
          <source>” Journal of the American Medical Informatics Association</source>
          , vol.
          <volume>22</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>1126</fpage>
          -
          <lpage>1131</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N. P.</given-names>
            <surname>Tatonetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. P.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Daneshjou</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Altman</surname>
          </string-name>
          , “
          <article-title>Datadriven prediction of drug effects and interactions,” Science Translational Medicine</article-title>
          , vol.
          <volume>4</volume>
          , no.
          <issue>125</issue>
          , pp.
          <fpage>125ra31</fpage>
          -
          <lpage>125ra31</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ward</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Wolverton</surname>
          </string-name>
          , “
          <article-title>Atomistic calculations and materials informatics: A review,” Current Opinion in Solid State and Materials Science</article-title>
          , vol.
          <volume>21</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>167</fpage>
          -
          <lpage>176</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>N.</given-names>
            <surname>Wilkins-Diehr</surname>
          </string-name>
          , “
          <article-title>Special issue: science gateways - common community interfaces to grid resources,” Concurrency and Computation: Practice and Experience</article-title>
          , vol.
          <volume>19</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>743</fpage>
          -
          <lpage>749</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Marru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gunathilake</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Herath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tangchaisin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pierce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mattmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gunarathne</surname>
          </string-name>
          , E. Chinthaka,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gardler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Slominski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Douma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Perera</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Weerawarana</surname>
          </string-name>
          , “
          <article-title>Apache Airavata: A framework for distributed applications and computational workflows</article-title>
          ,”
          <source>in Proceedings of the 2011 ACM Workshop on Gateway Computing Environments</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kacsuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Farkas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kozlovszky</surname>
          </string-name>
          , G. Hermann,
          <string-name>
            <given-names>A.</given-names>
            <surname>Balasko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Karoczkai</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Marton</surname>
          </string-name>
          , “
          <article-title>WS-PGRADE/gUSE generic DCI gateway framework for a large variety of user communities</article-title>
          ,
          <source>” Journal of Grid Computing</source>
          , vol.
          <volume>10</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>601</fpage>
          -
          <lpage>630</lpage>
          ,
          <year>Dec 2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Glatard</surname>
          </string-name>
          , M. E´ tienne Rousseau,
          <string-name>
            <given-names>S.</given-names>
            <surname>Camarasu-Pop</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Adalat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Beck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Das</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. F.</given-names>
            da
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Khalili-Mahani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Korkhov</surname>
          </string-name>
          , P.
          <article-title>-</article-title>
          <string-name>
            <surname>O. Quirion</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rioux</surname>
            ,
            <given-names>S. D.</given-names>
          </string-name>
          <string-name>
            <surname>Olabarriaga</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Bellec</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Evans</surname>
          </string-name>
          , “
          <article-title>Software architectures to integrate workflow engines in science gateways,” Future Generation Computer Systems</article-title>
          , vol.
          <volume>75</volume>
          , pp.
          <fpage>239</fpage>
          -
          <lpage>255</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wilde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hategan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Wozniak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Clifford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Katz</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Foster</surname>
          </string-name>
          , “
          <article-title>Swift: A language for distributed parallel scripting,” Parallel Computing</article-title>
          , vol.
          <volume>37</volume>
          , no.
          <issue>9</issue>
          , pp.
          <fpage>633</fpage>
          -
          <lpage>652</lpage>
          , Sep.
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>Chard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tuecke</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Foster</surname>
          </string-name>
          , “
          <article-title>Efficient and secure transfer, synchronization, and sharing of big data,” IEEE Cloud Computing</article-title>
          , vol.
          <volume>1</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>46</fpage>
          -
          <lpage>55</lpage>
          ,
          <year>Sept 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Adhikari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wilde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. F.</given-names>
            <surname>Freed</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T. R.</given-names>
            <surname>Sosnick</surname>
          </string-name>
          , “
          <article-title>Modeling large regions in proteins: Applications to loops, termini</article-title>
          , and folding,
          <source>” Protein Science</source>
          , vol.
          <volume>21</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>107</fpage>
          -
          <lpage>121</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11] J. Kr u¨ger,
          <string-name>
            <given-names>R.</given-names>
            <surname>Grunzke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gesing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Breuers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Brinkmann</surname>
          </string-name>
          , L.
          <string-name>
            <surname>de la Garza</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Kohlbacher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Kruse</surname>
            ,
            <given-names>W. E.</given-names>
          </string-name>
          <string-name>
            <surname>Nagel</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Packschies</surname>
            , R. M u¨llerPfefferkorn, P. Scha¨fer, C. Scha¨rfe,
            <given-names>T.</given-names>
            Steinke, T.
          </string-name>
          <string-name>
            <surname>Schlemmer</surname>
            ,
            <given-names>K. D.</given-names>
          </string-name>
          <string-name>
            <surname>Warzecha</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Zink</surname>
            , and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Herres-Pawlis</surname>
          </string-name>
          , “
          <article-title>The MoSGrid science gateway - a complete solution for molecular simulations</article-title>
          ,
          <source>” Journal of Chemical Theory and Computation</source>
          , vol.
          <volume>10</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>2232</fpage>
          -
          <lpage>2245</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E.</given-names>
            <surname>Afgan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Baker</surname>
          </string-name>
          , M. van den Beek et al., “
          <article-title>The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2016 update</article-title>
          ,”
          <source>Nucleic Acids Res</source>
          ., vol.
          <volume>44</volume>
          , no.
          <source>W1</source>
          , p.
          <fpage>W3</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gerow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Boyd-Graber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Blei</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Evans</surname>
          </string-name>
          , “
          <article-title>Measuring discursive influence across scholarship</article-title>
          ,
          <source>” Proceedings of the National Academy of Sciences</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y. N.</given-names>
            <surname>Babuji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chard</surname>
          </string-name>
          , and E. Duede, “
          <article-title>Enabling interactive analytics of secure data using cloud kotta</article-title>
          ,” in 8th Workshop on Scientific Cloud Computing, ser.
          <source>ScienceCloud '17</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>McLennan</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Kennell</surname>
          </string-name>
          , “
          <article-title>HUBzero: A platform for dissemination and collaboration in computational science</article-title>
          and engineering,” IEEE Des.
          <source>Test</source>
          , vol.
          <volume>12</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>48</fpage>
          -
          <lpage>53</lpage>
          , Mar.
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bicer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gursoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kettimuthu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. T.</given-names>
            <surname>Foster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. D.</given-names>
            <surname>Andrede</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F. D.</given-names>
            <surname>Carlo</surname>
          </string-name>
          , “
          <article-title>Real-time data analysis and autonomous steering of synchrotron light source experiments,” in 13th IEEE International Conference on e-Science (e-Science)</article-title>
          ,
          <source>Oct</source>
          <year>2017</year>
          , pp.
          <fpage>59</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>[17] “Libsubmit,” https://github.com/Parsl/libsubmit.</mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>T. G.</given-names>
            <surname>Armstrong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Wozniak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wilde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I. T.</given-names>
            <surname>Foster</surname>
          </string-name>
          , “
          <article-title>Compiler techniques for massively scalable implicit task parallelism,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser</article-title>
          .
          <source>SC '14</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>299</fpage>
          -
          <lpage>310</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>K.</given-names>
            <surname>Chard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tuecke</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Foster</surname>
          </string-name>
          , “
          <article-title>Efficient and secure transfer, synchronization, and sharing of big data,” IEEE Cloud Computing</article-title>
          , vol.
          <volume>1</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>46</fpage>
          -
          <lpage>55</lpage>
          ,
          <year>Sept 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Tuecke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ananthakrishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lidman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>McCollam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosen</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Foster</surname>
          </string-name>
          , “
          <article-title>Globus Auth: A research identity and access management platform</article-title>
          ,
          <source>” in 12th IEEE International Conference on eScience (e-Science)</source>
          ,
          <source>Oct</source>
          <year>2016</year>
          , pp.
          <fpage>203</fpage>
          -
          <lpage>212</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pitt</surname>
          </string-name>
          , “Swiftseq,” http://www.igsb.org/software/swiftseq.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>L.</given-names>
            <surname>Mart</surname>
          </string-name>
          ´ınez,
          <string-name>
            <given-names>R.</given-names>
            <surname>Andrade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. G.</given-names>
            <surname>Birgin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Mart</surname>
          </string-name>
          <article-title>´ınez, “PACKMOL: A package for building initial configurations for molecular dynamics simulations,”</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Comp</surname>
          </string-name>
          . Chemistry, vol.
          <volume>30</volume>
          , no.
          <issue>13</issue>
          , pp.
          <fpage>2157</fpage>
          -
          <lpage>2164</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Case</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. E.</given-names>
            <surname>Cheatham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Darden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gohlke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Merz</surname>
          </string-name>
          et al.,
          <source>“The Amber biomolecular simulation programs,” J. Comp. Chemistry</source>
          , vol.
          <volume>26</volume>
          , no.
          <issue>16</issue>
          , pp.
          <fpage>1668</fpage>
          -
          <lpage>1688</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>A.</given-names>
            <surname>Schleife</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kanai</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Correa</surname>
          </string-name>
          , “
          <article-title>Accurate atomistic firstprinciples calculations of electronic stopping</article-title>
          ,
          <source>” Phys. Rev. B</source>
          , vol.
          <volume>91</volume>
          , p.
          <fpage>014306</fpage>
          ,
          <year>Jan 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>E.</given-names>
            <surname>Deelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.-H.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Blythe</surname>
          </string-name>
          , James Gil et al.,
          <article-title>“Pegasus: A framework for mapping complex scientific workflows onto distributed systems,” Scientific Programming</article-title>
          , vol.
          <volume>13</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>219</fpage>
          -
          <lpage>237</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rocklin</surname>
          </string-name>
          , “
          <article-title>Dask: Parallel computation with blocked algorithms</article-title>
          and task scheduling,”
          <source>in Proc. 14th Python in Sci. Conf</source>
          .,
          <year>2015</year>
          , pp.
          <fpage>130</fpage>
          -
          <lpage>136</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Apache</given-names>
            <surname>Airflow</surname>
          </string-name>
          <string-name>
            <surname>Project</surname>
          </string-name>
          , “Apache Airflow,” https://airflow.incubator.apache.org/.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Spotify</surname>
          </string-name>
          , “Luigi,” https://github.com/spotify/luigi.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Ong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Medasani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Qu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kocher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brafman</surname>
          </string-name>
          , G. Petretto, G. Rignanese,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hautier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gunter</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Persson</surname>
          </string-name>
          , “
          <article-title>Fireworks: a dynamic workflow system designed for high-throughput applications</article-title>
          ,
          <source>” Concurrency and Computation: Practice and Experience</source>
          , vol.
          <volume>27</volume>
          , no.
          <issue>17</issue>
          , pp.
          <fpage>5037</fpage>
          -
          <lpage>5059</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>