<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Lessons Learned from Applying Big Data Paradigms to a Large Scale Scientific Workflow</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Silvina Caíno-Lores</string-name>
          <email>scaino@inf.uc3m.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Kropf</string-name>
          <email>peter.kropf@unine.ch</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrei Lapin</string-name>
          <email>andrei.lapin@unine.ch</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jesús Carretero</string-name>
          <email>jcarrete@inf.uc3m.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University Carlos III of Madrid, Av. Universidad</institution>
          ,
          <addr-line>30, Leganés, Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University Carlos III of Madrid, Av. Universidad</institution>
          ,
          <addr-line>30, Leganés, Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Neuchâtel</institution>
          ,
          <addr-line>Neuchâtel</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Neuchâtel</institution>
          ,
          <addr-line>Neuchâtel</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <abstract>
        <p>The increasing amount of data related to the execution of scienti c work ows has raised awareness of their shift towards parallel data-intensive problems. In this paper, we deliver our experience with combining the traditional highperformance computing and grid-based approaches for scienti c work ows, with Big Data analytics paradigms. Our goal was to assess and discuss the suitability of such dataintensive-oriented mechanisms for production-ready workows, especially in terms of scalability, focusing on a key element in the Big Data ecosystem: the data-centric programming model. Hence, we reproduced the functionality of a MPI-based iterative work ow from the hydrology domain, EnKF-HGS, using the Spark data analysis framework. We conducted experiments on a local cluster, and we relied on our results to discuss promising directions for further research.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>Scienti c work ows are key tools in many research areas
that rely on multiple, diverse, and distributed operations
over various datasets, usually yielding signi cant
computational complexity and data dependencies. Nowadays, the
increasing amount of input, intermediate, and even output
data related to the execution of work ows is shifting these
originally computationally intensive systems towards
parallel data-intensive problems.
c 2017 Copyright held by the owner/author(s).</p>
      <p>
        While current work ows rely on hundreds of gigabytes of
intermediate data [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], trends show that large scale
workows would have to address increasing data sizes, easily
reaching peta-scale [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. There has also been a rise of
scienti c many-task computing (MTC) work ows [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which are
capable to handle the huge data volume and the massive
computational requirements of these simulations.
      </p>
      <p>
        In this context, scienti c work ows face new performance
and scalability challenges in terms of data management,
workload distribution, load balance, and scheduling, to name
a few. Given the data-intensive nature of these problems,
recent works have suggested the opportunity of combining the
traditional high-performance computing (HPC) and
gridbased approaches with Big Data (BD) analytics and
highthroughput (HTC) paradigms [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. For example, typical BD
programming models, such as Apache Hadoop, have been
considered to substitute MPI parallelism induction
mechanisms, following a data-centric approach.
      </p>
      <p>
        Given the data-intensive nature of these problems,
recent works have suggested the opportunity of combining the
traditional high-performance computing (HPC) and
gridbased approaches with Big Data (BD) analytics and
highthroughput (HTC) paradigms [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. For example, typical BD
programming models, such as Apache Hadoop, have been
considered to substitute MPI parallelism induction
mechanisms, following a data-centric approach. Following this
trend, BD paradigms are increasingly seen as alternatives to
traditional HPC approaches for some major types of
scienti c applications, especially those with many loosely-coupled
tasks [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], or heterogeneous tasks with few interdependences
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>Our hypothesis is that BD techniques could be used to
scale-up scienti c work ows, although the architectural
differences between the analytics and scienti c worlds might
require novel approaches to achieve satisfactory results.
Therefore, in this work we aimed to assess the suitability of such
data-intensive-oriented mechanisms for production-ready
workows, especially in terms of scalability.</p>
      <p>
        In previous works [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] we showed that applying some of
these mechanisms could improve scalability in
parameterbased scienti c simulations. In particular, we focused on
increasing the addressable size and complexity of standalone
scienti c applications, executed as map-reduce-based
wrappers of the core implementation of their models.
      </p>
      <p>
        In this paper we address the suitability of such
dataintensive-oriented mechanisms for production-ready
workows. To study this, we reproduced the functionality of an
iterative work ow from the hydrology domain, EnKF-HGS
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], in Apache Spark 1.6.01, which is currently a major
representative of the data-centric analytics ecosystem. We
compared the original work ow and the Spark implementation
in a traditional cluster to evaluate scalability, and we
analysed the behaviour of the platform and the infrastructure as
the problem size increased.
      </p>
      <p>The goal of this paper is to present and discuss our
experience with the application of those paradigms, platforms and
infrastructures currently used in BD, to a scienti c
workow. The rest of this paper contains: background on the
BD techniques we relied on and their relation with scienti c
work ows, and a high-level description of the work ow from
the hydrology domain we considered as use case (Section
2); key architectural details of the proposed solution
implemented in Spark (Section 3); the preliminary evaluations we
conducted to study the behaviour of the resulting redesigned
work ow (Section 4); and relevant highlights and directions
for future work (Section 5).</p>
    </sec>
    <sec id="sec-2">
      <title>BACKGROUND</title>
      <p>
        Scienti c work ows are composed of heterogeneous and
coupled components that simulate di erent aspects of the
domain they model. These modules interact and exchange
signi cant volumes of data at runtime, hence making these
transfers e cient has a potential major impact in the overall
performance of the resulting application [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. As a
consequence, both the storage infrastructure and the logical le
system abstractions could a ect performance and
scalability, thus making data management a key aspect in work ow
design and implementation [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        Data-centric analytics and BD tools like Hadoop and Spark
are being explored to provide straightforward data
distribution and caching mechanisms in pleasingly-parallel
dataintensive HPC applications. The inherently parallel nature
of these tools has resulted in positive experimental results
showing their suitability for massively parallel workloads like
MTC-like work ows [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Nevertheless, challenges remain
with respect to work ows built with a pure HPC focus, like
those described in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which rely on MPI and traditional
storage infrastructures.
      </p>
      <p>
        However, this topic is still fairly new, and the
experience with applying these techniques is still limited. Previous
works have contributed with guidelines and methodological
approaches to make the design of scienti c work ows easier
and more e cient, with a user-centric and visual
perspective [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In this work, we focus not only in the design and
deployment of BD-inspired scienti c work ows, but also in
the performance and scalability issues they inherit from the
platform and infrastructure. We aim to provide insight on
the potential bene ts of redesigning HPC-oriented work ows
to BD platforms, while re ecting the technical and
performance issues that arise from these paradigms.
      </p>
      <p>
        Major examples of these resource-intensive work ows are
multi-scale data analysis applications. The hydrology
domain is a representative example of the former. One of
1Apache Spark 1.6.0 documentation is available at
http://spark.apache.org/docs/1.6.0/
the state of the art simulators in this domain is the
EnKFHGS work ow [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. EnKF-HGS is an iterative work ow that
implements an MPI version of the ensemble Kalman
lter (EnKF) technique for sequential data assimilation [
        <xref ref-type="bibr" rid="ref2 ref6">6,
2</xref>
        ]. EnKF-HGS runs an ensemble of model instantiations
{which we call realizations{ with di erent combinations of
input parameters and initial conditions.
      </p>
      <p>
        As shown in Figure 1, HGS allows the numerical
simulation of all the relevant surface water and groundwater
processes in a pre-alpine type of valleys. Hence, each HGS
simulation in the ensemble of realizations represents a
longrunning compute-intensive process, which comprises the
sequential execution of two proprietary simulation kernels:
GROK and HydroGeoSphere (HGS) [
        <xref ref-type="bibr" rid="ref1 ref14">14, 1</xref>
        ]. GROK is a
preprocessor that prepares the input les for HGS, which
makes GROK an I/O intensive application. HGS, on the
other hand, is an integrated hydrological modelling
simulator, which mainly relies on the CPU to solve di erential
equations. With the HGS simulation results, each model
realization is updated with the environmental eld
measurements and optimally weighted in order to achieve a higher
quality model prediction.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. SHIFTING TO A BIG DATA PARADIGM</title>
      <p>As described, the original work ow consisted of an MPI
implementation of an EnKF, which relied on two legacy
binaries to execute the simulation (GROK and HGS).
EnKFHGS operates with a set of realizations, which constitute
independent instantiations of the model, but with di erent
parameters. They are simulated independently, and the
output is gathered afterwards for further processing.</p>
      <p>Since the work ow is iterative, we selected the popular
data analysis tool Spark as representative of a BD
programming model and execution engine. The most relevant design
and implementation details of the nal implementation of
the work ow in Spark are described in the following
paragraphs. The procedures executed in the Spark driver process
are identi ed using letters (from A to D), while numbers (1
to 6) refer to tasks that are computed distributively in the
Spark executors.</p>
      <sec id="sec-3-1">
        <title>A. Data distribution</title>
        <p>The rst step is to load the necessary auxiliary les
that every executor will need to properly run its data
partition. This includes, for instance, the kernel
binaries. Spark guarantees that these les will be available
for the worker nodes in their current working directory.</p>
      </sec>
      <sec id="sec-3-2">
        <title>B. Input matrix composition</title>
        <p>Input data is read in the driver process in order to
initialize the base model, composed of two main
matrices, M1 and M2, in which each column c1;r and c2r
corresponds to an instantiation, r, of the model.
Additional data structures are created and initialised, and
the parameters of the simulation are obtained.</p>
      </sec>
      <sec id="sec-3-3">
        <title>C. Column distribution</title>
        <p>Both matrices are distributed by columns in order to
build the realization set, R. Each realization r is
composed of the corresponding columns from both
matrices, c1;r and c2;r, so that a data distribution process
is needed to create the distributed dataset that will be
transformed in the following stages and iterations. The
rationale behind distributing the workload this way is
that each realization can be simulated independently
from the others, without any further communication.
Additionally, we forced each partition to hold the data
for a single realization in order to induce ne-grained
parallelism. After realizations are distributed, the
followoing steps are executed for each realization:</p>
      </sec>
      <sec id="sec-3-4">
        <title>1. Data pre-processing</title>
      </sec>
      <sec id="sec-3-5">
        <title>2. Model simulation</title>
        <p>The GROK kernel writes the realization input data to
a local le. HGS will read the realizations from this
le in order to conduct the simulation of the model.
With the input les from GROK, HGS simulates the
model and writes its output for subsequent analysis. In
steps 1 and 2 we must ensure that both binaries will be
executed in the same node to exploit data locality. To
achieve this, we run GROK and HGS in the same map
function, which is an indivisible task in Spark. They
thus act as an inner pipeline within the work ow.</p>
      </sec>
      <sec id="sec-3-6">
        <title>3. Distributed post-processing</title>
        <p>The post-processing stage is partially distributed. First,
the output from each HGS execution is read in each
executor in order to create an updated realization set,
R0. With this information we create a distributed
matrix M10 , and conduct several distributed operations to
avoid gathering the whole matrix in the driver.</p>
      </sec>
      <sec id="sec-3-7">
        <title>4. Data analysis</title>
        <p>Further operations with auxiliary matrices are
executed in the driver in order to lter and randomize
the input for the following iteration. The goal of this
stage is to minimise the size of the dataset that needs
to be collected in a the driver prior the model update.
Note that to achieve this, signi cant data shu es must
be executed.</p>
        <p>Since not every stage of the analysis could be
distributed, there is a step in which we aggregate a nal
matrix that will be used to compute an update
matrix. This matrix is distributed afterwards, so we can
update the realizations without gathering the whole
dataset in a single node.</p>
      </sec>
      <sec id="sec-3-8">
        <title>5. Data update and caching</title>
        <p>The distributed update matrix is used to update every
realization in parallel. The resulting realization set is
persisted to the local storage of the nodes as a
faulttolerance measure, and the following iteration starts.</p>
      </sec>
      <sec id="sec-3-9">
        <title>6. Output persistence</title>
        <p>After every iteration is executed, the output is stored
to HDFS. This is executed in parallel, as every
partition is stored independently.
4.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>EVALUATION</title>
      <p>The goal of our evaluation is to assess the bene ts and
drawbacks of the application of the techniques discussed in
Sec. 2. We focused on absolute execution time and
speedup to analyse the e ects of the memory and virtualization
overheads of Spark.
4.1</p>
    </sec>
    <sec id="sec-5">
      <title>Experimental Setup</title>
      <p>In order to assess the performance and scalability of the
application, we selected a local cluster as baseline. The
speci cations and limitations of this testbed are described as
follows.</p>
      <p>This infrastructure comprised 11 slave nodes, with the
speci cations shown in Tab. 1. Each slave node holds 8GB
of RAM and two Intel Xeon E5405 @2.00GHz processors,
with four cores each. In addition, an auxiliary node was
necessary to host the driver process of the Spark
implementation, which required 7GB in the largest experiment
we conducted. This means that the container in charge of
running the driver would require 7GB plus a 10% memory
overhead (as con gured by default in the platform), 512MB
extra memory for heap space, and other overhead sources
like serialization bu ers. Since Spark adds signi cant
memory overhead to drivers and executors, we had to add a larger
node to bypass the memory constraints in the slave nodes.
As a consequence, we added a node with an overall amount
of 94GB of RAM and four Intel Xeon E7-4807 @1.87GHz
processors, with six cores each.</p>
      <p>Execution Models: Spark vs. MPI</p>
      <p>The objective of these experiments is to detect the e ects
that the execution models have on the performance of the
work ow execution. We analysed the MPI-based
implementation of EnKF-HGS and its data-centric version built in
Spark, both running in the local cluster formerly described.</p>
      <p>We allocated each Spark executor to one core in order
to fairly compare scalability against single-core MPI
processes. We run experiments increasing realization volumes
and executors number. We measured the absolute execution
time for a single execution (including the job launch time
required by Spark) and computed the speed-up achieved. The
results for those experiments are shown in Fig. 2, in which
(a) and (c) correspond to MPI, and (b) and (d) correspond
to Spark. Remarkably, Spark yields better execution times
for every experiment, and its speed-up is better the larger is
the experiment for a given number of workers. This might
be a result of the redesign process. However, the speed-up
in Spark for the largest experiment (i.e. 64 realizations on
64 executors) is lower than in the MPI case. The problem
in this case is that the 64 executors cannot be scheduled at
once due to their large memory requirements.</p>
      <p>The main conclusion from this experiment is that, while
the BD-inspired approach shows surprising performance
results, the memory overhead of the execution framework hurts
scalability, as less parallel executors can be allocated in the
same infrastructure. As a result, the slimmer MPI processes
seem more suitable for large scale execution of this work ow.
Another interesting aspect is related to the post-processing
stage of the work ow and its e ect on the overall execution
time. At some point, with the growing number of the parallel
executors, the post-processing computation becomes shorter
in time than the data transferring time. Consequently, the
post-processing stage starts to a ect the overall execution
time more than with a fewer number of the parallel
executors.</p>
    </sec>
    <sec id="sec-6">
      <title>CONCLUSIONS AND FUTURE WORKS</title>
      <p>Scienti c work ows are becoming data-intensive, and their
scale continues to grow with larger data volumes and
problem complexities. There is a rising interest in exploiting the
opportunity to leverage Big Data application models and
infrastructures to increase work ow scalability. Nevertheless,
the bene ts and drawbacks of these new tools have not been
fully studied yet.</p>
      <p>This work has explored the e ects these paradigms could
have in current work ows. We aimed to detect their
benets and drawbacks, and to extract from these knowledge a
series of lessons and future research topics. We focused this
work in two key elements to be compared: MPI and Apache
Spark. We experimented with a speci c work ow from the
hydrology domain, and analyzed its performance and
scalability implemented on MPI and Spark and executed in a
local cluster.</p>
      <p>The main conclusion from the experiments made is that,
while the BD-inspired approach shows surprising performance
results, the memory overhead of the execution framework
hurts scalability, as less parallel executors can be allocated
in the same infrastructure. As a result, the slimmer MPI
processes seem more suitable for large scale execution of this
work ow.</p>
      <p>While we considered this use case relevant for our
objectives in terms of complexity and size, further
experimentation with other work ows with di erent structures would be
necessary to corroborate our results and conclusions. In
addition, further analysis of the cost-performance trade-o of
applying Big Data paradigms to work ows should be
conducted, as well as a detailed analysis of the I/O-related
overhead speci c to the work ows.
6.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGEMENTS</title>
      <p>This work has been partially funded under the Spanish
Ministry of Economics and Competitiveness grant
TIN201341350-P, the COST Action IC1305 \Network for
Sustainable Ultrascale Computing Platforms" (NESUS), and the
FPU Training Program for Academic and Teaching Sta
FPU15/00422 by the Spanish Ministry of Education. We
gratefully acknowledge Wolfgang Kurtz (IBG-3,
Forschungszentrum JA~ ijlich GmbH) for providing the EnKF-HGS
simulator source code. We also acknowledge Oliver Schilling
(CHYN, University of Neucha^tel) for providing the
hydrological model use-case.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Brunner</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simmons</surname>
          </string-name>
          , C.T.:
          <article-title>Hydrogeosphere: A fully integrated, physically based hydrological model</article-title>
          .
          <source>Ground Water</source>
          <volume>50</volume>
          (
          <issue>2</issue>
          ),
          <volume>170</volume>
          {
          <fpage>176</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Burgers</surname>
            , G., van Leeuwen,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evensen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Analysis scheme in the ensemble Kalman lter</article-title>
          .
          <source>Monthly Weather Review</source>
          <volume>126</volume>
          (
          <issue>6</issue>
          ),
          <volume>1719</volume>
          {
          <fpage>1724</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Ca</surname>
            no-Lores,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fernandez</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <article-title>Garc a-</article-title>
          <string-name>
            <surname>Carballeira</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carretero</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A cloudi cation methodology for multidimensional analysis: Implementation and application to a railway power simulator</article-title>
          .
          <source>Simulation Modelling Practice and Theory</source>
          <volume>55</volume>
          ,
          <volume>46</volume>
          {
          <fpage>62</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Duro</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <article-title>GarcA~ a</article-title>
          , J.,
          <string-name>
            <surname>Isaila</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carretero</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wozniak</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          :
          <article-title>Flexible data-aware scheduling for work ows over an in-memory object store</article-title>
          .
          <source>In: Proceedings of IEEE/ACM CCGrid</source>
          <year>2016</year>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Etemadpour</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bomho</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lyons</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murray</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forbes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Designing and evaluating scienti c work ows for big data interactions</article-title>
          .
          <source>In: Big Data Visual Analytics (BDVA)</source>
          ,
          <year>2015</year>
          . pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          (Sept
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Evensen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Sequential data assimilation with a nonlinear quasi-geostrophic model using monte carlo methods to forecast error statistics</article-title>
          .
          <source>Journal of Geophysical Research: Oceans</source>
          <volume>99</volume>
          (
          <issue>C5</issue>
          ),
          <volume>10143</volume>
          {
          <fpage>10162</fpage>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Kurtz</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Hendricks Franssen,
          <string-name>
            <given-names>H.J.</given-names>
            ,
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.P.</given-names>
            ,
            <surname>Vereecken</surname>
          </string-name>
          , H.:
          <article-title>Joint assimilation of piezometric heads and groundwater temperatures for improved modeling of river-aquifer interactions</article-title>
          .
          <source>Water Resources Research</source>
          <volume>50</volume>
          (
          <issue>2</issue>
          ),
          <volume>1665</volume>
          {
          <fpage>1688</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Lapin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schiller</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kropf</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schilling</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brunner</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kapic</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Braun</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , Ma oletti, S.:
          <article-title>Real-time environmental monitoring for cloud-based hydrogeological modeling with hydrogeosphere</article-title>
          .
          <source>In: High Performance Computing and Communications</source>
          ,
          <source>2014 IEEE 6th Intl Symp on Cyberspace Safety and Security</source>
          ,
          <source>2014 IEEE 11th Intl Conf on Embedded Software and Syst (HPCC,CSS,ICESS)</source>
          ,
          <source>2014 IEEE Intl Conf on</source>
          . pp.
          <volume>959</volume>
          {
          <issue>965</issue>
          (Aug
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pacitti</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valduriez</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mattoso</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A survey of data-intensive scienti c work ow management</article-title>
          .
          <source>Journal of Grid Computing</source>
          <volume>13</volume>
          (
          <issue>4</issue>
          ),
          <volume>457</volume>
          {
          <fpage>493</fpage>
          (
          <year>2015</year>
          ), http://dx.doi.org/10.1007/s10723-015-9329-8
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Luckow</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mantha</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jha</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Pilot-abstraction: A valid abstraction for data-intensive applications on hpc, hadoop</article-title>
          and cloud infrastructures?
          <source>arXiv preprint arXiv:1501.05041</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>de Oliveira</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ogasawara</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , BaiA~ co,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Mattoso</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Scicumulus: A lightweight cloud middleware to explore many task computing paradigm in scienti c work ows</article-title>
          .
          <source>In: 2010 IEEE 3rd International Conference on Cloud Computing</source>
          . pp.
          <volume>378</volume>
          {
          <issue>385</issue>
          (
          <year>July 2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Raicu</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Many-task computing for grids and supercomputers</article-title>
          .
          <source>In: Many-Task Computing on Grids and Supercomputers</source>
          ,
          <year>2008</year>
          . MTAGS 2008. Workshop on. pp.
          <volume>1</volume>
          {
          <issue>11</issue>
          (Nov
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Szabo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sheng</surname>
            ,
            <given-names>Q.Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kroeger</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.:</surname>
          </string-name>
          <article-title>Science in the cloud: Allocation and execution of data-intensive scienti c work ows</article-title>
          .
          <source>Journal of Grid Computing</source>
          <volume>12</volume>
          (
          <issue>2</issue>
          ),
          <volume>245</volume>
          {
          <fpage>264</fpage>
          (
          <year>2014</year>
          ), http://dx.doi.org/10.1007/s10723-013-9282-3
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Therrien</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McLaren</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sudicky</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Panday</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A Three-dimensional Numerical Model Describing Fully-integrated Subsurface and Surface Flow</article-title>
          and
          <string-name>
            <given-names>Solute</given-names>
            <surname>Transport</surname>
          </string-name>
          .
          <source>Tech. rep. (</source>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Vahi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rynge</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juve</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mayani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deelman</surname>
          </string-name>
          , E.:
          <article-title>Rethinking data management for big data scienti c work ows</article-title>
          .
          <source>In: Big Data</source>
          , 2013 IEEE International Conference on. pp.
          <volume>27</volume>
          {
          <issue>35</issue>
          (Oct
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barbary</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nothaft</surname>
            ,
            <given-names>F.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sparks</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zahn</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franklin</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patterson</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perlmutter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Scienti c computing meets big data technology: An astronomy use case</article-title>
          .
          <source>In: Big Data (Big Data)</source>
          ,
          <source>2015 IEEE International Conference on</source>
          . pp.
          <volume>918</volume>
          {
          <issue>927</issue>
          (Oct
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Processing data-intensive work ows in the cloud (</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raicu</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Scienti c work ow systems for 21st century, new bottle or new wine?</article-title>
          <source>In: 2008 IEEE Congress on Services - Part I</source>
          . pp.
          <volume>467</volume>
          {
          <issue>471</issue>
          (
          <year>July 2008</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>