<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Framework for Data Stream Applications in a Distributed Cloud</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Civil Engineering and Computer Science Engineering University of Rome Tor Vergata</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>56</fpage>
      <lpage>63</lpage>
      <abstract>
        <p>The ever increasing diffusion of sensing and computing devices enables a new generation of data stream processing (DSP) applications that operate in a distributed Cloud environment. Despite this, most of the existing solutions, such as Apache Storm, are designed to run in a local cluster. In this paper we present our extension of Storm, which provides distributed monitoring, scheduling and management capabilities. Exploiting these new functionalities, the system can improve its performance and react to internal and external changes. Finally, we analyze open challenges of placing and adapting DSP applications.</p>
      </abstract>
      <kwd-group>
        <kwd>Data Stream Processing</kwd>
        <kwd>Adaptation</kwd>
        <kwd>Placement</kwd>
        <kwd>Apache Storm</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Matteo Nardelli</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        With the disruptive diffusion of sensing devices (e. g., smartphones, cars,
monitoring stations), the almost ubiquitous Internet connection, and the Fog
Computing [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] paradigm, urban environments are today permeated by an ever increasing
number of diffused and networked sensing and computing devices. All these
sensing devices continuously produce streams of data that can be collected by
distributed data stream processing (DSP) applications, to timely extract valuable
information about many fundamental aspects of the environment we live in
(e. g., urban mobility, public decision making, energy management). As data
increases, we cannot push it toward the core of Internet. To increase scalability
and reduce latency, a possible solution is to rely on distributed and near-edge
computation. Furthermore, determining the computational resources that should
host and execute each operator of the DSP application, i. e., solving the operator
placement problem, is challenging because the characteristics of computational
tasks are not known a-priori, the properties of the input streams change
continuously, and the load imposed has to be sustained for long provisioning times.
Therefore, we extended Storm [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], an open source DSP system, with policies
and mechanisms that allow to find a placement that optimizes a utility function
and to continuously adapt the placement when changes occur in the execution
environment.
      </p>
      <p>
        The main contributions of this paper are as follows: a) we describe how
our extension implements the MAPE (Monitor, Analyze, Plan, and Execute)
reference model for autonomic systems (Sect. 4); b) we show its benefits when
the placement is determined according to the distributed policy proposed by
Rizou et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] (Sect. 5); and c) we illustrate some of the open challenges for
DSP systems when they are executed in distributed environments (Sect. 6).
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>
        As technologies and needs evolve in time, the DSP paradigm has experienced
different generation of architectures [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], where the last one relies on Cloud-based
resources. Despite this, most DSP systems are still designed to run in a local
cluster, where the often homogeneous nodes are interconnected with negligible
network delays (e. g., [
        <xref ref-type="bibr" rid="ref14 ref15 ref17">14, 15, 17</xref>
        ]). These assumptions do not hold any more when
the DSP system runs in geographically distributed and dynamic environments,
where a great heterogeneity of devices are interconnected with not-negligible
network latencies. Storm, a framework of the last generation, is attracting increasing
interests. However, most of the proposed Storm extensions are all centralized
solutions (e. g., [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]), implicitly designed for clustered environments, which do
not scale well as the number of applications increases. Our extension, instead,
provides distributed monitoring, scheduling and management capabilities [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        A great variety of placement algorithms have been proposed in literature.
Lakshmanan et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] provide a comprehensive overview of them, but, as the
authors show, they differ each other on assumptions and optimization goals. Being
interested in a network-aware solution, in our previous work [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] we implemented
the Pietzuch’s algorithm [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], and here we evaluate the strategy proposed by
Rizou et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Both the solutions minimize network usage, however the authors
of [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] claim that their formulation has better convergence properties and works
better in a distributed environment than [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        Recently, another framework for large-scale processing is gaining interest:
Spark [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. It extends and improves the MapReduce approach (batch processing),
and, using the Spark Streaming module, can reduce the size of each batch and
process streams of data (micro-batch processing). This alternative is throughput
oriented, whereas Storm, which is a pure DSP system, can further minimize the
application latency, therefore is preferred in latency sensitive scenarios. Apache
Flink1 proposes a unified framework for batch and stream processing. Similarly
to Storm, Flink has been originally designed to run in a cluster environment and
shows the drawbacks we discuss in Sect. 4.1. DSP systems are also offered as
Cloud services. Google Cloud Dataflow2 provides a unified programming model
to process batch and streaming data on top of Google cloud platform. Amazon
offers Kinesis3, which resembles an evolved publish-subscribe system, suitable to
process near real-time streams of data. Both these Cloud-based services abstract
      </p>
      <sec id="sec-3-1">
        <title>1 https://flink.apache.org/ 2 https://cloud.google.com/dataflow/ 3 https://aws.amazon.com/kinesis/</title>
        <p>sources of words count words</p>
        <p>intermediate rank total rank
(WORD)
(WORD,COUNT)
spout
bolt
the underlying infrastructure, but it is reasonable to believe that they execute in
a centralized data center, conversely to the context investigated in this paper.
3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Apache Storm</title>
      <p>Storm4 is an open source and scalable DSP system maintained by the Apache
Software Foundation. It provides an abstraction layer where event-based
applications can be executed over a set of worker nodes interconnected by an overlay
network. A worker node is a generic computational resource, whereas the overlay
network comprises the logical links between these nodes. In Storm, an application
is represented by its topology, which is a directed acyclic graph with spouts and
bolts as vertices and streams as edges. A spout is a data source that feeds the
data into the system through one or more streams. A bolt is either a processing
element, which extracts valuable information from incoming data and generates
new outgoing streams, or a final information consumer. A stream is an unbounded
sequence of tuples, which are key-value pairs. We refer to spouts and bolts as
operators. Figure 1a shows an example of a DSP application. Storm uses three
types of entities with different grain to execute a topology. A task is an instance of
an operator in charge of a share of its incoming streams. An executor can execute
one or more tasks related to the same operator. A worker process is a Java process
that runs a subset of executors of the same topology. As represented in Fig. 1b,
there is a hierarchy among these entities: a group of tasks runs sequentially in
the executor, which is a thread within the worker process that serves as container
on the worker node. Besides the computational resources (i. e., worker nodes),
Storm includes two centralized components: Nimbus and ZooKeeper. Nimbus
coordinates the topology execution and defines the placement of its operators
on the available worker nodes. This assignment plan is communicated to all the
worker nodes through ZooKeeper, which is a shared memory service that enables
distributed coordination. Since each worker node can execute one or more worker
processes, a Supervisor component on the node starts and terminates worker
processes on the basis of the Nimbus decisions.</p>
      <sec id="sec-4-1">
        <title>4 http://storm.apache.org/</title>
        <p>4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Distributed Storm</title>
      <p>From Cluster to Distributed Cloud: A Gap to Close
Storm has been originally designed to run in a local cluster, where network
delays are negligible. If we deploy Storm in a distributed Cloud, it shows poor
performances, because of the assumption that data can quickly move between
computational nodes. We can summarize the limitations that Storm shows
in this new environment as follows: 1) it is unaware of QoS attributes (e. g.,
resource utilization, network delays) of computational and network resources;
2) its placement decision is static, therefore the system cannot adapt to internal
(i. e., application) and external (i. e., environmental) changes; and 3) if we create
a custom centralized scheduler that collects the QoS attributes for each node and
periodically evaluates the placement of each application, it will not scale well as
the number of applications and network resources increases. In a geographically
distributed environment, we would like to have a framework that considers network
delays and resource heterogeneity while determining placement decisions.
4.2</p>
      <p>
        Distributed Scheduling in Storm
We have extended the Storm architecture to run distributed, adaptive, and
QoSaware scheduling algorithms [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The newly introduced components, illustrated
in orange in Fig. 2, are: the AdaptiveScheduler, the QoSMonitor, and the
WorkerMonitor. We preserved the centralized scheduler, named BootstrapScheduler,
which defines the initial placement of the application. The AdaptiveScheduler is
the distributed scheduler that coordinates the MAPE control cycle. It executes on
each Supervisor together with the QoSMonitor, an infrastructure level monitoring
component. The WorkerMonitor is an application level monitor and runs on each
worker process. Exploiting the feedback control loop, the distributed scheduler
can react to internal and external changes of the operating conditions. In a single
loop iteration, it monitors the environment and the locally executed executors,
rke irto rke irto
ro onM ro onM
w W w W
      </p>
      <p>Supervisor
ASched QoSMonitor
rke irto
ro onM
w W
rke irto rke irto
ro onM ro onM
w W w W</p>
      <p>Supervisor
ASched QoSMonitor
rke irto
ro onM
w W
rke irto rke irto
ro onM ro onM
w W w W</p>
      <p>Supervisor
ASched QoSMonitor
rke irto
ro onM
w W
rke irto rke irto
ro onM ro onM
w W w W</p>
      <p>Supervisor
ASched QoSMonitor
rke irto
ro onM
w W</p>
      <p>Network
Nimbus BSched</p>
      <p>ZooKeeper
analyzes if there are candidate executors for a new reassignment, and, in positive
case, plans and executes the corresponding repositioning actions.</p>
      <p>
        Monitor. The AdaptiveScheduler acquires the information on computational
resources and on executors that run locally through the QoSMonitor and the
WorkerMonitors respectively. The QoSMonitor provides the QoS awareness to
each distributed scheduler, thus it is responsible of obtaining intra-node
information (i. e., utilization and availability) and inter-node information (i. e., network
delays). For the latter, it resorts on a network coordinates system [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] that provides
an accurate estimate of the delays between any two computational nodes without
the need of an exhaustive probing. The WorkerMonitor computes the exchanged
data rate for each executor that runs on the node.
      </p>
      <p>
        Analyze and Plan. A distributed scheduling policy drives these two phases.
Our previous work [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] relies on the Pietzuch’s placement algorithm [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In this
paper, we use the scheduling solution designed by Rizou et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], which places
the application minimizing the network usage (i. e., sum of bandwidth-delay
product for each application link). Implementing the Rizou’s algorithm within
the extended Storm requires just few changes. Basically, it needs to account
for the specific Storm application model, where a processing operator can be
instantiated in one or more executors and pinned operators are not modeled.
Furthermore, the algorithm can readily obtain QoS information (i. e., latency,
bandwidth) relying on the monitoring components.
      </p>
      <p>Execute. Finally, if a new assignment must take place, the executor is moved
to the new candidate node. The new assignment decision is shared with the
involved worker nodes through ZooKeeper. We note that in Storm an executor
reassignment does not preserve its state; thus, the executor is stopped on the
previous worker node and started on the new one.</p>
      <p>Thanks to the adaptation cycle, the distributed scheduler can manage changes
that may occur both in the infrastructure layer (e. g., a worker node appears or
fails) and application layer (e. g., data rate fluctuations).</p>
      <p>The source code of our extension is available at http://bit.ly/extstorm.
5</p>
    </sec>
    <sec id="sec-6">
      <title>Experimental Results</title>
      <p>
        We show the improvements and the self-adaptation capabilities of our distributed
scheduler equipped with the Rizou’s algoritm (named as dRizou) with respect
to the centralized and default EvenScheduler of Storm (named as cRR). For a
better evaluation, we also indicate the behaviour of dQoS, that is the distributed
scheduler equipped with the Pietzuch’s algoritm (further details in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]). dRizou
and dQoS place operators exploiting QoS attributes, whereas cRR uses a
roundrobin policy. The evaluation uses a cluster of 8 worker nodes (each can host
at most 2 worker processes) and 2 further nodes for Nimbus and ZooKeeper.
We emulated wide-area network latencies among the Storm nodes applying to
outgoing packets a Gaussian delay with mean and standard deviation in the
ranges [
        <xref ref-type="bibr" rid="ref12">12, 32</xref>
        ] ms and [
        <xref ref-type="bibr" rid="ref1 ref3">1, 3</xref>
        ] ms, respectively. The DSP application is composed of
a source, which generates 10 tuples/s, followed by a sequence of 5 operators before
Time (s)
      </p>
      <p>cRR
dRizou
dQoS
dRizou migrations
dQoS migrations
6000 7000
reaching the final consumer. The placement of source and consumer is fixed. The
other operators are unpinned and replicated (i. e., two executors are assigned
to each of them). Figure 3 shows the evolution of the application end-to-end
latency; on its bottom, we indicate the run-time reassignments performed by the
distributed schedulers (cRR does not intervene during the execution). We start
the application and, after 3240 s, we artificially increase the load on a subset of
three nodes using the Linux tool stress. The subset is composed by one worker
node running some application executors and two free worker nodes. This event
is represented in Fig. 3 with a vertical dotted line. As the distributed scheduler
(both dQoS and dRizou) perceives the change, it moves the application operators
on lightly loaded nodes. cRizou reduces the application latency with respect to
cRR of about 12.6 % (measured between 5000 s and the end of the experiment).
Furthermore, differently from dQoS, dRizou converges with a lower number of
reassignments, increasing the application availability.
6</p>
    </sec>
    <sec id="sec-7">
      <title>Open Challenges</title>
      <p>Although our extension enables the execution of the Storm, as a generic DSP
system, in a distributed environment, the peculiarities of Cloud computing require
an efficient management of scalability, elasticity and fault tolerance. With no
claim of completeness, we summarize some of the needed mechanisms.</p>
      <p>
        Stateful Migration: an operator is stateful if its behavior depends also
on its internal state. Therefore, moving a stateful operator requires an efficient
relocation of its internal, possibly extremely large, state across the network. In
literature, the general tendency is to use the strategy stop-move-play, which stops
the incoming streams, moves the operator and its state, and redirects the streams
to the new operator location (e. g., [
        <xref ref-type="bibr" rid="ref4 ref6">4, 6</xref>
        ]). Wu et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] improve this technique
by aggressively dividing the application-level state in computation slices, which
are asynchronously checkpointed to remote machines, enabling parallel state
migrations between nodes. However, most of the existing techniques do not fit
well in a latency sensitive scenario, because they do not explicitly consider QoS
attributes of communication links and computational nodes. A fast, live, and
QoS-aware migration strategy could bring important improvements to these
systems.
      </p>
      <p>
        Elastic Replication: the ability of the system to autonomously adapt the
number of replicas for each operator. This mechanism can increase non-functional
attributes of the applications (e. g., availability) with the penalty of a higher cost
and resource overhead. Bellavista et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] present a prototype that allows to
trade-off monetary cost and active replication. An alternative to active replication
is upstream backup, which achieves fault-tolerance using an upstream server
that stores a copy of the operator state. However, since this technique imposes
a higher recovery time, it is used as a second-class mechanism. For example,
Heinze et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] combine these two mechanisms to reduce the overall resource
consumption with respect to a recovery time threshold.
      </p>
      <p>
        Elastic Sharding: the ability of automatically scaling in and out the number
of shards for an operator based on the incoming load. Each shard of an operator
is in charge of a partition of its incoming stream; therefore, this mechanism
can increase the application scalability by handling a growing workload. As a
consequence, the system can acquire and release resources when needed, without
resorting in over- or under-provisioning (i. e., resource elasticity [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]). Increasing
the number of shards is critical for stateful operators, because the system needs to
preserve the consistency of the operations. In literature different works investigate
this issue. Some solutions define a-priory the maximum number of shards [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ],
expose some API to manually manage the state [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], or automatically determine
the optimal number of state partitions to be used [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>Solutions to the above mentioned issues are almost consolidated in a clustered
environment, however the emerging distributed Cloud scenario imposes a new
perspective. In a distributed environment, aside the number of replicas or shards,
the scheduler should also define their optimal placement with respect to some
QoS metrics (e. g., latency, bandwidth, reliability), considering the heterogeneity
of applications and resources. For example, replicas should be placed in different
availability zones; different shards of the same operator should let users experience
similar response times.
7</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusion</title>
      <p>The ever increasing diffusion of sensing and computing devices enables a new
generation of DSP systems. Starting from the major drawbacks of an existing
framework to the execution in distributed and dynamic environments, we
developed an extension of Apache Storm that provides distributed monitoring,
scheduling and management capabilities. The evaluation results showed that our
extension of Storm is suitable to operate in a distributed environment, where
QoS-awareness and adaptation capabilities can be truly beneficial to the
application performances. Finally, we highlighted some core mechanisms that can
improve performances of DSP systems when executed in a distributed Cloud.</p>
      <p>As future work, we will provide a formal definition of the placement problem
for DSP applications and design a new placement algorithm that better leverages
the potentialities of a distributed Cloud model.</p>
      <p>Acknowledgments. Thanks to the anonymous reviewers for the valuable
comments, to Valeria Cardellini, and to Gabriele Scolastri for the implementation of
the Rizou’s algorithm.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aniello</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baldoni</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Querzoni</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Adaptive online scheduling in Storm</article-title>
          .
          <source>In: Proc. of ACM DEBS '13</source>
          . pp.
          <fpage>207</fpage>
          -
          <lpage>218</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bellavista</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corradi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kotoulas</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reale</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Adaptive fault-tolerance for dynamic resource provisioning in distributed stream processing systems</article-title>
          .
          <source>In: EDBT</source>
          . pp.
          <fpage>85</fpage>
          -
          <lpage>96</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cardellini</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grassi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Lo</given-names>
            <surname>Presti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Nardelli</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Distributed QoS-aware Scheduling in Storm</article-title>
          .
          <source>In: Proc. of ACM DEBS '15</source>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Castro</surname>
            <given-names>Fernandez</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Migliavacca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Kalyvianaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Pietzuch</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Integrating scale out and fault tolerance in stream processing using operator state management</article-title>
          .
          <source>In: Proc. of ACM SIGMOD'13</source>
          . pp.
          <fpage>725</fpage>
          -
          <lpage>736</lpage>
          . ACM (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dabek</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cox</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaashoek</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morris</surname>
          </string-name>
          , R.:
          <article-title>Vivaldi: A decentralized network coordinate system</article-title>
          .
          <source>SIGCOMM Comput. Commun. Rev</source>
          .
          <volume>34</volume>
          (
          <issue>4</issue>
          ),
          <fpage>15</fpage>
          -
          <lpage>26</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gedik</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schneider</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirzel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>K.L.</given-names>
          </string-name>
          :
          <article-title>Elastic scaling for data stream processing</article-title>
          .
          <source>IEEE Trans. Parallel Distrib. Syst</source>
          .
          <volume>25</volume>
          (
          <issue>6</issue>
          ),
          <fpage>1447</fpage>
          -
          <lpage>1463</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Heinze</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aniello</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Querzoni</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jerzak</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Cloud-based data stream processing</article-title>
          .
          <source>In: Proc. of ACM DEBS '14</source>
          . pp.
          <fpage>238</fpage>
          -
          <lpage>245</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Heinze</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krahn</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jerzak</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , et al.:
          <article-title>An adaptive replication scheme for elastic data stream processing systems</article-title>
          .
          <source>In: Proc. of ACM DEBS '15</source>
          . pp.
          <fpage>150</fpage>
          -
          <lpage>161</lpage>
          . ACM (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hochreiner</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulte</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dustdar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lecue</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Elastic stream processing for distributed environments</article-title>
          .
          <source>Internet Computing, IEEE</source>
          <volume>19</volume>
          (
          <issue>6</issue>
          ),
          <fpage>54</fpage>
          -
          <lpage>59</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lakshmanan</surname>
            ,
            <given-names>G.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strom</surname>
          </string-name>
          , R.:
          <article-title>Placement strategies for internet-scale data stream systems</article-title>
          .
          <source>Internet Computing, IEEE</source>
          <volume>12</volume>
          (
          <issue>6</issue>
          ),
          <fpage>50</fpage>
          -
          <lpage>60</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pietzuch</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ledlie</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shneidman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roussopoulos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>Network-aware operator placement for stream-processing systems</article-title>
          .
          <source>In: Proc. of IEEE ICDE '06</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Rizou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Durr</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rothermel</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Solving the multi-operator placement problem in large-scale operator networks</article-title>
          .
          <source>In: Proc. of IEEE ICCCN</source>
          <year>2010</year>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Satyanarayanan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuster</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ebling</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fettweis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , et al.:
          <article-title>An open ecosystem for mobile-cloud convergence</article-title>
          .
          <source>IEEE Communications 53(3)</source>
          ,
          <fpage>63</fpage>
          -
          <lpage>70</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Toshniwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taneja</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shukla</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramasamy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , et al.:
          <source>Storm@Twitter. In: Proc. of ACM SIGMOD '14</source>
          . pp.
          <fpage>147</fpage>
          -
          <lpage>156</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Urbani</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Margara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jacobs</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voulgaris</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>AJIRA: a lightweight distributed middleware for MapReduce and stream processing</article-title>
          .
          <source>In: Proc. of IEEE ICDCS '14</source>
          . pp.
          <fpage>545</fpage>
          -
          <lpage>554</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>K.L.</given-names>
          </string-name>
          :
          <article-title>Chronostream: Elastic stateful stream computation in the cloud</article-title>
          .
          <source>In: IEEE Int'l Conf</source>
          .
          <article-title>ICDE 2015 (forthcoming</article-title>
          ) (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Zaharia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chowdhury</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franklin</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shenker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>Spark: Cluster computing with working sets</article-title>
          .
          <source>In: Proc. of USENIX HotCloud'10</source>
          . p.
          <volume>10</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>