<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploiting Stream Reasoning to Monitor multi-Cloud Applications</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Miglierina</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Balduini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Narges Shahmandi Hoonejani</string-name>
          <email>narges.shahmandi@mail.polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elisabetta Di Nitto</string-name>
          <email>elisabetta.dinitto@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Danilo Ardagna</string-name>
          <email>danilo.ardagna@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Politecnico di Milano</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This demo shows how we have used a stream reasoning mechanism, C-SPARQL, as the main component of a Monitoring Platform for multi-Clouds applications, that is, applications replicated or distributed on multiple Clouds. The C-SPARQL engine executes monitoring queries on data gathered both from application-level components deployed on the Cloud and from Cloud-level resources. We show how, through our Monitoring Platform, we can enable and disable monitoring queries on demand. Thus, we increase/decrease the monitoring granularity depending on the overall system status, and therefore limit the monitoring overhead when possible.</p>
      </abstract>
      <kwd-group>
        <kwd>multi-cloud applications</kwd>
        <kwd>cloud computing</kwd>
        <kwd>stream reasoning</kwd>
        <kwd>monitoring</kwd>
        <kwd>c-sparql</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Cloud Computing is a novel paradigm based on the idea that computation,
storage and communication resources are o ered as a service to any user. Users
can exploit the resulting services based on their needs, adopting a pay-per-use
cost model. This makes the adoption of Cloud Computing a very interesting
business opportunity, especially for those companies who are rapidly growing or
expecting to grow up in the future.</p>
      <p>Today research is investigating various fronteers of Cloud Computing, one of
which is the possibility of designing and managing an application on multiple
Clouds. This is an interesting problem as it could limit an issue called customer
lock-in, that is, the di culty for a user to change Clouds given the potentially
high costs of migration. Moreover, it could open up the possibility of signi cantly
increasing the availability of an application.</p>
      <p>In the MODAClouds project1, among the various issues, we focus on the
problem of how to monitor multi-Cloud applications in a scenario where
applications are replicated on di erent Clouds and monitoring data of various kinds
need to be collected and analyzed from these replicas.</p>
      <p>
        This demo, in particular, shows how we have used a stream reasoning
engine, that is, C-SPARQL [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], as the main component to execute monitoring
queries on data gathered both from application-level components deployed on
1 www.modaclouds.eu
the Cloud and from Cloud-level resources. We show how, through our monitoring
platform, we can enable and disable monitoring queries on demand thus
increasing/decreasing the monitoring capabilities of the system depending on its overall
status, and therefore limiting the overhead of monitoring when possible and its
cost (note that also the network tra c is charged by many Cloud providers).
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>MODAClouds Monitoring Platform</title>
      <p>The MODAClouds monitoring architecture is illustrated in Figure 1. Data
Collectors are in charge of gathering monitoring data from di erent data sources
where the application is running. In particular, they associate semantic meaning
to data, and send this information to the C-SPARQL engine in the RDF format.
C-SPARQL is an extension of SPARQL language that enables continuous queries
over streams of RDF data. Monitoring queries are evaluated by the engine
according to sliding time windows so that reasoning is performed on knowledge
evolving over time. The Knowledge Base contains the Ontology, which is a
formal speci cation of the common abstractions needed to represent and monitor
Cloud systems and applications, and the Permanent RDF Data, which contains
information about the deployed system. Any Observer can nally subscribe for
queries so to receive their results when available. Last, the Monitoring Manager
is responsible for enabling or disabling data collectors, registering or
unregistering queries and keeping the Knowledge Base up to date.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The Demo</title>
      <p>The objective of this demo is to show how the Monitoring Platform behaves when
monitoring a commercial Web application deployed on two di erent Clouds,
Amazon Web Services and Eucalyptus@iEAT (a private Cloud based on
Eucalyptus 3.0).</p>
      <p>In particular, monitoring focuses on the execution time on the server side.
The commercial application we use for the demo is Apache Open For Business
(Apache OFBiz), an open source enterprise resource planning (ERP) system. It
provides a suite of enterprise applications that integrate and automate many of
the business processes of an enterprise.</p>
      <p>We implemented Data Collectors by instrumenting the Control Servlet code,
which is in charge of handling incoming requests, and deployed it on the two
Clouds (located in the Ireland Amazon Region and Romania, respectively). Each
replica (which includes an application server and an Apache Derby DB) is
installed on a Virtual Machine in each Cloud. We deployed the C-SPARQL engine
on another Amazon VM (located in the Ireland Region but in di erent
availability zone). The replicas are stressed through a load injector (we use Apache
JMeter) running in Italy and sending requests to both of them.</p>
      <p>We registered two di erent queries on the C-SPARQL engine:
{ Query Q1 checks the average execution time of requests arriving to the whole
system, composed by the two OfBiz replicas. If this value, within a 60 sec
time window {which is updated with a 10 sec step{ is above a 5 sec threshold,
the query produces a Violation Event on its output stream. In this case we
say that the data on the input stream have violated the monitoring query.
{ Query Q2 works exactly as query Q1 (and lters the same input) except for
the fact that the average is computed for each replica and for each type of
request. Thus, the query provides ner grained results and allows the system
administrator to identify the Cloud installation that is not performing as
expected.</p>
      <p>The execution of Q2 requires a larger number of computational and bandwith
resources (and hence incurs also in higher cost since the incoming Amazon tra c
is charged). Indeed, in case of violations, Q1 sends a violation event every 10
sec, while Q2 sends an event every 10 sec for every type of request in violation.
We therefore con gured the Monitoring Platform to activate Q2 automatically
only in case query Q1 detects a violation.</p>
      <p>The Knowledge Base is stored in the VM containing the C-SPARQL engine,
and is fed with an ontology that contains both Monitoring Platform-speci c
concepts and those OfBiz concepts that are relevant for the monitoring activity.
The Knowldge Base also contains information about the system deployment.</p>
      <p>The Monitoring Manager registers Q1 and Q2, enables the two Data
Collectors and attaches a text viewer as Observer of the two queries so the user can
visualize the data that are streamed out of them.</p>
      <p>In the demo we load the system by means of the load injector so to send
requests uniformly to both VMs, increasing the rate linearly (up to 180
requests/minute reached after 10 minutes) so to cause a violation of Q1, the
activation of Q2 and then its violation. Q1 violations occur after 2 minutes, while
Q2 violations start immediately (after the 60 sec time window) in the two
deployments.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>The purpose of this demo is to show an application of stream reasoning on
monitoring of Cloud-based applications. While some Cloud providers o er their own
monitoring mechanisms, when considering monitoring of multi-Cloud
applications new and exible monitoring platforms are needed. Stream reasoning can
be at the core of such kinds of platform thanks to its ability to acquire large
quantities of owing data and reason on them in a exible way.</p>
      <p>Information produced by monitoring platforms vary depending on the QoS
constraints posed on the application under analysis. While in this demo we focus
on execution time, other kinds of metrics could be more interesting in other
cases, for instance the number of requests answered without errors, the average
CPU utilization, the number of accesses to data storage, etc. The possibility
of programming the behavior of the C-SPARQL engine through queries and
the use of ontology certainly o ers a way to support the selection, ltering,
correlation of such di erent types of data (that can be also semantically di erent
in di erent providers). Of course, the engine should also be connected to proper
Data Collectors able to provide the right data to reason on.</p>
      <p>
        Disabling and enabling queries depending on the status of the monitored
application is an important mean to increase the level of accuracy of monitoring
when needed and decreasing it when it would only imply a large overhead and
cost on the execution of the application itself. Issues to be considered in the
evolution of our platform include the de nition of guidelines for the development of
Data Collectors. The Data Collector we used in the demo has been hard coded
as part of the servlet of the OfBiz application. While this approach may be the
only solution in the worst case, less invasive approaches should be pursued in
general. A promising approach we are investigating concerns the adoption of
aspect-oriented programming [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] as a way to build monitoring code fragments
that are then weaved within the application code. This enables separation of
concerns still allowing the programmer to instrument the application code. In
the case of monitoring resources in a lower level than the application (e.g., the
operating system, the virtual machine, the application container, ....), data
collection can be achieved only by exploiting speci c (and often proprietary) APIs,
if any (e.g., Amazon CloudWatch). The interesting characteristics of our
Monitoring Platform is that we keep these issues completely decoupled from the logic
needed to analyze the data. This certainly simpli es the design, even in presence
of many special cases. Another important issue is how to support the designer in
the de nition of monitoring queries starting from the de nition of QoS
requirements and constraints typically stated during the design of an application. To
support this issue we are experimenting with the semi-automatic generation of
monitoring queries by adopting a model-driven approach.
      </p>
      <p>Acknowledgments. This research has been partially supported by the
European Commission, Grant no. FP7-ICT-2011-8-318484 MODAClouds project.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Barbieri</surname>
            ,
            <given-names>D. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Braga</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Ceri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Della Valle</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grossniklaus</surname>
          </string-name>
          , M.:
          <string-name>
            <surname>C-SPARQL</surname>
          </string-name>
          :
          <article-title>A continuous query language for RDF data streams</article-title>
          .
          <source>Int. Journal of Semantic Computing (4)</source>
          ,
          <issue>1</issue>
          ,
          <fpage>3</fpage>
          -
          <lpage>25</lpage>
          .
          <year>2010</year>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Elrad</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Filman</surname>
            ,
            <given-names>R. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bader</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          :
          <article-title>Aspect-oriented programming: Introduction. Commun</article-title>
          .
          <source>ACM (44)</source>
          ,
          <volume>10</volume>
          ,
          <fpage>29</fpage>
          -
          <lpage>32</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>