<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>IWSG</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>IoT-Hub: New IoT Data-Platform for Virtual Research Environments</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Rosa Filguiera</institution>
          ,
          <addr-line>Rafael Ferreira da Silva</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>13</volume>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>-This paper presents IoT-Hub a new scalable, elastic, efficient, and portable Internet of Things (IoT) data-platform based on microservices for monitoring and analysing largescale sensor data in real-time. IoT-Hub allows us to collect, process, and store large amounts of data from multiple sensors in distributed locations-which could be deployed as a backend for Virtual Research Environments (VRE) or Science Gateways. In the proposed data-platform, all required software, which involves a variety of state-of-the-art open-source middleware, is packed into containers and deployed in a cloud environment. As a result, the engineering and computational time and costs for deployment and execution is significantly reduced.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        The emergence of the Internet of Things (IoT) is introducing
a new era to the realm of computing and technology [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
proliferation of sensors and actuators that are embedded in
things enables these devices to understand the environments
and respond accordingly more than ever before. Additionally,
it opens unlimited possibilities to domain scientists and/or
data scientists for building models and analyses that turn
this sensation into big benefits to science and society.
Realtime processing of big data streams will gain importance as
embedded technology increases and we continue to generate
new types and methods of data analysis [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], particularly in
regard to IoT. However, this revolutionary spread of IoT
devices creates big challenges, such as choosing, deploying,
and managing adequate data-frameworks for data-intensive
computation in science, engineering, and many other fields.
      </p>
      <p>
        Virtual Research Gateways (VREs), also known as Science
Gateways [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], are web-tools accessible from anywhere. They
usually provide an integrated view of all available resources
with pervasive data access control, handle continuity between
sessions and support collaboration with shared data and
methods. VREs open up opportunities for sharing and comparing
both experiment data from experiments, observations, and
model runs and analytic interpretations of these data. They
are very popular in a variety of scientific communities (e.g.
seismology or astronomy) since they hide many technical and
management details whose use are not straightforward for
nonexperts. The connection to and between VREs and science
automation technologies has gained a lot of attention in the
last years. However, that is not the case for the VREs, IoT,
and all new middleware for emerging data-intensive analytics.
      </p>
      <p>In this paper, we present IoT-Hub, an integrated,
comprehensive, elastic, and portable data-platform based on
microservices. IoT-Hub combines the benefits of several
wellknown data-frameworks with Docker containers. The
current implementation of IoT-Hub includes a service-pipeline
composed by Apache Kafka, Apache Spark, Elasticsearch,
and Kibana middleware that enables automated gathering,
preprocessing, storing, and visualization of IoT streams in
a scalable, efficient, and robust manner. IoT-Hub acts as a
backend for VREs to run stream-based applications, deploying
cloud resources upon request. It reduces the engineering time
and effort (and possible human errors) required by scientists
or VRE administrators to build such complex systems. Our
hypothesis is that if we provide scientific communities with
portable and elastic platforms to interrogate the IoT data, it
will speed up scientific discoveries.</p>
      <p>
        We have demonstrated the feasibility of IoT-Hub via a
real use case application, which processes sensor data from
the British Geological Survey (BGS) environmental baseline
programme [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] (freely available online). IoT-Hub collects,
preprocesses, and stores in real-time time-series data from
several distributed locations and sensors types, and makes them
available to domain scientists (e.g. groundwater modelers) and
data scientists, so they can use it to build their models, make
predictions, and conduct analyses.
      </p>
      <p>This paper is structured as follows. Section II presents
background. Section III discusses IoT-Hub features. Section IV
presents the use case for testing the platform. We conclude
with a summary of achievements and outline future work.</p>
    </sec>
    <sec id="sec-2">
      <title>II. BACKGROUND AND RELATED WORK In this section, we provide a brief overview on the stateof-the-art encompassing VREs, IoT, and middleware for dataIntensive analytics.</title>
      <sec id="sec-2-1">
        <title>A. Virtual Research Environments (VREs)</title>
        <p>
          VREs can be defined as community-development set of
tools, applications, and data that is integrated via a portal or a
suite of applications, usually in a graphical user interface, that
is further customized to meet the needs of a specific
community [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. These tools sit behind the scenes and exploit a wealth
of resources residing on multiple computing infrastructures
and data providers (according to their policies). Some VREs
examples include:
        </p>
        <p>
          VERCE [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] is a data-intensive e-science environment to
enable innovative data analysis and data modeling
methods that fully exploit the increasing wealth of open data
generated by the observational and monitoring systems
of the global seismology community.
        </p>
        <p>
          MoSGrid [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] is a portal that offers an approach to carrying
out high-quality molecular simulations on distributed
compute infrastructures to scientists with all kinds of
background and experience levels.
        </p>
        <p>
          CyberSKA [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] is a collaborative portal which aims to
address the current and future needs of data-intensive
radio astronomy. A wide variety of tools and services that
have been developed and integrated with the CyberSKA
portal, including a distributed data management system,
a data access tool, remote visualization tools, and third
party applications.
        </p>
        <p>
          EFFORT [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] is an innovative platform to promote
persistent collaboration research in Rock Physics and
Volcanology. It organizes data from rock physics experiments and
volcano monitoring to open up opportunities for sharing
and comparing data, observations and model runs, and
analytical interpretation methods.
myExperiment [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] is a portal for collaboration and
sharing of workflows and experiments. In contrast to
systems that simply make workflows available, it
provides mechanisms to support the sharing of workflows
within and across multiple communities via a social web
approach.
        </p>
        <p>Having a closer look to the technologies, tools, systems,
and computing resources that are very often behind VREs’
backends, we can categorize them as follows:</p>
      </sec>
      <sec id="sec-2-2">
        <title>High Performance Computing solutions: aggregated com</title>
        <p>
          puting resources to perform high performance
computations (including processors, memory, disk, and operating
system) [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ];
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Distributed Computing Infrastructures: distributed sys</title>
        <p>tems characterized by heterogeneous networked
computers called to offer data processing facilities. This includes
high-throughput computing and cloud computing;</p>
      </sec>
      <sec id="sec-2-4">
        <title>Scientific workflow management systems (SWMS): sys</title>
        <p>
          tems enacting the definition and execution of
scientific workflows consisting of a list of tasks and
operations, the dependencies between the interconnected
tasks, control-flow structures, and the data resources to
be processed [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ];
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Data analytics frameworks and platforms: platforms and</title>
        <p>
          workbenches enabling scientists to execute analytic tasks.
Such platforms tend to provide their users with
implementations of algorithms and (statistical) methods for the
analytics tasks [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>These classes of solutions and approaches are not isolated,
rather they are expected to rely on each other to provide
VREs end users with easy to use, efficient, and effective
data processing facilities, e.g., SWMS rely on distributed
computing infrastructures to actually execute their constituent
tasks.</p>
        <p>The proposed framework (IoT-Hub) can be used as a VRE
backend, where scientists can simply inject and execute their
processing analyses (via VRE fronted) without putting effort
in operating the enabling technology. In order to meet this
goal, we have leveraged Docker containers, which allows us
to have an elastic computational environment based in loosely
coupled services, which are immediately portable. Docker
handles the packaging and execution of a container so that
it works identically across different machines, while exposing
the necessary interfaces for networking ports, volumes, and
so forth, allowing other users to reconstruct an equivalent
computational environment. Therefore, IoT-Hub can be
deployed on demand (as-a-service) reducing engineering time
and computational cost.</p>
      </sec>
      <sec id="sec-2-6">
        <title>B. Internet of Things (IoT): Big Data challenges</title>
        <p>
          The explosive increase in the number of devices connected
to the IoT and the exponential increase in data consumption
only reflect how the growth of big data perfectly overlaps with
that of IoT [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. And therefore, many architectural design
challenges have arisen for the delivery of big data services
based on the IoT. These challenges have been described in
detail in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. In this work, we have mainly focused in the
following ones:
        </p>
        <p>The number of IoT devices: With growth forecasted in
the number of connected “things” and expected to reach
billions world-wide, there will be masses of devices
which may be a data source, and which may be subject
to third party control;</p>
      </sec>
      <sec id="sec-2-7">
        <title>Risk of IoT device malfunction: With a great number</title>
        <p>of IoT devices and manufacturers it is reasonable to
assume there will be many occasions where IoT devices
malfunction in various ways;
Update frequency: Though some devices will produce
data reports at a low frequency there may be substantial
quantities of data streaming from more sophisticated
Internet connected things.</p>
        <p>Therefore, IoT-Hub has been designed to collect data
from a wide range of different IoT devices, geographically
distributed that stream data at different frequency ratios, and
could yield malfunction behaviors intermittently.</p>
        <p>
          Our proposed solution provides an IoT data-platform, which
support an ecosystem of third party application developers
(e.g.domain scientists and data scientists) to explore data
given the described challenges. IoT-Hub offers a degree of
flexibility making use of Docker and Docker-compose tools
for deploying services on demand in virtualized environments,
such as cloud systems. Previous works have targeted similar
environments, such as [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], which is a cloud-based autonomic
information system for delivering Agriculture-as-a-Service
(AaaS) through the use of cloud and big data technologies.
        </p>
      </sec>
      <sec id="sec-2-8">
        <title>C. Microservices &amp; Middleware</title>
        <p>
          With recent advances in cloud computing, virtualization,
containerization, continuous integration, and the DevOps
movement, deploying software solutions today is very
different from even just a few years ago. Today’s distributed
applications are built as a set of independently deployable
microservices distributed over clusters of commodity
hardware. The microservice term – also known as the microservice
architecture – refers to a new architectural style that structures
an application as a collection of loosely coupled services,
which implement business capabilities. So, we have built
IoT-Hub following microservice principles [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
        <p>The question that arises now is, which components should
be used in IoT-Hub to develop a high performance platform
to efficiently analyze IoT big data. For answering this
question, we developed a prototype platform in an elastic cloud
environment, and Falcon1, Apache Kafka2, Apache Spark3,
Elasticsearch4, Kibana5 and Docker6 have been initially
selected (see Table I). This selection could be easily extended
in the future for including additional data-frameworks, such
as Cassandra, Apache Flink, and Jupyter Notebooks.
Reliable, high-performance Python web
framework for building large-scale app
backends and microservices. It
encourages the REST architectural style with
minimal external dependencies, while
remaining highly effective.</p>
        <p>Distributed streaming platform that
allows for publishing and subscribing to
streams of records (topics) in a
faulttolerant way and process streams of
records as they occur.</p>
        <p>Fast and general engine for large-scale
data processing. Among other features,
it allows writing streaming jobs the
same way as writing batch jobs. It
supports Java, Scala and Python.</p>
        <p>A distributed, RESTful search and
analytics engine for performing and
combining many types of searches
structured, unstructured, geo or metric.</p>
        <p>An open source data visualization
plugin for Elasticsearch. It provides
visualization capabilities on top of the content
indexed on an Elasticsearch cluster. It
also supports remote I/O
A lightweight, stand-alone, executable
package of a piece of software that
includes everything needed to run it:
code, runtime, system tools, system
libraries, settings.</p>
        <p>Version
2.0
0.10.2.0
2.2.0
6.2.2
(oss)
6.2.2
(oss)
1.13.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>1https://falconframework.org/</title>
      <p>2https://kafka.apache.org/
3https://spark.apache.org/
4https://www.elastic.co/
5https://www.elastic.co/products/kibana
6https://www.docker.com/
Events / Data Streams</p>
      <p>Data Processing/
Data Cleaning</p>
      <p>Data Storage</p>
      <p>Data Analysis /
Visualizations
Location 1
SA SB
SC
Location 2
SA SB
SC
Location 3
SA SB
SC</p>
      <p>Falcon
data generation</p>
      <p>Data
Producers
(Domain
Scientists)</p>
      <p>Apache
Kafka</p>
      <p>Apache
Spark</p>
      <p>ElasticSearch</p>
      <p>Kibana</p>
      <p>Cloud Infrastructure
data preparation and storage</p>
      <p>applications</p>
      <p>Data
Engenieers</p>
      <p>Scientists
(Domain
Scientists /</p>
      <p>Data
Scientists)</p>
      <p>Fig. 1: IoT-Hub: Data-platform for gathering, quality
checking, storing, and visualizing environmental sensors streams.</p>
    </sec>
    <sec id="sec-4">
      <title>III. IOT-HU B FEATURES</title>
      <p>IoT-Hub integrates several middleware based on the
microservices architecture. Apache Kafka provides the
mechanism for ingesting real-time data streams and making them
available to downstream consumers in a parallel and
faulttolerant manner. Data in Apache Kafka is organized into topics
that are split into partitions for parallelism. A topic can be
viewed as an infinite stream where data is retained for a
configurable amount of time. Producers are applications that
publish stream of records to one or more topics. In our case,</p>
      <sec id="sec-4-1">
        <title>Apache Kafka streams events out to Apache Spark consumers</title>
        <p>which are subscribed to one or more topics for parsing their
content, all of which is done in near real-time.</p>
        <p>Spark Streaming API enables scalable, high-throughput,
fault-tolerant stream processing of live data streams. Data
can be ingested from many sources (e.g. Apache Kafka,
Flume, Twitter), but in this current version of IoT-Hub we
limited it to Apache Kafka. Spark Streaming API can be used
for processing the ingested data using complex algorithms
composed of high-level functions like map, reduce, join,
and window. The processed data can be published to yet
another Kafka topic for further consumption or it can be stored
as results in HDFS, databases, or dashboards. In this work, we
have selected Elasticsearch as temporary storage system. One
of the reasons for this choice is elasticsearch-hadoop provides
native integration between Elasticsearch and Apache Spark, in
the form of an RDD (Resilient Distributed Dataset).</p>
        <p>Kibana offers interactive visualizations (e.g. histograms,
line graphs, pie charts, sunbursts, etc.) and advanced time
series analysis on Elasticsearch data, by leveraging the full
aggregation capabilities of Elasticsearch.</p>
        <p>In order to have a portable, scalable, and elastic
dataplatform, we created a Docker cluster for each of the previous
middleware and connected them via docker-compose, since it
allows us to setup and run multi-container environments.
Figure 1 shows a visual description of the IoT-Hub components,
and the interactions from different roles (e.g. data producers,
data architects, data scientists, domain scientists) that we
anticipate. All Dockerfiles and docker-compose file
used to generate the IoT-Hub are available freely online in a
GitHub repository7 for allowing reproducibility and share our
approach among the scientific community.</p>
        <p>For our experiments, we have used the NSF-Chameleon
cloud8, using a CentOS7 image with 42 CPUS for deploying
our hub. Note that the proposed framework could be deployed
to any other Cloud system.</p>
        <p>IV. CASTE STUDY: ENVIRONMENTAL BASELINE</p>
        <p>MONITORING PROGRAMME</p>
        <p>
          To demonstrate the feasibility of IoT-Hub, we have used
the Environmental Baseline Monitoring programme [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], which
provides the perfect scenario for testing our platform with IoT
sensor data. The British Geological Survey (BGS), along with
partners from the Universities of Manchester, York,
Birmingham, Bristol, and Public Health England (PHE), is carrying
out a science-based environmental monitoring programme in
the areas of Lancashire [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] and Vale of Pickering [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. This
programme represents the first independent, integrated
monitoring study to characterize the environmental baseline in areas
subjected to close scrutiny in anticipation of the development
of a nascent UK shale-gas industry. The monitoring involves
ways of managing high volume, highly varied data, generated
by a range of IoT sensor data, including:
        </p>
        <p>Groundwater quality: The sensors installed in the
boreholes provide real-time measurements of water-quality
parameters: water level, temperature, pH, conductivity,
and dissolved gases (O2, CH4, CO2, Rn).</p>
        <p>Seismicity: The monitoring of background seismicity has
involved installation of a network of seismic stations in
the vicinity of the proposed shale gas wells. Real time
seismic data are being collected from the array of stations
to help characterize current levels of seismic activity. The
information captured in near real-time includes: station
code, station name, seismic data (for a single channel),
and timestamp.</p>
        <p>Air composition: The monitoring equipment measures
concentrations of ozone (O3), particulate matter (PM1,
PM2.5, PM4, and PM10), nitrogen oxides (NO, NO2
and NOx), methane (CH4), non-methane hydrocarbons
(NHMCs), hydrogen sulphide (H2S) and carbon dioxide
(CO2) as well as capturing meteorological information.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>7https://github.com/rosafilgueira/EMB datastreaming 8https://www.chameleoncloud.org/</title>
      <p>Lancashire - Groundwater sensors- boreholes from 1 to 5</p>
      <p>Fig. 2: Groundwater sensors from the areas of Lancashire
and Vale of Pickering. In total, we have 9 boreholes (sensors
are attached to boreholes) placed at geographically distributed
locations.</p>
      <p>We have initially focused on Groundwater quality sensor
data. However, very little work has to be done in IoT-Hub to
enable support to other sensors. This is discussed in Section V.
These groundwater sensors are attached to boreholes, which
are called emb1, emb2, : : : emb10. Figure2 shows the
locations of these boreholes. For simplicity, we have selected
emb2, emb3, and emb4 boreholes, but IoT-Hub supports
any number of boreholes and sensors.</p>
      <p>
        IoT-Hub collects in simulated real-time the water-quality
parameters described before, from sensors attached to the
selected boreholes (marked as 2, 3, and 4 in Figure 2).
Since we did not have direct access to these sensors, but
access to yearly compressed files instead, monthly datasets
were downloaded locally [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. These datasets are originally
in comma separated values (CSV) format and sensors provide
a single reading for every hour (every line corresponds to an
hour reading) per each day. In this work, we simulate the
setup where sensor data corresponding to one hour arrives
every 10 seconds to test the ability of IoT-Hub to deal with
high frequency data transmission. To achieve this, a feeder
script was implemented to send POST requests messages (with
a sensor reading) to Falcon web services every 10 seconds.
Then, Falcon was configured to act as a producer, publishing
streams to Apache Kafka by using the emb topic. Apache Kafka
ingests those streams in real-time and makes them available
to a Spark-Streaming application which acts as a consumer.
This application is subscribed to the emb topic, and stores
the data in Elasticsearch (see Listing 1) after performing a
quality check over the values received (e.g. if the data is
within the range specified by the sensor manufacturers). If
Once sensor data becomes available in Elasticsearch,
different applications can be run to analyze it. These applications
can vary from simple scripts to check the insertion of data, to
more complex machine learning analysis, such as the anomaly
detection that will be explained in the following subsection.
      </p>
      <p>Furthermore, IoT-Hub includes Kibana for exploring the
data visually (see Figure 3 as an example).
the values are within the predefined range, all the data is
stored into the corresponding fields and the normal label
is stored in the qc field. Otherwise, the data is stored and
annotated with the anomaly label instead. Note that Apache
Kafka offers a decoupling and buffering of the input streams.
Therefore producer-consumer need not to know about each
other. From multiple sources producers can write data to any
topic in kafka, and several consumers can be subscribed to the
same topic making each of them different processing analysis.
Since kafka persists data to disc, consumers can be shut down
for performing changes and when they are restarted, they will
retrieve all data from the time they were offline.</p>
      <p>Listing 1: Elasticsearch index for storing groundwater values.
i n d e x : e m b t e s t
t y p e : emb
f i e l d s :
s e n s o r i d &gt; t y p e : t e x t , d e s c r i p t i o n : I d o f t h e s e n s o r
d a t e &gt; t y p e : d a t e , d e s c r i p t i o n : Date (UTC)
t i m e &gt;t y p e : d a t e , d e s c r i p t i o n : Time (UTC)
s e c &gt; t y p e : i n t e g e r , d e s c r i p . : M i c r o s i e m e n s p e r c e n t i m e r
ph &gt; t y p e : f l o a t , d e s c r i p t i o n : PH
w a t e r l e v e l &gt; t y p e : f l o a t , d e s c r i p t i o n : Water l e v e l aOD
w at er t em p &gt; t y p e : f l o a t , d e s c r i p t i o n : Water t e m p e r a t u r e
t d g &gt; t y p e : i n t e g e r , d e s c r i p t i o n : T o t a l d i s s o l v e d g a s
qc &gt; t y p e : t e x t , d e s c r i p t i o n : Q u a l i t y c o n t r o l
(
S
m)
u
/
c
S
S
E
)900
/Smc700
CE(C400
200</p>
    </sec>
    <sec id="sec-6">
      <title>2) Minimal or no parameter tuning.</title>
      <p>3) Both globalized and localized AD in evolving time
series.</p>
      <p>The criteria address two main problems in the AD domain.
First, the monitoring of the network for the detection of
anomalies without a-priori knowledge. Any selected method
should be able to perform well both in global and local AD.
Detecting localized anomalies can be utilized as a warning
system to prevent device failures. As a consequence, this
delivers more consistent up-time, improved data quality, and
thereby results in a more robust system. Second, the
monitoring of global trends that change over time and detect any
unusual seasonal patterns. This can help in improving the
understanding about the nature of seasonal patterns or changes
in research environments.</p>
      <p>
        Due to the aforementioned reasons, the algorithm selected
is Twitter’s Seasonal Hybrid Extreme Studentized Deviate
(SH-ESD) AD algorithm [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] that uses robust statistics with a
focus on analyzing long term and short trends in time series.
The underlying algorithm employs time series decomposition
to detect both global and local anomalies. A combination of
piecewise approximation to extract the trend of a time series
and an ESD test for anomalies is performed to accommodate a
more localized AD. What is of particular interest is the ability
to detect both local and global anomalies or seasonal changes
and identify when a different pattern emerges in a continuous
data flow.
      </p>
      <p>The only algorithm parameter that was pre-configured was
the maximum anomaly (set to 0.2), m, that controls the
algorithm’s upper bound of suspected anomalies. The algorithm
ran for each of the different measurements. For brevity, in
Figure 4 only the most interesting AD case is shown. In this
case, the first (left) and the second run (right) of the algorithm
after 15 minutes are shown. The first figure shows normal
deviation and no anomalies are detected. In the second run,
the algorithm picks up the anomalies that show a small spike
before the sensor stops provides readings of value zero.
Fig. 3: Kibana screenshot for filtering the normal values, and
counting the different values of the water level parameter.
845.2
845
)
cm844.7
/
S
u
(C844.5
E
S
844.2
SEC AD</p>
      <sec id="sec-6-1">
        <title>A. Anomaly Detection</title>
        <p>An anomaly detection (AD) algorithm has been
implemented to periodically interrogate the data that is
automatically collected, processed, and stored by IoT-Hub. The AD
algorithm is called every 15 minutes to review all the data
transmitted in that period. The IoT-Hub presents new
challenges in an AD context due to its continuous data streaming.</p>
        <p>Therefore, we have established a three criteria that an
algorithm must fulfill in IoT-Hub:
1) Robustness in seasonal changes (i.e., weather patterns).</p>
        <p>Fig. 4: IoT-Hub: AD results in emb3, Lancashire sensor for
February 2017. First run (left), second run (right). Detected
anomalies are shown as turquoise dots.</p>
        <p>The performance of the method is promising in a real world
scenario for AD. However, as can be seen by the high number
of anomalies there is always the caveat of too many false
positives, which is something to be avoided in real world
scenarios. The parameter selection has also to be evaluated
periodically. Careful consideration of algorithms that fulfill the
above criteria with a focus on precision rather than recall have
to be chosen in order to be able to deliver robust results.</p>
        <p>In the future, we plan to create a warning system (e.g., via
a VRE fronted) that makes use of IoT-Hub for running the
described AD analysis to interpret sensors in the field. This
system will send out alerts to domain scientists subscribed to
these alerts, as well as to those in charge of deploying and
maintaining the sensors in the field.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>V. CONCLUSIONS AND FUTURE WORK</title>
      <p>In this paper, we have presented IoT-Hub, which delivers
specialized data-framework to exploit the IoT in a scalable,
efficient, and robust manner reducing the engineering time
and computational cost. We have demonstrated the
feasibility of the proposed solution by using the Environmental</p>
      <sec id="sec-7-1">
        <title>Baseline Monitoring programme. Data from different sensors</title>
        <p>are collected, preprocessed, and stored in real-time using
different microservices middleware. To test the capacity of
IoT-Hub for running complex data-analytics tasks, we have
implemented an anomaly detection algorithm, which queries
data from Elasticsearch and detects the anomalies of each of
the water-quality parameters. All the middleware that forms
IoT-Hub has been containerized, which enables flexible and
agile development, and deployment in cloud-based
infrastructures.</p>
        <p>The current version of IoT-Hub has been pre-configured
for working with groundwater sensors. To extend this work to
other sensors, it only would require to: (1) create a new Kafka
topic to produce and consume new datasets; (2) modify the
Spark-Streaming application to consume and check the data
from this new topic; and (3) create a new Elasticsearch index.</p>
        <p>One of the main uses of IoT-Hub could be to act as a
backend for VREs or Scientific Gateways for running
dataintensive applications and deploying cloud resources upon
request.</p>
        <p>As future work, we plan to include more middleware in
IoT-Hub, such as Cassandra database (for high performance
operations and handling massive datasets), an RDF repository
(to store and harvest RDF data), SparQL Endpoint (to query
a knowledge base via the SPARQL language), a job
submission system (to submit applications to distributed computing
infrastructures), and Jupyter Notebook (to offer an interactive
computational environment).</p>
        <p>
          Additionally, we also plan to create a warning system that
makes use of IoT-Hub for running the described AD analysis
to interpret sensors in the field, and leverage IoT-Hub
capabilities to process near real-time logs from scientific workflow
executions [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ].
        </p>
        <p>Acknowledgments. This work was carried out when the lead author
was with the British Geological Survey. It was funded under the
Scottish Informatics and Computer Science Alliance with the Postdoctoral
and Early Career Researcher Exchanges fellowship, partially funded
by DOE under Contract DESC0012636, “ Panorama - Predictive
Modeling and Diagnostic Monitoring of Extreme Science
Workflows”, and by “ DARE -Delivering Agile Research Excellence on</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gubbi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Buyya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Marusic</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Palaniswami</surname>
          </string-name>
          , “
          <article-title>Internet of things (iot): A vision, architectural elements, and future directions,” Future Gener</article-title>
          .
          <source>Comput. Syst.</source>
          , vol.
          <volume>29</volume>
          , no.
          <issue>7</issue>
          , pp.
          <fpage>1645</fpage>
          -
          <lpage>1660</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Atkinson</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Parsons</surname>
          </string-name>
          , “
          <article-title>The digital-data challenge,” in The DATA Bonanza - Improving Knowledge Discovery for Science, Engineering</article-title>
          and Business,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Atkinson</surname>
          </string-name>
          et al., Eds. Wiley,
          <year>2013</year>
          , ch. 1, pp.
          <fpage>5</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Candela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Castelli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Pagano</surname>
          </string-name>
          , “
          <article-title>Virtual research environments: An overview and a research agenda,” Data Science Journal</article-title>
          , vol.
          <volume>12</volume>
          , pp.
          <fpage>GRDI75</fpage>
          -
          <lpage>GRDI81</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>[4] “Instrumenting the earth</article-title>
          ,” http://www.bgs.ac.uk/Sensors/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Catlett</surname>
          </string-name>
          , “
          <article-title>Teragrid: A foundation for us cyberinfrastructure</article-title>
          ,”
          <source>in Proceedings of the 2005 IFIP International Conference on Network and Parallel Computing (NPC'05)</source>
          ,
          <year>2005</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>1</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Atkinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Carpene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Casarotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Claus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Filgueira</surname>
          </string-name>
          et al., “
          <article-title>Verce delivers a productive e-science environment for seismology research,” in 2015 IEEE 11th International Conference on e-Science (e-Science)(E-SCIENCE)</article-title>
          , vol.
          <volume>00</volume>
          ,
          <year>2015</year>
          , pp.
          <fpage>224</fpage>
          -
          <lpage>236</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7] L.
          <string-name>
            <surname>de la Garza</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Krger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Schrfe</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Rttig</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Aiche</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Reinert</surname>
            , and
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Kohlbacher</surname>
          </string-name>
          , “
          <article-title>From the desktop to the grid: conversion of knime workflows to guse,”</article-title>
          <source>in Proc. IWSG</source>
          <year>2013</year>
          ,
          <article-title>ser</article-title>
          .
          <source>CEUR-WS</source>
          , vol.
          <volume>993</volume>
          ,
          <year>2013</year>
          , p.
          <fpage>9</fpage>
          . [Online]. Available: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>993</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Kiddle</surname>
          </string-name>
          et al.,
          <article-title>“Cyberska: An on-line collaborative portal for dataintensive radio astronomy</article-title>
          ,” in
          <source>2011 ACM Workshop on Gateway Computing Environments (GCE '11)</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Filgueira</surname>
          </string-name>
          et al.,
          <article-title>“escience gateway stimulating collaboration in rock physics and volcanology</article-title>
          ,” in 2014 IEEE 10th International Conference on e-Science - Volume
          <volume>01</volume>
          ,
          <string-name>
            <surname>ser.</surname>
          </string-name>
          E-SCIENCE '
          <volume>14</volume>
          ,
          <year>2014</year>
          , pp.
          <fpage>187</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>D. De Roure</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Goble</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Stevens</surname>
          </string-name>
          , “
          <article-title>The design and realisation of the myExperiment Virtual Research Environment for social sharing of workflows,” Future Generation Computer Systems</article-title>
          , vol.
          <volume>25</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>561</fpage>
          -
          <lpage>567</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>G.</given-names>
            <surname>Hager</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Wellein</surname>
          </string-name>
          ,
          <article-title>Introduction to High Performance Computing for Scientists and Engineers</article-title>
          , 1st ed. Boca Raton, FL, USA: CRC Press, Inc.,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>I. J.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , E. Deelman,
          <string-name>
            <given-names>D. B.</given-names>
            <surname>Gannon</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Shields</surname>
          </string-name>
          ,
          <article-title>Workflows for e-Science: scientific workflows for grids</article-title>
          . Springer Publishing Company, Incorporated,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          , E. Pacitti,
          <string-name>
            <given-names>P.</given-names>
            <surname>Valduriez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Mattoso</surname>
          </string-name>
          , “
          <article-title>A survey of data-intensive scientific workflow management</article-title>
          ,
          <source>” J. Grid Comput.</source>
          , vol.
          <volume>13</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>457</fpage>
          -
          <lpage>493</lpage>
          , Dec.
          <year>2015</year>
          . [Online]. Available: http://dx.doi.org/10.1007/s10723-015-9329-8
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>C.-W. Tsai</surname>
            ,
            <given-names>C.-F.</given-names>
          </string-name>
          <string-name>
            <surname>Lai</surname>
          </string-name>
          , H.-C.
          <article-title>Chao, and</article-title>
          <string-name>
            <given-names>A. V.</given-names>
            <surname>Vasilakos</surname>
          </string-name>
          , “
          <article-title>Big data analytics: a survey</article-title>
          ,
          <source>” Journal of Big Data</source>
          , vol.
          <volume>2</volume>
          , no.
          <issue>1</issue>
          , p.
          <fpage>21</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Yaqoob</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. A. T.</given-names>
            <surname>Hashem</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. I. A.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Imran</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. V.</given-names>
            <surname>Vasilakos</surname>
          </string-name>
          , “
          <article-title>The role of big data analytics in internet of things,” Comput</article-title>
          . Netw., vol.
          <volume>129</volume>
          , no.
          <issue>P2</issue>
          , pp.
          <fpage>459</fpage>
          -
          <lpage>471</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16] “
          <article-title>Iot big data framework architecture</article-title>
          ,” https://www.gsma.com/iot/ wp-content/uploads/2016/11/CLP.25-
          <fpage>v1</fpage>
          .0.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Gill</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Chana</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Buyya</surname>
          </string-name>
          , “
          <article-title>Iot based agriculture as a cloud and big data service: The beginning of digital india</article-title>
          ,
          <source>” JOEUC</source>
          , vol.
          <volume>29</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>I. Nadareishvili</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mitra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>McLarty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and M.</given-names>
            <surname>Amundsen</surname>
          </string-name>
          , Microservice Architecture: Aligning Principles, Practices, and
          <string-name>
            <surname>Culture</surname>
          </string-name>
          , 1st ed.
          <source>O'Reilly Media</source>
          , Inc.,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19] “
          <article-title>Environmental baseline monitoring in the lancashire</article-title>
          ,” http://www.bgs.ac.uk/research/groundwater/shaleGas/monitoring/ lancsDataSummary.html.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20] “
          <article-title>Environmental baseline monitoring in the vale of pickering</article-title>
          ,” http://www.bgs.ac.uk/research/groundwater/shaleGas/monitoring/ vopDataSummary.html.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>O.</given-names>
            <surname>Vallis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hochenbaum</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Kejariwal</surname>
          </string-name>
          , “
          <article-title>A novel technique for long-term anomaly detection in the cloud</article-title>
          ,
          <source>” in 6th USENIX Conference on Hot Topics in Cloud Computing (HotCloud'14)</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>E.</given-names>
            <surname>Deelman</surname>
          </string-name>
          et al.,
          <string-name>
            <surname>“</surname>
            <given-names>PANORAMA</given-names>
          </string-name>
          :
          <article-title>An approach to performance modeling and diagnosis of extreme scale workflows</article-title>
          ,”
          <source>International Journal of High Performance Computing Applications</source>
          , vol.
          <volume>31</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>4</fpage>
          -
          <lpage>18</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>