<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>© Dmitry Namiot</string-name>
          <email>dnamiot@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lomonosov Moscow State University</institution>
          ,
          <addr-line>Moscow</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Proceedings of the XVII International Conference «Data Analytics and Management in Data Intensive Domains» (DAMDID/RCDL'2015)</institution>
          ,
          <addr-line>Obninsk</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>132</fpage>
      <lpage>137</lpage>
      <abstract>
        <p>Data persistence for time series is an old and in many cases traditional task for databases. In general, the time series is just a sequence of data elements. The typical use case is a set of measurements made over a time interval. Much of the data generated by sensors, in a machine to machine communication, in Internet of Things area could be collected as time series. Time series are used in statistics, mathematical and finance. In this paper, we provide a survey of data persistence solutions for time series data. The paper covers the traditional relational databases, as well as NoSQL-based solutions for time series data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>According to the classical definition, a time series is
simply a sequence of numbers collected at regular
intervals over a period of time. More generally, a time
series is a sequence of data points (not necessarily
numbers). And typically, time series consisting of
successive measurements made over a time interval.</p>
      <p>So, time series exist in any domain of applied
science and engineering which involves temporal
measurements. By this reason, data persistence
mechanisms for time series are among oldest tasks for
databases.</p>
      <p>Let us start from the relational databases. At the first
hand, the table design looks simple. We can create a
table with a timestamp as a key column. Each new
measurement will simply add a new row. And columns
will describe our measurements (attributes). For the
different time series we can add series ID column too:</p>
      <sec id="sec-1-1">
        <title>CREATE TABLE TS AS ( ts_time TIMESTAMP ts_id INT,</title>
      </sec>
      <sec id="sec-1-2">
        <title>NOT NULL PRIMARY KEY,</title>
        <p>ts_value FLOAT</p>
        <p>In the real machine to machine (M2M) or Internet of
Things (IoT) application, we will have more than one
sensor. So, more likely, our application should support
many time series simultaneously. Of course, in any
practical system the whole set of attributes is limited.
But for the each individual measurement we could have
(potentially) any subset from this limited list. It is the
typical use case for Machine to Machine (M2M) or
Internet of Things (IoT) applications. Devices in our
system will provide data asynchronously (it is the most
practical use case). So, each row in our table will have
many empty (null-value) columns. This decision (one
row per measurement) leads to the very inefficient use
of disk space. Also, it complicates the future processing.</p>
        <p>
          Now let us discuss the possible operations. For time
series (for measurements) the main operation in data
space is adding new data (INSERT statement in SQL).
Updating or deleting data is completely uncommon for
time series (measurements). And reading has got some
special moments too. Obviously, the reading of data is
closely related to processing methods. The challenge in
a database of evolving time series is to provide efficient
algorithms and access methods for query processing,
taking into consideration the fact that the database
changes continuously as new data become available [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>
          Many (most of) algorithms for time series data
mining actually always work with only part of the data.
And for the streamed data the sliding window is the
most natural choice. Let us discuss some well-known
techniques. Random Sampling lets us sampling the
stream at periodic intervals. In this approach, we
maintain a sample called the “reservoir,” from which a
random sample can be generated. As the data stream
flows, every new element has a certain probability of
replacing an old element in the reservoir [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Instead of
sampling the data stream randomly, we can use the
sliding window model to analyze stream data [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The
basic idea is that rather than running computations on
all of the data seen so far, or on some sample (as the
above-mentioned random sampling), we can make
decisions based only on recent data. So, any element for
analysis, arrived at some time t will be declared as
expired at time t+w, where w is the window “size”. In
many practical use cases, we can assume (e.g. sensing in
IoT applications) that the only recent events may be
important.
        </p>
        <p>A histogram approach partitions the data into a set
of contiguous buckets.</p>
        <p>In a titled frame model, we use different
granularities for time frames. The most recent time is
registered at the finest granularity; the most distant time
is registered at a coarser granularity. And so on.</p>
        <p>
          We present this short survey for highlighting the fact
that time series data mining algorithms actually almost
always work with only part of the data. It is a common
use case for time series processing. In terms of SQL,
SELECT statement with some complex condition is
uncommon for time series processing. What we need to
read for most of the algorithms is some limited portion
(window) of recent constantly updated (added) data. In
the same time, we need the full log of data (e.g., for the
verification, audit, billing, etc.). Probably, it is the
perfect example of so-called lambda architecture [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. It
is illustrated in Figure 1.
        </p>
      </sec>
      <sec id="sec-1-3">
        <title>Fig, 1. Lambda architecture [5]</title>
        <p>On this picture we have data source with constantly
updated data. Most of the data should be processed in
the real time. It is especially true for Internet of Things
and M2M systems, where time series processing is the
main source behind control actions and conclusions
(alerts). In the same time, we could still have some
processing without the strong limitations for
time-todecision. And database could save processed data for
queries from users and applications. It would be
convenient to have such a processing as a part of a
database system.</p>
        <p>The rest of the paper is organized as follows. The
section 2 is devoted to time series support in relational
databases. In section 3 we describe NoSQL solutions for
time series.
2 Time Series and relational databases</p>
        <p>
          In this section, we would like to discuss time series
persistence (and processing, of course) and “traditional”
databases. As the first example, we have tested
TokuDB as an engine for time series data [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. This
engine uses a Fractal Tree index (instead of more
known B-tree). A Fractal Tree index is a tree data
structure. Any node can have more than two sub-nodes
(children) [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Like a B-tree, it that keeps data always
sorted and allows fast searches and sequential access.
But unlike a B-tree, a Fractal Tree index has buffers at
each node. Buffers allow changes to be stored in
intermediate locations. Buffers lets schedule disk writes
so that each writing operations deals with a large block
of data [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. The simple database related analogue is a
transactions monitor. This optimization lets perform
fast data writing (INSERT operations we are mostly
interested for time series). Also, these “local” buffers
can be used during replications. Figure 2 illustrates the
benchmark for INSERT operations (provided by
TokuDB).
        </p>
        <p>.</p>
        <p>Of course, the “traditional” relation systems are easy
to maintain, they could be cheaper to host and it could
be easier (cheaper) to find developers. So, the
maintenance is the biggest advantage.</p>
        <p>
          As the second position in this section we would like
to mention time series-related SQL extension in Vertica
database [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Vertica provides so-called time series
analytics as an extension for SQL. As we have
mentioned above, input records for time series data
usually appear at non-uniform intervals. It means they
might have gaps (missed values). Vertica provides
socalled gap-filling functionality. This option fills in
missing data points using an interpolation scheme.
Secondly, Vertica allows developers to use event-based
windows to break time series data into windows that
border on significant events within the data. SQL
extension proposed by Vertica could be used as a useful
example of language-based support for time series in
SQL databases. Here is a typical example:
        </p>
        <p>SELECT item, slice_time, ts_first_value(price,
'const') price FROM ts_test WHERE price_time
BETWEEN timestamp '2015-04-14 09:00' AND
timestamp '2015-04-14 09:25' TIMESERIES slice_time
AS '1 minute' OVER (PARTITION BY item ORDER
BY price_time) ORDER BY item, slice_time, price;</p>
        <p>This request should fill missed data from 09:00 till
09:25 with 1 minute step in the returned snapshot.</p>
        <p>Vertica provides additional support for time series
analytics with the following SQL extensions:</p>
        <p>The SELECT … TIMESERIES clause supports
gap-filling and interpolation computation.</p>
        <p>TS_FIRST_VALUE and TS_LAST_VALUE are
time series aggregate functions that return the value at
the start or end of a time slice, respectively, which is
determined by the interpolation scheme.</p>
        <p>TIME_SLICE is a (SQL extension) date/time
function that aggregates data by different fixed-time
intervals and returns a rounded-up input TIMESTAMP
value to a value that corresponds to the start or end of
the time slice interval.</p>
        <p>
          The similar examples are so-called
windowfunctions in PostgreSQL [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. A window function
performs a calculation across a set of table rows that are
somehow related to the current row. This is comparable
to the type of calculation that can be done with an
aggregate function. But unlike regular aggregate
functions, use of a window function does not cause rows
to become grouped into a single output row — the rows
retain their separate identities. Behind the scenes, the
window function is able to access more than just the
current row of the query result. The typical example (a
moving average for three rows) is illustrated below:
        </p>
      </sec>
      <sec id="sec-1-4">
        <title>SELECT id_sensor, name_sensor, temperature, avg(temperature) OVER (ORDER BY id_sensor ROWS BETWEEN 1 PRECEDING AND 1 FOLLOWING)</title>
        <p>FROM temperature_table;</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3 Time Series in NoSQL systems</title>
      <p>NoSQL as one of the basic principles proclaimed
rejection of a universal model of data. The data model
must meet the required processing methods. The second
basic principle is the lack of the dedicated programming
access tools (layers). Data access API is a part of
NoSQL systems, and it presents one of the important
elements for the final selection of a data model.</p>
      <p>
        By our opinion, NoSQL solutions for time series
could be described as “best practices” in using NoSQL
stores for time series. Let us see, for example, the
architecture for OpenTSDB [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. It is one of the popular
NoSQL solutions for time series.
      </p>
      <p>
        OpenTSDB is a set of so-called Time Series
Daemon (TSD) and command line utilities. Each TSD
(and they are independent) is just a wrapper for access
to HBase database [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Each TSD is independent. Each
TSD uses the HBase to store and retrieve time series
data. TSD itself supports a set of protocols for access to
data.
      </p>
      <p>The second dimension of this architecture solution is
the optimized schema for data. In this case, the schema
is highly optimized for fast aggregations of similar time
series. This schema is actually almost a de-facto
standard for presenting time-series in a so-called
BigTable data model. In OpenTSDB, a time series data
point consists of a metric name, a timestamp, a value
and a set of tags (key-value pairs). So, for example,
suppose we have a metric (name) data.test, a key name
is host, a key value is host1. In this case, each row in
master table looks so:
data.test Time Value host host1
here Time is a timestamp, Value is a measured
value. And all APIs will use the similar format for data
writing – there is no schema definition. Set of keys lets
present so-called multivariate time series.</p>
      <p>OpenTSDB handles things a bit differently by
introducing the idea of 'tags'. Each time series still has a
'metric' name, but it's much more generic, something
that can be shared by many unique time series. Instead,
the uniqueness comes from a combination of tag
key/value pairs that allows for flexible queries with very
fast aggregations. Every time series in OpenTSDB must
have at least one tag. The underlying data schema will
store all of the tag’s time series next to each other so
that aggregating the individual values is very fast and
efficient. OpenTSDB was designed to make these
aggregate queries as fast as possible.</p>
      <p>
        OpenTSDB follows to one of the commonly used
patterns for time series data persistence in
columnoriented databases like HBase. The basic data storage
unit in HBase is a cell. Each cell is identified by the
Row ID, column-family name, column name and the
version. Each cell can have multiple versions of data. At
the physical level, each column family is stored
continuously on disk and the data are physically sorted
by Row ID, column name and version [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The version
dimension is used by HBase for time-to-live (TTL)
calculations. Column families may be associated with a
TTL value (length). So, HBase will automatically delete
rows once the expiration time is reached. For time series
data, this feature lets automatically delete old (obsolete)
measurements, for example. The possible schemes for
time series data are:
      </p>
      <p>a) The row key is constructed as a combination of a
timestamp and sensor ID. Each column is the offset of
the time for the timestamp in the row. E.g., the
timestamp is one hour, the column offset is two
minutes. Each cell contains the values for all sensor’s
(defined by Sensor ID) measurements at the moment
timestamp + offset. As the format for cell’s data, we can
use JSON or even commas separated values. As a
variation, we can combine all data in a row into a binary
object (blob).</p>
      <p>b) The row key is constructed as a combination of a
timestamp and sensor ID. Each column corresponds to
one measurement (metric) and contains values for all
time offsets.</p>
      <p>
        KairosDB [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] is a rewrite of the original
OpenTSDB and uses Cassandra as a data store.
      </p>
      <p>There are several patterns for storing time series data
in Cassandra. When writing data to Cassandra, data is
sorted and written sequentially to disk. When retrieving
data by row key and then by range, you get a fast and
efficient access pattern due to minimal disk seeks.</p>
      <p>The simplest model for storing time series data is
creating a wide row of data for each measurement. E.g.:</p>
      <p>SensorID, {timestamp1, value1}, {timestamp2,
value2} … {timestampN, valueN}</p>
      <p>Cassandra can store up to 2 billion columns per row.
For high frequency measured data, we can add a shard
interval to a row key. The solution is to use a pattern
called row partitioning by adding data to the row key to
limit the amount of columns you get per device. E.g.,
instead of some generic name like smart_meter1 we can
use smart_meter1_day (e.g. smart_meter_20150414).</p>
      <p>
        Another common pattern for time series data is
socalled rolling storage. Imagine we are using this data for
a dashboard application, and we only want to show the
last 10 temperature readings. Older data is no longer
useful, so they can be purged eventually. With many
other databases, we would have to setup a background
job to clean out older data. With Cassandra, we can take
advantage of a feature called expiring columns to have
our data quietly disappear after a set amount of seconds
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        What is really interesting and new for NoSQL
solutions is the growing support for SenML [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
SenML is defined by a data model for measurements
and simple meta-data about measurements and devices.
The data in SenML is structured as a single object with
attributes. The object contains an array of entries
(measurements). Each entry is an object that has
attributes such as a unique identifier for the sensor, the
time the measurement was made, and the current value
(Figure 4). Serializations for this data model are defined
for JSON and XML.
      </p>
      <p>
        Geras DB [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] uses SenML as data format. Another
important feature for any time series database is MQTT
support. MQTT is a popular connectivity protocol for
Machine-to-Machine and Internet of Things
communications [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. It was designed as an extremely
lightweight publish/subscribe messaging transport.
Sensors as data sources may use MQTT, so a time series
database should be able to acquire data right from
MQTT [
        <xref ref-type="bibr" rid="ref21 ref22">21, 22</xref>
        ].
      </p>
      <p>
        Druid is an open-source analytics data store
designed for OLAP queries on time series data (trillions
of events, petabytes of data). Druid provides
costeffective and always-on real-time data ingestion,
arbitrary data exploration, and fast data aggregation
[
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Druid is a system built to allow fast ("real-time")
access to large sets of seldom-changing data. It
provides:
      </p>
      <p>the column-based storage format for partially nested
data structures;</p>
      <p>the hierarchical query distribution with intermediate
pruning;
indexing for quick filtering;
realtime ingestion (ingested data is immediately
available for querying);</p>
      <p>the fault-tolerant distributed architecture that doesn’t
lose data.</p>
      <p>Data is ingested by Druid directly through its
realtime nodes, or batch-loaded into historical nodes from a
deep storage facility. Real-time nodes accept
JSONformatted data from a streaming datasource.
Batchloaded data formats can be JSON, CSV, or TSV.
Realtime nodes temporarily store and serve data in real time,
but eventually push the data to the deep storage facility,
from which it is loaded into historical nodes. Historical
nodes hold the bulk of data in the cluster.</p>
      <p>Real-time nodes chunk data into segments, and they
are designed to frequently move these segments out to
deep storage. To maintain cluster awareness of the
location of data, these nodes must interact with Mysql to
update metadata about the segments, and with Apache
ZooKeeper to monitor their transfer.</p>
      <p>Figure 5 illustrates Druid architecture.</p>
      <p>
        SciDB’s [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] native multi-dimensional array data
model is designed for ordered, highly dimensional,
multifaceted data. SciDB’s data is never overwritten,
allowing you to record and access data corrections and
updates over time. SciDB is designed to efficiently
handle both dense and sparse arrays providing dramatic
storage efficiencies as the number of dimensions and
attributes grows. Math operations run directly on the
native data format. Partitioning data in each coordinate
of an array facilitates fast joins and access along any
dimension, thereby speeding up clustering, array
operations and population selection.
      </p>
      <p>
        BlinkDB [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] supports a slightly constrained set of
SQL-style declarative queries and provides approximate
results for standard SQL aggregate queries, specifically
queries involving COUNT, AVG, SUM and
PERCENTILE and is being extended to support any
User-Defined Functions (UDFs). Queries involving
these operations can be annotated with either an error
bound or a time constraint, based on which the system
selects an appropriate sample to operate on. For
example:
      </p>
      <p>SELECT avg(Temperature) from Table where
SensorID=1 WITHIN 2 seconds</p>
      <p>
        SAP HANA’s [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] column-oriented in-memory
structures have been extended to provide efficient
processing of series data. SAP HANA provides:
series property aspect of tables;
built-in Special SQL functions for working with
series data;
      </p>
      <p>analytic functions: special SQL functions for
analyzing series data;</p>
      <p>storage support: advanced techniques for storing
equidistant data using dictionary encoding</p>
      <p>By adding Series Data descriptors to column tables,
users can identify which columns contain series data,
period information, hints on how to handle missing
timestamps, and so on. By explicitly telling HANA
about time series data, it can more efficiently store and
manage this data to increase performance and decrease
the memory footprint through improved compression.</p>
      <p>TABLESAMPLE allows ad-hoc random samples
over column tables so it is easy, for example, to
calculate a result from a defined percentage of the data
in a table.</p>
      <p>There are examples for built-in functions:
SERIES_GENERATE – generate a complete series
SERIES_DISAGGREGATE – move from coarse
units (e.g., day) to finer (e.g., hour)</p>
      <p>SERIES_ROUND – convert a single value to a
coarser resolution</p>
      <p>SERIES_PERIOD_TO_ELEMENT – convert a
timestamp in a series to its offset from the start</p>
      <p>SERIES_ELEMENT_TO_PERIOD – convert an
integer to the associated period .</p>
      <p>Analytical functions are:</p>
      <p>CORR – Pearson product-moment correlation
coefficient</p>
      <p>CORR_SPEARMAN - Spearman rank correlation
LINEAR_APPROX - Replace NULL values by
interpolating adjacent non-NULL values</p>
      <p>MEDIAN - Compute median value</p>
      <p>
        InfluxDB [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] is open-source, distributed, time
series database with no external dependencies.
InfluxDB is targeted at use cases for DevOps, metrics,
sensor data, and real-time analytics. The key moments
behind InfluxDB are:
      </p>
      <sec id="sec-2-1">
        <title>SQL like query language.</title>
        <p>HTTP based API.</p>
        <p>Database managed retention policies for data.
Built-in management interface.</p>
        <p>On the fly aggregation.</p>
        <p>SQL-like query with aggregation by time looks so:
SELACT mean (value) FROM T GROUP BY
time(5m).</p>
        <p>The lack of external dependencies makes Influx very
attractive from the practical point of view. The opposite
approach is the above-mentioned Druid with almost full
of Apache stack (Zookeeper, etc.)</p>
        <p>
          From cloud-based solutions for time series data, we
can mention Blueflood [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]. Blueflood uses Cassandra
as the data store.
        </p>
        <p>As per the classical definition, Big Data could be
described via so-called 3V: Variety, Velocity and
Volume. By our opinion, for time series databases the
key factor is Velocity. NoSQL solution for time series
should be selected in case of high frequency
measurements. And in case of NoSQL solutions for
time series, Cassandra is the preferred choice.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Kontaki</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papadopoulos</surname>
            ,
            <given-names>A. N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Manolopoulos</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Adaptive similarity search in streaming time series with sliding windows</article-title>
          .
          <source>Data &amp; Knowledge Engineering</source>
          ,
          <volume>63</volume>
          (
          <issue>2</issue>
          ),
          <fpage>478</fpage>
          -
          <lpage>502</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Chatfield</surname>
            <given-names>C.</given-names>
          </string-name>
          <article-title>The analysis of time series: an introduction</article-title>
          . - CRC press,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Han</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kamber</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pei</surname>
            <given-names>J</given-names>
          </string-name>
          .
          <article-title>Data mining: Concepts and techniques</article-title>
          . - Morgan Kaufmann,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Fan</surname>
            , Wei,
            <given-names>and Albert</given-names>
          </string-name>
          <string-name>
            <surname>Bifet</surname>
          </string-name>
          .
          <article-title>"Mining big data: current status, and forecast to the future."</article-title>
          <source>ACM SIGKDD Explorations Newsletter 14.2</source>
          (
          <year>2013</year>
          ):
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Lambda</given-names>
            <surname>Architecture: Design Simpler</surname>
          </string-name>
          , Resilient, Maintainable and Scalable Big Data Solutions http://www.infoq.com/articles/lambdaarchitecture-scalable
          <article-title>-big-data-solutions</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Bartholomew</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <source>MariaDB Cookbook</source>
          . Packt Publishing Ltd.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gibbons</surname>
            ,
            <given-names>P. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mowry</surname>
            ,
            <given-names>T. C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Valentin</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2002</year>
          , June). Fractal prefetching B+
          <article-title>- trees: Optimizing both cache and disk performance</article-title>
          .
          <source>In Proceeding of the 2002 ACM SIGMOD international conference on Management of data</source>
          (pp.
          <fpage>157</fpage>
          -
          <lpage>168</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Bender</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Farach-Colton</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Fineman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Fogel,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Kuszmaul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ;
            <surname>Nelson</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          (
          <year>June 2007</year>
          ).
          <article-title>"Cache-Oblivoius streaming B-trees"</article-title>
          .
          <source>Proceedings of the 19th Annual ACM Symposium on Parallelism in Algorithms</source>
          and Architectures (CA: ACM Press):
          <fpage>81</fpage>
          -
          <lpage>92</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>[9] TOKUDB® VS. INNODB FLASH MEMORY http://www.tokutek.com/tokudb-formysql/benchmarks-vs-innodb-flash/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Lamb</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuller</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varadarajan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vandiver</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doshi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Bear</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>The vertica analytic database: C-store 7 years later</article-title>
          .
          <source>Proceedings of the VLDB Endowment</source>
          ,
          <volume>5</volume>
          (
          <issue>12</issue>
          ),
          <fpage>1790</fpage>
          -
          <lpage>1801</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Obe</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsu</surname>
            <given-names>L. S.</given-names>
          </string-name>
          <article-title>PostgreSQL: up and running</article-title>
          . - "
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          ,
          <source>Inc."</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Wlodarczyk</surname>
            ,
            <given-names>T. W.</given-names>
          </string-name>
          (
          <year>2012</year>
          , December).
          <article-title>Overview of time series storage and processing in a cloud environment</article-title>
          .
          <source>In Proceedings of the 2012 IEEE 4th International Conference on Cloud Computing Technology and Science (CloudCom)</source>
          (pp.
          <fpage>625</fpage>
          -
          <lpage>628</lpage>
          ). IEEE Computer Society.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>[13] OpenTSDB http://opentsdb.net</mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>George</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>HBase: the definitive guide. "</article-title>
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          ,
          <source>Inc.".</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Han</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Stroulia</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2012</year>
          ,
          <article-title>September). A three-dimensional data model in hbase for large time-series dataset analysis</article-title>
          .
          <source>In Maintenance and Evolution of Service-Oriented and Cloud-Based Systems (MESOCA)</source>
          ,
          <year>2012</year>
          IEEE 6th International Workshop on the (pp.
          <fpage>47</fpage>
          -
          <lpage>56</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Goldschmidt</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jansen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koziolek</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doppelhamer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Breivold</surname>
            ,
            <given-names>H. P.</given-names>
          </string-name>
          (
          <year>2014</year>
          , June).
          <article-title>Scalability and Robustness of Time-Series Databases for Cloud-Native Monitoring of Industrial Processes</article-title>
          .
          <source>In Cloud Computing (CLOUD)</source>
          ,
          <year>2014</year>
          IEEE 7th International Conference on (pp.
          <fpage>602</fpage>
          -
          <lpage>609</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17] Planet Cassandra http://planetcassandra.org
          <article-title>/getting-started-withtime-series-data-modeling/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Jennings</surname>
            , Cullen,
            <given-names>Jari</given-names>
          </string-name>
          <string-name>
            <surname>Arkko</surname>
            , and
            <given-names>Zach</given-names>
          </string-name>
          <string-name>
            <surname>Shelby</surname>
          </string-name>
          .
          <article-title>"Media types for sensor markup language (SENML)." (</article-title>
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Geras</surname>
            <given-names>DB</given-names>
          </string-name>
          http://1248.io/geras.php Retrieved: Feb,
          <year>2015</year>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Hunkeler</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Truong</surname>
            ,
            <given-names>H. L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Stanford-Clark</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2008</year>
          ,
          <article-title>January)</article-title>
          .
          <article-title>MQTT-S-A publish/subscribe protocol for Wireless Sensor Networks</article-title>
          .
          <source>In Communication systems software and middleware and workshops</source>
          ,
          <year>2008</year>
          . comsware
          <year>2008</year>
          . 3rd international conference on (pp.
          <fpage>791</fpage>
          -
          <lpage>798</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Namiot</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sneps-Sneppe M. On</surname>
          </string-name>
          IoT Programming //International Journal of Open Information Technologies.
          <article-title>-</article-title>
          <year>2014</year>
          . - . 2. - .
          <volume>10</volume>
          . - p.
          <fpage>25</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Sneps-Sneppe</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Namiot</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2012</year>
          , April).
          <article-title>About M2M standards and their possible extensions</article-title>
          .
          <source>In Future Internet Communications (BCFIC)</source>
          ,
          <year>2012</year>
          2nd Baltic Congress on (pp.
          <fpage>187</fpage>
          -
          <lpage>193</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>[23] Druid http://druid.io</mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>[24] Druid Whitepaper http://static.druid.io/docs/druid.pdf</mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Stonebraker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brown</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poliakov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Raman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2011</year>
          , January).
          <source>The architecture of SciDB. In Scientific and Statistical Database Management</source>
          (pp.
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          ). Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mozafari</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Panda</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milner</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madden</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Stoica</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2013</year>
          , April).
          <article-title>BlinkDB: queries with bounded errors and bounded response times on very large data</article-title>
          .
          <source>In Proceedings of the 8th ACM European Conference on Computer Systems</source>
          (pp.
          <fpage>29</fpage>
          -
          <lpage>42</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Färber</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>May</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehner</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Große</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Müller</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rauhe</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Dees</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>The SAP HANA Database--An Architecture Overview</article-title>
          .
          <article-title>IEEE Data Eng</article-title>
          . Bull.,
          <volume>35</volume>
          (
          <issue>1</issue>
          ),
          <fpage>28</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Leighton</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cox</surname>
            ,
            <given-names>S. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Car</surname>
            ,
            <given-names>N. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stenson</surname>
            ,
            <given-names>M. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vleeshouwer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Hodge</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>A Best of Both Worlds Approach to Complex, Efficient, Time Series Data Delivery</article-title>
          .
          <source>In Environmental Software Systems. Infrastructures, Services and Applications</source>
          (pp.
          <fpage>371</fpage>
          -
          <lpage>379</lpage>
          ). Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Blueflood</surname>
          </string-name>
          :
          <article-title>A new Open Source Tool for Time Series Data at Scale https://developer</article-title>
          .rackspace.com/blog/bluefloodannouncement/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>