<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cloud Data Management: A Short Overview and Comparison of Current Approaches</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Siba Mohammad</string-name>
          <email>siba.mohammad@iti.uni-</email>
          <email>siba.mohammad@iti.unimagdeburg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sebastian Breß</string-name>
          <email>sebastian.bress@st.ovgu.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eike Schallehn</string-name>
          <email>eike@iti.cs.uni-magdeburg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Otto-von-Guericke University</institution>
          ,
          <addr-line>Magdeburg</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <abstract>
        <p>Users / Applications To meet the storage needs of current cloud applications, new data management systems were developed. Design decisions were made by analyzing the applications workloads and technical enviQrounemrye Lnatn.gIutawgeas realized that traditional Relational Database Management Systems (RDBMSs) with their centralized architecture, strong consistency, and relational model do not t the elasticity and scalability requirements of theDcilsoturidb.uted Processing System Di erent architectures with a variety of data partitioning schemes and replica placement strategies were developed. As for the data model, the key-value pairs wiSthtruitcstuvarerida tDiaotnas System were adopted for cloud storage. The contribution of this paper is to provide a comprehensible overview of key-characteristics of current solutions and outline the problems they do and do not address. This paper should sDeirsvtreibausteadn Setnotrraygep Soyinsttefmor orientation of future research regarding new applications in cloud computing and advanced requirements for data management.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Data management used within the cloud or o ered as a
service from the cloud is an important current research eld.
However, there are only few publications that provide a
survey and compare the di erent approaches. The contribution
of this paper is to provide a starting point for researchers and
developers who want to work on cloud data management.
Cloud computing is a new technology that provides resources
as an elastic pool of services in a pay-as-you-go model [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
Whether it is storage space, computational power, or
software, customers can get it over the internet from one of the
cloud service providers. Big players in the market, such as
Google [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], Amazon [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], Yahoo! [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and Hadoop [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], de ned
the assumptions for cloud storage systems based on
analyzing the technical environment and applications workload.
First, a data management system will work on a cluster of
storage nodes where components failure is the normal
situation rather than the exception. Thus, fault tolerance and
recovery must be built in. The system must be portable
      </p>
      <p>
        Users / Applications
across heterogeneous hardware and software platforms. It
will store tera bytes of data, thus parameters of I/O
operations and block sizes must be adapted to large sizes. Based
on these assumptions the requirements for a cloud DBMS
are: elasticity, scalability, fault tolerance and self
manageability [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The rest of the paper is organized as follows.
First, we provide an overview of the architecture and the
family tree of the cloud storage systems. Then, we discuss
how di erent systems deal with the trade-o in the
Consistency, Availability, Partition tolerance (CAP) theorem.
After that, we discuss di erent schemes used for data
partitioning and replication. Then, we provide a list of the cloud
data models.
2.
      </p>
      <p>ARCHITECTURE OVERVIEW</p>
      <p>
        There are two main approaches to provide data
management systems for the cloud. In the rst approach, each
customer gets an own instance of a Database Management
System (DBMS), which runs on virtual machines of the
service provider [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] as illustrated in Figure 1. The DBMS
supports full ACID requirements with the disadvantage of
loosing scalability. If an application requires more
computing or storage resources than the maximum allocated for an
instance, the customer must implement partitioning on the
application level using a di erent database instance for each
partition [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Another solution is on demand assignment of
resources to instances. Amazon RDS is an example of
relational database services, that supports MySQL, Oracle, and
SQL Server.
      </p>
      <p>In the second approach, data management is not provided
as a conventional DBMS on a virtualized platform, but as
a combination of interconnected systems and services that
can be combined according to application needs. Figure 2
illustrates this architecture. The essential part is the
dis</p>
      <sec id="sec-1-1">
        <title>Users / Applications</title>
      </sec>
      <sec id="sec-1-2">
        <title>Query Language</title>
      </sec>
      <sec id="sec-1-3">
        <title>Distributed Processing System</title>
      </sec>
      <sec id="sec-1-4">
        <title>Structured Data System</title>
      </sec>
      <sec id="sec-1-5">
        <title>Distributed Storage System</title>
        <p>tributed storage system. It is usually internally used by the
cloud services provider and not provided as a public service.
It is responsible for providing availability, scalability, fault
tolerance, and performance for data access. Systems in this
layer are divided in three categories:</p>
        <p>Distributed File Systems (DFS), such as Google's File
System (GFS).</p>
        <p>Cloud based le services, such as Amazon's Simple
Storage Service (S3).</p>
        <p>
          Peer to peer le systems, such as Amazon's Dynamo.
The second layer consists of structured data systems and
provides simple data models such as key-value pairs, which
we discuss in Section 6. These systems support various APIs
for data access, such as SOAP and HTTP. Examples of
systems in this layer are Google's Bigtable, Cassandra, and
SimpleDB. The third layer includes distributed processing
systems, which are responsible for more complex data
processing, e.g, analytical processing, mass data
transformation, or DBMS-style operations like joins and aggregations.
MapReduce [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] is the main processing paradigm used in
this layer.
        </p>
        <p>
          The nal layer includes query languages. SQL is not
supported. However, developers try to mimic SQL syntax for
simplicity. Most query languages of cloud data management
systems support access to one domain, key space, or table,
i.e., do not support joins [
          <xref ref-type="bibr" rid="ref2 ref4">4, 2</xref>
          ]. Other functionalities, such
as controlling privileges and user groups, schema creation,
and meta data access are supported. Examples of query
languages for cloud data are HiveQL, JAQL, and CQL. The
previous components complement each other and work
together to provide di erent sets of functionalities. One
important design decision that was made for most cloud data
query languages is not supporting joins or aggregations.
Instead, the MapReduce framework is used to perform these
operations to take advantage of parallel processing on di
erent nodes within a cluster. Google pioneered this by
providing a MapReduce framework that inspired other systems.
        </p>
        <p>For more insight into connections and dependencies
between these systems and components, we provide a family</p>
      </sec>
      <sec id="sec-1-6">
        <title>Users / Applications</title>
        <p>RDS</p>
        <sec id="sec-1-6-1">
          <title>HadoopDB</title>
          <p>Relational Cloud StoragHei vSeervice</p>
        </sec>
        <sec id="sec-1-6-2">
          <title>Cassandra</title>
        </sec>
      </sec>
      <sec id="sec-1-7">
        <title>Relational DBMS</title>
        <sec id="sec-1-7-1">
          <title>HBase</title>
        </sec>
        <sec id="sec-1-7-2">
          <title>Hadoop MR </title>
          <p>HDFS Dynamo
S3</p>
        </sec>
        <sec id="sec-1-7-3">
          <title>Google MR GFS</title>
        </sec>
        <sec id="sec-1-7-4">
          <title>File system</title>
        </sec>
        <sec id="sec-1-7-5">
          <title>SimpleDB</title>
        </sec>
        <sec id="sec-1-7-6">
          <title>Bigtable</title>
          <p>use the system
use concepts </p>
        </sec>
        <sec id="sec-1-7-7">
          <title>DBMS</title>
          <p>tree of cloud storage systems as illustrated in Figure 3. We
use a solid arrow to illustrate that a system uses another one
such as Hive using HDFS. We use a dotted arrow to
illustrate that a system uses some aspects of another system like
the data model or the processing paradigm. An example of
this is Cassandra using the data model of Bigtable. In this
family tree, we cover a range of commercial systems, open
source projects, and academic research as well. We start
on the left side with distributed storage systems GFS and
HDFS. Then, we have the structured storage systems with
API support such as Bigtable. Furthermore, there are
systems that support a simple QL such as SimpleDB. Next, we
have structured storage systems with support of MapReduce
and simple QL such as Cassandra and HBase. Finally, we
have systems with sophisticated QL and MapReduce
support such as Hive and HadoopDB. We compare and classify
these systems based on other criteria in the coming sections.
An overview is presented in Figure 4.
3.</p>
          <p>
            CONSISTENCY, AVAILABILITY,
PARTITION TOLERANCE (CAP) THEOREM
Tightly related to key features of cloud data management
systems are discussions on the CAP theorem [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]. It states
that consistency, availability, and partition tolerance are
systematic requirements for designing and deploying
applications for distributed environments. In the cloud data
management context these requirements are:
          </p>
          <p>Consistency: includes all modi cations on data that
must be visible to all clients once they are committed.
G
F
S</p>
          <p>H
D
F</p>
          <p>S
file</p>
          <p>file
chunk</p>
          <p>chunk
X
X</p>
          <p>X
X</p>
          <p>D
y
n
a
m
o
X
X
Key
space
set of
items</p>
          <p>X
X</p>
          <p>S
i
m
p
l
e
D
B
X
X
X
X
X</p>
          <p>Property
ACID
BASE
SCLA
Tunable
Consistency
Hash
Range
List
Composite
H
i
v
e</p>
          <p>R
D
S
X</p>
          <p>X
item</p>
          <p>X
domain
item
chunk
block</p>
          <p>Availability: means that all operations on data, whether
read or write, must end with a response within a
speci ed time.</p>
          <p>Partition tolerance: means that even in the case of
components' failures, operations on the database must
continue.</p>
          <p>The CAP theorem also states that developers must make
trade-o decisions between the three con icting requirements
to achieve high scalability. For example, if we want a data
storage system that is both strongly consistent and partition
tolerant, the system has to make sure that write operations
return a success message only if data has been committed
to all nodes, which is not always possible because of
network or node failures. This means that its availability will
be sacri ced.</p>
          <p>In the cloud, there are basically four approaches for DBMSs
in dealing with CAP:
Atomicity, Consistency, Isolation, Durability (ACID):
With ACID, users have the same consistent view of
data before and after transactions. A transaction is
atomic, i.e., when one part fails, the whole transaction
fails and the state of data is left unchanged. Once a
transaction is committed, it is protected against crashes
and errors. Data is locked while being modi ed by
a transaction. When another transaction tries to
access locked data, it has to wait until data is unlocked.
Systems that support ACID are used by applications
that require strong consistency and can tolerate its
affects on the scalability of the application as already
discussed in Section 2.</p>
          <p>Basically Available, Soft-state, Eventual consistency
(BASE):</p>
          <p>The system does not guarantee that all users see the
same version of a data item, but guarantees that all
of them get a response from the systems even if it
means getting a stale version. Soft-state refers refers
to the fact, that the current status of a managed
object can be ambiguous, e.g. because there are several
temporarily inconsistent replicas of it stored.
Eventually consistent means that updates will propagate
through all replicas of a data item in a distributed
system, but this takes time. Eventually, all replicas
are updated. BASE is used by applications that can
tolerate weaker consistency to have higher availability.
Examples of systems supporting BASE are SimpleDB
and CouchDB.</p>
          <p>Strongly Consistent, Loosely Available (SCLA): This
approach provides stronger consistency than BASE.
The scalability of systems supporting SCLA in the
cloud is higher compared to those supporting ACID.
It is used by systems that choose higher consistency
and sacri ce availability to a small extent. Examples
of systems supporting SCLA are HBase and Bigtable.
Tunable consistency: In this approach, consistency is
congurable. For each read and write request, the user
decides the level of consistency in balance with the level
of availability. This means that the system can work in
HadoopDB</p>
          <p>RDS
A
BASE
CouchDB
SimpleDB
Dynamo
S3</p>
          <p>HBase</p>
          <p>Hive</p>
          <p>Bigtable</p>
          <p>Tunable 
Consistency P</p>
          <p>Cassandra</p>
          <p>
            PNUTS
high consistency or high availability and other degrees
in between. An example of a system supporting
tunable consistency is Cassandra [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ] where the user
determines the number of replicas that the system should
update/read. Another example is PNUTS [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ], which
provides per record time-line consistency. The user
determines the version number to query at several points
in the consistency time-line. In Figure 5, we classify
di erent cloud data management systems based on the
consistency model they provide.
4.
          </p>
          <p>PARTITIONING TECHNIQUES</p>
          <p>
            Partitioning, also known as sharding, is used by cloud
data management systems to achieve scalability. There is a
variety of partitioning schemes used by di erent systems on
di erent levels. Some systems partition data on the le level
while others horizontally partition the key space or table.
Examples of systems partitioning data on the le level are
the DFSs such as GFS and HDFS which partition each le
into xed sized chunks of data. The second class of systems
which partition tables or key space uses one of the following
partitioning schemes [
            <xref ref-type="bibr" rid="ref18 ref6">18, 6</xref>
            ] :
List Partitioning A partition is assigned a list of discrete
values. If the key of the inserted tuple has one of these
values, the speci ed partition is selected . An example
of a system using list as the partitioning scheme is
Hive [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ].
          </p>
          <p>Range Partitioning The range of values belonging to one
key is divided into intervals. Each partition is assigned
one interval. A partition is selected if the key value of
the inserted tuple is inside a certain range. An example
of a system using range partitioning is HBase.</p>
          <p>Hash Partitioning The output of a hash function is
assigned to di erent partitions. The hash function is
applied on key values to determine the partition. This
scheme is used when data does not lend itself to list
and range partitioning. An example of a system using
hash as the partitioning scheme is PNUTS.</p>
          <p>There are some systems that use a composite partitioning
scheme. An example is Dynamo, which uses a composite of
hash and list schemes (consistent hashing). Some systems
allow partitioning data several times using di erent
partitioning schemes each time. An example is Hive, where each
table is partitioned based on column values. Then, each
partition can be hash partitioned into buckets, which are stored
in HDFS.</p>
          <p>
            One important design consideration to make is whether
to choose an order-preserving partitioning technique or not.
Order preserving partitioning has an advantage of better
performance when it comes to range queries. Examples of
systems using order preserving partitioning techniques are
Bigtable and Cassandra. Since most partitioning methods
depend on random position assignment of storage nodes,
the need for load balancing to avoid non uniform
distribution of data and workloads is raised. Dynamo [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ] focuses on
achieving a uniform distribution of keys among nodes
assuming that the distribution of data access is not very skewed,
whereas Cassandra [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] provides a load balancer that
analyzes load information to lighten the burden on heavily
loaded nodes.
          </p>
          <p>REPLICATION TECHNIQUES</p>
          <p>
            Replication is used by data management systems in the
cloud to achieve high availability. Replication means storing
replicas of data on more than one storage node and probably
more than one data center. The replica placement strategy
a ects the e ciency of the system [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. In the following we
describe the replication strategies used by cloud systems:
Rack Aware Strategy: Also known as the Old Network
Topology Strategy. It places replicas in more than one
data center on di erent racks within each data center.
Data Center Aware Strategy: Also known as the New
Network Topology Strategy. In this strategy, clients
specify in their applications how replicas are placed
across di erent data centers.
          </p>
          <p>Rack Unaware Strategy: Also known as the Simple
Strategy. It places replicas within one data center using a
method that does not con gure replica placement on
certain racks.</p>
          <p>Replication improves system robustness against node
failures. When a node fails, the system can transparently read
data from other replicas. Another gain of replication is
increasing read performance using a load balancer that directs
requests to a data center close to the user. Replication has a
disadvantage when it comes to updating data. The system
has to update all replicas. This leads to very important
design considerations that impact availability and consistency
of data. The rst one is to decide whether to make
replicas available during updates or wait until data is consistent
across all of them. Most systems in the cloud choose
availability over consistency. The second design consideration is
to decide when to perform replica con icts resolution, i.e.,
during writes or reads. If con ict resolution is done during
write operations, writes could be rejected if the system can
not reach all replicas or a speci ed number of them within
a speci c time. Example of that is the WRITE ALL
operation in Cassandra, where the write fails if the system could
not reach all replicas of data. However, some systems in the
cloud choose to be always writeable and push con ict
resolution to read operations. An example of that is Dynamo
which is used by many Amazon services like the shopping
cart service where customer updates should not be rejected.
6.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>DATA MODEL</title>
      <p>Just as di erent requirements compared to conventional
DBMS-based applications led to the previously described
di erent architectures and implementation details, they also
led to di erent data models typically being used in cloud
data management. The main data models used by cloud
systems are:
Key-value pairs It is the most common data model for
cloud storage. It has three subcategories:</p>
      <p>
        Row oriented: Data is organized as containers
of rows that represent objects with di erent
attributes. Access control lists are applied on the
object (row) or container (set of rows) level [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
An example is SimpleDB.
      </p>
      <p>
        Document Oriented: Data is organized as a
collection of self described JSON documents.
Document is the primarily unit of data which is
identi ed by a unique ID. Documents are the unit
for access control [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Example of a cloud data
management system with document oriented data
model is CouchDB.
      </p>
      <p>
        Wide column: In this model, attributes are grouped
together to form a column family. Column family
information can be used for query optimization.
Some systems perform access control and both
disk and memory accounting at the column
family level [
        <xref ref-type="bibr" rid="ref17 ref3 ref8">8, 17, 3</xref>
        ]. An example of that is Bigtable.
Systems of wide column data model should not be
mistaken with column oriented DB systems. The
former deals with data as column families on the
conceptual level only. The latter is more on the
physical level and stores data by column rather
than by row.
      </p>
      <p>Relational Model (RM) The most common data model
for traditional DBMS is less often used in the cloud.
Nevertheless, Amazon's RDS supports this data model
and PNUTS a simpli ed version of it.
7. SUMMARY AND CONCLUSION</p>
      <p>The cloud with its elasticity and pay-as-you-go model is
an attractive choice for outsourcing data management
applications. Cloud service providers, such as Amazon and
Microsoft, provide relational DBMSs instances on virtual
machines. However, the cloud technical environment,
workloads, and elasticity requirements lead to the development
of new breed of storage systems. These systems range from
highly scalable and available distributed storage systems
with simple interfaces for data access to fully equipped DBMSs
that support sophisticated interfaces, data models, and query
languages.</p>
      <p>Cloud data management systems faced with the CAP
theorem trade-o provide di erent levels of consistency ranging
from eventual consistency to strict consistency. Some
systems allow users to determine the level of consistency for
each data input/output request by determining the number
of replicas to work with or the version number of the data
item. List, range, hash, and composite partitioning schemes
are used to partition data to achieve scalability. With
partitioning comes the need for load balancing with two
basic methods: uniform distribution of data and workloads,
and analyzing load information. With data partitioned and
distributed over many nodes, taking into consideration the
possibility of node and network failures, comes the need for
replication to achieve availability. Replication is done on
the partition level using rack aware, data center aware, and
rack unaware placement strategies. Cloud data management
systems support relational data model, and key-value pairs
data model. The key-value pairs is widely used with di
erent variations: document oriented, wide column, and row
oriented.</p>
      <p>As outlined throughout this paper, several typical
properties of traditional DBMS, such as advanced query languages,
transaction processing, and complex data models, are mostly
not supported by cloud data management systems. Some
of them, because they are simply not required for current
cloud applications. Others, because their implementation
would lead to losing some of the important advantages, e.g.,
scalability and availability, of current cloud data
management approaches. On the one hand, providing features of
conventional DBMS appears to lead toward worthwhile
research directions and is currently addressed in ongoing
research. On the other hand, advanced requirements, which
are completely di erent may arise from future cloud
applications, e.g., interactive entertainment, on-line role playing
games, and virtual realities, have interesting characteristics
of continuous, collaborative, and interactive access patterns,
sometimes under real-time constraints. In a similar way,
new ways of human-computer interaction in real-world
environments addressed, for instance, in ubiquitous
computing and augmented reality are often very data-intensive and
sometimes require expensive processing, which could be
supported by cloud paradigms. Nevertheless, neither traditional
DBMS nor cloud data management can currently su ciently
support those applications.</p>
    </sec>
    <sec id="sec-3">
      <title>ACKNOWLEDGMENTS</title>
      <p>We would like to thank the Syrian ministry of higher
education for partially funding this research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Amazon</surname>
            <given-names>RDS</given-names>
          </string-name>
          FAQ. http://aws.amazon.com/rds/faqs/. [online; accessed 10-March-2012].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Cassandra</given-names>
            <surname>Query Language Documentation</surname>
          </string-name>
          . http://caqel.deadcafe.org/cql-doc.
          <source>[online; accessed 25-July-2011].</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>[3] HBase: Bigtable-like structured storage for Hadoop HDFS</article-title>
          . http://wiki.apache.org/hadoop/hbase. [online; accessed 01-March-2012].
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Jaql</given-names>
            <surname>Overview</surname>
          </string-name>
          . http://www.almaden.ibm.com/cs/projects/jaql/.
          <source>[online; accessed 31-August-2011].</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Armbrust</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fox</surname>
          </string-name>
          , R. Gri th,
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Joseph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. H.</given-names>
            <surname>Katz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Konwinski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Patterson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rabkin</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Stoica</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Zaharia</surname>
          </string-name>
          .
          <article-title>Above the clouds: A berkeley view of cloud computing</article-title>
          .
          <source>Technical report</source>
          , EECS Department, University of California, Berkeley,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Avi</surname>
          </string-name>
          <string-name>
            <surname>Silberschatz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Henry F.</given-names>
            <surname>Korth. Database System Concepts Fifth Edition. McGraw-Hill</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Borthakur</surname>
          </string-name>
          .
          <article-title>The Hadoop Distributed File System: Architecture and Design</article-title>
          .
          <source>The Apache Software Foundation</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>F.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghemawat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. C.</given-names>
            <surname>Hsieh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Burrows</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chandra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fikes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Gruber</surname>
          </string-name>
          .
          <article-title>Bigtable: A distributed storage system for structured data</article-title>
          .
          <source>In In Proceedingss of the 7th Conference on Usenix Symposium on Operating Systems Design and Implementation</source>
          - Volume
          <volume>7</volume>
          , pages
          <fpage>205</fpage>
          {
          <fpage>218</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B. F.</given-names>
            <surname>Cooper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ramakrishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Silberstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bohannon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.-A.</given-names>
            <surname>Jacobsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Puz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weaver</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Yerneni</surname>
          </string-name>
          . Pnuts:
          <article-title>Yahoo!'s hosted data serving platform</article-title>
          .
          <source>Proc VLDB Endow</source>
          .,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Curino</surname>
          </string-name>
          , E. Jones,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Popa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Malviya</surname>
          </string-name>
          , E. Wu,
          <string-name>
            <given-names>S.</given-names>
            <surname>Madden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Balakrishnan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Zeldovich</surname>
          </string-name>
          .
          <article-title>Relational Cloud: A Database Service for the Cloud</article-title>
          .
          <source>In 5th Biennial Conference on Innovative Data Systems Research</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Das</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A. E. Abbadi.</surname>
          </string-name>
          <article-title>ElasTraS: An Elastic, Scalable, and Self Managing Transactional Database for the Cloud</article-title>
          .
          <source>Technical report</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghemawat</surname>
          </string-name>
          .
          <source>Mapreduce: simpli ed data processing on large clusters. Commun. ACM</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G.</given-names>
            <surname>DeCandia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hastorun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jampani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Kakulapati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lakshman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pilchin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sivasubramanian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vosshall</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Vogels</surname>
          </string-name>
          . Dynamo:
          <article-title>Amazon's highly available key-value store</article-title>
          .
          <source>SIGOPS Oper. Syst. Rev.</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fox</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. D.</given-names>
            <surname>Gribble</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chawathe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Brewer</surname>
          </string-name>
          , and P. Gauthier.
          <article-title>Cluster-based scalable network services</article-title>
          .
          <source>SIGOPS Oper. Syst. Rev.</source>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hewitt. Cassandra The De nitive Guide</surname>
          </string-name>
          .
          <string-name>
            <given-names>O</given-names>
            <surname>Reilly Media</surname>
          </string-name>
          , Inc,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>J. L. J. Chris Anderson</surname>
            and
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Slater. CouchDB The De nitive Guide</surname>
          </string-name>
          .
          <source>OReilly Media</source>
          , Inc,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lakshman</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Malik</surname>
          </string-name>
          .
          <article-title>Cassandra: a decentralized structured storage system</article-title>
          .
          <source>SIGOPS Oper. Syst. Rev.</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S. B. e. a. Lance</given-names>
            <surname>Ashdown</surname>
          </string-name>
          , Cathy Baird.
          <source>Oracle9i Database Concepts</source>
          .
          <source>Oracle Corporation</source>
          .,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>R. H. Prabhakar</given-names>
            <surname>Chaganti. Amazon SimpleDB Developer Guide</surname>
          </string-name>
          <article-title>Scale your application's database on the cloud using Amazon SimpleDB</article-title>
          . Packt Publishing Ltd,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Thusoo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Sarma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Shao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Chakka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Anthony</surname>
          </string-name>
          , H. Liu,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wycko</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Murthy</surname>
          </string-name>
          .
          <article-title>Hive: a warehousing solution over a map-reduce framework</article-title>
          .
          <source>Proc. VLDB Endow</source>
          .,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>