<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Big Data Analytics Technologies and Platforms: a brief review</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ticiana L. Coelho da Silva</string-name>
          <email>ticianalc@ufc.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Regis P. Magalha˜es</string-name>
          <email>regismagalhaes@ufc.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Igo R. Brilhante</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jose´ Antonio F. de Maceˆdo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Arau´ jo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paulo A. L. Rego</string-name>
          <email>pauloalr@ufc.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aloisio Vieira Lira Neto</string-name>
          <email>aloisio.lira@prf.gov.br</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brazilian Federal Highway Police</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Federal University of Ceara ́</institution>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>25</fpage>
      <lpage>32</lpage>
      <abstract>
        <p>A plethora of Big Data Analytics technologies and platforms have been proposed in the last years. However, in 2017, only 53% of companies are adopting such tools. It seems that the industry is not convinced about Big Data promises or maybe choosing the right technology/platform requires indepth knowledge about the capabilities of all these tools. Before deciding the right technology or platform to choose from, the organizations have to investigate the application/algorithm needs and the advantages and drawbacks of each technology/platform. In this paper, we aim at helping organizations in the selection of technologies/platforms more appropriate to their analytic processes by offering a short-review according to some categories of Big Data problems as processing (streaming and batch), storage, data integration, analytics, data governance, and monitoring.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        <xref ref-type="bibr" rid="ref7">Herodotou et al. 2011</xref>
        ,
        <xref ref-type="bibr" rid="ref2">Begoli and Horey 2012</xref>
        ,
        <xref ref-type="bibr" rid="ref11">Singh and Reddy 2015</xref>
        ], together with
the authors’ development experience with Big Data problems, advocate that all Big Data
technologies should require fault tolerance, scalability, elasticity, distributed architecture,
generic storage and processing of large volumes of data in the order of terabytes or even
petabytes. Besides, a Big Data platform with an ecosystem of services and technologies
should also provide resource management, data governance, and monitoring. In this work,
we only refer to the technologies and platforms regarding such features.
      </p>
      <p>
        Notably, we aim at comparing different Big Data technologies and analytics
platforms according to the following categories: processing (streaming and batch),
storage, data integration, analytics, governance, and monitoring. There exist several
papers that compare big data technologies, to name a few [
        <xref ref-type="bibr" rid="ref8">Inoubli et al. 2018</xref>
        ,
        <xref ref-type="bibr" rid="ref10">Sakr 2016</xref>
        ,
        <xref ref-type="bibr" rid="ref11">Singh and Reddy 2015</xref>
        ]. However, they do not address big data platform analytics. On
the other hand, this study aims to help organizations in the selection of platforms more
suitable to their analytic processes. Since, typically, before deciding on the right
technology or platform to choose from, the user/organization investigates what the
application/algorithm needs are and what each technology/platform may provide. It is worth
to mention that our focus is not comparing Big Data technologies and platforms for
different applications, like Cloud Computing and Internet of Things, but to compare them
according to categories of Big Data problems.
      </p>
      <p>The remainder of the paper is structured as follows: Section 2 and 3 provide an
overview of the relevant Big Data technologies and platforms, respectively, from the
stateof-the-art works. Moreover, these sections present a comparison of such technologies and
platforms based on some categories of problems. Finally, Section 4 draws final
considerations and research challenges.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Big Data Technologies</title>
      <p>
        A plethora of Big Data technologies have been proposed [
        <xref ref-type="bibr" rid="ref1">Alexandrov et al. 2014</xref>
        ,
        <xref ref-type="bibr" rid="ref15">White 2012</xref>
        ,
        <xref ref-type="bibr" rid="ref3">Borkar et al. 2011</xref>
        ,
        <xref ref-type="bibr" rid="ref17">Zaharia et al. 2010</xref>
        ]. In this section, we briefly describe
some of those technologies and provide a comparison between them according to some
categories of problems.
      </p>
    </sec>
    <sec id="sec-3">
      <title>2.1. Overview of Big Data Technologies</title>
      <p>
        This section presents an overview of the most widely used and recently discussed Big
Data technologies [
        <xref ref-type="bibr" rid="ref10">Sakr 2016</xref>
        ,
        <xref ref-type="bibr" rid="ref11">Singh and Reddy 2015</xref>
        ]. To this end, we examine the main
features of YARN/Hadoop, Spark, Flink and Hyracks/ASTERISK.
      </p>
      <p>
        In many Big Data scenarios, Apache Hadoop has become the data and
computational de facto standard for sharing and accessing data and computational resources
[
        <xref ref-type="bibr" rid="ref14">Vavilapalli et al. 2013</xref>
        ]. Hadoop is a scalable open-source computation framework that
allows the partitioning of computation processes across many host servers which are not
necessarily high-performance computers [
        <xref ref-type="bibr" rid="ref15">White 2012</xref>
        ]. It has two main components: a
MapReduce execution engine and a distributed file system (DFS) called HDFS – Hadoop
Distributed FileSystem. The advantages of Hadoop mainly lie in its high flexibility,
scalability, low-cost, and reliability for managing and efficiently processing a large volume
of structured and unstructured datasets, as well as providing job schedules for balancing
data, resource and task loads. Hadoop evolved to YARN – Yet Another Resource
Negotiator, whose architecture decouples the programming model from the resource management
infrastructure and delegates many scheduling functions to per-application components
[
        <xref ref-type="bibr" rid="ref14">Vavilapalli et al. 2013</xref>
        ].
      </p>
      <p>
        Apache Spark [
        <xref ref-type="bibr" rid="ref18">Zaharia et al. 2016</xref>
        ] is a unified engine for distributed data
processing. It has a programming model similar to MapReduce but extends it with a
datasharing abstraction called Resilient Distributed Datasets, or RDDs. Using this extension,
Spark can capture a wide range of processing workloads that previously needed
separate engines, including SQL, streaming, machine learning, and graph processing. Spark
[
        <xref ref-type="bibr" rid="ref17">Zaharia et al. 2010</xref>
        ] was also designed to overcome the disk I/O limitations and improve
the performance of earlier systems. The main feature of Spark is its ability to perform
in-memory computations. It allows the data to be cached in memory, thus eliminating the
YARN’s disk overhead limitation for iterative tasks.
      </p>
      <p>
        Apache Flink [
        <xref ref-type="bibr" rid="ref4">Carbone et al. 2015</xref>
        ] is an open-source stream and batch
processing framework for distributed and high-performing applications originated from
[
        <xref ref-type="bibr" rid="ref1">Alexandrov et al. 2014</xref>
        ] project. It is built on the philosophy that many classes of data
processing applications, including real-time analytics, continuous data pipelines,
historical data processing, and iterative algorithms can be expressed and executed as pipelined
fault-tolerant data flows. Flink can run as a completely independent framework, or on top
of HDFS and YARN. It leverages in-memory storage for improving the performance of
the runtime execution. The main novelties of Flink in comparison to previous Big Data
technologies: a distributed data flow runtime that exploits pipelined streaming
execution for batch and stream workloads; exactly-once state consistency through lightweight
checkpointing; native iterative processing; and a sophisticated window semantics,
supporting out-of-order processing.
      </p>
      <p>
        Hyracks/ASTERIX [
        <xref ref-type="bibr" rid="ref3">Borkar et al. 2011</xref>
        ] is a partitioned-parallel software platform
designed to run data-intensive computations on large shared-nothing clusters. Hyracks
includes a collection of operators that can be used to assemble data processing jobs
without needing to write Map and Reduce code. Moreover, it also provides a Yarn compatible
layer to run existing MapReduce jobs. The Hyracks presents a scalable information
management system that supports the storage, querying, and analysis of large collections of
semi-structured nested data objects. Hyracks provides performance gains over
MapReduce through its more flexible user model, while also being a more efficient
implementation than Hadoop for MapReduce jobs for a variety of data-intensive use cases. Hyracks
also achieves fault recovery performance gains over Hadoop by offering a less pessimistic
approach to fault handling.
      </p>
    </sec>
    <sec id="sec-4">
      <title>2.2. Comparison of Big Data Technologies</title>
      <p>Companies using the Big Data technologies are usually facing challenges like: (i) dealing
with the storage of heterogeneous sources such as structured, unstructured and
semistructured data; (ii) the need to discover knowledge from large and heterogeneous datasets
by not only applying SQL queries, but also performing complex machine learning
algorithms or graph computations; (iii) continuously receiving streams of data that must be
continuously processed in order of milliseconds for (near) real-time analytics. Based on
these challenges, we present a comprehensive discussion on how those frameworks are
able or not to provide support for streaming and batch processing, generic storage and
data analytics.</p>
      <p>Batch Processing. This kind of data processing is intimately related to long-time
running computation over a large volume of data, all at once, over a period. It is typically
performed in tasks of ETL (Extract, Transform and Load), data aggregation, training and
updating machine learning models. Hadoop was broadly adopted in batch processing due
to its MapReduce implementation for distributing the data processing within a computing
cluster with many nodes. Hyracks also performs batch data processing. However, Spark
has become the main adopted engine for large data processing by a variety of companies1,
since it brings fast in-memory data processing capability, which overcomes the Hadoop
reading and writing overheads.</p>
      <p>Streaming Processing. In stream processing, the data is processed and the results
produced strictly within specific time constraints (often in the order of milliseconds and
sometimes microseconds depending on the application and the user requirements). For
instance, Spark Streaming receives live input data streams and divides the data into
microbatches, which are processed by the Spark engine and used to generate the final stream
of results in batches. Micro-batching allows handling a stream as a sequence of small
batches or chunks of data. However, it can introduce considerable overhead in the form
of scheduling tasks. On the other hand, Flink can deliver all of the advantages of buffering
with none of the task-scheduling overhead. Flink can also perform well on real-time or
near-real-time scenarios, where insights from data should be available at nearly the same
moment of data generation.</p>
      <p>Generic Storage. HDFS can store a diverse mix of structured, unstructured and
semistructured data. Hyracks can consume the data from HDFS, and it also provides
the data storage AsterixDB to ingest, store, index, query, and analyze mass quantities of
data using a flexible data model (ADM). Spark supports integration with a wide variety of
file systems, including HDFS, MapR File System, Cassandra, Amazon S3, or the
implementation of a custom solution. Flink enables the integration of heterogeneous data sets,
ranging from strictly structural relational data, unstructured text data and semi-structured
data. It also works with HDFS and connects to various other data storage systems. Flink
and Spark do not provide a primary storage solution.</p>
      <p>
        Data Analytics. YARN/Hadoop actively supports several top-level projects to create
development tools and to manage its data flow and processing such as Giraph, Pig, Hive,
Mahout, and HBase. Spark also supports a wide range of applications, including ETL,
Machine Learning (MLib), Stream Processing (Spark Streaming), and Graph computation
(GraphX). Flink’ stack offers libraries with high-level APIs for different use cases:
Complex Event Processing (CEP), Machine Learning (FlinkML), and Graph Analytics (Gelly).
The software stack of Hyracks system is composed of various interfaces for analytics as
well, like SQL (Hivesterix), XQuery (Apache VXQuery), and Graph (Pregelix). Even
Hyracks can efficiently execute complex distributed data-flow operations and express full
relational algebras. It also exposes low-level APIs and requires a machine learning (ML)
expert to reformulate its algorithms as dataflow operators [
        <xref ref-type="bibr" rid="ref13">Sparks et al. 2013</xref>
        ].
      </p>
      <p>Finally, Table 1 summarizes our discussion. It provides a short-comparison of Big
Data technologies capabilities according to the categories of Big Data problems analyzed.
```````T`ec`hn`olo`gy`</p>
      <p>Category</p>
      <p>Processing Type
Generic Storage
Data Analytics</p>
      <p>Hadoop</p>
      <p>Batch</p>
      <p>HDFS
SQL, ML,</p>
      <p>Graph</p>
      <p>Spark</p>
      <p>Mini-batch
no primary storage</p>
      <p>ETL, ML,</p>
      <p>Graph</p>
      <p>Flink
Streaming,</p>
      <p>Batch
no primary storage</p>
      <p>ML, CEP,</p>
      <p>Graph</p>
      <p>Hyracks</p>
      <p>Batch</p>
      <p>AsterixDB
SQL, XQuery,</p>
      <p>Graph</p>
    </sec>
    <sec id="sec-5">
      <title>3. Big Data Platforms</title>
      <p>
        A Big Data platform is an ecosystem of services and technologies that needs to perform
analysis on voluminous, complex and dynamic data. Thus, scaling up the hardware
platform becomes imminent and choosing the right hardware/software technologies becomes
a crucial decision if the user’s requirements are to be satisfied in a reasonable amount of
time [
        <xref ref-type="bibr" rid="ref11">Singh and Reddy 2015</xref>
        ]. A set of Big Data platforms has recently emerged,
including Big Data Europe (BDE) [
        <xref ref-type="bibr" rid="ref9">Jabeen et al. 2017</xref>
        ], Hortonworks 2 and Cloudera3. In this
section, we briefly describe these Big Data platforms and provide a comparison between
them according to some categories of Big Data problems.
      </p>
    </sec>
    <sec id="sec-6">
      <title>3.1. Overview</title>
      <p>
        BDE platform [
        <xref ref-type="bibr" rid="ref9">Jabeen et al. 2017</xref>
        ] developed a computing infrastructure for handling
large volumes of data in a variety of formats. It addresses the requirements of
simplifying use, easing deployment, managing heterogeneity and improving scalability, and
facilitates the execution and integration of Big Data frameworks and tools like Hadoop,
Spark, Flink and many others. The authors have decided to use Docker as packaging
and deployment methodology as well as managing the variety of underlying hardware
resources efficiently alongside the varying software requirements. BDE allows performing
a variety of Big Data flow tasks such as message passing (via Kafka, Flume), storage (via
Hive, Cassandra), analysis (via Spark, Flink) or publishing (via GeoTriples). Moreover,
the platform is open-source and completely free.
      </p>
      <p>Hortonworks Data Platform (HDP) is an open-source modern data architecture
that delivers immediate value by slashing storage costs as it integrates Yarn into its data
center, and by optimizing Enterprise Data Warehouse costs by offloading low-value
computing tasks such as ETL to Yarn. Yarn allows HDP to integrate all data processing
engines across the community and commercial ecosystem to deliver consistent shared
services and resources across the platform. Ambari is an intuitive Web UI and a robust
REST API that makes HDP management simpler, consistent and secure. Furthermore,
HDP is a complete solution offering not just data processing and management, but the
enterprise capabilities to match the demands of an enterprise spanning security, governance,
and operations.</p>
      <p>Cloudera was the first company to develop and distribute Apache Hadoop-based
software, and it has made data analytics on Big Data more convenient and accessible
to anyone interested. It integrates Hadoop with more than a dozen other critical open
source projects. Cloudera created a functionally advanced system that helps to perform
2https://br.hortonworks.com/
3https://www.cloudera.com/
end-to-end Big Data workflows. Different projects compose Cloudera ecosystem for a
variety of Big Data tasks: streaming processing (via Spark), message passing (via Kafka,
Flume), storage (via Accumulo, Hive, Pig, HBase), analysis (via Flink, Impala), searching
(via Cloudera Search) or providing an extensible and productive web GUI for users (via
HUE).</p>
    </sec>
    <sec id="sec-7">
      <title>3.2. Comparison of Big Data Platforms</title>
      <p>Following we present a set of recurrent problems usually faced by organizations that
might become more complex when dealing with Big Data: (i) integrate different Big
Data sources and provide a transparent view to the users; (ii) manage and protect the
organization’s data assets in order to guarantee generally understandable, correct, complete
and secure corporate data; (iii) monitor the data, resources and applications to review and
evaluate the health and performance of the whole system. Each challenge can be
summarized in one of the following categories of problems: Data Integration, Data Governance,
and Monitoring Services. In what follows, we provide a comparison between the Big Data
platforms mentioned in this work, and what they provide to deal with such problems.</p>
      <p>Data Integration. Data integration involves combining data from different
sources and providing users with a unified view of them. HDP has partnered with
Talend, a powerful and versatile open source solution for Big Data integration that natively
supports Hadoop, including connectors for HDFS, HBase, Pig, Sqoop, and Hive without
having to write any code. Talend also supports Cloudera Navigator. Another alternative
for HDP is Oracle Data Integrator (ODI). A user can create a flow from sources to targets
of different technologies, including relational databases, applications, XML, JSON, Hive
tables, HBase, HDFS files, and so on. BDE platform goes further than HDP and Cloudera
by comprising a Semantic Data Lake – a repository provided for processing and analysis
the datasets in their original formats – named Ontario. Ontario builds a Semantic Layer
on top of the Data Lake, which is responsible for mapping data into existing Semantic
vocabularies/ontologies. A successful mapping process, termed Semantic Lifting, provides
a view over the whole data. In this way, data can be extracted, queried or analyzed from
the heterogeneous sources in the lake as if it was in a single format using a high-level
query language. Another relevant component is Semagrow, a SPARQL query processing
system that federates multiple remote endpoints.</p>
      <p>
        Data Governance. Data Governance is a system of decision rights and
accountabilities for information-related processes, executed according to agreed-upon models
which describe who can take what actions with what information, and when, under what
circumstances, using what methods [
        <xref ref-type="bibr" rid="ref6">Data Governance Institute 2018</xref>
        ]. [
        <xref ref-type="bibr" rid="ref12">Soares 2012</xref>
        ]
expands this definition by including policies regarding the optimization, privacy, and
monetization of Big Data. Governing Big Data systems can be complex. Securing datasets
consistently across multiple repositories can be extremely error-prone. Cloudera
Navigator Data Management component is a fully integrated data management and security
tool for the Hadoop that has been designed to meet compliance, data governance, and
auditing needs of global enterprises. HDP uses Apache Atlas and Apache Ranger, which
combine data classification with security policy enforcement. Apache Atlas was created
as part of the Hadoop Data Governance initiative, and it offers the ability to view the
cross-component lineage, providing a complete view of the data movement through some
parsing engines such as Apache Storm, Kafka, Falcon, and Hive. Apache Ranger provides
      </p>
      <p>XXXXXXPXlatXforXmX</p>
      <p>Category</p>
      <p>Data Integration
Data Governance</p>
      <p>Monitoring</p>
      <p>HDP
Talend, ODI
Atlas, Ranger</p>
      <p>Ambari</p>
      <p>BDE
Ontario, Semagrow</p>
      <p>No support
Prometheus, ELK stack</p>
      <p>Cloudera</p>
      <p>Talend
Cloudera Navigator
Cloudera Manager
centralized security management for Hadoop. By integrating Atlas and Ranger, HDP
allows companies to implement dynamic, runtime access policies that pro-actively prevent
violations. BDE does not delve much into data governance since it does not address issues
such as data privacy, sharing, and rights.</p>
      <p>Monitoring. Monitoring is the process of proactively reviewing and evaluating
what has been monitored (as data, resources or applications). Monitoring software helps
to measure and track the data usually using dashboards, alerts, and reports. Cloudera
Manager provides many features for monitoring the health and performance of the clusters
components (hosts, service daemons) as well as the performance and resource demands
of the jobs running on clusters. BDE distinguishes between resource monitoring and
status monitoring. The former allows to follow up the health of a server or a component
in the platform (CPU usage, memory usage, network I/O and disk utilization) while the
latter offers insight in the status of a specific application. For resource monitoring, the
tools Docker Stats, cAdvisor, Prometheus, InfluxDB, and Grafana can be useful at BDE
platform. For status monitoring, BDE supports docker built-in logging and ELK stack.
As part of HDP, Apache Ambari allows to plan, install and securely configure clusters of
computers, by making it easier to provide ongoing cluster maintenance and management.</p>
      <p>Finally, a summary of this section is presented in Table 2, that provides a
shortreview of Big Data platforms according to the categories of Big Data problems analyzed.</p>
    </sec>
    <sec id="sec-8">
      <title>4. Conclusion</title>
      <p>This paper surveys various Big Data technologies and platforms that are currently
available and discusses their capabilities. A comparison between different technologies based
on some important Big Data problems has been made. In addition, we also compare
different Big Data platforms based on their support for data integration, data governance,
and monitoring. By providing this guideline, we aim at helping organizations in the
selection of technologies/platforms more appropriate to their Big Data problems. A future
work consists of an empirical evaluation of these technologies/platforms by using
different Big Data scenarios/applications. Moreover, we intend to compare them according to
other categories of Big Data problems, such as how these platforms/technologies manage
and integrate different data analysis outputs and algorithms.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>This work has been supported by FUNCAP SPU 8789771/2017 research project and
CAPES fellowship.</p>
      <p>Dresner Advisory Services (2017). Big Data Analytics Market Study.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Alexandrov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergmann</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ewen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freytag</surname>
            ,
            <given-names>J.-C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hueske</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heise</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kao</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leich</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leser</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markl</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , et al. (
          <year>2014</year>
          ).
          <article-title>The stratosphere platform for big data analytics</article-title>
          .
          <source>The VLDB Journal</source>
          ,
          <volume>23</volume>
          (
          <issue>6</issue>
          ):
          <fpage>939</fpage>
          -
          <lpage>964</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Begoli</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Horey</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Design principles for effective knowledge discovery from big data</article-title>
          .
          <source>In Joint ICSA and ECSA</source>
          , pages
          <fpage>215</fpage>
          -
          <lpage>218</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Borkar</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grover</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Onose</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Vernica</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Hyracks: A flexible and extensible foundation for data-intensive computing</article-title>
          .
          <source>In ICDE</source>
          , pages
          <fpage>1151</fpage>
          -
          <lpage>1162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Carbone</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Katsifodimos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ewen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markl</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haridi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Tzoumas</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Apache flink: Stream and batch processing in a single engine</article-title>
          .
          <source>Bulletin of the IEEE Computer Society TCDE</source>
          ,
          <volume>36</volume>
          (
          <issue>4</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dolan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunlap</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellerstein</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Welton</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Mad skills: new analysis practices for big data</article-title>
          .
          <source>Proceedings of the VLDB Endowment</source>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          ):
          <fpage>1481</fpage>
          -
          <lpage>1492</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Data</given-names>
            <surname>Governance Institute</surname>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Definitions of data governance</article-title>
          . http:// www.datagovernance.com/adg_data_governance_definition/. Accessed:
          <fpage>2018</fpage>
          -05-01.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Herodotou</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lim</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borisov</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cetin</surname>
            ,
            <given-names>F. B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Babu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Starfish: a self-tuning system for big data analytics</article-title>
          .
          <source>In CIDR</source>
          , pages
          <fpage>261</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Inoubli</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aridhi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mezni</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maddouri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Nguifo</surname>
            ,
            <given-names>E. M.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>An experimental survey on big data frameworks</article-title>
          .
          <source>Future Generation Computer Systems.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Jabeen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Archer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scerri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Versteden</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ermilov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mouchakis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Big data europe</article-title>
          .
          <source>In EDBT/ICDT Workshops.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Sakr</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <source>Big data 2.0 processing systems: a survey</source>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Reddy</surname>
            ,
            <given-names>C. K.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>A survey on platforms for big data analytics</article-title>
          .
          <source>Journal of Big Data</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Soares</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Big data governance: an emerging imperative</article-title>
          . Mc Press.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Sparks</surname>
            ,
            <given-names>E. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Talwalkar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kottalam</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franklin</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M. I.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kraska</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Mli: An api for distributed machine learning</article-title>
          .
          <source>In ICDM.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Vavilapalli</surname>
            ,
            <given-names>V. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murthy</surname>
            ,
            <given-names>A. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Douglas</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Graves</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lowe</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seth</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al. (
          <year>2013</year>
          ).
          <article-title>Apache hadoop yarn: Yet another resource negotiator</article-title>
          .
          <source>In Proceedings of the 4th Symposium SOCC.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>White</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Hadoop: The definitive guide</article-title>
          . ”
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          , Inc.”.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>G.-Q.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Data mining with big data</article-title>
          .
          <source>IEEE TKDE</source>
          ,
          <volume>26</volume>
          (
          <issue>1</issue>
          ):
          <fpage>97</fpage>
          -
          <lpage>107</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Zaharia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chowdhury</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franklin</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shenker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Stoica</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Spark: Cluster computing with working sets</article-title>
          .
          <source>HotCloud</source>
          ,
          <volume>10</volume>
          (
          <fpage>10</fpage>
          -10):
          <fpage>95</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Zaharia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xin</surname>
            ,
            <given-names>R. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wendell</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Armbrust</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dave</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meng</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venkataraman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franklin</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          , et al. (
          <year>2016</year>
          ).
          <article-title>Apache spark: a unified engine for big data processing</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>59</volume>
          (
          <issue>11</issue>
          ):
          <fpage>56</fpage>
          -
          <lpage>65</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>