<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Large-scale Computing Frameworks: Experiments and Guidelines</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michele Baglioni</string-name>
          <email>michele.baglioni1@studenti.unipg.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabrizio Montecchiani</string-name>
          <email>fabrizio.montecchiani@unipg.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mario Rosati</string-name>
          <email>mario.rosati@e4-analytics.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Big data engines</institution>
          ,
          <addr-line>Apache Spark, Dask, Ray, Kubernetes</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dipartimento di Ingegneria, Università degli Studi di Perugia</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>E4 Analytics</institution>
          ,
          <addr-line>67039 Sulmona (AQ)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Large-scale computing frameworks are key technologies to fulfill the computational requirements of massive data analysis. In particular, while Apache Spark has emerged as de facto standard for big data analytics after Hadoop's MapReduce, tools such as Dask and Ray can greatly boost the performance of Python applications in distributed environments. The goal of this paper is to study the performance of these three frameworks on a common playground. We focus on cloud-native architectures, which merge the benefits of big data and cloud computing. We refrain from considering high-level features such as ML models, we instead consider simple data processing operations, common ingredients of more complex pipelines. As a byproduct of our experiments, we ofer a set of guidelines for the development of cloud-native data processing applications.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Massive data processing pipelines running on single-node machines or large clusters can benefit
of computing frameworks that hide low-level operations needed to parallelize the workload.
In this regard, so-called Big Data engines play a crucial role as they make it simple to write
applications that transparently scale with the underlying infrastructure. For instance, Apache
Spark1 has emerged as the de facto standard for big data analytics after Hadoop’s MapReduce [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
It ofers an open-source unified analytics engine for large-scale data processing, with transparent
fault-tolerance. Spark embraces key principles such as in-memory computing, data locality
and lazy evaluation, in order to achieve great speedups and strong scalability. While Spark
has a powerful set of programming interfaces for multiple languages, other projects aim for
a seamless and transparent scaling of Python applications. Among them, Dask2 is a newer
open-source library to scale Python code, providing a familiar user interface [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Indeed, it
mirrors the APIs of other popular libraries such as Pandas, scikit-learn and NumPy. Similarly,
CEUR
Workshop
Proceedings
Ray3 is an open-source compute framework to scale Python applications, introduced only about
5 years ago, with a focus on reinforcement learning, deep learning, hyperparameters tuning, and
model serving [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We point the reader to the oficial documentations of these three frameworks
for additional information and details. Also, some design features will be discussed in Section 2.
      </p>
      <p>
        The growth of these technologies has been accompanied by the growth of cloud-native
applications, which are increasingly studied in research and adopted in industry. Cloud-native
applications consists of several processes that run in isolated containers, which are typically
spread over the nodes of a cluster. In this respect, Kubernetes4 is becoming the standard
technology to automate software deployment, scaling, and management through containers
orchestration [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Given the popularity of Spark, Dask, and Ray, and the importance of making informed
choices in terms of technology, the goal of this paper is to study the performance of these
three frameworks on a common playground. We use a cloud-native environment based on
Kubernetes, and we experiment standard data processing pipelines fed by a large dataset
consisting of numerical values. In particular, while we are well aware of the diferences and
peculiarities in terms of high-level APIs ofered by the considered frameworks, our focus is on
standard preprocessing operations that are typically performed in any data-processing pipeline.
We do not consider high-level features such as ML/DL models as they require diverse datasets
and specialized hardware (such as suitable GPUs), to be conveniently evaluated. As a byproduct
of our experiments, we ofer a set of guidelines for the development of cloud-native
dataprocessing applications. To the best of our knowledge, this is the first paper that experimentally
compares these three frameworks. The only similar comparison we are aware of is restricted to
Dask and Spark, and it focuses on neuroimaging pipelines [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Paper structure. A detailed description of the experimental setting can be found in Section 2.
Results and guidelines are discussed in Section 3 and Section 4, respectively. Section 5 concludes
the paper with future research directions.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Experimental Setting</title>
      <sec id="sec-3-1">
        <title>2.1. Infrastructure</title>
        <p>In this section we describe the infrastructure and the benchmark used to run the experiments.
Hardware. In terms of hardware, the experiments were conducted on a Dell PowerEdge
XE8545 server, designed to take advantage of the industry’s most advanced technologies. The
server is equipped with two sockets and has a 4U format, with two 128-core AMD EPYC
processors, NVIDIA A100 Tensor Core GPUs and 1024 GB of RAM. As for storage, it has 4
Dell Ent NVMe CM6 disks of 2TB each. See also Table 1 for full references. Furthermore, it
is equipped with NVLink, PCIe Gen 4.0 and NVMe SSD to improve I/O and data processing
performance. These technologies speed up data transfer between the CPU and GPU, improving</p>
        <sec id="sec-3-1-1">
          <title>3https://www.ray.io/ 4https://kubernetes.io/</title>
          <p>the overall eficiency of the system. The server is certified by NVIDIA for its performance,
manageability, security and scalability.</p>
          <p>Processor
Memory</p>
          <p>GPUs
Operating System</p>
          <p>Storage</p>
          <p>Dual AMD EPYC 7713, 64C, 2.8 GHz
1024 GB (16x64GB @ 3200 MT/s)
4 x NVIDIA A100 SXM4 80 GB</p>
          <p>Linux Ubuntu Server 20.04 LTS
4 x Dell Ent NVMe CM6 RI 1.92TB 2.1.8</p>
          <p>
            Kubernetes. We deployed a bare-metal single-node Kubernetes cluster on the previously
described server. We remark that the scheduling and resource allocation process is a critical
aspect for such experiments [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]. To face this issue, we employed Volcano5, an add-on for running
HPC workloads on Kubernetes. In particular, diferent from Kubernetes, Volcano features gang
scheduling6, which ensures that each job will run at “full speed”, that is, it will start to run only
when all its tasks are ready to be deployed. The adopted scheduler configuration is reported in
Listing 1.
          </p>
          <p>Listing 1: Volcano scheduler configuration.
1 a p i V e r s i o n : v1
2 d a t e :
3 v o l c a n o − s c h e d u l e r . c o n f : |
4 a c t i o n s : ” enqueue , a l l o c a t e , b a c k f i l l ”
5 t i e r s :
6 − p l u g i n s :
7 − name : p r i o r i t y
8 − name : gang
9 − name : c o m p l i a n c e
10 k i n d : ConfigMap
11 m e t a d a t a :
12 c r e a t i o n T i m e s t a m p : ” 2 0 2 3 − 0 2 − 1 5 T09 : 1 6 : 4 7 Z ”
13 name : v o l c a n o − s c h e d u l e r − c o n f i g m a p
14 namespace : v o l c a n o − s y s t e m</p>
          <p>For each framework a specific queue has been defined, as reported in Listing 2, in order to
decouple the workloads of the diferent frameworks; see also Fig. 1 for a schematic illustration.</p>
          <p>Listing 2: Volcano queue definition.
1 a p i V e r s i o n : s c h e d u l i n g . v o l c a n o . sh / v 1 b e t a 1
2 k i n d : Queue
3 m e t a d a t a :
4 name : t e s t − s p a r k</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>5https://volcano.sh/en 6https://volcano.sh/en/docs/plugins/</title>
          <p>Finally, to ensure reproducibility, Table 2 provides all the versions and configurations used
for the entire software stack.</p>
          <p>Storage. Read and write operations are fundamental in the context of data processing.
Therefore, it is essential to configure the system to operate in the best conditions. In particular, we
implemented a scalabale object storage system via MinIO7 within the Kubernetes cluster. This
system consists of 4 physical disks for a total of 7 TiB of available space.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Benchmark</title>
        <p>
          Dataset. The dataset was generated synthetically by sampling random values in the
range [
          <xref ref-type="bibr" rid="ref1">1, 103</xref>
          ] from a uniform probability distribution. The dataset has 1.6 ⋅ 109 rows and
a schema, represented in Table 3, where there are three columns named a, b, and c. The first
column contains integers, while the other two columns contain decimal numbers represented
as doubles.
        </p>
        <p>
          Format. As for the format in which to store the dataset, we used Apache Parquet [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. It is a
selfdescribing data format that embeds the schema into the data itself. It supports eficient encoding
and compression schemes that help lower data storage costs and maximize the efectiveness of
data queries. Parquet has additional advantages, such as storing data in compressed form using
Snappy [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], a method developed by Google and designed for space and query eficiency. The
dataset was divided into 160 blocks. The archived dataset has a size of 13.8 GB, and, once loaded
into memory, it occupies about 39 GB. Dividing the dataset into blocks has the advantage of
facilitating the parallelization of data processing operations on a cluster, as each block can be
processed independently of the others. The block count of 160 equals the least common multiple
of the number of worker nodes in all configurations evaluated. This ensures that, in all cluster
configurations we evaluated, the workers can perform their read operations concurrently.
Operations. At high-level, the benchmark consists of reading data from the distributed
storage, manipulating the data with elementary functions, and finally writing the data back in
the storage; see Fig. 2 for an illustration. The pipeline is therefore composed of the following
• Read: the data are loaded into the main memory, performing a read operation from the
distributed storage.
• Sort: the data items are sorted in ascending order based on the value of a specific column.
• Shufle: the data items are randomly shufled.
• Create new column: a new column of is created from other existing columns, for example
to combine two columns or extract information from one column.
• Transform of a column: an existing column is transformed to make it more suitable for
subsequent processing.
• Sum between two columns: the values of two columns are added together, generating a
new column to store the result.
• Weighted average: the weighted average of the data items is calculated on the basis of
certain weights assigned to each item.
• Math function: the following math function is applied to the values of multiple existing
columns: cos(1) ∗ arctan(2) + log(3)
• Write: Processed data is written as a new object into the distributed storage.
Applications. In order to support reproducibility of the experiments, we adopted the
Kubernetes Operator Pattern [
          <xref ref-type="bibr" rid="ref4 ref9">4, 9</xref>
          ], which makes it possible to integrate domain knowledge into
Kubernetes’ orchestration process. In order to implement this pattern, we first define new types
of resources that the Kubernetes API can manage (called CRDs); next, we add the operator,
which is a software component running inside the cluster in order to manage the entire lifecycle
of the CRDs. In particular, the operator interacts with the Kubernetes API and reacts to creation,
modification, or removal of custom resources (CRs). Therefore, for each of the three frameworks,
the respective operators were installed via Helm chart in the Kubernetes cluster. Fig. 3 shows
how a Spark application is deployed and launched in the Kubernetes cluster. As it will be
clarified later, each app is executed with a number  of workers, with  ranging between 1 and
32. Listing 3 shows one of the YAML files used to define Spark applications. The applications
are deployed and launched similarly for the other frameworks. The only exception is that Dask
does not provide a direct integration with the Volcano scheduler, hence we relied on a Volcano
job to launch Dask applications.
        </p>
        <p>Listing 3: Definition of a Spark Application.
cpu : 32</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Experimental Results</title>
      <p>Metrics and procedure. In order to assess the performance of the experimented frameworks,
besides the runtime, we considered both speedup and eficiency, briefly recalled below. Let
be the number of available processors, let   be the time taken to process the workload with
 ∈ [1,  ]</p>
      <p>
        processors, then the metrics are defined as follows (a linear speedup and constant
eficiency are ideal targets, see, e.g., [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]).
      </p>
      <p>
        Each framework has been evaluated with a number of workers equal to 2 , with  ranging in
[
        <xref ref-type="bibr" rid="ref5">0, 5</xref>
        ]. The results of the various executions are extracted from the logs of the respective master
node’s pod. For the sake of reliability, the results are averaged over 10 executions of the same
application with the same workers configuration.
      </p>
      <p>Overall performance.</p>
      <p>We being by remarking that, given a pipeline of operations, each
framework adopts a specific strategy to break the workload into smaller tasks and to assign the

(1)
(2)
speedup () =
eficiency () =

1


()</p>
      <p>tasks to the workers. In particular, triggering the execution of single operations to measure each
single run time would hinder such optimization strategies and invalidate the results. Therefore,
we measure the run time of the whole pipeline, such that each framework can fully exploit its
own optimization strategy. The run time includes the time needed to read the data, which will
also be evaluated separately in the next paragraph.</p>
      <p>Table 4 reports the results of the experiments. In terms of overall execution time, Ray appears
to be the fastest framework, with the best performance in all worker configurations except in
the single worker setting, in which Dask is only a few seconds faster. Between Dask and Spark,
we see that Dask outperforms Spark up to 2 workers, while Spark performs better between
4 and 32 workers. In terms of speedup and eficiency, Spark exhibits very good performance
up to 16 workers, showing an eficiency drop with 32 workers. The Ray framework shows
better scalability, with a smaller drop in terms of eficiency. On the other hand, Dask exhibits
low scalability compared to the other two frameworks, its eficiency rapidly decreases as the
number of workers increases.</p>
      <p>Reading performance. In this paragraph we analyze the performance of the frameworks in
terms of reading performance, that is, the time taken to read the data from the object storage
and to load the data into the main memory. Table 5 reports the results of our experiments.
Again, for the single worker scenario, Dask is the fastest framework, and it remains the fastest
also with 2 workers. With 4 workers Spark becomes the fastest framework, and it remains
the fastest also with 8 workers, together with Ray. With 16 and 32 workers, Ray is the fastest
framework, with Spark slightly slower. In terms of speedup, again Ray and Spark outperform
Dask, which also exhibits the worse behavior in terms of eficiency.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Discussion, Limitations and Guidelines</title>
      <p>
        Discussion. We begin by briefly discussing possible relations between the design principles
behind the three frameworks and the observed performance. Spark has been designed to handle
large amounts of data and complex operations, using distributed data structures to keep data in
memory and minimize disk access operations. In particular, its ability to optimize operations
through its distributed execution engine, dividing the operations to be performed into tasks and
distributing them eficiently on the nodes of the cluster, makes it able to scale very well with
configurations with high availability of resources, while it appears less eficient with few workers.
Ray is a system designed to handle high-speed, low-latency distributed processing operations.
To achieve this goal, eficient communication between nodes and an actor-based programming
model are used, which allows operations to be distributed asynchronously, minimizing workers’
downtime. This architecture makes it very eficient in all configurations. Dask uses a data
partitioning strategy based on blocks of variable size, which allows to better adapt to the size of
the dataset and to minimize data movements between nodes. With few workers, Dask’s data
partitioning strategy proves to be particularly efective, as it allows you to make the most of
the available resources, minimizing execution times. However, as the number of worker nodes
increases, coordinating distributed operations becomes more complex, and Dask may show
limits in terms of scalability compared to Spark and Ray. In particular, Dask can sufer from
increased overhead for coordinating distributed operations and managing distributed memory,
which can slow performance on large clusters.
Limitations. Our experiments where performed on a highly optimized infrastructure,
however, the performance of the frameworks are afected by the use of containers. For instance,
it was shown that Kubernetes pods may deteriorate data locality, and make a worse usage of
memory and CPU [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Also, our experiments did not measure the resiliency of the three
frameworks, nor their performance when some tasks are aborted or delayed. Our results should not
be generalized to larger clusters, in particular with multiple racks in which network bottlenecks
may arise. Similarly, much larger datasets, in the order of Terabytes, may lead to diferent
behaviours of the frameworks. Also, we only considered simple data processing operations,
while all the frameworks have advanced libraries to train and use ML/DL models. Since the
performance of these advanced features strongly depend on their implementation and on the
availability of specialized hardware, we cannot draw any conclusion about model training from
our experiments.
      </p>
      <p>Guidelines. Based on our findings and with the above limitations in mind, Ray appears
to be the most eficient framework over all configurations. On the other hand, Dask is the
fastest framework with only one worker, but it is less scalable than Spark and Ray. Spark
performs somewhere in between, with performance improving as nodes increase. In the case of
read-intensive applications, Dask seems a good choice if only few nodes are available, while
Spark or Ray would be recommended with a large number of workers. Besides performance,
the ecosystem is an important aspect to consider when choosing a technology over another.
In this respect, Spark stands out for its reliability, as it is certainly a well-tested framework,
with an established community of developers ofering a wide range of resources, including
documentation and reusable code. On the other hand, Ray allows you to run Spark and Dask
code on the infrastructure managed by the Ray Core, thus making it possible to use particular
features of other frameworks or existing legacy code. Another key aspect when staring out a
project with a new technology is the required learning curve. In this regard, Dask’s learning
curve is rather shallow for Python developers, with the possibility of using the same syntax of
the most common libraries, such as Pandas and Numpy.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Future Work</title>
      <p>We plan to extend our experiments by considering larger datasets and more advanced features
ofered by the three frameworks. In particular, we would like to compare their performance in
terms of model training. To this aim, we plan to equip our cluster with suitable GPUs, which
represent the current standard technology in this field. This type of comparison should involve
diferent ML/DL models and hence multiple datasets, for instance, including images and text.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Salloum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dautov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. X.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>Big data analytics on apache spark</article-title>
          ,
          <source>Int. J. Data Sci. Anal</source>
          .
          <volume>1</volume>
          (
          <year>2016</year>
          )
          <fpage>145</fpage>
          -
          <lpage>164</lpage>
          . URL: https://doi.org/10.1007/s41060-016-0027-9. doi:
          <volume>10</volume>
          .1007/ s41060-016-0027-9.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rocklin</surname>
          </string-name>
          ,
          <article-title>Dask: Parallel computation with blocked algorithms and task scheduling, in: SciPy, scipy</article-title>
          .org,
          <year>2015</year>
          , pp.
          <fpage>126</fpage>
          -
          <lpage>132</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Moritz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Nishihara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tumanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Liaw</surname>
          </string-name>
          , E. Liang,
          <string-name>
            <given-names>M.</given-names>
            <surname>Elibol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          , W. Paul,
          <string-name>
            <given-names>M. I.</given-names>
            <surname>Jordan</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Stoica</surname>
          </string-name>
          ,
          <article-title>Ray: A distributed framework for emerging AI applications</article-title>
          , in: OSDI, USENIX Association,
          <year>2018</year>
          , pp.
          <fpage>561</fpage>
          -
          <lpage>577</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ibryam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Huß</surname>
          </string-name>
          , Kubernetes Patterns:
          <article-title>Reusable Elements for Designing Cloud-native</article-title>
          <string-name>
            <surname>Applications</surname>
          </string-name>
          ,
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          ,
          <year>2019</year>
          . URL: https://books.google.it/books?id=
          <fpage>Ax53wgEACAAJ</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dugré</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Hayot-Sasson</surname>
          </string-name>
          ,
          <string-name>
            <surname>T. Glatard,</surname>
          </string-name>
          <article-title>A performance comparison of dask and apache spark for data-intensive neuroimaging pipelines</article-title>
          ,
          <source>in: WORKS@SC, IEEE</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>40</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guitart</surname>
          </string-name>
          ,
          <article-title>Fine-grained scheduling for containerized HPC workloads in kubernetes clusters</article-title>
          , in: HPCC/DSS/SmartCity/DependSys 2022, IEEE,
          <year>2022</year>
          , pp.
          <fpage>275</fpage>
          -
          <lpage>284</lpage>
          . doi:
          <volume>10</volume>
          .1109/
          <string-name>
            <surname>HPCC-DSS-SmartCity-DependSys57074</surname>
          </string-name>
          .
          <year>2022</year>
          .
          <volume>00068</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Sharma</surname>
          </string-name>
          , U. Marjit, U. Biswas,
          <article-title>Eficiently processing and storing library linked data using apache spark and parquet</article-title>
          ,
          <source>Information Technology and Libraries</source>
          <volume>37</volume>
          (
          <year>2018</year>
          )
          <fpage>29</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Janet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Balakrishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. R.</given-names>
            <surname>Prasad</surname>
          </string-name>
          ,
          <article-title>Optimizing data movement within cloud environment using eficient compression techniques</article-title>
          ,
          <source>in: 2016 International Conference on Information Communication and Embedded Systems (ICICES)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICICES.
          <year>2016</year>
          .
          <volume>7518896</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Henning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wetzel</surname>
          </string-name>
          , W. Hasselbring,
          <article-title>Reproducible benchmarking of cloud-native applications with the kubernetes operator pattern</article-title>
          ,
          <source>in: SSP</source>
          , volume
          <volume>3043</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEURWS.org,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Eager</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zahorjan</surname>
          </string-name>
          , E. Lazowska,
          <article-title>Speedup versus eficiency in parallel systems</article-title>
          ,
          <source>IEEE Transactions on Computers</source>
          <volume>38</volume>
          (
          <year>1989</year>
          )
          <fpage>408</fpage>
          -
          <lpage>423</lpage>
          . doi:
          <volume>10</volume>
          .1109/12.21127.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , B. Han,
          <string-name>
            <given-names>Y</given-names>
            .
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>A comparative performance study of spark on kubernetes</article-title>
          ,
          <source>J. Supercomput</source>
          .
          <volume>78</volume>
          (
          <year>2022</year>
          )
          <fpage>13298</fpage>
          -
          <lpage>13322</lpage>
          . URL: https://doi.org/10.1007/s11227-022-04381-y. doi:
          <volume>10</volume>
          .1007/ s11227-022-04381-y.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>