<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>SEBD</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Maximizing Eficiency in Existing Data Preparation Pipelines</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Angelo Mozzillo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>First Year PhD Sudent, ICT Doctorate at DBGroup, University of Modena and Reggio Emilia</institution>
          ,
          <addr-line>Modena</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>31</volume>
      <fpage>02</fpage>
      <lpage>05</lpage>
      <abstract>
        <p>Data preparation involves transforming, cleaning, and converting raw data into a usable format for further analysis. This process can be time-consuming and resource-intensive. Typically, data to be analyzed is placed in two-dimensional tabular structures called DataFrames. These are the de facto standards for data science and machine learning tasks to store and process large amounts of structured data. Pandas is the most commonly used API for manipulating DataFrames in Python due to its popularity and comprehensive functionality. However, it is important to note that Pandas has some limitations, such as being single-core and non-distributed, which can impact its eficiency and performance for large datasets. Several libraries have been developed to expand the functionality of Pandas by utilizing multi-core and distributed computing capabilities. Therefore, the choice of library can significantly impact the eficiency and performance of the data preparation pipeline, and it is important to consider the specific requirements of the project when selecting a library. In this paper, I will discuss my primary contributions during the first months of my PhD program, which mainly focused on creating an open-source framework to evaluate the performance of seven Python libraries on five datasets of diferent sizes. The primary objective is to identify the best combination of these libraries to create eficient data preparation pipelines that deal with the needs of users who frequently encounter several libraries claiming to perform similar tasks. Finally, I will describe some ideas concerning the future directions of this research.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Data Preparation</kwd>
        <kwd>Big Data</kwd>
        <kwd>Data Science Pipeline</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Data preparation is an important process that involves exploring, combining, cleaning, and
transforming raw data into curated datasets that can be used for various purposes [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. This
process is essential for ensuring that the data is accurate, reliable, and suitable for the intended
use. In simpler terms, data preparation can be defined as the set of pre-processing operations
that are the pre-requisite for an efective data science task. These operations are aimed at
transforming the raw data into a structured format that is more suitable for analysis and
decision-making.
      </p>
      <p>
        One of the most common data structures used for storing and manipulating data in the data
preparation process is the DataFrame. A DataFrame is a two-dimensional data structure that
consists of rows and columns [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Pandas is the most widely used library for operating on
DataFrames [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]; it is considered the standard for all data preparation operations. It provides
an API with a wide range of functions for carrying out the various stages of data preparation.
These functions allow users to perform tasks such as data cleaning, merging, aggregation, and
transformation to ensure that the data is of high quality and suitable for further analysis.
      </p>
      <p>Despite its many advantages, Pandas has significant limitations when dealing with Big Data.
The main drawbacks of Pandas are:
• Memory: Since Pandas loads the full dataset into memory, the size of the dataset that can
be analyzed is constrained by the memory capacity of the system.
• Speed: Pandas is a single-threaded library by design, so working with large datasets might
make it slow. Certain tasks, including grouping and sorting, can take a very long time to
complete.
• Scalability: Pandas cannot readily grow over diferent nodes or a cluster since they are
not optimized to fully exploit distributed or parallel systems.</p>
      <p>
        Over the years, in response to the need to process large amounts of data eficiently, several
libraries have been developed to overcome this limitation. There has been limited research
conducted in the field of optimizing data preparation pipelines in DataFrame systems. Previous
studies have mainly focused on evaluating diferent libraries on data processing tasks. First,
Petersohn et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] proposed a technique that improves the eficiency and scalability of data
processing in parallel DataFrame systems. Second, Watson et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] introduced a data science
benchmark specifically designed to evaluate the performance of systems handling data
processing and analytics tasks. Furthermore, it is important to note that the existing literature often
refers to individual articles presenting specific libraries, rather than providing a comprehensive
comparison of these libraries in the context of data preparation pipelines [
        <xref ref-type="bibr" rid="ref3 ref7 ref8">3, 7, 8</xref>
        ]. Although
these articles help to provide valuable insights into the capabilities of individual libraries, a
comprehensive evaluation and comparison of libraries in the context of a complete data
preparation pipeline remains limited. The DFBen framework, introduced in the 2 section, aims to fill
this gap by conducting a systematic evaluation and comparison of various libraries in terms of
their performance and efectiveness in executing data preparation pipelines.
      </p>
      <p>DFBen focuses on investigating various crucial aspects of these pipelines, enabling researchers
and practitioners to make informed decisions. Firstly, the investigation of lazy evaluation and
its impact on overall performance is essential, as it ofers advantages such as reduced
memory consumption and improved processing eficiency. Secondly, scalability is a critical factor,
considering the exponential growth of data volumes. Assessing the libraries’ ability to scale
eficiently and utilize computational resources is crucial for real-world implementation.
Moreover, the choice of I/O formats in data preparation plays a significant role, as diferent formats
have distinct characteristics related to storage eficiency, data compression, and compatibility
with downstream processing. Lastly, the setup and configuration requirements of each library
significantly influence the ease of adoption and integration into existing pipelines.</p>
      <p>In Section 3, I present future research activities, including transitioning to a FaaS
(Function-asa-Service) framework and utilizing machine learning models to determine the ideal combination
of libraries based on input dataset features. These innovations promise to enhance the eficiency
and accuracy of data preparation pipelines.</p>
    </sec>
    <sec id="sec-2">
      <title>2. DFBen: DataFrame Evaluation Framework</title>
      <p>DFBen is an open-source project that aims to compare various commercial libraries, such as
Pandas1, Dask2, Spark3, Modin4, Polars5, Vaex6, and Rapids7, that utilize DataFrames for data
preparation tasks. In order to determine how much the libraries have improved on the standard
set by Pandas, this project conducts a comparison of the most commonly used functions grouped
into four stages of data preparation. To thoroughly assess the libraries’ potential, tests were
conducted on a number of distinct datasets with diferent sizes, schemas, and types.</p>
      <p>To identify the most efective libraries for constructing data preparation pipelines, I evaluate
their performance in four key areas. First, I assess the performance of individual functions used
for data preparation. Next, I test libraries in the four key stages of the data preparation pipeline:
Input/Output (I/O), Exploratory Data Analysis (EDA), Data Transformation, and Data Cleaning.
By looking at the performance of the libraries in each of these areas, I can identify the best
ones for each stage of the pipeline. Lastly, I analyze how well libraries perform across the full
data preparation pipeline. In addition, I investigate the impact of scalability on the performance
of these libraries. This enables me to assess their efectiveness in handling large datasets and
complex pipelines.</p>
      <p>The project’s open-source nature enables seamless incorporation of additional libraries
and preparators through a Library Class that encapsulates the library with the implemented
preparators, as well as a Dataset Class that allows for the inclusion of additional datasets in the
configuration file. Moreover, the pipelines are managed using a JSON file, enabling the easy
addition and modification of existing pipelines. In its current state, DFBen is a helpful tool for
anyone looking to compare the many DataFrame librariesavailable and select the one that best
ifts their needs and hardware capabilities in terms of performance and scalability.</p>
      <sec id="sec-2-1">
        <title>2.1. Framework Design</title>
        <p>
          Data preparation involves multiple stages and is typically performed using preparators, which
are responsible for small-scale preparation steps. In line with Naumann et al.’s (2020) Data
Preparation: A Survey of Commercial Tools [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], similar to data preparation tools, libraries often
ofer a range of preparator implementations for sequential application or the creation of data
preparation pipelines. The mentioned paper involved the selection of 40 commonly preparators,
which I subsequently implemented within the DFBen framework.
        </p>
        <p>To ensure optimal pipeline design, I select one of the top three Kaggle8 most voted notebooks,
specifically tailored to the pre-selected dataset. This approach leverages the collective expertise
of the data science community and ensures that the pipeline design is based on best practices
1https://pandas.pydata.org/.
2https://www.dask.org/.
3https://spark.apache.org/.
4https://modin.readthedocs.io/en/stable/.
5https://pandas.pydata.org/.
6https://vaex.io/docs/.
7https://rapids.ai/.
8https://www.kaggle.com/.
and state-of-the-art techniques. To provide an evaluation with diferent levels of granularity, as
described in Section 2.3, I organized the preparators into four key stages:
• I/O: this stage includes the functions that handle the input and output of data in various
formats, such as databases, CSV or JSON files and web APIs;
• EDA: here, the functions are grouped that deal with the exploratory analysis of data
characteristics for better understanding and detecting any error or anomaly;
• Data Transformation: at this stage, functions deal with transformation, such as data
normalization, deduplication, categorical data encoding, data aggregation, and other
techniques that make the data more suitable for analysis;
• Data Cleaning: the final stage involves data cleaning, which may include handling missing
or erroneous data, correcting outliers, eliminating outliers, and other techniques that
ensure data quality and reliability of analysis results.</p>
        <p>
          The datasets selection for this study, as described in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], was driven by their suitability for
evaluating the performance and capabilities of the libraries in various contexts. The first dataset,
120 years of Olympic history-athletes and results presents a comprehensive historical record
of all Olympic Games held between 1896 and 2016. This dataset is ideal for testing a lot of
preparatrs since its popularity on kaggle. The second dataset, NYC Yellow Taxi encompasses
detailed information about every taxi trip taken in New York City. To evaluate the scalability of
the libraries, two versions of this dataset are utilized: one containing data from 2015, and the
other encompassing data from 2009 to 2015. This allows for a comprehensive examination of
the libraries’ ability to handle large volumes of real-world data. To evaluate type inference and
handling of skewed and sparse string datasets, two additional datasets are included alongside
the mentioned datasets. The Loan Data dataset contains valuable information about loans issued
by a financial institution. The California State Patrol dataset encompasses information about
trafic stops conducted by the California Highway Patrol.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Implementation</title>
        <p>As described in Section 2.1, I used 50 preparators for my study. To implement these preparators,
I relied on functions defined in the libraries as well as custom functions designed to cover all
possible cases. To ensure optimal isolation, repeatability, and eficient management of machine
configurations, the code is executed within a Docker 9 container for each library.</p>
        <p>The tests are run on 3 machine configurations representing 3 diferent types of use: PC
(8 cores-16GB RAM), Workstation (16 cores-64GB RAM) and Server (24 cores-180 GB RAM).
In all configurations, the Tesla T4 GPU with 16GB RAM used by Rapids library is included.
In this way, users can test the framework in diferent contexts and verify its performance
on machines with diferent hardware configurations. Finally, the management of DFBen’s
configurations is entrusted to JSON files that can be easily configured by users, who can
customize the framework’s settings and the functions to be executed according to their own
needs and preferences.
103
()se
m
i
T
e
g
a
r
e
v
A102
load_dataset
fil_nan
sort</p>
        <p>edit
cast_columns_types</p>
        <p>Method</p>
        <p>Pipeline Step
Input</p>
        <p>EDA
data_cleaning</p>
        <p>output
data_transformation</p>
        <p>Step
Pipeline Full</p>
        <p>Library
dask
vaex
polars
pandas
pandas20
modin_dask
modin_ray
spark
rapids
dask
vaex
polars
pandas</p>
        <p>Frpaamndeaws2o0rk modin_dask
modin_ray
spark
rapids</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Evaluation</title>
        <p>
          For evaluating the performance, I track the execution time in diferent cases [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. It is important
to note that some libraries utilize lazy evaluation [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], which is a programming technique that
delays the execution of a computation or evaluation until it is actually needed. This is in contrast
to eager evaluation where computations are executed immediately regardless of whether the
result is actually used or not.
        </p>
        <p>Lazy evaluation can lead to more eficient use of resources as unnecessary computations are
avoided. When comparing libraries like Pandas, which do not use lazy evaluation, with those
that do, it is necessary to calculate the execution time accurately in order to have a fair and
objective comparison. In order to do that, the evaluation framework consists of three types of
tests:
• Core, is responsible for forcing the execution of lazy libraries and evaluating the
performance of individual operations. This allows an accurate assessment of the eficiency of
each library’s function and the identification of any performance bottlenecks.
• Pipeline-Step, executes a data preparation pipeline and calculates the execution time at
the conclusion of each stage. This method shows how each library’s function performs at
each stage of the pipeline and shows any ineficiencies or room for improvement.
• Pipeline-Full, component runs the entire data preparation pipeline. This approach provides
a comprehensive evaluation of each library’s performance and allows comparison of the
overall eficiency of completing the entire data preparation process.</p>
        <p>Figure 1 displays the initial result of the three types of evaluation explained above, which
have been obtained by running the tool on the Loan Dataset using a machine configuration
of 24 cores and 180 GB of RAM. However, it is important to note that these results are only in
the preliminary stage, and the analysis of the data is still ongoing. Further explorations will
be needed to provide more detailed information, considering factors beyond eficiency. The
efectiveness of the libraries involves various aspects that need to be evaluated. Therefore, it will
be crucial to consider the specific requirements of the dataset and employ a refined selection
process to identify the libraries that best meet the desired criteria.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Future Directions</title>
      <p>
        During the first few months of my PhD project, I mainly focused on creating an open-source
framework that enables the comparison of seven libraries for the construction of data preparation
pipelines. DFBen provides an opportunity for future research directions. In the next year, I plan
to work on two main research activities:
1. The first one can be enabled by the tool with the automatic selection of the most efective
combination of libraries for constructing pipelines. To address this, the tool will gather
information about the characteristics of the dataset and the required pipeline and, by
taking into account the execution times of individual functions, will be able to suggest
the proper library. This would remove the dificulty of manually selecting the appropriate
one for a given task.
2. The second research activity can be the transition of DFBen to a FaaS framework. One
reason for this is that implementing a data preparation pipeline requires iterating the
steps described in Section 2.1 several times. When resources are limited, a small sample
of the reference dataset may sufice for testing, but this may not always be representative,
and it may not show certain record types that would lead to errors. Frameworks such as
Spark currently dominate this context [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], but in the implementation phase, it may not
be necessary to use and configure a framework such as this due to the iterative nature of
the process. FaaS has the potential to revolutionize the way data transformation pipelines
are implemented by making it easier to manage and scale individual functions in response
to events or triggers. However, this approach also requires careful consideration of issues
such as security, data privacy, and performance optimization. [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]
      </p>
      <p>To achieve these objectives, I want to study and test various methods and approaches for
creating data transformation pipelines based on FaaS. Additionally, I will explore novel
approaches for incorporating machine learning to construct pipelines that combine libraries, thus
enhancing their efectiveness and eficiency. By doing so, I believe that this research will make
a valuable contribution to the field of data preparation pipelines and help advance the field of
data science as a whole.
I wish to express my gratitude to my tutor, Prof. Sonia Bergamaschi and my co-tutor, Giovanni
Simonini for their guidance and support throughout my PhD journey, as well as to my colleagues
at the DBGroup who have been a constant source of encouragement. In addition, I would like
to express my sincere appreciation to Leonardo for funding my research grant.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hameed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <article-title>Data preparation: A survey of commercial tools</article-title>
          ,
          <source>ACM SIGMOD Record</source>
          <volume>49</volume>
          (
          <year>2020</year>
          )
          <fpage>18</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Gagliardelli</surname>
          </string-name>
          , G. Papadakis, G. Simonini,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          , T. Palpanas,
          <article-title>Generalized supervised meta-blocking</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>15</volume>
          (
          <year>2022</year>
          )
          <fpage>1902</fpage>
          -
          <lpage>1910</lpage>
          . URL: https://www. vldb.org/pvldb/vol15/p1902-gagliardelli.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Petersohn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Macke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Xin</surname>
          </string-name>
          , W. Ma,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Mo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Hellerstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Joseph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Parameswaran</surname>
          </string-name>
          ,
          <article-title>Towards scalable dataframe systems</article-title>
          , arXiv preprint arXiv:
          <year>2001</year>
          .
          <volume>00888</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>W. McKinney</surname>
          </string-name>
          <article-title>, pandas: a foundational python library for data analysis and statistics, Python for high performance and scientific computing 14 (</article-title>
          <year>2011</year>
          )
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Petersohn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Durrani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Melik-Adamyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Joseph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Parameswaran</surname>
          </string-name>
          ,
          <article-title>Flexible rule-based decomposition and metadata independence in modin: a parallel dataframe system</article-title>
          ,
          <source>Proceedings of the VLDB Endowment</source>
          <volume>15</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Watson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S. V.</given-names>
            <surname>Babu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ray</surname>
          </string-name>
          ,
          <article-title>Sanzu: A data science benchmark</article-title>
          ,
          <source>in: Procedings International Conference on Big Data (Big Data)</source>
          , IEEE,
          <year>2017</year>
          , pp.
          <fpage>263</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Breddels</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Veljanoski</surname>
          </string-name>
          ,
          <article-title>Vaex: big data exploration in the era of gaia</article-title>
          ,
          <source>Astronomy &amp; Astrophysics</source>
          <volume>618</volume>
          (
          <year>2018</year>
          )
          <article-title>A13</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rahmany</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Zin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Sundararajan</surname>
          </string-name>
          ,
          <article-title>Comparing tools provided by python and r for exploratory data analysis</article-title>
          ,
          <source>IJISCS (International Journal of Information System and Computer Science</source>
          )
          <volume>4</volume>
          (
          <year>2020</year>
          )
          <fpage>131</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bloss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hudak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Young</surname>
          </string-name>
          ,
          <article-title>Code optimizations for lazy evaluation</article-title>
          ,
          <source>Lisp and Symbolic Computation</source>
          <volume>1</volume>
          (
          <year>1988</year>
          )
          <fpage>147</fpage>
          -
          <lpage>164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Werner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kuhlenkamp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Klems</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tai</surname>
          </string-name>
          ,
          <article-title>Serverless big data processing using matrix multiplication as example</article-title>
          ,
          <source>in: Procedings International Conference on Big Data (Big Data)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>358</fpage>
          -
          <lpage>365</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Sreekanti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. C.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schleier-Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Faleiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Hellerstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tumanov</surname>
          </string-name>
          , Cloudburst:
          <article-title>Stateful functions-as-a-service</article-title>
          , arXiv preprint arXiv:
          <year>2001</year>
          .
          <volume>04592</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H.</given-names>
            <surname>Shafiei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Khonsari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mousavi</surname>
          </string-name>
          ,
          <article-title>Serverless computing: a survey of opportunities, challenges, and applications</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>54</volume>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>