<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On Testing a Data Access Architecture Based on Micro-services</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michał Bodziony</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Filip Szóstak</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Konrad Tarnacki</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robert Wrembel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IBM Poland Software Lab Kraków</institution>
          ,
          <country country="PL">Poland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Poznan University of Technology</institution>
          ,
          <addr-line>Poznań</addr-line>
          ,
          <country country="PL">Poland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <abstract>
        <p>Architectures that integrate or transport data between systems are more and more frequently built on microservices, i.e., software components that are independent on each other but communicate over a network to provide some services. They ofer advantages like increased flexibility of the architecture and scalability, but the price to pay for it are increased communication costs. In this paper we evaluate the performance of three architectures (varying in complexities) based on micro-services and relate their eficiency (data reading time and throughput) to a base-line architecture (client-server). The paper presents conclusions from a joint project between IBM Software Lab Kraków and Poznan University of Technology.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;data integration</kwd>
        <kwd>data movement</kwd>
        <kwd>data source connector</kwd>
        <kwd>micro-service</kwd>
        <kwd>container</kwd>
        <kwd>Docker</kwd>
        <kwd>Kubernetes</kwd>
        <kwd>performance</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction and motivation</title>
      <p>For over six decades, Data integration (DI) has been an active research domain [1, 2], with the common
objective of providing end-users with a unified view on heterogeneous and typically distributed data. DI
research has led to the development of a few standard DI architectures, namely: federated [3], mediated
[4], data warehouse [5], lambda [6], data lake [7], data lakehouse [8], polystore [9], data mesh [10], and
recently - data spaces [11, 12]. In each of these architectures, an integration layer is an indispensable
component that facilitates the transportation of data from source storage systems into an integrated
destination system. This layer is implemented through sophisticated software that executes DI processes.
A DI process typically involves a sequence of tasks to ingest data from diverse sources and pre-process
them into formats suitable for analytical and machine learning applications. A variety of commercial
and open-source DI tools are available on the market [13].</p>
      <p>In recent years, architectures based on micro-services are frequently adopted to facilitate the
integration and transportation of large data volumes across heterogeneous systems [14, 15, 16, 17, 18]. Such
novel architectures present a viable alternative to monolithic approaches, which typically are based on
data silos [19, 20].</p>
      <p>Micro-services [21] are software components that are independent on each other and communicate
over a network to provide some services, like access to distributed and heterogeneous data sources,
data pre-processing. This modularity allows for independent development, deployment, management,
extensions with new functionalities (e.g., by deploying a new service), and scaling of individual services.
Each service can be managed independently, based on its specific needs (e.g., resource allocation policy).
By decomposing an integration process into a set of fine-grained services, each service is responsible
for a specific integration task (e.g., data ingestion, cleaning, transformation). Thus, micro-services ofer
alternative and more flexible solutions to many of the limitations of monolithic approaches. Moreover,
diferent services can be built using the most appropriate technology stack (e.g., a programming language,
data storage system) for their specific function. The loosely coupled nature of micro-services makes a
system more fault resistant, i.e., if one service fails, it is less likely to bring down the entire integration
pipeline. Micro-services communicate with each other over a network, typically using APIs or message
brokers.</p>
      <p>Data integration by means of micro-services ofers a few advantages as compared to standard
centralized data integration approaches. First, each micro-service is responsible for making available
data from its own DS. Second, it may encapsulate data access and privacy protection rules. Third, a
particular service can be scaled up or down (depending on a data volume being processed) independently
on other components of the whole integrated system. Thus, micro-services allow to build a modular
DI architecture, where each implemented module is a small and independent service that executes a
particular data integration tasks.</p>
      <p>Micro-services share the same design principle with data mesh and data spaces, which is
decentralization. While micro-services changed the technology for building applications by breaking down
monoliths, data mesh and data spaces aimed at breaking data silos into domain-specific sources. For
this reason, micro-services are particularly attractive to building data mesh and data spaces, where
every independent data source is encapsulated in one or more specialized services, each of which ofers
a particular data set and ofers specific data processing means.</p>
      <p>From a technological perspective, micro-services cooperate with modern specialized software like
Airbyte and Apache Arrow Flight, which provide means for eficient transportation of large data
volumes. The increasing popularity of novel DI architectures and software tools to implement them
creates demand for benchmarks, which could evaluate their performance. Organizations require tools
that not only assist in selecting the best technologies, but also help identify issues with solutions that
have already been implemented.</p>
      <p>
        Unfortunately, the micro-services architecture pays the price of lower performance due to time
overhead. The overhead is caused by: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) increased network communication between services, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) data
format conversions between services (e.g., from the Airbyte format to the columnar format of Arrow
Flight, and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) time consumed by instantiating containers and PODs (outlined in Section 2).
      </p>
      <p>In this short paper we report our findings on the performance of a simple data mesh architecture,
implemented by means of micro-services, as reported in Section 4. The performance evaluation was
done within a joint project run by IBM Software Lab Kraków and Poznan University of Technology.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Technologies</title>
      <p>In this section we outline modern technologies for building DI architectures based on micro-services.
We applied these technologies in the evaluation described in this paper. The technologies include:
Airbyte, Arrow, Arrow Flight, Cloud Pak for Data, and Docker.</p>
      <p>Apache Airbyte is an alternative tool for data extraction and transportation [22]. The heart of Airbyte
is its large library of connectors - it ofers over 600 pre-built connectors for data source and destination
systems. The tool includes also API for building new connectors.</p>
      <p>Apache Arrow is a platform for software development that provides means for transporting large
data volumes between systems. It uses a standardized in-memory columnar format for representing
table-like data [23, 24]. The columnar format storage allows eficient data processing by applying Single
Instruction Multiple Data [25] and compression [26]. For this reason, it is frequently used in analytical
systems (like data warehouses) and architectures for data integration and data transportation. If systems
are located on the same node, Arrow provides out-of-the-box zero-copy data sharing that allows one
system to make available data for usage directly to another system, by means of shared memory.</p>
      <p>Data in the Arrow columnar format can be transported in bulk in a client-server architecture using
Apache Arrow Flight, which uses the Google Remote Procedure Call (gRPC) mechanism [27]. To further
increase data throughput, Arrow Flight groups data into structures called RecordBatches and uses
parallel transportation (it was reported that Arrow Flight was able to achieve 20 Gb/s throughput per
core [28]). Both technologies can be integrated with another component of a DI architecture, namely
IBM Cloud Pak for Data.</p>
      <p>Cloud Pak for Data [29, 30] is an ecosystem of software services for integrating, storing, managing,
preparing for analysis and analyzing data. It is available either as an on-premise software or as a service
managed on IBM Cloud. An important feature of this software is its extensibility with the set of data
sources that can be accessed by Cloud Pak for Data. To this end, a development toolkit called Connector
SDK is used for creating connectors to new data sources [31]. The SDK generates skeleton files for
developing connectors in Java and wraps them within an Arrow Flight server.</p>
      <p>Micro-services are developed with the support of containerization that is a lightweight form of
virtualization, where an application’s executable code, all the needed libraries, and other files are
packaged into a unit, called a container. Containers represent separate components of a DI architecture,
like a database, a data ingestion component, a data wrangling component, an analytical application.
Two primary technologies are used for containerization, namely Docker and Kubernetes. Docker [32] is
used for building and running individual containers, whereas Kubernetes is an orchestration tool that
automates the deployment, scaling, and management of containerized applications [33]. To manage
containers, Kubernetes wraps them into units called PODs.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Related research</title>
      <p>Benchmarking data processing technologies and architectures, like on-line transaction processing
(OLTP) databases, on-line analytical systems (OLAP) - mostly data warehouses, and stream processing
systems for years has been an active field of research and technological development. These eforts
resulted in several benchmarks recognized by industry and research. The core standard includes the
family of TPC benchmarks (tpc.org) for evaluating OLTP, OLAP, big data, the Internet of Things,
and Artificial Intelligence solutions. There are also methods for evaluating specialized systems, like
spatio-temporal data warehouses (e.g., [34, 35]) or graph databases (e.g., [36, 37, 38]).</p>
      <p>In the context of the experimental evaluation of data integration/ data transportation architectures
based on micro-services few works has been published so far.</p>
      <p>
        [39] presented a benchmark for evaluating big data solutions that apply containerization. Performance
metrics included processing time, throughput, resource usage, and latency. The data from HiBench
(a big data benchmark suite) were used. Paper [40] presented the performance evaluation of Arrow
Flight on a high-performance computing cluster, parameterized by the degree of parallelism and data
volume. The evaluation is based on two metrics, namely: throughput and the total time required to
query a dataset. The NYC Taxi dataset was used in the tests. [41] presented a system for genomics data
processing on computational clusters, based on Arrow - for data storage and Arrow Flight - for data
transportation. The evaluation included: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) runtime of data shufling, i.e., transferring data between
a cluster nodes, duplicate removal, and overall data pre-processing; (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) cluster scalability w.r.t. the
number of nodes; (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) the impact of memory size on performance; and (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) the throughput of Arrow Flight
w.r.t. the number of Flight connections to the cluster. The standard genomic benchmarking set NA12878
was used in the evaluation.
      </p>
      <p>
        Despite the aforementioned eforts in testing architectures based on micro-services, to the best of our
knowledge there is no evaluation that would clearly answered questions: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) what is the time overhead
for deploying components in the DI architectures, like Docker containers and Kubernetes PODs, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
what is the time spent on reading data varying in size, and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) what is the throughput and its stability
w.r.t. varying data size. The outcomes from the project reported in this paper aim at answering these
questions.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Tested architecture</title>
      <p>
        For the purpose of this project, a simple architecture was built, with the following components used as
building blocks: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) Airbyte, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) Arrow Flight - serving as a layer for moving data in a columnar format,
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) Docker and Kubernetes, (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) IBM Cloud Pak for Data, (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) data sources - a relational database PostgreSQL,
a non-relational storage MongoDB, and Parquet files.
      </p>
      <p>As illustrated in Figure 1, four architectural variants were tested. Variant A served as the simplest
baseline architecture, where the Cloud Pak for Data client established a direct connection to a data
source using a native connector. Specifically, Psycopg2 was used for PostgreSQL, PyMongo for MongoDB,
and Pandas for Parquet.
variant A
variant B
variant C
variant D</p>
      <p>native
connector
Docker container</p>
      <p>Airbyte
connector
POD
Docker container</p>
      <p>Airbyte
connector
POD
Docker container</p>
      <p>Airbyte
connector</p>
      <p>Cloud Pak
for Data
Cloud Pak
for Data
Cloud Pak
for Data
port
900n</p>
      <p>POD
Docker container</p>
      <p>Arrow Flight
server</p>
      <p>Arrow
Flight
API</p>
      <p>Cloud Pak
for Data</p>
      <p>In variant B, the Airbyte source connector was used to connect to the data source by means of a
dedicated connector for each of the three aforementioned data sources. The connectors were instantiated
within a Docker container. In variant C, the connectors were instantiated in a Kubernetes POD [42].
Finally, in variant D, Arrow Flight was used as an intermediate layer between the connector and Cloud
Pak for Data. The variant represents a typical data transportation architecture for large data volumes.</p>
      <p>By extending the complexity of the architectures from B to D (by adding new components) we were
able to measure time overhead introduced by each of these new components.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Experimental evaluation</title>
      <p>The experiments were run on a workstation with: Intel Core i7-7700HQ 2.8 GHz, 16 GB of DDR4 RAM
2400 MHz, and an SSD disk, on Ubuntu 22.04 LTS.</p>
      <p>Each experiment was repeated 10 times. To mitigate the impact of outliers, the minimum and
maximum values were discarded, and the average was computed from the remaining 8 measurements.
This average is presented in the subsequent charts. The batch size for Arrow Flight was equal to 100000
rows, which we found the most eficient for the whole evaluation.</p>
      <p>
        Test data came from the the TPC-DS benchmark [43]. Specifically, the data were taken from table
store_sales (23 attributes of types integer or decimal, average record length = 164 B). The following four
data volumes were used in the experiments: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) 1000 records (0.2 MB), 500000 records (78 MB), 1000000
(156 MB), 2000000 records (313 MB), and 4000000 records (625 MB).
      </p>
      <p>The overall measured time for variant A comprised connecting to the data source, reading data, and
returning the data as a stream of rows. Variant B introduced the additional overhead of container
creation, variant C added POD creation, and variant D further included the creation of two PODs and
data translation to the Arrow Flight format.</p>
      <p>
        The goal of these experiments was to assess: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) the time overhead for deploying additional
components of the architectures, like Docker containers and Kubernetes PODs, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) time spent on reading data
varying in size, and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) throughput and its stability w.r.t. varying data size.
      </p>
      <p>The source codes of all the implemented components are available at GitHub: https://github.com/
Cheriit/airbyte-arrow-integration.</p>
      <sec id="sec-5-1">
        <title>5.1. Time overhead for deploying components</title>
        <p>To quantify the deployment time of additional architectural components, a query selecting a single
record from the smallest dataset was executed from Cloud Pak for Data. The overall query execution
times, from initiation to result retrieval, for three distinct data sources are presented in Figure 2.</p>
        <p>The results indicated that the architecture in variant A of the tested architecture was the most
eficient, likely due to the C implementation (thus eficient) of its native connectors, contrasting with
the multi-language implementation of the Airbyte connectors. Moreover, variants B, C, and D sufered
a performance penalty from the mandatory conversion of ingested data to the JSON format required by
Airbyte. The total measured time of variant C was further increased by POD deployment and connection
setup. Finally, the performance of variant D was additionally impacted by the deployment of an Arrow
Flight server in a separate POD and the subsequent data conversion from the row-oriented Airbyte
format to the columnar Arrow format.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Reading data of varying sizes</title>
        <p>In this experiment, four data volumes difering in size were read by means of the four tested architectures.
Figure 3 shows the results where PostgreSQL served as the data source. Similarly as in the previous
experiment, variant A is the fastest, and variant D is the slowest.</p>
        <p>Figure 4 shows the results where MongoDB served as the data source. Again, variant A is the fastest,
and variant D is the slowest. Notice, that similar characteristics was observed for Parquet serving as a
data source.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Throughput</title>
        <p>
          This experiment aimed at: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) measuring the throughput ofered by the four tested architectures and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
verifying whether the throughput remains stable when a transmitted data volume increases (in case of
memory leaks and malfunctioning communications between the components such a behavior could be
observed).
        </p>
        <p>The obtained results for PostgreSQL are shown in Figure 5. Variant A ofers the highest throughput
and variant D - the lowest. Moreover, the throughput remains stable for the tested data volumes.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Summary and conclusions</title>
      <p>A data integration architecture built with the support of micro-services has a few advantages (see Section
1) but its main drawback is additional time overhead added by multiple components of the architecture,
like Docker containers and Kubernetes PODs, establishing connections between the components, and data
format conversions. The experiments showed that native connectors (in the client-server architecture)
ofered the best performance. The more complex the micro-services architecture is the less eficient it
is. The findings from this project have been included into the IBM Knowledge Base.</p>
      <p>
        The project allowed us to identify paths for future works, which include: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) eficient compression
techniques for data in the Arrow format, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) caching techniques for data returned by deployment
commands of the components (e.g., discovery), (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) techniques for re-usability of PODs.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>The author(s) have not employed any Generative AI tools.
and Analytical Processing of Big Data (DOLAP) @EDBT/ICDT, volume 3931, CEUR-WS.org, 2025.
[13] Gartner data integration tools reviews and ratings, 2025. URL: https://www.gartner.com/reviews/
market/data-integration-tools, accessed Jan, 2026.
[14] D. Chia, Scaling data pipelines on kubernetes, 2020. URL: https://airbyte.com/blog/
scaling-data-pipelines-kubernetes, accessed Jan, 2026.
[15] Cloud Native Computing Foundation annual survey, 2022. URL: https://www.cncf.io/reports/
cncf-annual-survey-2022/, accessed Jan, 2026.
[16] S. Singh, C. H. Muntean, S. Gupta, Resilient microservices: an investigation into istio efectiveness
in kubernetes, Cluster Computing 29 (2026).
[17] S. Zeng, H. Zhang, Z. Yan, CASLO: joint scaling and deployment for microservices leveraging
context-aware SLO assignment, Journal of Network and Computer Applications 247 (2026).
[18] A. C. M. Saucedo, G. Rodríguez, F. G. Rocha, R. P. dos Santos, Migration of monolithic systems to
microservices: A systematic mapping study, Information &amp; Software Technology 177 (2025).
[19] O. Bruns, L. Söhn, T. Tietz, J. J. Steller, E. Posthumus, T. Schrade, H. Sack, Gotta catch’em all: From
data silos to a knowledge graph, in: Satellite Events @ Semantic Web (ESWC), volume 15344 of
Lecture Notes in Computer Science, Springer, 2024.
[20] M. Sienkiewicz, From data silos to data mesh: A case study in financial data architecture, in: Int.</p>
      <p>Conf. on Database and Expert Systems Applications (DEXA), volume 16046 of LNCS, Springer,
2025.
[21] IBM Developer. Build software with fine-grained, loosely coupled services, 2025. URL: https:
//developer.ibm.com/depmodels/microservices/, accessed Jan, 2026.
[22] Airbyte. Hundreds of connectors out-of-the-box, 2025. URL: https://airbyte.com/connectors,
accessed Jan, 2026.
[23] W. McKinney, Introducing apache arrow flight: A framework for fast data transport, 2019. URL:
https://arrow.apache.org/blog/2019/10/13/introducing-arrow-flight/, accessed Jan, 2026.
[24] What is Apache Arrow? Capabilities &amp; benefits, 2018. URL: https://www.dremio.com/resources/
guides/apache-arrow/, dremio; accessed Jan, 2026.
[25] J. Hildebrandt, J. Pietrzyk, A. Krause, D. Habich, W. Lehner, Partition-based SIMD processing
and its application to columnar database systems, Datenbank-Spektrum 23 (2023). URL: https:
//doi.org/10.1007/s13222-022-00431-0.
[26] M. Kuschewski, D. Sauerwein, A. Alhomssi, V. Leis, Btrblocks: Eficient columnar compression for
data lakes, ACM on Management of Data 1 (2023).
[27] gRPC documentation, 2021. URL: https://grpc.io/docs/, accessed Jan, 2026.
[28] It’s time to replace ODBC &amp; JDBC, 2019. URL: https://www.dremio.com/blog/
is-time-to-replace-odbc-jdbc/, dremio; accessed Jan, 2026.
[29] IBM Documentation. Overview of Cloud Pak for Data, 2025. URL: https://www.ibm.com/docs/en/
cloud-paks/cp-data/5.3.x?topic=overview-cloud-pak-data, accessed Jan, 2026.
[30] IBM Cloud Pak for Data, ???? URL: https://www.ibm.com/products/cloud-pak-for-data.
[31] IBM Cloud Pak for Data - Connector SDK, 2025. URL: https://github.com/IBM/cp4d-connector-sdk,
accessed Jan, 2026.
[32] A safer container ecosystem, for everyone, 2026. URL: https://www.docker.com/, accessed Jan,
2026.
[33] Production-grade container orchestration, 2026. URL: https://kubernetes.io/, accessed Jan, 2026.
[34] C. Düntgen, T. Behr, R. H. Güting, Berlinmod: a benchmark for moving object databases, VLDB</p>
      <p>Journal 18 (2009).
[35] V. Pandey, A. Kipf, T. Neumann, A. Kemper, How good are modern spatial analytics systems?,</p>
      <p>VLDB Endowment 11 (2018).
[36] Graph benchmark. Comparing and understanding graph databases, ????. URL: https://
graphbenchmark.com/, accessed Jan, 2026.
[37] Open graph benchmark, ????. URL: https://ogb.stanford.edu/, accessed Jan, 2026.
[38] Graph Data Council, ????. URL: https://ldbcouncil.org/, accessed Jan, 2026.
[39] S. Ceesay, A. Barker, B. Varghese, Plug and play bench: Simplifying big data benchmarking using
containers, in: Int. Conf. on Big Data (BigData), IEEE, 2017.
[40] T. Ahmad, Benchmarking Apache Arrow Flight - a wire-speed protocol for data transfer, querying
and microservices, in: Benchmarking in the Data Center: Expanding to the Cloud (BID@PPOPP),
ACM, 2022.
[41] T. Ahmad, C. Ma, Z. Al-Ars, H. P. Hofstee, Communication-eficient cluster scalable genomics
data processing using Apache Arrow Flight, in: Int. Symp. on Parallel and Distributed Computing
(ISPDC), IEEE, 2022.
[42] Kubernetes documentation. PODs, 2025. URL: https://kubernetes.io/docs/concepts/workloads/
pods/, accessed Jan, 2026.
[43] TPC Benchmark DS - standard specification 3.2.0, 2021. URL: https://www.tpc.org/TPC_
Documents_Current_Versions/pdf/TPC-DS_v3.2.0.pdf, accessed Jan, 2026.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Timakum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Song</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Song, DOLAP: A 25 year journey through research trends and performance (invited talk)</article-title>
          ,
          <source>in: Int. Workshop on Design, Optimization, Languages and Analytical Processing of Big Data (DOLAP)</source>
          , volume
          <volume>3653</volume>
          , CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Wrembel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abelló</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Song</surname>
          </string-name>
          ,
          <article-title>DOLAP data warehouse research over two decades: Trends and challenges</article-title>
          ,
          <source>Information Systems</source>
          <volume>85</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Elmagarmid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rusinkiewicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sheth</surname>
          </string-name>
          , Management of Heterogeneous and Autonomous Database Systems, Morgan Kaufmann Publishers, ISBN 1-55860-216-X,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Wiederhold</surname>
          </string-name>
          ,
          <article-title>Mediators in the architecture of future information systems</article-title>
          ,
          <source>Computer</source>
          <volume>25</volume>
          (
          <year>1992</year>
          )
          <fpage>38</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Vaisman</surname>
          </string-name>
          , E. Zimányi,
          <source>Data Warehouse Systems - Design and Implementation</source>
          ,
          <string-name>
            <surname>Second Edition</surname>
          </string-name>
          ,
          <source>Data-Centric Systems and Applications</source>
          , Springer,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gillet</surname>
          </string-name>
          , É. Leclercq,
          <string-name>
            <given-names>N.</given-names>
            <surname>Cullot</surname>
          </string-name>
          ,
          <article-title>Lambda+, the renewal of the lambda architecture: Category theory to the rescue</article-title>
          ,
          <source>in: Int. Conf. Advanced Information Systems Engineering (CAiSE)</source>
          ,
          <source>LNCS 12751</source>
          , Springer,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Hai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Koutras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Quix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jarke</surname>
          </string-name>
          ,
          <article-title>Data lakes: A survey of functions and systems</article-title>
          ,
          <source>IEEE Trans. Knowl. Data Eng</source>
          .
          <volume>35</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Errami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hajji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. A. E.</given-names>
            <surname>Kadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Badir</surname>
          </string-name>
          ,
          <article-title>Spatial big data architecture: From data warehouses and data lakes to the lakehouse</article-title>
          ,
          <source>Journal of Parallel and Distributed Computing</source>
          <volume>176</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Chirkova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Gadepally</surname>
          </string-name>
          , T. G. Mattson,
          <article-title>Enabling query processing across heterogeneous data models: A survey</article-title>
          ,
          <source>in: IEEE Int. Conf. on Big Data</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dehghani</surname>
          </string-name>
          , Data Mesh:
          <article-title>Delivering Data-Driven Value at Scale</article-title>
          ,
          <source>O'Reilly</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Ayala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bilalli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gómez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mazón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Romero</surname>
          </string-name>
          ,
          <article-title>Challenges to enforce data quality in data spaces</article-title>
          ,
          <source>in: Int. Workshop on Design, Optimization, Languages and Analytical Processing of Big Data (DOLAP) @EDBT/ICDT</source>
          , volume
          <volume>3931</volume>
          , CEUR-WS.org,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Morejón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Berenguer</surname>
          </string-name>
          , L. de Espona,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tomás</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mazón</surname>
          </string-name>
          ,
          <article-title>Exploring content-based catalogs for enhanced discovery services in data spaces</article-title>
          ,
          <source>in: Int. Workshop on Design, Optimization</source>
          , Languages
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>