<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On Three Missing Pieces in the Data Integration Puzzle (Position Paper)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Robert Wrembel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Poznan University of Technology</institution>
          ,
          <addr-line>Poznań</addr-line>
          ,
          <country country="PL">Poland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Data integration (DI) for years has been among the most frequently researched topics. A common goal of DI is to make heterogeneous and distributed data available for an end user in a unified format that is suitable for analyses. Research and development works resulted in a few standard DI architectures. In all of them data are moved from data sources (DSs) into an integrated system by means of an integration layer. This layer runs DI processes that are complex workflows composed of multiple tasks responsible for extracting data from DSs and making them available for analytics. Methods for developing DI processes have been researched and developed for decades, but designing and managing DI processes is still dificult and time costly. Moreover, the still open problems include: (1) the optimization of DI processes, (2) managing user-defined functions in DI processes, especially available as black-boxes, and (3) discovering and managing broken data lineage in a DI pipeline. In this position paper we formulate the three research hypotheses concerning the aforementioned problems and show how these hypotheses can be verified to develop the missing solutions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;data integration</kwd>
        <kwd>data integration processes</kwd>
        <kwd>user-defined functions</kwd>
        <kwd>data lineage</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction and motivation</title>
      <p>tion architectures (i.e., federated, mediated, or data mesh).</p>
      <p>
        DI processes are managed by a dedicated software, called
For years, the widespread of complex, data-driven sys- a DI engine (ETL engine in DW architectures) [11, 12, 13].
tems has been observed (e.g., in medicine, agriculture, Methods for designing DI processes have been
resmart cities). They produce huge volumes of highly het- searched and developed for decades (see [14, 15]). Despite
erogeneous data (a.k.a. big data). These data need to be these substantial number of works, designing and
manintegrated to feed various analytical and machine learn- aging DI processes is still dificult and time costly [ 16].
ing (ML) applications. Consequently, data integration (DI) Moreover, the still open fundamental research
probtechniques are core components of information systems. lems include: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) the optimization of DI processes, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
DI architectures and processes are among very frequently managing user-defined functions (UFDs) in DI processes,
researched topics [1, 2, 3, 4]. and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) managing data lineage in a DI pipeline. In this
      </p>
      <p>
        DI aims at consolidating data from multiple distributed position paper we discuss possible solutions to these
proband heterogeneous data sources (DSs), to deliver these lems. This discussion is motivated by our cooperation
data to an end user in a common format. Research and with the IT sector and with the banking sector.
development works resulted in a few standard DI
architectures, namely: federated [5] and mediated [6], data
warehouse (DW) [7], data lake (DL) [8], data lake house 2. Related research
(DLH) [9], and (7) data mesh [10]. In all of these
architectures, data are moved from DSs into an integrated 2.1. Optimization of DI processes
system by means of an integration layer. This layer is Any DI process has to finish its work within a given
implemented by a sophisticated software, which runs the time window, typically a few hours, in order to make a
so-called DI processes. DW/DL/DLH available for users. Because a DI process (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
      </p>
      <p>
        DI processes range from simple to complex workflows moves large volumes of heterogeneous data between DSs
of dozens to thousands of tasks. These tasks are responsi- and a DW/DL/DLH and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) executes complex data
proble for extracting data from DSs, transforming data into cessing algorithms, its execution is time costly. Therefore,
a common model and data structures, cleaning data, and providing solutions for eficient execution of DI processes
loading them into either a central repository (i.e., a DW, (a.k.a. DI process optimization) is of high importance,
DL, or DLH) or making them available in virtual integra- since it is crucial for the whole DI architecture. Big data
Published in the Proceedings of the Workshops of the EDBT/ICDT 2025 (with their volume and complexity) further increase the
Joint Conference (March 25-28, 2025), Barcelona, Spain dificulty of DI process optimization. In order to reduce
$ robert.wrembel@cs.put.poznan.pl (R. Wrembel) the execution time of a DI process, a few classes of
solu0000-0©002012-56C0op3y7ri-gh5t7©12802(5Rfo.r Wthisrpeamperbbeyli)ts authors. Use permitted under tions have been proposed by industry and research.
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCreEatUiveRCoWmmoornksLsichenospeAPttrribouctioene4d.0iInntgersna(tiConEal U(CCRB-YW4.0S)..org) Industry approaches are based on the following
optimization techniques: scaling-up or scaling-out a data pets of a logic tailored to a specific and non-typical, data
integration server or adding a specialized hardware like processing problem. Thus, UDFs extend the functionality
FPGAs [17, 18], parallel processing of DI tasks, and mov- of a DI architecture with tailored tasks. UDFs represent
ing the execution of some DI tasks into a DS - this tech- code written in multiple programming languages,
externique is commonly called push-down [19, 20]. nal to a DI design environment and a DI engine. Typically,
      </p>
      <p>
        Research approaches draw upon the following ideas. such UDFs are treated by a DI engine as black-boxes
(furThe first class of solutions is based on changing the ther called black-box UDFs - BBUDFs), i.e., code snippets
order of DI tasks (a.k.a. task reordering), so that a whose internal logic and performance characteristics are
new reordered process is more eficient that the original unknown.
one [21, 22, 23, 24]. These methods are computationally Multiple works addressed the usage of UDFs in data
too expensive (exponential) [21] - the search space of engineering. They can be categorized as: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) explicit code
all possible valid orders of tasks in such processes is annotations by a programmer [33, 34, 35, 36], (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) static
too large to be fully searched. As a consequence, the code analysis and debugging [33, 37, 38], (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) eficient
applicability of these methods is limited to rather simple compilation and execution [39, 40, 41, 42], (4) UDFs in
DI processes. Moreover, none of the aforementioned database queries [43, 44, 45, 46], and (5) UDFs in data
approaches supports the optimization of DI processes pipelines [36, 47, 48]. Most of these solutions assume that
with UDFs. The reordering of operators is based on the UDFs are treated as white-boxes (their implementation
semantics of the operators, which is well known and code is accessible), which prevents their applications to
understood for traditional relational operators. On the BBUDFs.
contrary, the semantics of UDFs is frequently unknown. Using a BBUDF in a DI process makes such a process
      </p>
      <p>A special class of reordering is the push-down tech- dificult to optimize and manage (e.g., impact analysis).
nique. The principle of this technique is to move data First, because the internal logic of such a BBUDF is
unprocessing from a DI engine into a DS server for execu- known. Second, its performance characteristics have to
tion [19, 20, 25]. However, push-down so far has been be learned, which is a time consuming task. Third, it
researched and developed mainly for relational DSs. To causes data lineage (see Section 2.3) extremely
challengthe best of our knowledge, push-down for a few non- ing. For this reason, discovering the internal logic and
relational DSs is available in Informatica and Spark. In- performance characteristics of BBUDFs is of paramount
formatica is able to push-down filters, joins, and aggre- importance for building eficient DI processes with data
gations into a Hadoop cluster. Spark query programming lineage.
library allows to push-down filter predicates and
aggregations to Parquet, ORC, and a few relational databases 2.3. Data lineage in DI processes
connected to via a JDBC driver [26].</p>
      <p>
        Despite these solutions are available in software tools, Data lineage is a set of techniques that document the
still open problems concern among others: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) eficient lifecycle of data, including source information and any
implementation of a pushed-down operation in a DS, data transformations that have been applied within any
given its functionality and internal features, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) an over- DI process [49]. These techniques are crucial for data
all method for assessing profitability of push-down, (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) management, governance, and compliance, since they
extending push-down to key-value, column-family, doc- provide insights into where data came from and how
ument, and graph storage. they were processed along a data pipeline. Data lineage
      </p>
      <p>The second class of solutions uses parallel processing techniques have been researched for decades. They can
of either individual tasks within a process or portions be categorized according to a few taxonomies.
of the process [27, 28, 29, 30, 31]. A challenge in this The first one divides them into annotation-based and
approach is to figure out the most eficient parallelization inversion-based. In the annotation-based techniques,
scheme for a given DI task, or a subset of tasks, or the input data are annotated with information about their
whole DI process [32], especially when DI processes use origin and how they were processed along a data pipeline.
UDFs. Task parallelization has been researched quite The annotations are propagated throughout the whole
extensively and a few sound solutions to this challenge pipeline to be included in final results [ 50, 51, 52, 53,
have been proposed. 54]. The inversion-based techniques invert queries,
trying to infer the original data that produced the result
2.2. UDFs in DI processes of a given query [55, 56, 57]. To this end, the so-called
inversion functions are used.</p>
      <p>Big data integration processes use not only predefined DI The second taxonomy divides lineage techniques into
tasks available in DI design, development, and manage- tuple-level and value-level. The tuple-level lineage
supment tools [12], but also require the deployment of user ports tracking the processing of individual data rows, e.g.,
defined functions . UDFs allow to implement code snip- [58, 55]. The value-level lineage supports tracking the
lineage of specific values within rows, e.g., [56].</p>
      <p>The third taxonomy organizes lineage into eager and
lazy. In the eager technique, a lineage information is built
for all output data while executing a query (analysis) [59,
60, 61, 62, 63, 64]. In the lazy technique, this information
is constructed after a query was executed, e.g., [55].</p>
      <p>The lineage information is typically represented either
by IDs that are attached to source tuples (and in some
cases their individual attributes), e.g., [59, 61, 62] or by
additional structures that relate source and final (after
processing) data, e.g., [64].</p>
      <p>More advanced solutions using annotation were
presented in [53, 65, 66] - these solutions are known as the
semiring provenance model. It leverages the
mathematical structure of semirings to represent and analyze a
lineage information.
3. The missing pieces
non-relational data sources will allow to increase
performance of DI processes (i.e., reduce their execution
time). Based on the developed execution cost models and
implementation of code snippets, it will be possible to
eficiently push-down typical DI tasks [ 19] into the most
popular non-relational DSs.</p>
      <p>H2: we foresee that it will be possible to label BBUDFs
with known performance classes. First, these classes
will be built by ML techniques from performance
characteristics (e.g., RAM usage, CPU usage) of known UDFs.</p>
      <p>Second, performance characteristics of a given BBUDF
will be collected and used to assign the BBUDF to one
of the already known performance classes. Since similar
code snippets expose similar performance characteristics
[67], at some extent, a given performance class might
represent a similar internal logic of code snippets. This
approach will allow to reason (with a given probability)
about an expected BBUDF performance and (in some
cases) about its internal logic. As a consequence, it might
be further possible to decide whether a given BBUDF
could be reordered (in the spirit of [21]) and/or
parallelized (in the spirit of [68]).</p>
      <p>H3: we foresee to be able to develop techniques for
discovering broken lineage links between data and data
objects (e.g., tables, views, materialized views, stored
procedures and functions). The techniques should ofer
the discovery of the broken lineage links from a given
data object onward as well as from a given data object
backwards. Every discovered link will be accompanied
by a probability of its connection with another object. To
this end, we envisage using ML and statistical modeling.</p>
      <sec id="sec-1-1">
        <title>From the analysis of the state of the art on managing</title>
        <p>DI processes, we draw the following conclusions. First,
while the push-down technique has been researched
mainly for relational DSs, to the best of our knowledge,
the applicability and eficiency of this technique on big
data sources still faces mulitple open issues. Second, the
optimization of DI processes with UDFs has not yet been
researched enough to provide advanced and acceptable
solutions. This problem becomes especially challenging
for UDFs made available as black-boxes. For BBUDFs,
a still open research question is how to reason about
performance and internal logic of such UDFs. Third, the
existing data lineage techniques (both from research and
industry) do not handle cases when lineage links are 4. Possible solutions
broken by temporary objects.</p>
        <p>
          In this context, the main goal of the paper is to point In order to verify hypotheses H1-H3 (and hopefully to
out to research methods for managing DI processes prove them) we foresee to start researching and
developwith the focus on: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) eficient push-down methods on ing the solutions outlined in this section.
non-relational DSs - as the method of increasing
performance of DI processes, (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) learning the performance 4.1. Hypothesis H1
characteristics of black-box UDFs and, if possible, their
internal logic - as the method of optimizing the execu- To prove H1 it is necessary to conduct research on the
tion of DI processes with BBUDFs, (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) discovering and most popular non-relational DSs i.e., key-value,
columnmanaging data lineage in DI processes in the presence family, text, and graph. To this end, an execution cost
of temporary objects and UDFs, which broke lineage - model for each typical tasks being pushed-down (e.g.,
with the aim of providing additional functionality to data ifltering, value transformation, structure transformation,
governance. In this paper we formulate the following data anonymization) should be formulated for each type
research hypotheses (with their verification methods of a given data source. This would allow us to design
detailed in Section 4). implementation code skeletons for tasks pushed-down
        </p>
        <p>
          H1: we foresee that it will be possible to build a push- into these storage systems.
down optimizer that will be able to: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) assess whether There are two fundamental research questions on the
a given DI task profits from being pushed-down and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) push-down technique. Q1: whether a given DI task
propose eficient (alternative) implementations of a given pushed-down will cause an overall system performance
task pushed-down into a particular DS. As a consequence, improvement? Q2: given that the answer for Q1 is
posiwe envisage that the push-down technique applied to tive, how to eficiently implement a given task in a given
DS? In order to answer Q1, one could train a binary clas- techniques applied to the known UDFs in the same class.
sifier on various types of data describing DSs, DI tasks, Being able to apply parallelization to a given BBUDF is a
and performance characteristics, as discussed below. major step towards its performance improvement.
        </p>
        <p>
          Features of a DS should include among others: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) a The solution could be based on machine learning
DS type, its producer and software version, (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) the sup- (ML) techniques. To support the solution, the following
port for indexes, if any, (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) the type of a query optimizer, components are foreseen: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) the repository of
perforif any, (4) the support for parallel processing, (5) current mance characteristics of known (white-box) UDFs, which
deployment architecture - centralized or distributed, (6) include CPU and RAM usage as well as I/O and elapsed
physical parameters of hardware (e.g., the number of processing time, all represented as time series (TSs), (
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
CPUs, main memory size, disk type). classification algorithms for TSs, and (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) similarity
mea
        </p>
        <p>
          Features of a DI task executed in a DI engine should sures for TSs. We assume that classes of performance
include among others: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) task type (e.g., filtering, joining, characteristics of known UDFs have been built on TSs
aggregating, window function), (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) performance charac- obtained from excessive experiments in parallel and
centeristics of the task (CPU time, elapsed processing time, tralized computing architectures. The white-box UDF
RAM usage, I/O usage) for various data volumes and characteristics are labeled to reflect their performance
types - they will be collected by means of experiments. classes with parallelization and without. Based on these
        </p>
        <p>
          Features of a pushed-down DI task should include characteristics, various classifiers will be built to classify
among others: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) task type (e.g., filtering, joining, aggre- BBUDFs based on their performance characteristics.
gating, window function), (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) performance characteristics When a new BBUDF is made available in the system,
of the task (CPU time, elapsed processing time, RAM us- its performance characteristics are collected by means
age, I/O usage) for various data volumes and types and of experiments. Having collected the characteristics, the
various set ups of a DS. classification algorithms assign the BBUDF a label
rep
        </p>
        <p>
          Performance characteristics of each DS should in- resenting one of the existing performance classes. We
clude CPU time, elapsed processing time, RAM usage, I/O anticipate two alternative scenarios of TSs classification,
usage, for various workload types of DI processes. These namely: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) a few independent classifiers and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) an
encharacteristics will be collected and stored in a repository semble of classifiers (a.k.a. Time Series Forests).
for DI processes without and with pushed-down tasks. Our prior preliminary work on this problem [67]
        </p>
        <p>
          In order to answer Q2 we envisage the following ap- showed that popular classification algorithms, namely:
proaches. The fist one could be based on a pre-defined kNN, ROCKET [71], and HIVE-COTE [72] produced
library of code templates. For a given DI task, there promising results on TSs classification. This approach
will be a few code alternatives available for each DS. The can be substantially extended by enlarging the set of
classelection of a given code template will be based on the sifiers with two that recently have gained high popularity
decision of another classifier, learned on performance in the research community, i.e., multi-layer perceptron
characteristics of these code templates in a given DS. The [73] and CatBoost [74]. Moreover, the models can be
second approach could utilize a genetic algorithm to build based on much larger experiments and for much
evolve a given code template from the library. As in the broader types of BBUDFs, implementing operations like
classical genetic approach to producing code snippets, those listed in [19] but on non-relational data.
each code template would be evaluated by means of a fit- Notice that in order to classify TSs, a similarity
meaness score, based on how eficiently it performs a specific sure is needed. In [67], we used the dynamic time
warptask. The third solution could apply Large Technology ing (DTW) similarity measure [75]. TSs being compared
Models (LTMs) for the generation of code snippets, e.g., may difer in amplitude, length, and phase. For this
rea[69, 70] of pushed-down tasks. In this solution one could son it is necessary to normalize them before comparing
use the aforementioned features of data sources, DI tasks, [76, 77]. We foresee extensive validation not only DTW
and their performance characteristics to prompt an LTM. but also other advanced similarity measures, like
motifs [78], shapelets [79], locality-sensitive hashing [80],
4.2. Hypothesis H2 possibly combined with DTW, as inspired by [81, 82].
To conclude, our previous works on discovering the
inProving H2 is based on the assumption that knowing ternals of BBUDFs was based on a try and error approach
the performance class of a BBUDF one can get some [83], but it turned out to be applicable only to a small class
insights (with a certain level of probability) on the kind of of simple DI tasks (implementing filtering, projection,
filoperations being executed and performance (for example tering+projection), as in general the problem transforms
w.r.t. a data volume) of the BBUDF, by analyzing other to the Boolean Satisfiability Problem (belonging to the
(known) UDFs belonging to the same class. Moreover, NP-complete class) [84]. For this reason, in our opinion
understanding how to parallelize the BBUDF can also be the only sound solution to reason about the internals
ifgured out in some cases by analyzing the parallelization of BBUDFs should be based on ML techniques.
4.3. Hypothesis H3
Verifying H3 could be based on the relational data model
due to the fact that: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) it contains a well defined and rich
set of commands in the Data Definition Language and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
it is one of the most popular data models, and (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) broken
lineage has not been studied for this model. Since it is
possible to discover only verisimilar broken lineage links,
two alternative discovery approaches are foreseen. The
ifrst one is based on ML, whereas the second one - on
statistical modeling.
        </p>
        <p>
          The ML approach uses two diferent probabilistic
classifiers, where classes represent data objects and links
between them represent possible lineage links. Each link
is labeled with a probability of one object being connected
to another object. This approach has already been
veriifed on a small pilot project [ 85] and for some simple
scenarios real broken links were discovered. However, this
approach still needs to be substantially extended to: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
handle more complex broken lineage links, with longer
chains of dependencies between objects, (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) improve
the obtained prediction quality (e.g., measure F1), (
          <xref ref-type="bibr" rid="ref3">3</xref>
          )
be learned on substantially larger and versatile database
schemas, and (4) provide advanced and clear
visualizations of lineage graphs and discovered broken links.
        </p>
        <p>The statistical approach assumes applying the
Hidden Markov Model (HMM) that by learning the patterns in
data transformations can automatically generate lineage
graphs of data objects. In this application of the HMM,
states represent diferent data sets and structures of data
objects (e.g., tables, views, materialized views),
observations represent transformations applied to data and data
objects, transition probability represents the probability
of moving from one hidden state to another, and
emission probability represents the likelihood of observing
a particular transition. Thus, the model will discover
and represent the sequence of data and data object states
and transformations that occurred in a given DI process
from a DS, through a staging area to a destination system,
which in our case will be a data warehouse.</p>
        <p>The model can be learned and tested on historical
lineage links acquired from code repositories (e.g., GitHub)
and recognized database benchmarks (e.g., tpc.org). At
this stage of our research, we envisage using the
popular Baum-Welch algorithm to tune the parameters of the
HMM and the Viterbi algorithm for discovering the most
likely sequence of transitions from one state to another
[86].</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>5. Conclusions</title>
      <p>From the analysis of the state of the art on managing DI
processes, we draw the following conclusions.
• Push-down has been researched mainly for
relational DSs. To the best of our knowledge, the
applicability and eficiency of this technique on
big data sources reveals multiple open issues.
• The optimization of DI processes with UDFs has
not yet been researched enough to provide
advanced and acceptable solutions. This problem
becomes especially challenging for UDFs made
available as black-boxes. For BBUDFs, a still open
research question is how to reason about
performance and internal logic of such UDFs.</p>
      <p>
        The fact, that we have previously been
cooperating with IBM Software Lab Kraków (Poland) on
pilot projects on: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) developing techniques for
push-down on NoSQL DSs [87, 88] and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
learning about BBUDFs [67], shows that solutions to
these problems are of interest and value for the
IT sector.
• From our cooperation with the financial sector
in Poland, we received a feedback clearly
indicating that solutions for discovering data lineage
links and visualizing them for large data
repositories (like data warehouses) is of significant
importance. Legacy systems typically do not maintain
data lineage and there is an evident need to build
and visualize such links in these systems. For
example, a small DW in one of the companies
in the financial sector includes over 11000 stored
procedures and functions, with over 2 million
of codes and over 5000 of (materialized) views.
Building and visualizing lineage even for such a
small legacy DW is challenging. The problem
is aggravated with the presence of temporary
objects and UDFs that break data lineage.
Unfortunately, the existing data lineage techniques
(both from research and industry) do not handle
cases when lineage links are broken by temporary
objects.
      </p>
      <sec id="sec-2-1">
        <title>The aforementioned conclusions motivate us to address these three research problems in this position paper. We consider them as important missing solutions in the data integration research.</title>
        <p>on Design, Optimization, Languages and Analyti- Mauroux, Networking and storage: The next
comcal Processing of Big Data (DOLAP), volume 3653, puting elements in exascale systems?, IEEE Data
CEUR-WS.org, 2024. Engineering Bulletin 43 (2020).
[4] J. Zhu, Y. Mao, L. Chen, C. Ge, Z. Wei, Y. Gao, Fu- [19] Product documentation: Infosphere information
sionquery: On-demand fusion queries over multi- server 11.3, https://www.ibm.com/docs/en/iis/11.
source heterogeneous data, VLDB Endowment 17 3?topic=jobs-processing-data, accessed Feb, 2025,
(2024). ????
[5] A. Bouguettaya, B. Benatallah, A. Elmargamid, In- [20] Informatica, White paper: How to Achieve Flexible,
terconnecting Heterogeneous Information Systems, Cost-efective Scalability and Performance through
Kluwer Academic Publishers, ISBN 0792382161, Pushdown Processing, 2007.</p>
        <p>1998. [21] A. Simitsis, P. Vassiliadis, T. K. Sellis, State-space
[6] P. Brezany, A. M. Tjoa, H. Wanek, A. Wöhrer, Me- optimization of ETL workflows, IEEE TKDE 17
diators in the architecture of grid information sys- (2005).
tems, in: Int. Conf. on Parallel Processing and Ap- [22] A. Simitsis, P. Vassiliadis, T. Sellis, Optimizing etl
plied Mathematics (PPAM), volume 3019 of LNCS, processes in data warehouses, in: ICDE, 2005.</p>
        <p>Springer, 2003. [23] N. Kumar, P. S. Kumar, An eficient heuristic for
[7] S. A. Errami, H. Hajji, K. A. E. Kadi, H. Badir, Spatial logical optimization of ETL workoflws, in: VLDB
big data architecture: From data warehouses and Workshop on Enabling Real-Time Business
Intellidata lakes to the lakehouse, Journal of Parallel and gence, 2010.</p>
        <p>Distributed Computing 176 (2023). [24] R. Halasipuram, P. M. Deshpande, S.
Padmanab[8] R. Hai, C. Koutras, C. Quix, M. Jarke, Data lakes: A han, Determining essential statistics for cost based
survey of functions and systems, IEEE Transactions optimization of an ETL workflow, in: EDBT, 2014.
on Knowledge and Data Engineering 35 (2023). [25] X. Yu, M. Youill, M. E. Woicik, A. Ghanem, M.
Ser[9] A. A. Harby, F. H. Zulkernine, Data lakehouse: A afini, A. Aboulnaga, M. Stonebraker, Pushdowndb:
survey and experimental study, Information Sys- Accelerating a DBMS using S3 computation, CoRR
tems 127 (2025). abs/2002.05837 (2020).
[10] Z. Dehghani, Data Mesh: Delivering Data-Driven [26] B. Konieczny, What’s new in Apache Spark</p>
        <p>Value at Scale, O’Reilly, ISBN 1492092398, 2022. 3.1 - predicate pushdown for JSON, CSV and
[11] M. Hameed, F. Naumann, Data preparation: A Apache Avro, https://www.waitingforcode.com/
survey of commercial tools, SIGMOD Record 49 apache-spark-sql/what-new-apache-spark-3.
(2020). 1-predicate-pushdown-json-csv-apache-avro/
[12] Gartner, Magic quadrant for data integration tools, read, accessed Feb, 2025, 2021.</p>
        <p>2022. [27] H. Herodotou, H. Lim, G. Luo, N. Borisov, L. Dong,
[13] IBM InfoSphere Information Server, https://www. F. B. Cetin, S. Babu, Starfish: A self-tuning system
ibm.com/information-server, accessed Dec 2024, for big data analytics, in: Biennial Conf. on
Inno???? vative Data Systems Research (CIDR), volume 11,
[14] S. M. F. Ali, R. Wrembel, From conceptual design to 2011.</p>
        <p>performance optimization of ETL workflows: cur- [28] X. Liu, C. Thomsen, T. B. Pedersen,
Mapreducerent state of research and open problems, The VLDB based dimensional ETL made easy, VLDB
EndowJournal 26 (2017). ment 5 (2012).
[15] A. Simitsis, S. Skiadopoulos, P. Vassiliadis, The [29] X. Liu, C. Thomsen, T. B. Pedersen, ETLMR: A
history, present, and future of ETL technology (in- highly scalable dimensional ETL framework based
vited), in: Int. Workshop on Design, Optimization, on mapreduce, Trans. Large-Scale Data- and
Languages and Analytical Processing of Big Data Knowledge-Centered Systems 8 (2013).
(DOLAP), volume 3369 of CEUR Workshop Proceed- [30] X. Liu, N. Iftikhar, An ETL optimization framework
ings, CEUR-WS.org, 2023. using partitioning and parallelization, in: ACM
[16] 2020 state of data science, https://know. SAC, 2015.</p>
        <p>anaconda.com/rs/387-XNW-688/images/ [31] S. M. F. Ali, J. Mey, M. Thiele, Parallelizing
Anaconda-SODS-Report-2020-Final.pdf ; ac- user–defined functions in the etl workflow using
cessed Dec 2024, 2020. orchestration style sheets, Int. Journal of Applied
[17] M. Owaida, G. Alonso, L. Fogliarini, A. Hock-Koon, Mathematics and Computer Science 29 (2019).</p>
        <p>P.-E. Melet, Lowering the latency of data processing [32] S. M. F. Ali, R. Wrembel, Towards a cost model to
pipelines through fpga based hardware acceleration, optimize user-defined functions in an ETL
workVLDB Endowment 13 (2019). lfow based on user-defined performance metrics,
[18] A. Lerner, R. Hussein, A. Ryser, S. Lee, P. Cudré- in: European Conf. on Advances in Databases and</p>
        <p>Information Systems (ADBIS), LNCS 11695, 2019. T. Stoltmann, U. Leser, Versatile optimization of
[33] F. Hueske, M. Peters, M. J. Sax, A. Rheinländer, udf-heavy data flows with sofa, in: SIGMOD, ACM,
R. Bergmann, A. Krettek, K. Tzoumas, Opening 2014.
the black boxes in data flow optimization, VLDB [48] C. Yan, Y. Lin, Y. He, Predicate pushdown for data
Endowment 5 (2012). science pipelines, SIGMOD 1 (2023).
[34] F. Hueske, M. Peters, A. Krettek, M. Ringwald, [49] What is data lineage?, https://www.ibm.com/topics/
K. Tzoumas, V. Markl, J.-C. Freytag, Peeking data-lineage, accessed Dec 2024, ????
into the optimization of data flow programs with [50] Y. R. Wang, S. E. Madnick, A polygen model for
mapreduce-style udfs, in: ICDE, 2013. heterogeneous database systems: The source
tag[35] P. Große, N. May, W. Lehner, A study of partitioning ging perspective, in: Int. Conf. on Very Large Data
and parallel UDF execution with the SAP HANA Bases (VLDB), 1990.
database, in: Conf. on Scientific and Statistical [51] L. Chiticariu, W.-C. Tan, G. Vijayvargiya, DBNotes:
Database Management (SSDBM), 2014. a post-it system for relational databases based on
[36] A. Rheinländer, A. Heise, F. Hueske, U. Leser, F. Nau- provenance, in: SIGMOD, 2005.
mann, Sofa: An extensible logical optimizer for udf- [52] J. N. Foster, T. J. Green, V. Tannen,
Annoheavy data flows, Information Systems 52 (2015). tated XML: queries and provenance, in: ACM
[37] P. Holanda, M. Raasveldt, M. Kersten, Don’t keep SIGMOD-SIGACT-SIGART symposium on
Princimy udfs hostage - exporting udfs for debugging pur- ples of Database Systems (PODS), 2008.
poses, in: Simpósio Brasileiro de Banco de Dados, [53] P. Senellart, Provenance and Probabilities in
Rela2017. tional Databases, SIGMOD Record 46 (2018).
[38] Y. Huang, Z. Wang, C. Li, Udon: Eficient debugging [54] D. Dosso, S. B. Davidson, G. Silvello, Data
proveof user-defined functions in big data systems with nance for attributes: attribute lineage, in: USENIX
line-by-line control, SIGMOD 1 (2023). Conf. on Theory and Practice of Provenance, TAPP,
[39] V. Simhadri, K. Ramachandra, A. Chaitanya, R. Gu- USENIX Association, 2020.</p>
        <p>ravannavar, S. Sudarshan, Decorrelation of user [55] Y. Cui, J. Widom, Practical lineage tracing in data
defined function invocations in queries, in: Int. warehouses, in: Int. Conf. on Data Engineering
Conf. on Data Engineering (ICDE), 2014. (ICDE), 2000.
[40] A. Crotty, A. Galakatos, K. Dursun, T. Kraska, [56] A. Woodruf, M. Stonebraker, Supporting
fineC. Binnig, U. Çetintemel, S. Zdonik, An architec- grained data lineage in a database visualization
ture for compiling udf-centric workflows, VLDB environment, in: Int. Conf. on Data Engineering
Endowment 8 (2015). (ICDE), 1997.
[41] K. Ramachandra, K. Park, Blackmagic: Automatic [57] M. Yamada, H. Kitagawa, T. Amagasa, A. Matono,
inlining of scalar udfs into SQL queries with froid, Augmented lineage: traceability of data analysis
VLDB Endowment 12 (2019). including complex UDF processing, The VLDB
[42] M. E. Schüle, J. Huber, A. Kemper, T. Neumann, Journal 32 (2023).</p>
        <p>Freedom for the sql-lambda: Just-in-time-compiling [58] P. Buneman, S. Khanna, W. C. Tan, Why and where:
user-injected functions in postgresql, in: Int. Conf. A characterization of data provenance, in: Int. Conf.
on Scientific and Statistical Database Management on Database Theory (ICDT), volume 1973 of LNCS,
(SSDBM), ACM, 2020. 2001.
[43] E. Friedman, P. M. Pawlowski, J. Cieslewicz, [59] O. Benjelloun, A. D. Sarma, A. Y. Halevy,
Sql/mapreduce: A practical approach to self- M. Theobald, J. Widom, Databases with uncertainty
describing, polymorphic, and parallelizable user- and lineage, VLDB Journal 17 (2008).
defined functions, VLDB Endowment 2 (2009). [60] O. Benjelloun, A. D. Sarma, A. Y. Halevy, J. Widom,
[44] K. Ramachandra, K. Park, K. V. Emani, A. Halver- ULDBs: Databases with uncertainty and lineage, in:
son, C. A. Galindo-Legaria, C. Cunningham, Froid: Int. Conf. on Very Large Data Bases (VLDB), ACM,
Optimization of imperative programs in a relational 2006.</p>
        <p>database, VLDB Endowment 11 (2017). [61] D. Bhagwat, L. Chiticariu, W. C. Tan, G.
Vijay[45] M. Sichert, T. Neumann, User-defined operators: vargiya, An annotation management system for
Eficiently integrating custom algorithms into mod- relational databases, VLDB Journal 14 (2005).
ern databases, VLDB Endowment 15 (2022). [62] L. Chiticariu, W. C. Tan, G. Vijayvargiya, Dbnotes:
[46] K. Chasialis, T. Palaiologou, Y. Foufoulas, A. Simit- a post-it system for relational databases based on
sis, Y. E. Ioannidis, Qfusor: A UDF optimizer plugin provenance, in: Int. Conf. on Management of Data
for SQL databases, in: Int. Conf. on Data Engineer- (SIGMOD), 2005.</p>
        <p>ing (ICDE), IEEE, 2024. [63] B. Glavic, G. Alonso, Perm: Processing
prove[47] A. Rheinländer, M. Beckmann, A. Kunkel, A. Heise, nance and data on the same data model through
query rewriting, in: Int. Conf. on Data Engineering [77] Á. B. Hernández, M. S. Pérez, S. Gupta, V.
Muntés(ICDE), 2009. Mulero, Using machine learning to optimize
paral[64] F. Psallidas, E. Wu, Smoke: Fine-grained lineage at lelism in big data applications, Future Generation
interactive speed, VLDB Endowment 11 (2018). Computer Systems 86 (2018).
[65] Y. Amsterdamer, D. Deutch, V. Tannen, Provenance [78] A. Mueen, Time series motif discovery: dimensions
for aggregate queries, in: ACM SIGMOD-SIGACT- and applications, Wiley Interdisciplinary Reviews:
SIGART Symposium on Principles of Database Sys- Data Mining and Knowledge Discovery 4 (2014).
tems (PODS), 2011. [79] J. Grabocka, N. Schilling, M. Wistuba, L.
Schmidt[66] T. J. Green, G. Karvounarakis, V. Tannen, Prove- Thieme, Learning time-series shapelets, in: ACM
nance semirings, in: ACM SIGACT-SIGMOD- SIGKDD Int. Conf. on Knowledge Discovery and
SIGART Symposium on Principles of Database Sys- Data Mining (KDD), 2014.</p>
        <p>tems (PODS), 2007. [80] M. Datar, N. Immorlica, P. Indyk, V. S. Mirrokni,
[67] M. Bodziony, B. Ciesielski, A. Lehnhardt, R. Wrem- Locality-sensitive hashing scheme based on
pbel, On reasoning about black-box UDFs by clas- stable distributions, in: ACM Symposium on
Comsifying their performance characteristics, in: Har- putational Geometry, 2004.
nessing Opportunities: Reshaping ISD in the post- [81] M. Ceccarello, J. Gamper, Fast and scalable mining
COVID-19 and Generative AI Era (ISD), 2024. of time series motifs with probabilistic guarantees,
[68] S. M. F. Ali, R. Wrembel, Framework to optimize VLDB Endowment 15 (2022).
data processing pipelines using performance met- [82] A. Charane, M. Ceccarello, J. Gamper, Shapelets
rics, in: Inf. Conf. on Big Data Analytics and Knowl- evaluation using silhouettes for time series
classifiedge Discovery (DaWaK), LNCS 12393, 2020. cation, in: Int. Workshop on Design, Optimization,
[69] F. Monti, F. Leotta, J. Mangler, M. Mecella, Languages and Analytical Processing of Big Data
S. Rinderle-Ma, Nl2processops: Towards llm- (DOLAP), volume 3653 of CEUR Workshop
Proceedguided code generation for process execution, in: ings, 2024.</p>
        <p>Business Process Management Forum (BPM), vol- [83] M. Bodziony, H. Krzyzanowski, L. Pieta, R.
Wremume 526 of LNBIP, Springer, 2024. bel, On discovering semantics of user-defined
func[70] F. Mu, L. Shi, S. Wang, Z. Yu, B. Zhang, C. Wang, tions in data processing workflows, in: BiDEDE @
S. Liu, Q. Wang, Clarifygpt: A framework for ACM SIGMOD/PODS Conference, ACM, 2021.
enhancing llm-based code generation via require- [84] S. Malik, L. Zhang, Boolean satisfiability from
thements clarification, ACM Software Engineering 1 oretical hardness to practical success,
Communca(2024). tions of the ACM 52 (2009).
[71] A. Dempster, F. Petitjean, G. I. Webb, ROCKET: [85] M. Grocholewski, T. Gruszczyński, Technologia
linexceptionally fast and accurate time series classifi- eage dla obiektów i kodu w bazach danych:
koncation using random convolutional kernels, Data cepcja, implementacja, weryfikacja
eksperymenMining and Knowledge Discovery 34 (2020). talna (in polish); Lineage technique for objects and
[72] M. Middlehurst, J. Large, M. Flynn, J. Lines, code in databases: concept, implementation, and
A. Bostrom, A. J. Bagnall, HIVE-COTE 2.0: a new experimental evaluation, 2024. Master thesis at
Pozmeta ensemble for time series classification, Ma- nan University of Technology.</p>
        <p>chine Learning 110 (2021). [86] J. P. Coelho, T. M. Pinho, J. Boaventura-Cunha,
Hid[73] M. Kulyabin, P. A. Constable, A. E. Zhdanov, I. O. den Markov Models: Theory and Implementation
Lee, D. A. Thompson, A. K. Maier, Attention to the using MATLAB, Data-Centric Systems and
Applielectroretinogram: Gated multilayer perceptron for cations, CRC Press, 2021.</p>
        <p>ASD classification, IEEE Access 12 (2024). [87] M. Bodziony, S. Roszyk, R. Wrembel, On evaluating
[74] L. Zhang, D. Jánosík, Enhanced short-term load performance of balanced optimization of ETL
proforecasting with hybrid machine learning models: cesses for streaming data sources, in: Int. Workshop
Catboost and xgboost approaches, Expert Systems on Design, Optimization, Languages and
Analytiwith Applications 241 (2024). cal Processing of Big Data (DOLAP), volume 2572,
[75] D. J. Berndt, J. Cliford, Using dynamic time warp- CEUR-WS.org, 2020.</p>
        <p>ing to find patterns in time series, in: Workshop [88] M. Bodziony, R. Morawski, R. Wrembel,
EvaluKnowledge Discovery in Databases, AAAI Press, ating push-down on NoSQL data sources:
exper1994. iments and analysis paper, in: BiDEDE @ ACM
[76] S. Pumma, W. Feng, P. Phunchongharn, SIGMOD/PODS Conference, ACM, 2022.</p>
        <p>S. Chapeland, T. Achalakul, A runtime
estimation framework for ALICE, Future Generation
Computer Systems 72 (2017).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Wrembel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abelló</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Song</surname>
          </string-name>
          ,
          <article-title>DOLAP data warehouse research over two decades: Trends and challenges</article-title>
          ,
          <source>Information Systems</source>
          <volume>85</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Siddiqi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kern</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Boehm, SAGA: A scalable framework for optimizing data cleaning pipelines for machine learning applications</article-title>
          ,
          <source>SIGMOD</source>
          <volume>1</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Timakum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Song</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Song, DOLAP: A 25 year journey through research trends and performance (invited talk)</article-title>
          ,
          <source>in: Int. Workshop</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>