<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>with Exasol</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christoph Großmann</string-name>
          <email>christoph.grossmann@online.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Johannes Schildgen</string-name>
          <email>johannes.schildgen@oth-regensburg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Computer Science and Mathematics, Regensburg University of Applied Sciences</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>M. Leyer, J. Wichmann (Eds.): Proceedings of the LWDA 2023 Workshops: BIA</institution>
          ,
          <addr-line>DB, IR, KDML and WM. Marburg</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Machine Learning, SQL</institution>
          ,
          <addr-line>Database, Exasol</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <abstract>
        <p>This paper introduces a novel, extendable, no-code framework for integrating machine-learning algorithms into SQL using the Exasol database. The framework combines the strengths of the highperformance, parallel-processing analytical Exasol database with the flexible and sophisticated machine learning algorithms of the Python library Scikit-Learn, while providing a seamless integration into SQL. This paper explores the technical background, the concept, and the implementation of the framework. The CREATE MODEL command for creating a machine learning model and the PREDICT function for prediction using a pre-trained model are discussed in detail. The main contributions of the framework are its seamless integration into SQL, scalability, and leveraging of existing database infrastructure. An overview of related work is also given.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>CEUR
Workshop
Proceedings
employee as our data source. Additionally, we specify the prediction target or label salary and
the features position and the birthyear used to determine the label.</p>
      <p>CREATE MODEL "model" ON employees PREDICT (salary) USING ("position", birthyear);</p>
      <p>We can use the created model to predict the salaries of employees. We again use the position
and birthyear of the employee table as features and predict the missing salary entries. The
prediction result is shown in Table 1. This example is elaborated on further in the following.
SELECT name, "position", birthyear, PREDICT "model" USING ("position", birthyear)
FROM employees WHERE salary IS NULL;</p>
      <p>Our framework facilitates eficient resource utilization by capitalizing on Exasol’s parallel
processing capabilities and ETL pipeline, enabling scalability to handle large datasets and
complex analytical workloads. Furthermore, in-database exploratory data analysis is simplified
by the availability of no-code ML functionality. Finally, this integration eliminates the need
for data movement between diferent systems, reducing latency, and enhancing the overall
eficiency of the analytics workflow.</p>
      <p>This paper provides an overview of the technical background, related work, the concept, and
the implementation of the framework. It closes with a discussion and a conclusion.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Technical Background</title>
      <p>
        Exasol is a proprietary distributed relational analytical database management system. It runs on
the Linux-based operating system (OS) ExaCluster OS, which provides a runtime environment
and a storage layer for the database. Exasol being a cluster of nodes allows it to execute queries
in parallel and makes it cloud-ready. Furthermore, Exasol uses column-oriented storage and
in-memory processing. For data unfit to be stored in the database, Exasol provides the file
system BucketFS [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. We chose Exasol since it provides the necessary tools for extending
the database and SQL for ML, and opportunities for improving ML processes using database
features. The tools needed for our framework are a query rewriter and a way to execute code
written in a scripting language, preferably Python, inside the SQL pipeline. The opportunities
for improving ML processes are massively parallel processing using the parallel processing
infrastructure of Exasol clusters and optimization through automatic query optimization.
      </p>
      <p>
        The features of the Exasol database our framework relies upon are introduced in the following.
The aforementioned BucketFS is a plain file system and can be accessed using a HTTPS interface.
Data stored in BucketFS is replicated over all nodes of a cluster. Eventual consistency is
guaranteed [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. Our framework uses BucketFS for ML model storage. Script-language container
are the basis for extending the database for ML. They are Docker containers and contain a
complete Linux installation with all packages required to execute code in scripting languages
like Python, R, or Java. A set of pre-built containers is distributed by Exasol. Nevertheless, it is
possible to build a custom container. Script-language containers are stored in BucketFS [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ].
User-defined function ( UDF) scripts provide the interface for extending the SQL pipeline with
the script languages provided by script-language containers [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. These scripts are executed
through SQL, pass their input data to a program written in another language and executed in an
instance of the currently active script-language container, and then pass the results back to the
database. Since UDF scripts are executed within the SQL pipeline, they can make use of database
parallelization [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. UDF scripts already make it possible to extend the Exasol database with
ML. Scripting programs combine SQL with the scripting language Lua. Thus, they can execute
multiple successive SQL statements and provide control structures [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Preprocessor scripts are
query rewriters that analyze and rewrite all SQL statements before they are processed. Thus,
they can convert unsupported SQL constructs into statements supported by the SQL parser.
They can be seen as specialized scripting programs [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        We chose Python since it is a popular language for ML providing many popular libraries like
Scikit-Learn, PyTorch, and TensorFlow [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]. This decision does not limit our framework to
Python. Support for ML libraries written in other script languages can be added in the future.
The framework currently integrates ML algorithms of the Scikit-Learn library due to its ease of
use, performance, and standardized API [
        <xref ref-type="bibr" rid="ref16 ref16">16, 16</xref>
        ]. Exasol provides Python libraries for accessing
the database [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and BucketFS [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Related Work</title>
      <p>
        Exasol’s developers and community provide information and many examples for creating UDF
scripts for ML and data analysis tasks [
        <xref ref-type="bibr" rid="ref20 ref21 ref22">20, 21, 22</xref>
        ]. These scripts each only handle one specific
use case, while our framework provides a generic solution. Furthermore, Exasol provides an
extension to use pre-trained ML models via the Transformers API [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
      </p>
      <p>
        There are several other approaches to integrate ML into database systems. Among these,
our framework stands out through its focus on smooth SQL integration. The approach closest
to our framework is the Apache MADlib analytics library, which uses user-defined functions
and aggregates to implement in-database ML algorithms [
        <xref ref-type="bibr" rid="ref24 ref25">24, 25</xref>
        ]. Many well-known database
vendors have solutions for integrating ML into the database like Oracle [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] and IBM [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. But
there are also many diferent approaches by the scientific community. Schule et al. propose a
complete ML pipeline using recursive tables while training models on GPUs. [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] Makrynioti
et al. introduce sql4ml, a framework for translating objective functions written in SQL into an
equivalent TensorFlow graph [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. Dolmatova et al. introduce relational matrix algebra (RMA),
which seamlessly integrates linear-algebra operations into the relational model [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. Kersten
et al. propose SciQL, a SQL-based query language with both tables and arrays as first-class
citizens [
        <xref ref-type="bibr" rid="ref31 ref32">31, 32</xref>
        ]. Apart from these approaches, other approaches that start with a high-level
statistical programming language and aim to build a parallel processing infrastructure using
database systems exist [
        <xref ref-type="bibr" rid="ref33 ref34 ref35 ref36">33, 34, 35, 36</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>4. Concept</title>
      <p>An important part of our framework is the convenient, well-integrated syntax for handling ML
models. ML models are handled as database objects stored in system tables and with support
for DDL commands. These commands include CREATE, RENAME, DROP, ALTER, and REPLACE.
The CREATE command creates a new model and trains it. The RENAME command renames an
existing model and all associated files. The DROP command deletes an existing model and all
associated files. The ALTER command allows for changing the parameters of an existing model.
The REPLACE command replaces an existing model with the newly trained one. In addition to
these commands, we introduce three commands unique to ML models: IMPORT, RETRAIN, and
PREDICT. The IMPORT command creates the metadata for an ML model that already exists in
BucketFS. The RETRAIN command retrains the specified ML model with the updated data in the
source table or view. An error is thrown, if the source table or view is missing columns needed
for the training of the model. The PREDICT function uses a previously trained ML model and
the specified input data to predict values.</p>
      <p>In the following, we present the syntax of the CREATE and the PREDICT command. Additionally,
examples are given for better understanding. These examples use the “employees” table shown
in Table 2. The table contains the name, position, year of birth, and salary of diferent employees.
Some of the employee salaries are NULL and thus unknown. We will create an ML model to
predict these values.
name</p>
      <sec id="sec-5-1">
        <title>Jacob Taylor</title>
      </sec>
      <sec id="sec-5-2">
        <title>Emma Anderson</title>
      </sec>
      <sec id="sec-5-3">
        <title>Daniel Young</title>
      </sec>
      <sec id="sec-5-4">
        <title>Ava Thompson</title>
      </sec>
      <sec id="sec-5-5">
        <title>Emily Wilson</title>
      </sec>
      <sec id="sec-5-6">
        <title>John Anderson</title>
        <p>⋮
position</p>
      </sec>
      <sec id="sec-5-7">
        <title>Software Engineer</title>
      </sec>
      <sec id="sec-5-8">
        <title>Software Engineer</title>
      </sec>
      <sec id="sec-5-9">
        <title>Sales Associate</title>
      </sec>
      <sec id="sec-5-10">
        <title>Sales Associate</title>
      </sec>
      <sec id="sec-5-11">
        <title>Software Engineer</title>
        <p>Sales Associate
⋮
birthyear</p>
        <sec id="sec-5-11-1">
          <title>4.1. Model Creation</title>
          <p>The syntax for creating an ML model using our framework is shown in Figure 1. The name
determines the unique object identifier of the model. This identifier is needed for all further
interactions with the model, like using it for predictions. The source identifier determines
which table or view is used as the input for training the model. Thus, the table or view has to
exist and preferably contain data. If the table or view is empty, the model has to be retrained after
the data is inserted. The column specifiers in the PREDICT clause determine which columns of
the source table or view are the labels of the model and thus contain the values to be predicted.
The column specifiers in the USING clause determine which columns of the source table or
view are the features of the model and thus contain the values that can be used to predict
(
,
=</p>
          <p>,
name</p>
          <p>ON
source</p>
          <p>PREDICT
USING
(</p>
          <p>)
,
column</p>
          <p>WITH
key
value
the labels. The WITH clause allows for setting additional parameters using key-value pairs.
These parameters can be used to determine the output type of the model, to specify the ML
algorithm to be used, and to pass additional settings to the algorithm. Examples of output types
are classification and regression. In case no output type is determined using the WITH clause,
regression is assumed if all labels are of the data type DOUBLE PRECISION (or its aliases DOUBLE,
FLOAT, NUMBER, and REAL). Otherwise, classification is assumed as the output type.</p>
          <p>As an example, we create a model "sal", which uses the employee table as its source.
The label to predict is the salary and the features are the position and the birthyear of
employees. Furthermore, we specify the model to use the 'DecisionTreeRegressor' function,
which determines the output to be a regression. Additionally, we specify a maximum depth of
64 for the created decision tree. The query to create the specified model is the following:
CREATE MODEL "sal" ON employees PREDICT (salary) USING ("position", birthyear)
WITH 'Function' = 'DecisionTreeRegressor', 'max_depth' = 64;</p>
          <p>The source table for this model contains some NULL values in the salary column. For training,
only tuples without NULL values in labels are used. After the training, the model is stored in
BucketFS for future use.</p>
        </sec>
        <sec id="sec-5-11-2">
          <title>4.2. Prediction</title>
          <p>The syntax for the variadic function PREDICT is shown in Figure 2. The name corresponds to
the identifier of an already existing ML model. The output of the prediction is one set of labels
for each input row. These labels correspond to the trained labels, having the same name and
a compatible data type. The column parameter list determines which columns serve as the
features of the model. The number and position of these features have to match the number
and position of the features used in the training step. The data types of the features have
to be compatible with the features used for training, while the name of the features is of no
importance. Furthermore, a prediction does not have to use the same table that was used for
training.</p>
          <p>As an example, we use the previously trained ML model to predict the salary of employees,
for whom this information is missing. We again use the employee table as source as well as
name
position and birthyear as features. Since salary is a currency value, we format it to have
two decimal places by casting it as a DECIMAL(14,2). Furthermore, we rename the result of the
prediction to pred_salary to avoid duplicate column names. The query for this prediction is
the following:
SELECT name, "position", birthyear, salary AS original_salary,
PREDICT "sal" USING ("position", birthyear) FROM employees;</p>
          <p>The result of the query is shown in Table 3. For each employee, this information is predicted
based on the data of employees with valid salary information. To persist the prediction, INSERT,
CREATE TABLE AS, or UPDATE queries can be used.
name</p>
        </sec>
      </sec>
      <sec id="sec-5-12">
        <title>Jacob Taylor</title>
      </sec>
      <sec id="sec-5-13">
        <title>Emma Anderson</title>
      </sec>
      <sec id="sec-5-14">
        <title>Daniel Young</title>
      </sec>
      <sec id="sec-5-15">
        <title>Ava Thompson</title>
      </sec>
      <sec id="sec-5-16">
        <title>Emily Wilson</title>
      </sec>
      <sec id="sec-5-17">
        <title>John Anderson</title>
        <p>⋮
position</p>
      </sec>
      <sec id="sec-5-18">
        <title>Software Engineer</title>
      </sec>
      <sec id="sec-5-19">
        <title>Software Engineer</title>
      </sec>
      <sec id="sec-5-20">
        <title>Sales Associate</title>
      </sec>
      <sec id="sec-5-21">
        <title>Sales Associate</title>
      </sec>
      <sec id="sec-5-22">
        <title>Software Engineer</title>
        <p>Sales Associate
⋮</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Implementation</title>
      <p>
        The implementation of the framework works with both the single-node “Community Version”
[
        <xref ref-type="bibr" rid="ref37">37</xref>
        ] as well as proprietary cluster versions. To avoid a library version mismatch within the
default script-language container, the Exasol script-language container version 8.0.0 is used.
      </p>
      <sec id="sec-6-1">
        <title>5.1. Available Algorithms</title>
        <p>Our framework currently supports five algorithms of the Python library Scikit-Learn. These
supported algorithms are listed in Table 4. All parameters of the algorithms are supported
by our framework. In the future, other algorithms of the Scikit-Learn framework and other
frameworks, even ones written in other programming languages, will be supported by our
framework.</p>
      </sec>
      <sec id="sec-6-2">
        <title>5.2. Framework Layers</title>
        <p>To process commands interacting with ML models, our framework employs several layers.
These layers are visualized in Figure 3. The first layer processing incoming queries consists of
preprocessor scripts. This layer converts the custom SQL syntax of the framework to scripting
programs, which are the second layer of the framework. The scripting programs handle the
metadata of the model and call the UDF scripts, which are the third, final layer. UDF scripts
handle the calling of the actual ML functionality provided by the Python library Scikit-Learn.
Furthermore, the UDF scripts handle the storage and loading of models to and from BucketFS.
In the following, the implementation is discussed in further detail.</p>
        <p>Queries
Preprocessor Scripts</p>
        <p>Scripting Programs</p>
        <p>UDF Scripts</p>
        <p>BucketFS
Metadata</p>
        <p>Scikit-Learn</p>
        <sec id="sec-6-2-1">
          <title>Python</title>
          <p>For each algorithm supported by the framework, two UDF scripts need to be created. The
ifrst UDF script handles model creation and training. For this purpose, the script takes the
name of the model, settings for the algorithm, a list of features, and a label. The number of
labels is only restricted in the current implementation of the framework. The script passes the
settings, features, and labels to the algorithm, starts the training of the model, and finally stores
the model in BucketFS. Since the currently implemented algorithms cannot process character
strings as input or output, it is necessary to map these to integers. This is handled by the UDF
scripts and a mapping dictionary. As an example, let us assume the key 'Software Engineer'
is mapped to the integer value 1. On prediction, each input instance of 'Software Engineer'
would also be mapped to 1. In the case of classification, each output instance of 1 would be
mapped to 'Software Engineer'. Mapping dictionaries are created before model creation and
stored alongside the model in BucketFS. In a future version of the framework, this mapping
functionality will be replaced by in-database mapping tables.</p>
          <p>As an example for creating a model using UDF scripts, we use the statement created by the
preprocessor when processing the following CREATE MODEL statement.</p>
          <p>CREATE MODEL "sal" ON employees PREDICT (salary) USING ("position", birthyear)
WITH 'Function' = 'DecisionTreeRegressor', 'max_depth' = 64;</p>
          <p>Since UDF scripts are ML-function-specific, the function to be used has to be determined
before the execution. In our case, the function parameter set to 'DecisionTreeRegressor'
means that the decision-tree-regressor function of the Scikit-Learn library is selected. The
preprocessor script rewrites the CREATE MODEL statement into the following statement.
SELECT ML.sklearn_tree_DecisionTreeRegressor_train</p>
          <p>('sal', '{"model_params":{"max_depth":64}}', "position", birthyear, salary)
FROM employees WHERE salary IS NOT NULL ORDER BY RANDOM();</p>
          <p>The second UDF script handles prediction. The parameters of the script are the name of the
model, settings, the row identifier, and the features used for predicting labels. The features
passed to the prediction script have to match the number, position, and data type of the features
which were used to create the model. The script loads the model and all associated mapping
dictionaries from BucketFS and passes the features to the model for prediction. The predicted
labels are then combined with the internal row identifiers by position and the set is returned.
An important restriction of Exasol is that no other expression can be present in the SELECT
clause when calling a UDF script that emits a table. To solve this problem, we use common
table expressions.</p>
          <p>As an example for prediction with a pre-trained model using UDF scripts, we use the
statement created by the preprocessor when processing the following statement using the PREDICT
function.</p>
          <p>SELECT name, "position", birthyear, salary AS original_salary,
PREDICT "sal" USING ("position", birthyear) FROM employees;</p>
          <p>The prediction UDF script is determined using the stored model metadata. The preprocessor
script rewrites the previous statement into the following statement.</p>
          <p>WITH pred AS (SELECT ML.sklearn_tree_DecisionTreeRegressor_predict
('sal', '', ROWID, "position", birthyear) FROM employees)
SELECT e.name, e."position", e.birthyear, e.salary AS original_salary, p.label
AS salary FROM employees e JOIN pred p ON e.ROWID = p.identifier GROUP BY IPROC();</p>
          <p>As already discussed, scripting programs combine the handling of metadata with the execution
of ML functionality. The metadata of the framework could theoretically also be managed inside
of UDF scripts, but this would necessitate the use of a database connector. This would defeat
the purpose of executing ML in-database since an external connection to the database is needed.
Scripting programs take the parameters extracted by the preprocessor scripts, as input. The
preprocessor script takes incoming queries containing the custom syntax of our framework,
splits them into tokens, and extracts the parameters of clauses of statements. In case a new ML
model is created, the scripting programs choose the ML function to be used according to the
settings the user provided in the WITH clause. If multiple ML functions fit the given settings,
the function with the lowest priority value is selected. When training an ML model, all settings
relevant to the model are passed to the UDF script. In case an existing ML model is needed, the
scripting programs determine the function used to create the model through the metadata of
the model. When executing predictions, the scripting programs use either Exasol’s internal row
ID of the source table ROWID or the ROW_NUMBER function as the row identifier for data passed to
the prediction UDF script. The GROUP BY IPROC() clause groups the rows by the node they are
stored on. Thus, each row is processed locally on the node it is stored on and only the results
are transmitted over the network.</p>
        </sec>
      </sec>
      <sec id="sec-6-3">
        <title>5.3. Tracked Metadata</title>
        <p>The metadata of the framework is stored in two tables. The ML.Algorithm table contains
information about the algorithms integrated into the framework, The information contained
about algorithms includes their algorithm type, their output type, the module or library it is
contained in, and the function it references to. The ML.Model table contains information about
all created ML models created by the user. The information about models includes the algorithm
used to train the model, its name, the source table or view, the features used during training,
the labels used during training, and the settings used during training. This information is used
for PREDICT or RETRAIN statements, for example.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Discussion</title>
      <p>Our framework is an extendable, no-code integration of ML into SQL while employing Exasol’s
distributed, parallel processing capabilities and ETL pipeline in addition to Scikit-Learn’s flexible
and sophisticated ML algorithms. The main contributions of the proposed framework lie in its
seamless integration into SQL, scalability, and leveraging of existing database infrastructure.
The integration of ML into SQL also benefits users familiar with SQL, such as data scientists,
analysts, and database administrators. The framework enables them to leverage their existing
SQL skills, making the transition to advanced analytics and ML more accessible. Furthermore,
it is also possible to export and import models, since all ML models of the framework are stored
in BucketFS.</p>
      <p>However, certain restrictions need to be addressed. Firstly, the prediction phase currently
uses UDF scripts. In future work, the prediction step will be changed to use preprocessor
scripts. Other future work includes replacing the mapping directories with mapping tables in
the database. Furthermore, enabling more than one possible label is also future work. The WITH
clause is currently restricted to exclusively textual values. Moreover, when using a model the
version of the used libraries has to match the versions of the libraries used for creating the
model. This can be achieved by using the same script-language container that was used for the
model creation. Currently, the user of the framework has to activate the correct script-language
container for each model. The automation of this process is also future work. Future work also
includes the extension of the framework with additional algorithms of the Scikit-Learn library
and other ML libraries written in Python or other programming languages. Future directions
for our framework include distributed training, incremental model training, sample weights,
model statistics, explainability functions, and data preparation.</p>
      <p>
        In comparison to other approaches, our framework stands out through its smooth SQL
integration. Furthermore, our framework has the advantage of employing the well-established
ML library Scikit-Learn. However, by employing third-party libraries, our framework has a
disadvantage compared to approaches implementing ML algorithms directly in SQL. Examples
of approaches like this are Apache MADlib [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] and Oracle Machine Learning [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. No eficiency
and speed comparisons between these approaches and our framework have been done yet.
      </p>
      <p>In comparison to traditional ML approaches involving data movement between databases and
separate analytics platforms, the proposed framework ofers advantages in terms of reduced data
transfer, improved performance, and enhanced scalability. These advantages are all achieved by
using the database as the singular platform for data storage and analysis.</p>
    </sec>
    <sec id="sec-8">
      <title>7. Conclusion</title>
      <p>This paper introduced a novel, extendable, no-code framework to integrate ML into SQL with
Exasol. This framework bridges the gap between traditional SQL-based analytics and ML,
empowering users to perform advanced analytical tasks directly within the database environment.
We introduced Exasol, a high-performance, parallel-processing analytical database, and its
features relevant to the framework. These features include scripting programs, preprocessor
scripts, UDF scripts, script language containers, and the file system BucketFS. Furthermore, we
discussed related work in the form of other approaches for ML with Exasol and other
frameworks for in-database ML. The concept for the framework was introduced while discussing
the syntax of the CREATE MODEL command for creating a new ML model and the PREDICT
function for prediction using a pre-trained model in detail. The implementation of the framework
consists of three layers: preprocessor scripts, scripting programs, and UDF scripts. Each layer
provides a part of the complete functionality to translate incoming queries and execute the
required ML functionality. Additionally, metadata for ML models is tracked in tables. The
current restrictions of the framework and solutions were discussed. In comparison to other
frameworks, our framework stands out with its seamless integration into SQL but is probably
outshone regarding eficiency by frameworks re-implementing ML directly in the database. The
main contributions of our framework are its seamless integration into SQL, scalability, and
leveraging of existing database infrastructure.</p>
      <p>
        The framework was initially created during the master’s thesis “Extending SQL for Machine
Learning” [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ]. The source code is freely available at https://github.com/christoph-grossmann/
Exasol_DB_ML_Framework.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Verbraeken</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wolting</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Katzy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kloppenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Verbelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Rellermeyer</surname>
          </string-name>
          ,
          <article-title>A survey on distributed machine learning</article-title>
          ,
          <source>ACM Comput. Surv</source>
          .
          <volume>53</volume>
          (
          <year>2020</year>
          ). URL: https: //doi.org/10.1145/3377454. doi:
          <volume>10</volume>
          .1145/3377454.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Sarker</surname>
          </string-name>
          ,
          <article-title>Machine learning: Algorithms, real-world applications and research directions</article-title>
          ,
          <source>SN Computer Science</source>
          <volume>2</volume>
          (
          <year>2021</year>
          )
          <article-title>160</article-title>
          . URL: https://doi.org/10.1007/s42979-021-00592-x. doi:
          <volume>10</volume>
          . 1007/s42979- 021- 00592- x.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , G. Chen,
          <string-name>
            <given-names>H. V.</given-names>
            <surname>Jagadish</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. C.</given-names>
            <surname>Ooi</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.-L. Tan</surname>
          </string-name>
          ,
          <article-title>Database meets deep learning</article-title>
          ,
          <source>ACM SIGMOD Record</source>
          <volume>45</volume>
          (
          <year>2016</year>
          )
          <fpage>17</fpage>
          -
          <lpage>22</lpage>
          . URL: https://doi.org/10.1145%
          <fpage>2F3003665</fpage>
          . 3003669. doi:
          <volume>10</volume>
          .1145/3003665.3003669.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , Exasol website,
          <year>2023</year>
          . URL: https://www.exasol.com/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , Exasol white papers,
          <year>2023</year>
          . URL: https://www.exasol.com/resource-type/ white-papers/.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , BucketFS client,
          <year>2023</year>
          . URL: https://github.com/exasol/bucketfs-client.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , BucketFS database access,
          <year>2023</year>
          . URL: https://docs.exasol.com/db/latest/database_ concepts/bucketfs/database_access.htm.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , Exasol script languages,
          <year>2023</year>
          . URL: https://github.com/exasol/ script-languages-release.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , Adding new packages to existing
          <source>script languages</source>
          ,
          <year>2023</year>
          . URL: https: //docs.exasol.com/db/latest/database_concepts/udf_scripts/adding_new_packages_ script_languages.htm.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kozachuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Graupmann</surname>
          </string-name>
          ,
          <article-title>Bring your language to your data with EXASOL, in: Datenbanksysteme für Business</article-title>
          ,
          <source>Technologie und Web</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , UDF scripts,
          <year>2023</year>
          . URL: https://docs.exasol.com/db/latest/database_concepts/udf_ scripts.htm.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , Scripting programs,
          <year>2023</year>
          . URL: https://docs.exasol.com/db/latest/database_ concepts/scripting.htm.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , SQL preprocessor,
          <year>2023</year>
          . URL: https://docs.exasol.com/db/latest/database_concepts/ sql_preprocessor.htm.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nagpal</surname>
          </string-name>
          , G. Gabrani,
          <article-title>Python for data analytics, scientific</article-title>
          and
          <source>technical applications</source>
          ,
          <source>2019 Amity International Conference on Artificial Intelligence (AICAI)</source>
          (
          <year>2019</year>
          ). doi:
          <volume>10</volume>
          .1109/ AICAI.
          <year>2019</year>
          .
          <volume>8701341</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Raschka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Patterson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Nolet</surname>
          </string-name>
          ,
          <article-title>Machine learning in Python: Main developments and technology trends in data science, machine learning</article-title>
          ,
          <source>and artificial intelligence, Information</source>
          <volume>11</volume>
          (
          <year>2020</year>
          ). URL: https://www.mdpi.com/2078-2489/11/4/193. doi:
          <volume>10</volume>
          .3390/info11040193.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , E. Duchesnay, Scikit-Learn:
          <article-title>Machine Learning in Python</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <year>2011</year>
          )
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , Pyexasol,
          <year>2023</year>
          . URL: https://docs.exasol.com/db/latest/connect_exasol/drivers/ python/pyexasol.htm.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          ,
          <string-name>
            <surname>Exasol BucketFS Utils Python</surname>
          </string-name>
          ,
          <year>2023</year>
          . URL: https://github.com/exasol/ bucketfs-utils-python.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , Exasol BucketFS Python library,
          <year>2023</year>
          . URL: https://github.com/exasol/ bucketfs-python.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          ,
          <source>Data science with Exasol</source>
          ,
          <year>2023</year>
          . URL: https://github.com/exasol/ data-science-examples.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          ,
          <article-title>Data science</article-title>
          and
          <source>UDFs examples</source>
          ,
          <year>2023</year>
          . URL: https://docs.exasol.com/db/latest/ advanced_analytics/advancedexamples.htm.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          ,
          <source>Machine learning example in R and Python</source>
          ,
          <year>2022</year>
          . URL: https://exasol.my.site.com/ s/article/Machine-Learning-
          <article-title>Example-in-R-and-Python?language=en_US.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          , Exasol transformers extension,
          <year>2023</year>
          . URL: https://github.com/exasol/ transformers-extension.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>J. M. Hellerstein</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Ré</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Schoppmann</surname>
            ,
            <given-names>D. Z.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Fratkin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Gorajek</surname>
            ,
            <given-names>K. S.</given-names>
          </string-name>
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Welton</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>The MADlib analytics library: Or MAD skills, the SQL</article-title>
          ,
          <source>Proceedings of the VLDB Endowment</source>
          <volume>5</volume>
          (
          <year>2012</year>
          )
          <fpage>1700</fpage>
          -
          <lpage>1711</lpage>
          . URL: https://doi.org/10. 14778/2367502.2367510. doi:
          <volume>10</volume>
          .14778/2367502.2367510.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>The</given-names>
            <surname>Apache Software Foundation</surname>
          </string-name>
          , Apache MADlib community artifacts,
          <year>2023</year>
          . URL: https: //github.com/apache/madlib-site/tree/asf-site/community-artifacts.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Oracle</surname>
          </string-name>
          ,
          <source>Machine learning in Oracle database</source>
          ,
          <year>2023</year>
          . URL: https://www.oracle.
          <source>com/ artificial-intelligence/database-machine-learning/.</source>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>IBM</given-names>
            <surname>Corporation</surname>
          </string-name>
          ,
          <article-title>In-database machine learning Db2 database</article-title>
          ,
          <year>2023</year>
          . URL: https: //www.ibm.com/docs/en/db2/11.5
          <article-title>?topic=content-in-database-machine-learning&amp;mhsrc= ibmsearch_a&amp;mhq=machinelearningdatabase.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>M.</given-names>
            <surname>Schule</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lang</surname>
          </string-name>
          , M. Springer, A. Kemper,
          <string-name>
            <given-names>T.</given-names>
            <surname>Neumann</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Gunnemann, In-Database Machine Learning with SQL on GPUs</article-title>
          ,
          <source>in: 33rd International Conference on Scientific and Statistical Database Management</source>
          ,
          <string-name>
            <surname>SSDBM</surname>
          </string-name>
          <year>2021</year>
          ,
          <article-title>Association for Computing Machinery</article-title>
          , New York, NY, USA,
          <year>2021</year>
          , p.
          <fpage>2536</fpage>
          . URL: https://doi.org/10.1145/3468791.3468840. doi:
          <volume>10</volume>
          . 1145/3468791.3468840.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>N.</given-names>
            <surname>Makrynioti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ley-Wild</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vassalos</surname>
          </string-name>
          ,
          <article-title>Machine Learning in SQL by Translation to Tensorflow, in: Proceedings of the Fifth Workshop on Data Management for End-To-End Machine Learning</article-title>
          ,
          <source>DEEM '21</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2021</year>
          . URL: https://doi.org/10.1145/3462462.3468879. doi:
          <volume>10</volume>
          .1145/3462462.3468879.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>O.</given-names>
            <surname>Dolmatova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Augsten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Böhlen</surname>
          </string-name>
          ,
          <article-title>A relational matrix algebra and its implementation in a column store</article-title>
          ,
          <source>in: Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, ACM</source>
          ,
          <year>2020</year>
          . URL: https://doi.org/10.1145%
          <fpage>2F3318464</fpage>
          .3389747. doi:
          <volume>10</volume>
          .1145/3318464.3389747.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kersten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ivanova</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Nes, SciQL, a query language for science applications</article-title>
          ,
          <source>in: Proceedings of the EDBT/ICDT 2011 Workshop on Array Databases, AD '11</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2011</year>
          , p.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          . URL: https://doi.org/10.1145/ 1966895.1966896. doi:
          <volume>10</volume>
          .1145/1966895.1966896.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kersten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ivanova</surname>
          </string-name>
          , N. Nes,
          <article-title>SciQL: Bridging the gap between science and relational DBMS</article-title>
          ,
          <source>in: Proceedings of the 15th Symposium on International Database Engineering Applications</source>
          , IDEAS '11,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2011</year>
          , p.
          <fpage>124</fpage>
          -
          <lpage>133</lpage>
          . URL: https://doi.org/10.1145/2076623.2076639. doi:
          <volume>10</volume>
          .1145/ 2076623.2076639.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Khamis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. Q.</given-names>
            <surname>Ngo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Olteanu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schleich</surname>
          </string-name>
          , AC/DC: In-Database Learning Thunderstruck,
          <source>in: Proceedings of the Second Workshop on Data Management for End-ToEnd Machine Learning</source>
          ,
          <source>DEEM'18</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2018</year>
          . URL: https://doi.org/10.1145/3209889.3209896. doi:
          <volume>10</volume>
          .1145/3209889.3209896.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>J. V. D'silva</surname>
            ,
            <given-names>F. D.</given-names>
          </string-name>
          <string-name>
            <surname>Moor</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Kemme</surname>
          </string-name>
          , AIDA
          <article-title>- abstraction for advanced in-database analytics</article-title>
          ,
          <source>Proceedings of the VLDB Endowment</source>
          <volume>11</volume>
          (
          <year>2018</year>
          )
          <fpage>1400</fpage>
          -
          <lpage>1413</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Mlog:
          <article-title>Towards declarative in-database machine learning</article-title>
          ,
          <source>Proceedings of the VLDB Endowment</source>
          <volume>10</volume>
          (
          <year>2017</year>
          )
          <fpage>1933</fpage>
          -
          <lpage>1936</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Naughton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <article-title>Towards linear algebra over normalized data</article-title>
          ,
          <year>2017</year>
          . doi:
          <volume>10</volume>
          .14778/3137628.3137633.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>Exasol</surname>
          </string-name>
          ,
          <source>Exasol community edition</source>
          ,
          <year>2023</year>
          . URL: https://docs.exasol.com/db/latest/get_ started/communityedition.htm.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>C.</given-names>
            <surname>Großmann</surname>
          </string-name>
          ,
          <string-name>
            <surname>Extending</surname>
            <given-names>SQL</given-names>
          </string-name>
          <article-title>for Machine Learning</article-title>
          ,
          <source>Master thesis, Ostbayerische Technische Hochschule Regensburg</source>
          ,
          <year>2023</year>
          . doi:https://doi.org/10.35096/othr/pub- 6059.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>