<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of Par</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1007/978-981-99-7108-4_18</article-id>
      <title-group>
        <article-title>First Experiences on the Application of Lakehouses in Industrial Practice</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jan Schneider</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arnold Lutsch</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoph Gröger</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Holger Schwarz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bernhard Mitschang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute for Parallel and Distributed Systems, University of Stuttgart</institution>
          ,
          <addr-line>Universitätsstraße 38, 70569 Stuttgart</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Robert Bosch GmbH</institution>
          ,
          <addr-line>Borsigstraße 4, 70469 Stuttgart</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>8</volume>
      <fpage>213</fpage>
      <lpage>222</lpage>
      <abstract>
        <p>In recent years, so-called lakehouses have emerged as a new type of data platform that intends to combine characteristics of data warehouses and data lakes. Although companies started to employ the associated concepts and technologies as part of their analytics architectures, little is known about their practical medium- and long-term experiences as well as proven architectural decisions. Additionally, there is only limited knowledge about how lakehouses can be utilized efectively in an industrial context. Hence, it remains unclear under which circumstances lakehouses represent a viable alternative to conventional data platforms. To address this gap, we conducted a case study on a real-world industrial case, in which manufacturing data needs to be managed and analytically exploited. Within the scope of this case, a dedicated analytics department has been testing and leveraging a lakehouse approach for several months in a productive environment with high data volumes and various types of analytical workloads. The paper at hand presents the results of our within-case analyses and focuses on the industrial setting of the case as well as the architecture of the utilized lakehouse. This way, it provides preliminary insights on the application of lakehouses in industrial practice and refers to useful architectural decisions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Data Lakehouse</kwd>
        <kwd>Data Platform</kwd>
        <kwd>Platform Architecture</kwd>
        <kwd>Data Analytics</kwd>
        <kwd>Case Study</kwd>
        <kwd>Industry Experience</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        analyses and are typically less suited for many types of
advanced analytics [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9, 10</xref>
        ], such as data mining.
      </p>
      <p>
        With the growing range of capabilities for data acquisi- Data lakes attempt to close this gap by enabling the
tion and the current advances in the field of analytics [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], storage of raw data in their original formats, without
data is becoming an increasingly important asset for en- requiring the data to be transformed into a pre-defined
terprises of all business fields. For example, in the in- schema before it can be loaded onto the data platform.
dustrial sector, data from the shop floor can be exploited For this purpose, data lakes typically employ highly
scalwith data mining and machine learning techniques for able and cost-efective storage systems like distributed
eficiently orchestrating manufacturing processes, pre- ifle systems or object storages. However, in comparison
dicting the quality of products and scheduling the main- to data warehouses, this increased flexibility comes at
tenance of machines [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Similarly, this also applies to the cost of less comfortable data management and
analorganizations from other business fields, such as health- ysis capabilities. In summary, it can be concluded that
care [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and agriculture [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. data warehouses typically represent a reasonable choice
      </p>
      <p>
        Data platforms constitute the technical foundation for for use cases in which the analysis questions are already
all kinds of analytics applications within enterprises, as known in advance, while exploratory workloads with
they are capable of storing and managing huge amounts unknown analysis questions are more appropriately
supof data for analytical purposes and thus support data ported by data lakes. Consequently, enterprises often
collection, processing and analysis [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Traditional data need to operate both types of data platforms in parallel
warehouses [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and the more modern data lakes [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] repre- and either exchange or replicate the data between them
sent the two most popular types of data platforms. Origi- in order to be able to serve all kinds of analytical
worknally, they were designed for diferent kinds of analytics loads. This commonly results in complex architectures,
applications and hence tend to show rather opposing high operational costs and slow analytical processes [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
characteristics [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]: While conventional data warehouses In order to address these issues, eforts have recently
share many similarities with relational databases and been made to develop so-called lakehouses, which
repare primarily utilized for reporting and Online Analyti- resent another variant of data platform that intends to
cal Processing (OLAP) workloads, their proprietary data combine desirable characteristics of data warehouses and
formats and rigid data models impede explorative data data lakes. This way, lakehouses are supposed to serve
all kinds of analytical workloads by a single data
plat35th GI-Workshop on Foundations of Databases (Grundlagen von Daten- form. In literature, multiple diferent definitions for
lakebanken), May 22-24, 2024, Herdecke, Germany. houses exist [
        <xref ref-type="bibr" rid="ref9">9, 11, 12</xref>
        ] and apparently, there is currently
$ jan.schneider@ipvs.uni-stuttgart.de (J. Schneider)
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License no final consensus on how lakehouses can be
characAttribution 4.0 International (CC BY 4.0).
terized. Nevertheless, most authors seem to agree that
data lakes, which are based on highly scalable storage
systems and have been enhanced for additional data
warehousing capabilities with the help of certain
specialized frameworks, can generally be referred to as
lakehouses [13, 14, 15, 16, 17].
      </p>
      <p>The most popular representatives of such frameworks 2. The Manufacturing Case
are the open-source projects Delta Lake1, Apache Hudi2
and Apache Iceberg3. All of these frameworks provide
libraries for various popular data query and processing
engines, such as Apache Spark4 and Apache Flink5, and
enhance them for additional data warehousing
capabilities when working on data that resides on distributed
ifle systems or object storages. This includes features for
abstracting stored data as tables with relational
characteristics, ensuring ACID properties and enabling eficient
batch and stream processing [13, 14, 12]. Consequently,
data can be flexibly stored in open formats and in a
directly accessible manner, but still be conveniently
processed and queried. This allows such lakehouses to cover
large portions of the typical analytical workloads of data
warehouses and data lakes.</p>
      <p>In literature, several works can be found that
propose and discuss diferent lakehouse implementations
based on these frameworks, such as for the domains of
healthcare [18, 19, 20], biomedical research [21], network
management [22], IT security [23] and geospatial
analytics [24]. However, these descriptions often lack details
regarding the chosen architectural decisions and the
extent to which these decisions have proven over time,
especially in terms of aspects such as data organization,
data modelling and data flow. Furthermore, to the best
of our knowledge, there are yet no works available that
present industrial real-world implementations of
lakehouses for the manufacturing sector. As a consequence,
it remains unclear under which circumstances and with
which architectures lakehouses may represent a viable
alternative to conventional data warehouses and data
lakes for enterprises in this field. In an efort to address
this gap, we conducted a case study on a real-world in- Table 1
dustrial case in which an analytics department developed, Characteristics of the investigated manufacturing case.
tested and leveraged a lakehouse approach for the man- Source Data:
agement and analytical exploitation of manufacturing Source Systems:
data from the shop floor with a volume in the magnitude Data Types:
of terabytes. This lakehouse has been in use for several Data Volumes:
months in a productive environment with various analyt- Analytical
ical workloads and hence represents a suitable candidate Workloads:
for investigating architectural decisions. Analytics Types:</p>
      <p>Section 2 explains the methodological approach of our Users:
study and introduces the industrial setting of the case.</p>
      <sec id="sec-1-1">
        <title>Case studies are generally recognized as an appropriate</title>
        <p>research method for complex topics and topics where
the context needs to be taken into consideration [25].
Both applies to the research field of data platforms, as the
construction of such platforms constitutes complex tasks
that are subject to rapid developments and innovations,
while architectural decisions and practical experiences
of enterprises likely depend on contextual aspects, such
as the domain and size of the enterprise, the available
data volumes and the analytical use cases.</p>
        <p>In the scope of our case study for the manufacturing
case, we conducted interviews with one solution
architect and one data engineer of the responsible analytics
department, who both have been involved in the
development of the corresponding lakehouse. The interviews
lasted between 45 and 60 minutes and followed a
semistructured approach, for which questions had already
been prepared in advance, but were spontaneously
supplemented by follow-up questions during the interviews.
The asked questions were related to the context of the
manufacturing case, including the available data sources
and the requirements for the data analyses that are
supposed to be performed, as well as various architectural
aspects of the developed lakehouse (cf. Section 3).
Afterwards, the transcribed answers of the interview
participants were structured and analyzed with the help of
qualitative coding techniques.</p>
        <p>Table 1 summarizes important characteristics of the
investigated case, including details about the source data
and the intended analytical use cases.</p>
        <p>Machine and Sensor Data
Manufacturing Execution System
Structured, Unstructured
Terabytes
Reporting, OLAP, Machine Learning,
Near-realtime Reporting
Descriptive, Diagnostic, Predictive
Business Users, Data Analysts,</p>
        <p>Data Scientists</p>
      </sec>
      <sec id="sec-1-2">
        <title>Section 3 then presents the first results of our within-case analyses by discussing the architecture of the data platform and pointing to interesting architectural decisions. Finally, Section 4 concludes our work.</title>
      </sec>
      <sec id="sec-1-3">
        <title>1https://delta.io, accessed: 30.04.2024</title>
        <p>2https://hudi.apache.org, accessed: 30.04.2024
3https://iceberg.apache.org, accessed: 30.04.2024
4https://spark.apache.org, accessed: 30.04.2024
5https://flink.apache.org, accessed: 30.04.2024</p>
      </sec>
      <sec id="sec-1-4">
        <title>This manufacturing case is situated at a large-scale, globally operating manufacturer, which develops and produces technical components of high volume. Along the shop floor, manufacturing machines and sensors col</title>
        <p>Event Hub
Confluent Kafka</p>
        <p>Data Stream
Data</p>
        <p>Stream
Manufacturing</p>
        <p>Data</p>
        <p>Extract &amp; Load</p>
        <p>Harmonization</p>
        <p>Delivery
Serialized Data</p>
        <p>Delta Lake</p>
        <p>Tabulated Data</p>
        <p>Delta Lake</p>
        <p>Harmonized Data</p>
        <p>Delta Lake
Ingested Images
Various Formats</p>
        <p>Application Data</p>
        <p>Delta Lake
Machine Learning
Prepared Images
Various Formats</p>
        <p>Queries
Queries
Data</p>
        <p>Stream
Data Science</p>
        <p>Tools</p>
        <p>Reporting
OLAP
Near-realtime
Reporting
Machine Learning</p>
        <p>Cloud Object Storage</p>
        <p>Azure Data Lake Storage Gen2
Batch
Processing</p>
        <p>Stream
Processing</p>
        <p>Relational
Dataset
(Table)</p>
        <p>Data
Files</p>
        <p>Zone
lect various types of data, reflecting the quality of the ifed with both the lakehouse approach in general, as well
workpieces and the condition of the machines. This data as the architecture they have developed, since it allowed
is primarily made available in a Manufacturing Execu- them to meet their requirements in terms of performance
tion System (MES) and ingested as data stream into a and the required range of analytical workloads. We
therelakehouse-like data platform. Besides structured mea- fore assume that their architectural decisions regarding
surement values of the machines and sensors, this data the lakehouse have practically proven their suitability at
also includes graphical images of particularly error-prone least in the medium term.
parts of the produced technical components. A dedicated In the context of this work, lakehouses can largely be
analytics department constisting out of multiple solution regarded as data lakes that consist of a distributed file
architects and data engineers is involved in this case and system or an object storage and have been enhanced for
responsible for the development and operation of the data additional data warehousing capabilities with the help
platform. The goal of this data platform is the collection, of specific frameworks, such as Delta Lake, Apache Hudi
management, preparation and analysis of the generated and Apache Iceberg [12]. Therefore, we utilized the Data
manufacturing data in order to enable self-service ana- Lake Architecture Framework by Giebler et al. [26] as a
lytics for business users and data analysts. This includes reference to guide and structure our within-case analysis
traditional reporting and OLAP workloads, but also near- for the architecture of the data platform. However, we
realtime reporting, since some analysis results are sup- limited ourselves to the aspects Data Organization, Data
posed to be displayed on dashboards along the shop floor. Modeling, Data Flow, Data Storage and Infrastructure, as
In addition, data scientists pursue to train machine learn- these were most strongly covered during the interviews.
ing models from the available image data, which should Figure 1 illustrates the current architecture of the
lakeenable the automatically detection of faulty workpieces house that is applied in the investigated manufacturing
with the help of image classification techniques. case. Here, the lakehouse is depicted in the center, while
the data sources and analytical workloads are shown on
the left- and right-hand side. The lakehouse consists of
3. The Lakehouse Architecture multiple zones [27], which separate data of diferent
granularity, quality and purpose. The individual aspects of
As part of the case study, we examine the architecture of this architecture according to the Data Lake Architecture
the lakehouse data platform that is employed in the man- Framework are discussed in the following sections.
ufacturing case and, in particular, focus on interesting
architectural decisions that were made during its
development, for example with respect to aspects such as data
organization, data modeling and data flow.</p>
        <p>During the interviews, both participants confirmed to
us that after several months of operation, they are
satisthe Machine Learning Zone, where they can be used by preferred over continuously running stream processing
data scientists for machine learning experiments. jobs for economic reasons, b) that ingested raw data is
stored and managed as serialized JSON strings in tables
3.3. Data Modeling and not in raw text files, c) that data modelling is carried
out only informally and that de-normalization techniques
According to the participants of the interviews, data are applied in order to increase the query performance
modeling is carried out only informally, meaning that at the expense of higher costs for storage space and d)
data models are created ad-hoc depending on the struc- that zone models appear to be a suitable technique for
ture of the data that is provided by the data sources data organization within lakehouses.
and the requirements of the analytical use cases. Hence, In future work, we plan to compare this case with
widely researched modelling approaches, such as normal several other real-world cases from diferent domains
forms [28], multi-dimensional modeling [29] or the Data in terms of architectural similarities, the motivational
Vault concept [30] are not explicitly applied. Moreover, factors for enterprises to utilize lakehouses, practical
the data is intentionally de-normalized, as this enables a experiences and encountered challenges. This way, we
higher query performance at the cost of increased storage want to further expand the findings of our work and
space, which represents a reasonable trade-of for cloud become capable of generalizing them.
environments due to the lower costs for storage space in
comparison to computational resources.
3.4. Data Organization
The lakehouse architecture of the manufacturing case
applies a zone model for organizing the data of diferent
granularity, quality and application-specificity within
the lakehouse. In particular, this zone model defines
four zones: A Raw Zone, which stores the raw and only
slightly processed data, a Harmonization Zone, in which
the relational, consolidated and unified data resides, a
Delivery Zone for pre-aggregated, application-specific
data and a Machine Learning Zone that holds data that
is relevant for machine learning activities. These zones
can be roughly mapped to the Raw Zone, Harmonized
Zone, Delivery Zone and Explorative Zone of the Zone
Reference Model (ZRM), which was originally proposed
by Giebler at al. [27] for the data organization within data
lakes. According to the participants of the interviews,
both the use of a zone model in general, as well as the
zones that were specifically selected for this case have
proven their suitability. Therefore, it can be concluded
that zone models for data lakes appear also to be relevant
in the context of lakehouses and may be a suitable choice
for organizing the data in these kind of data platforms.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Conclusion</title>
      <p>This paper presented a real-world case in which a
lakehouse has been developed and leveraged for the
management and analysis of manufacturing data in
industrial practice. In the scope of our study, we particularly
focused on the architecture of the lakehouse, as well
as the industrial setting and underlying goals. In this
course, some interesting architectural decisions could
be observed: Our study revealed that in the investigated
case a) the periodic execution of stream processing jobs is</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Zhang</surname>
          </string-name>
          , R. Zhong,
          <article-title>Big Data Analytics for Intelligent Manufacturing Systems: A Review</article-title>
          ,
          <source>Journal of Manufacturing Systems</source>
          <volume>62</volume>
          (
          <year>2022</year>
          )
          <fpage>738</fpage>
          -
          <lpage>752</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.jmsy.
          <year>2021</year>
          .
          <volume>03</volume>
          .005.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dogan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Birant</surname>
          </string-name>
          ,
          <article-title>Machine learning and data mining in manufacturing</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>166</volume>
          (
          <year>2021</year>
          )
          <article-title>114060</article-title>
          . doi:
          <volume>10</volume>
          .1016/j. eswa.
          <year>2020</year>
          .
          <volume>114060</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rehman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Naz</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Razzak</surname>
          </string-name>
          ,
          <article-title>Leveraging big data analytics in healthcare enhancement: trends, challenges and opportunities</article-title>
          ,
          <source>Multimedia Systems</source>
          <volume>28</volume>
          (
          <year>2022</year>
          )
          <fpage>1339</fpage>
          -
          <lpage>1371</lpage>
          . doi:
          <volume>10</volume>
          .1007/ s00530-020-00736-8.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Javaid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Haleem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Suman</surname>
          </string-name>
          ,
          <source>Understanding the potential applications of Artificial Intelligence in Agriculture Sector, Advanced Agrochem 2</source>
          (
          <year>2023</year>
          )
          <fpage>15</fpage>
          -
          <lpage>30</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.aac.
          <year>2022</year>
          .
          <volume>10</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Gröger</surname>
          </string-name>
          , Industrial Analytics - An
          <string-name>
            <surname>Overview</surname>
          </string-name>
          , it - Information
          <source>Technology</source>
          <volume>64</volume>
          (
          <year>2022</year>
          )
          <fpage>55</fpage>
          -
          <lpage>65</lpage>
          . doi:
          <volume>10</volume>
          . 1515/itit-2021-0066.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>W. H.</given-names>
            <surname>Inmon</surname>
          </string-name>
          ,
          <article-title>Building the data warehouse, fourth edition</article-title>
          , 4th ed. ed., Wiley, Indianapolis, Ind.,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Giebler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gröger</surname>
          </string-name>
          , E. Hoos,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schwarz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mitschang</surname>
          </string-name>
          ,
          <article-title>Leveraging the Data Lake: Current State and Challenges, in: Big Data Analytics and Knowledge Discovery</article-title>
          , volume
          <volume>11708</volume>
          of Lecture Notes in Computer Science, Springer International Publishing, Cham,
          <year>2019</year>
          , pp.
          <fpage>179</fpage>
          -
          <lpage>188</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -27520-4_
          <fpage>13</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nambiar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mundra</surname>
          </string-name>
          ,
          <article-title>An Overview of Data Warehouse and Data Lake in Modern Enterprise Data Management</article-title>
          ,
          <source>Big Data and Cognitive Computing</source>
          <volume>6</volume>
          (
          <year>2022</year>
          )
          <article-title>132</article-title>
          . doi:
          <volume>10</volume>
          .3390/bdcc6040132.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Armbrust</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ghodsi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Xin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zaharia</surname>
          </string-name>
          , Lake-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>