<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Datenbank-Spektrum</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1007/s13222-019-00318-7</article-id>
      <title-group>
        <article-title>Use of Domain Knowledge to Support Industrial Data Analytics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Peter Reimann</string-name>
          <email>peter.reimann@gsame.uni-stuttgart.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Industrial Data Analytics, Domain-specific Data Characteristics, Metadata Modeling, Domain Knowledge</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Graduate School of Excellence advanced Manufacturing Engineering, University of Stuttgart</institution>
          ,
          <addr-line>Nobelstr. 12, Stuttgart</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>19</volume>
      <issue>2019</issue>
      <fpage>137</fpage>
      <lpage>148</lpage>
      <abstract>
        <p>Industrial Data Analytics refers to data analyses across diferent phases of the industrial product life cycle. The specific characteristics of available industrial data often pose challenges for common data management and data analysis methods. This paper gives an overview on the projects of the research group ICT Platform for Manufacturing at the Graduate School of Excellence advanced Manufacturing Engineering (GSaME) of the University of Stuttgart. These projects are related to Industrial Data Analytics and ofer approaches to addressing the domain-specific data characteristics. Relevant research areas are metadata management, use of domain knowledge to improve data preparation, the management of machine learning (ML) models, and approaches to meta-learning and automated machine learning (AutoML). In addition, this paper details on two specific research contributions. Firstly, it discusses a metadata model that facilitates a democratized access to data in virtual product development projects. The second contribution is an approach to exploit domain knowledge during data preparation in order to address two of the most important challenging data characteristics in industrial data: a multi-class imbalance and a data bias that is due the high variety of underlying products.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR</p>
      <p>
        ceur-ws.org
1. Introduction
Industrial Data Analytics refers to problems and solution
approaches to data management, data provision, and
data analytics across diferent phases of the industrial
product life cycle [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This paper gives an overview on
the research topics of the research group ICT Platform
for Manufacturing at the Graduate School of Excellence
advanced Manufacturing Engineering (GSaME) of the
University of Stuttgart. This research group deals with
both application-oriented and fundamental research in
the area of Industrial Data Analytics. It examines data
and their potential for data analysis in various phases of
a product life cycle, e.g., for analyzing simulation data
in the product development phase [
        <xref ref-type="bibr" rid="ref3">2</xref>
        ], sensor data from
test benches in the production phase [
        <xref ref-type="bibr" rid="ref4">3, 4</xref>
        ], or data from
the product usage phase describing the configurations of
sold products [5].
      </p>
      <p>The specific characteristics of available industrial data
pose challenges for common data management and data
analysis methods [6, 7, 8, 9]. For instance, data may come
in diverse and heterogeneous formats and be contained in
isolated data silos across diferent organizational units of
a company. This makes it dificult or even impossible to
over, the high product diversity increases the number
and complexity of patterns and correlations contained in
learning (AutoML) [18, 19].</p>
      <p>
        After giving an overview on related research projects
of the group ICT Platform for Manufacturing in Section 2,
this paper details on two specific contributions in
Sections 3 and 4. The first contribution is a metadata model
that connects data from heterogeneous and previously
The second major contribution is an approach to exploit
domain knowledge from a taxonomy during data
preparation in order to address two of the most important
challenging data characteristics and kinds of bias in
industrial data: a multi-class imbalance and a data bias that
is due the high variety of underlying products [
        <xref ref-type="bibr" rid="ref6">10, 14</xref>
        ].
acquire relevant data for a particular analysis [
        <xref ref-type="bibr" rid="ref3">2</xref>
        ]. More- isolated data sources in virtual product development [13].
      </p>
      <p>CRM
…
&amp; tr
sageU uppoS
2
7
3
MDA/PDA
6
CAQ</p>
    </sec>
    <sec id="sec-2">
      <title>Product Life</title>
    </sec>
    <sec id="sec-3">
      <title>Cycle</title>
      <p>4
5
CPS
MES
CAP
SCM</p>
      <sec id="sec-3-1">
        <title>Wiki</title>
        <p>D
lvee rP
opem tcduo 1
n
t</p>
      </sec>
      <sec id="sec-3-2">
        <title>Process</title>
      </sec>
      <sec id="sec-3-3">
        <title>Simulation PDM</title>
      </sec>
      <sec id="sec-3-4">
        <title>Product</title>
      </sec>
      <sec id="sec-3-5">
        <title>Simulation CAD</title>
      </sec>
      <sec id="sec-3-6">
        <title>Digital</title>
      </sec>
      <sec id="sec-3-7">
        <title>Fabric</title>
        <sec id="sec-3-7-1">
          <title>Metadata Model to Describe CAE Data and Work</title>
          <p>
            1 Activities in Virtual Product Development Projects
2. Research Group ICT Platform
for Manufacturing
and configurations of ML tools and software to perform,
e.g., data collection, data preprocessing, model training,
and model deployment [21]. One contribution of this
Figure 1 gives an overview on the projects and research project is a method to structure the collabaration among
topics related to Industrial Data Analytics in the research diferent stakeholdes, e.g., data scientists, IT experts,
busigroup ICT Platform for Manufacturing. Furthermore, the ness analysts, engineers or other domain experts during
ifgure assigns the projects to the phases of a typical prod- development projects for ML solutions [22]. In
addiuct life cycle, from which the projects mainly acquire and tion, AssistML is a novel concept that enhances the ML
analyze data. The first project deals with metadata mod- model management platform of the previous project by
eling in the area of virtual product development [
            <xref ref-type="bibr" rid="ref3">2</xref>
            ]. The approaches to meta-learning [17]. This way, AssistML
second project focuses on the production phase and how automates the discovery and recommendation of ML
soto identify quality issues in complex assembly products, lutions for a given use case, and it makes this task feasible
e.g., truck engines [
            <xref ref-type="bibr" rid="ref4">3</xref>
            ]. The major contributions of these for non-experts such as citizen data scientists.
two projects are discussed in Sections 3 and 4. The fith project deals with hybrid approaches and
          </p>
          <p>The major outcome of the third project is a platform their application to fault detection and fault diagnosis in
to manage machine learning (ML) models, e.g., classi- a production line [15]. Hybrid approaches combine
dataifcation or regression models [ 16]. Data scientists of a driven models with knowledge-based models, such as
company may develop ML models for specific use cases physics-based simulation models. These diferent kinds
and upload them in the platform. The ML models are of data-driven and knolwedge-driven models enhance
then associated with appropriate metadata to further de- and complement each other. This is particularly useful
scribe them. This metadata covers, amongst others, life in cases when pure data-driven solutions are not
adecycle information of the ML model, e.g., whether it is quate due to a lack of data. A major contribution of this
in the training or application phase or whether it has project is PUSION, a generic and automated framework
already been retired and is thus not used anymore [20]. for decision fusion in classification ensembles that may
Moreover, the metadata covers semantic information [16]. be composed of diverse data-driven or knowledge-driven
This for instance includes information about the domain- models [18]. By combining the decisions of these diverse
specific use case, e.g., fault detection or fault diagnosis, or classification models via decision fusion algorithms, the
about the machine in a production line for which the ML overall prediction accuracy may be enhanced.
model has been developed. It helps other stakeholders The focus of the next project is related to methods
ifnd appropriate models for their specific use cases and for data-driven prediction of product failures during the
thus efectively facilitates reusability of ML models. product usage phase [5]. Here, diferent kinds of data</p>
          <p>The contributions of the fourth project ofer struc- set shifts may occur, i.e., changes of the statistical data
tured methods to specify, configure and select whole ML distribution over time [23]. This means that both the
solutions. These ML solutions constitute combinations decision boundaries of classification patterns and the
ltua taa
x d
tnoeC teaM
s
e
rko iiit</p>
          <p>v
W tc</p>
          <p>A
s
r
taaD itnena
o</p>
          <p>C
ltua taa
x d
tnoeC teaM
factory:</p>
        </sec>
        <sec id="sec-3-7-2">
          <title>Stuttgart date: 04.03.20</title>
          <p>②
Product
Planning</p>
          <p>④
Product
Simulation</p>
          <p>⑥
Prototyping
①
Product
Specification</p>
          <p>③
Virtual
Prototype</p>
          <p>⑤
Simulation</p>
          <p>Result</p>
          <p>⑦
Physical
Prototype
…
⑧
Product
Testing
owner:
Mrs. X
…
dimensions:
3
voxels:
30x40x50
…
…
…
material:
paper
…
⑨
Test
Results
…
…
statistical distribution of these patterns may change. This or CSV-formated files. In addition, diferent data are
has to be reflected via adequate approaches to detect data contained in isolated data silos across individual
organiset shifts and to adapt classification models to the new zational units that are often not willing to share their data
statistical distribution if necessary. with other stakeholders, even from the same company.</p>
          <p>Finally, one project concerns data quality of text data. Altogether, this makes it nearly impossible for domain
exIt introduces the QUALM concept for continuous data perts, i.e., development engineers, to acquire and explore
quality measurement and improvement at diferent steps the data they need for a particular data analysis.
of text analysis pipelines [12]. QUALM data quality indi- Ziegler et al. [13] thus propose a metadata model to
cators quantify text characteristics, e.g., the number of describe all related CAx data and to address the
aboveabbreviations or spelling mistakes, and give hints how mentioned challenges. This metadata model not only
these may afect the quality of analysis results. Corre- describes data, but it ofers a connected view on data,
sponding QUALM modifiers use these text characteristics metadata, and work activities in virtual product
develto enhance the text quality. An example of such a modi- opment projects. Figure 2 shows an idealized example
ifer is an approach to select the best-fitting training data of an instance of the metadata model. It covers several
for an analysis task based on the similarity between this data containers [27, 28] (blue in the figure) that abstract
training data and the input data [24]. In addition, QUALM from heterogeneous data formats and point to the
unofers a hybrid method for information extraction, which derlying data sources or files via URIs. Furthermore, the
exploits both structured and unstructured data sources to metadata explicitly describes the work activities (yellow)
yield more relevant information from these sources [25]. that are carried out by development engineers in
virtual product development projects. These work activities
are connected to the data containers the activities
con3. Metadata Model for Virtual sume and produce. So, the whole metadata describes full
Product Development Projects and connected views of project workflows including the
activities and the data. The grey elements shown in
FigProduct development projects in companies are mainly ure 2 represent metadata that further describe either data
virtual and digitized thanks to several computer-aided containers or work activities via contextual information.
systems, e.g., for Computer-Aided Design (CAD), A major benefit of this metadata model is that product
Computer-Aided Engineering (CAE), and Computer- development engineers easily understand it, because it
Aided Testing (CAT) [26] These CAx systems produce a is based on the project workflows and work activities
huge amount of data that ofer additional opportunities these engineers carry out in their daily work life. This
for data analysis, e.g., to gain insights how to improve facilitates a democratized data access, so that product
dea product design. However, several challenges hinder velopment engineers may easily find the data associated
exploiting the full potential for data analysis [13]. In par- to the work activities in development projects they are
ticular, diferent CAx systems store their data in various familiar with. It facilities an expert-led data exploration
heterogeneous formats, e.g., proprietary 2D or 3D ge- and a subsequent data analysis to gain sophisticated
inometry files, plain text, images, videos, XML documents, sights from the underlying data. The metadata model
after SPH</p>
          <p>s
p
u
o
r
G
t
c
u
d
o
r
P</p>
        </sec>
        <sec id="sec-3-7-3">
          <title>Classes</title>
          <p>Heterogeneous
Segmentation
according to
Product
Hierarchy</p>
          <p>(SPH)</p>
        </sec>
      </sec>
      <sec id="sec-3-8">
        <title>Product Hierarchy</title>
        <sec id="sec-3-8-1">
          <title>Domain Knowledge</title>
          <p>4. Approach to Exploit Domain</p>
          <p>Knowledge for Data Preparation
Hirsch et al. [11] introduce a challenging use case of
data-driven identification of quality issues in complex
assembly products, e.g., truck engines. The use cases
constitutes a multi-class classification problem, where each
class corresponds to one of 84 engine components, e.g.,
cylinders, fuel injectors, or turbochargers. The problem
is to train a classification model that is able to identify
one of these multiple classes and engine components that
are the cause of a particular quality issue.</p>
          <p>With 1050 data instances, the data set of this use
cases is of rather low size and contains several kinds of
noise [11]. Nevertheless, literature comprises techniques
which are able to deal with these two data
characteristics. However, such data-driven methods usually still
show poor prediction performance, as they are not able
to address two additional kinds of domain-specific data
ance, i.e., the class labels occur in an imbalanced way in
the data [32]. Here, many learning algorithms tend to
ignore the patterns of class labels that are
underrepresented. The second challenging data characteristic results
from the fact that companies ofer</p>
          <p>heterogeneous product
groups with a high product variety. Here, the class
patterns, i.e., decision boundaries in the feature space of a
particular class usually difer across individual product
the class patterns of seldom product groups that are
underrepresented in data.</p>
          <p>
            Hirsch et al. propose an approach to data preparation
that efectively addresses both challenges arising from a
multi-class imbalance and from heterogeneous product
groups [
            <xref ref-type="bibr" rid="ref6">10, 14</xref>
            ]. Figure 3 shows the major steps of this
approach. It divides the whole data set  into several
subsets   ⊆  . After this data preparation, a classification
model is trained for each of the data subsets   .
          </p>
          <p>The first step, Segementation according to Product
Hierarchy (SPH), uses domain knowledge from a product
hierarchy or product taxonomy to divide the data set into
one subset   for each product group. As only data of
one particular product group is contained in each
subset   , this subset contains a significantly less number
of class patterns, i.e., usually only one pattern for each
remaining class in   . In addition, these class patterns
are more evenly distributed in each subset. So,
learndistinguish all class patterns, even those that have
previously been underrepresented in the whole data set</p>
          <p>The second step, Class Partitioning according to
Imbalance (CPI), further divides some of the subsets   to
mines the degree of class imbalance in each subset  
resulting from SPH using the Gini coeficient as an
imbalance metric. If the value of the Gini coeficient as higher
a subset  
than a threshold, e.g., 30 %, the subset   is divided into</p>
          <p>+ containing only data instance of majority
classes and a subset   − for instances of minority classes.</p>
          <p>Here, CPI uses a quantile approach to determine the point
of intersection between majority and minority classes.
for classification ensembles [ 30], e.g., Random Forest [31], ing algorithms have much less problems to identify and
characteristics [11]. The first one is a
multi-class imbal- address multi-class imbalance. Therefore, CPI first
deter</p>
          <p>
            Hirsch et al. [
            <xref ref-type="bibr" rid="ref6">10</xref>
            ] apply their approach to the data of
the above-mentioned uses case for a data-driven
identification of quality issues in assembled truck engines and
discuss the evaluation results. In addition, the authors
prove the generality of their approach by applying it to
several synthetic data sets that show varying data and
class distributions [14]. In both evaluations, they
compare their approach exploiting domain knowledge with
a data-driven baseline that applies Random Forest and a
feature selection technique to the whole data set  . They
show that their approach leads to an average increase of
classification accuracy between 4 and 13 %-points.
Furthermore, it leads to a reduction of the number of rework
steps needed to repair faulty truck engines.
          </p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Gröger</surname>
          </string-name>
          ,
          <string-name>
            <surname>Industrial Analytics -- An Overview</surname>
          </string-name>
          , it - Information
          <source>Technology</source>
          <volume>64</volume>
          (
          <year>2022</year>
          )
          <fpage>55</fpage>
          -
          <lpage>65</lpage>
          . doi:10.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>1515/itit-2021-0066.</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          , et al.,
          <article-title>A Graph-based Approach to Manage CAE Data in a Data Lake</article-title>
          ,
          <string-name>
            <surname>Procedia</surname>
            <given-names>CIRP</given-names>
          </string-name>
          93 (
          <year>2020</year>
          )
          <fpage>496</fpage>
          -
          <lpage>501</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.procir.
          <year>2020</year>
          .
          <volume>04</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.</given-names>
            <surname>Hirsch</surname>
          </string-name>
          , et al.,
          <article-title>Analytical Approach to Support Fault Diagnosis and Quality Control in End-Of-Line Testing</article-title>
          ,
          <string-name>
            <surname>Procedia</surname>
            <given-names>CIRP</given-names>
          </string-name>
          72 (
          <year>2018</year>
          )
          <fpage>1333</fpage>
          -
          <lpage>1338</lpage>
          . doi:10.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          1016/j.procir.
          <year>2018</year>
          .
          <volume>03</volume>
          .024.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>V.</given-names>
            <surname>Hirsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Reimann</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. Mitschang,</surname>
          </string-name>
          <article-title>ExAcknowledgments ploiting Domain Knowledge to Address MultiClass Imbalance and a Heterogeneous Feature This work was supported by the German Research Foun- Space in Classification Tasks for Manufacturing dation (DFG) within the Excellence Initiative II and by Data</article-title>
          ,
          <source>PVLDB</source>
          <volume>13</volume>
          (
          <year>2020</year>
          )
          <fpage>3258</fpage>
          -
          <lpage>3271</lpage>
          . doi:
          <volume>10</volume>
          .14778/ the Ministry of Science,
          <source>Research and Arts of the State of 3415478.3415549</source>
          .
          <article-title>Baden-Wurttemberg within the sustainability</article-title>
          support of [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Hirsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Reimann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mitschang</surname>
          </string-name>
          ,
          <article-title>Data-Driven the projects of the Excellence Initiative II. Further thanks Fault Diagnosis in End-of-Line Testing of Comgo to the colleagues that have been part of the research plex Products, in: Proc. of the 6th International group ICT Platform for Manufacturing from 2017 to 2022: Conference on Data Science and Advanced AnaVitali Hirsch</article-title>
          , Cornelia Kiefer, Marco Spieß,
          <article-title>Alejandro lytics (DSAA)</article-title>
          , IEEE, Washington,
          <string-name>
            <surname>D.C.</surname>
          </string-name>
          , USA,
          <year>2019</year>
          . Gabriel Villanueva Zacarias, Christian Weber, Yannick doi:
          <volume>10</volume>
          .1109/DSAA.
          <year>2019</year>
          .
          <volume>00064</volume>
          .
          <string-name>
            <surname>Wilhelm</surname>
            , and Julian Ziegler. [12]
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Kiefer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Reimann</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Mitschang</surname>
          </string-name>
          , QUALM: Ganzheitliche Messung und Verbesserung der Datenqualität in der Textanalyse,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>