<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a Holistic Data Preparation Tool</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Valerie Restat</string-name>
          <email>valerie.restat@fernuni-hagen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Meike Klettke</string-name>
          <email>meike.klettke@uni-rostock.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Uta Störl</string-name>
          <email>uta.stoerl@fernuni-hagen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Hagen</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Rostock</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Data-driven systems and machine learning-based decisions are becoming increasingly important and are having an impact on our everyday lives. The prerequisite for this is good data quality, which must be ensured by preprocessing the data. However, a number of challenges arise in the process. These include the results of the process in terms of data quality, e.g., combating bias and ensuring fairness, and the preprocessing process itself. Here, human involvement and the lack of intelligent solutions and applications for domain experts without in-depth IT knowledge play a major role. This paper summarizes these challenges and provides an overview of the current state of the art. It proposes the design of a holistic tool, along with the necessary tasks to overcome these challenges and to support data preprocessing.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;data preparation</kwd>
        <kwd>data quality</kwd>
        <kwd>data preprocessing</kwd>
        <kwd>data wrangling</kwd>
        <kwd>data cleaning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        With the increasing amount of data, the quality of data
is decreasing [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, data-driven systems
increasingly influence our everyday lives and support us in
making decisions [2]. This ranges from search and
recommendation services to medical diagnosis, hiring and loans
decisions [
        <xref ref-type="bibr" rid="ref1">1, 3</xref>
        ]. Thereby, the quality of the decisions
depends on the quality of the data [4], which makes
ensuring data quality a key issue in big data management
and an important aspect of almost every data-driven
project [5, 6]. In order to ensure the quality of the data,
it is necessary to preprocess them. This preprocessing
possesses a number of challenges that will have to be
solved by the data management community in the future.
      </p>
      <p>In this paper, we highlight the challenges currently
encountered in data preprocessing, with regard to the
requirements for the results of the process and the process
itself. To address these challenges, we propose the design
of a holistic tool and present several tasks that need to
be addressed for this purpose in future studies.</p>
      <p>The rest of the paper is structured as follows.
Section 2 describes the challenges facing data preprocessing.
Section 3 depicts the current state of the art. Section 4
presents the necessary tasks that must be addressed in
future work to achieve a holistic tool to support data
preprocessing. Finally, section 5 summarizes the paper.</p>
      <p>© 2022 Copyright for this paper by its authors. Use permitted under Creative
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAstthribouptionP4r.0oIncteerenadtiionnagl s(CC(CBYE4U.0)R.-WS.org)</p>
    </sec>
    <sec id="sec-2">
      <title>2. Challenges of Data</title>
    </sec>
    <sec id="sec-3">
      <title>Preprocessing</title>
      <p>Data preprocessing, also referred to as data wrangling,
data engineering or data preparation, is necessary to
ensure data quality. This includes among other steps data
profiling, data cleaning, data transformation [ 7], and data
integration. Especially data cleaning is important to
improve machine learning based solutions, as shown in [8]
and [9]. Thus, we focus mostly on this task in the
paper. Data cleaning consists mainly of two components:
Error detection and error repairing [10], whereby a
distinction can be made between classical approaches and
those based on machine learning [7]. Error detection is
done e.g. by integrity constraints like functional
dependencies or denial constraints, error repairing is done e.g.
by domain experts or by using reference data sets [10].
Error repairing is often more dificult because the use of
domain experts is time-consuming and expensive, and
reference data sets are usually not available in suficient
quality [10]. Hence, data cleaning is a continuous
process, the evaluation of which is essential for good quality
data [10]. The aspects and challenges of data quality are
described in the following section 2.1. Subsequently, the
challenges related to the preprocessing process itself are
presented in section 2.2.</p>
      <sec id="sec-3-1">
        <title>2.1. Data Quality</title>
        <sec id="sec-3-1-1">
          <title>Data of good quality is a prerequisite for data analysis</title>
          <p>and machine learning, which is why the results of data
preprocessing must be ensured in terms of quality [11].</p>
          <p>Although the definition of data quality varies in the
literature, it is undisputed that data quality depends on many
diferent factors and does not only concern accuracy.</p>
          <p>In [12], diferent data quality aspects and definitions
from 1985 to 2009 were studied and 40 dimensions were
identified, including timeliness, currency, accuracy and
completeness, to name the most referenced.</p>
          <p>These are also reflected in [ 11], in which a hierarchical
data quality framework was formulated from the
perspective of data users. They identified the following five
dimensions:
In addition to the challenges posed by data quality,
practitioners face many challenges in the data preprocessing
workflow itself. In [ 14], the results of a user survey of
data analysts and infrastructure engineers show that data
cleaning is time-consuming and needs human
involvement. Participants described data cleaning as an
itera• Availability: This refers to timeliness, also men- tive approach in which evaluation is not automated but
tioned in [12], and accessibility (also a part of the merely ad-hoc. In addition, a discrepancy emerged
beFAIR principles Findable, Accessible, Interoperable tween the data analysts and the infrastructure engineers,
and Re-usable [13]). who define data quality diferently. While the engineers
• Usability: This includes documentation, credibil- tended to look at errors in data at the syntactic level,
ity and metadata. e.g., incorrect schemas or inconsistencies, that require a
precise definition of the errors, analysts tended to
inves• Reliability: Elements of this dimension include tigate semantic errors through domain expertise, which
the aforementioned aspects of accuracy and com- is more dificult to translate into clear rules.
pleteness, as well as integrity, consistency and A variety of tools exist to automate data cleaning.
auditability. In [15], the authors examined multiple tools using
real• Relevance: This refers to the fitness of data, which world data sets. The results of the examination show
plays a particularly important role in terms of that there is no single dominant tool, as most tools are
fairness, which will be discussed later in greater only suitable for a certain type of error, while a variety
detail. of diferent errors occur in real-world data. A holistic
• Presentation Quality: The last dimension includes multi-tool strategy improved the results, but still failed
readability and structure. These are crucial to to achieve acceptable error coverage. The authors thus
arrive at a valid description of the data to enhance also emphasize the need for human involvement. Further,
users’ understanding of these data. they highlight the need for real-world data sets for the
development and testing of new approaches as well as
advances in combination of data cleaning tools.</p>
          <p>Data cleaning tools were also investigated in [16] and
a survey was conducted for this purpose. The results
show that most tools already require a preprocessed and
cleaned data set as input, e.g. with uniform delimiters
or the same number of fields per row. It is also shown
here that human involvement is necessary, both domain
knowledge and IT-knowledge are required for the usability
of the tools. As in [15], the lack of intelligent solutions is
addressed and the need for automated data preparation
tasks and pipelining is highlighted.</p>
          <p>
            As mentioned earlier, the increasing amount of data
generates additional challenges to data preprocessing,
especially in terms of volume and variety, also shown
in [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. It is noted here that due to the ever-increasing
amount of data, quality is compromised and also many
data cleaning tools do not scale suficiently. In terms
of variety, the challenges not only arise from diverse
data formats but also from the variety of diferent errors.
          </p>
          <p>Moreover, repairing the data is dificult due to diferent
constraints and can also possibly lead to new errors. The
need for domain experts is underlined here as well.</p>
          <p>Another important aspect is to create metadata during
preprocessing and to track and document data
provenance [17]. The latter is especially important for
reproducibility. The terminology is not standardized [18], so
we adopt ACM’s definition [ 19]. It states that an
exper</p>
          <p>Another important aspect, which has also gained
increasing attention in recent years is the evaluation of data
quality in terms of bias and fairness. If a prediction uses
a data set that is not representative of the entire
population for which predictions are being made, there is a bias
in the data [2]. This goes hand in hand with the
aforementioned fitness of the data. In addition, prejudices,
preconceptions, and various historical perceptions may
be contained in data [2]. This problem is amplified when
biased data are used for algorithms whose consequently
biased output is in turn used for further predictions [2].</p>
          <p>Bias can also be introduced into the data through
preprocessing. This can be caused by several variants during
the data cleaning or data filtering step, such as an
accidental introduction of bias by methods for missing value
imputation, as shown in [3]. The example presented here
considers a form that ofers a binary gender choice, and
the option of not specifying gender. Suppose about half of
the respondents were women and half men, but women
were more likely than men not to report their gender
and some respondents would also identify themselves as
non-binary. If mode imputation were applied now, all
unspecified values would be set as male, thus skewing the
distribution and even excluding non-binary individuals.</p>
          <p>There are various other examples of this, making it clear
that ways to improve data quality and control bias are
an important aspect of data preprocessing [3].
iment is reproducible if a diferent team with the same stream machine learning model. HoloClean [24] is a
measurement procedure and setup can obtain the mea- system for automatic error repair that uses, in addition to
surements under the same operational conditions. the data set to be cleaned, integrity constraints, external</p>
          <p>In [20], the iterative nature of data preprocessing is data, and matching dependencies as input to create a
mentioned as a particular challenge for the reproducibil- probabilistic model, which suggests a cleaned data set
ity of such pipelines. Furthermore, it is described that it based on statistical learning and probabilistic inference.
is important to understand data lineage in order to en- HoloDetect [25], an error detection system, is related to
sure reproducibility. This also includes the identification this. In addition to the data set to be cleaned, a
trainof errors and the possibilities of a rollback, using data ing data set and optional denial constraints are required
lineage. as input. Through data augmentation, the training data</p>
          <p>
            The importance of reproducibility is also emphasized set is expanded and a machine learning model is trained
in [21], and with it the need for new algorithms to be that classifies whether a cell is faulty or not. Raha [
            <xref ref-type="bibr" rid="ref2">26</xref>
            ]
comparable. It also states that the reproducibility of data and Baran [
            <xref ref-type="bibr" rid="ref3">27</xref>
            ] are systems for error detection and error
preprocessing in particular has so far received little at- repairing, respectively. They do not require
configuratention, compared to other areas. So the importance of a tion and thus reduce human involvement. Instead, the
well-defined data preprocessing process is highlighted. configuration is done automatically. Raha is based on
clustering to train a binary classifier that predicts for
each cell whether it is clean or dirty. Baran can
subse3. State of the Art quently be used for error repairing. In an optional ofline
phase, external sources with value-based corrections can
In this chapter, we describe the current state of the art be used to pre-train error corrector models. In the online
in data preprocessing tools. A number of tools already phase, the error corrector models are updated and used
exist, some of which have been studied in [16]: Altair to generate potential corrections where a binary
classiPPMrroeenppaaarrraachttiiooDnna42,,taTSaAPbrPleepAaaugriaPlterieoDpn5a1,t,aTPaPalxreeanptdaaDrSaaetitlofanSP3er,reSvpAiacrSeatDDioaanttaa6
vifemraalpicdrhaeitdnioiecntlsecaairfpniaitbniigslitpainelast.afcIottrurmealliwecsoitrohrnedcaattisaocnha.enmaTlaFyXstios[daen2t8de]cdtiasetraaand Trifacta Wrangler7. Initially, 42 commercial tools rors and suggest possible fixes. Further, the schema and
were selected, of which these seven were chosen based on its diferent versions can be used to analyze the evolution
following criteria: Tools specific to the data preprocess- of the data.
ing task, comprehensive coverage of 40 features, guides However, even if especially Raha and Baran reduce
and documentation, availability of a trial version, a GUI, human involvement, it is mostly needed in the other
and customer support. The evaluation of the seven tools tools and IT knowledge is partly necessary in addition
based on three data sets and 40 features, classified into to domain knowledge. Moreover, none of these systems
the following six categories: data discovery, data valida- can be considered a holistic tool for data preprocessing.
tion, data structuring, data enrichment, data filtering, and They focus only on a certain aspect of data cleaning (e.g.,
data cleaning. The results show that there is no tool that only error detection) and only on certain types of errors
can cover all these features. As described in section 2.2, (e.g., domain value violations). The evaluation is mostly
all tools require a pre-preprocessing as well, and the au- done on only one or a few aspects of data quality (mostly
thors describe data preprocessing as a mainly manual accuracy), fairness for example is not considered in any
task performed by domain experts with data engineering of the systems, neither is reproducibility. Moreover, they
knowledge. are only suitable for tabular data; semi-structured or
un
          </p>
          <p>
            Other systems for data cleaning are described also in structured data cannot be used. Yet there is an increasing
literature, some of which are mentioned in [22]. Boost- number of formats like JSON or text data. The data
qualClean [23] detects and repairs domain value violations. ity of semi-structured and unstructured data need to be
As input, it requires a relational table, libraries of func- better explored in the future [
            <xref ref-type="bibr" rid="ref5">29</xref>
            ].
tions for detecting and repairing errors, and a
userspecified classifier training procedure. The system relies
on boosting to maximize the performance of a
down4. Holistic Data Preparation Tool
1https://www.altair.com/monarch/
2https://www.paxata.com/self-service-data-prep/
3https://www.sap.com/germany/products/
database-data-management.html
4https://www.sas.com/en_us/software/data-preparation.html
5https://www.tableau.com/products/prep
6https://www.talend.com/products/data-preparation/
7https://www.trifacta.com/products/why-trifacta/
          </p>
        </sec>
        <sec id="sec-3-1-2">
          <title>In line with the challenges described in section 2, we now</title>
          <p>outline tasks that the community will have to address
in the future. We propose the design of a holistic tool
to support domain experts in data preprocessing, which
must overcome the challenges mentioned.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>4.1. Data Quality</title>
        <sec id="sec-3-2-1">
          <title>In section 2.1, we described the challenges that arise in connection with data quality. Here we see the following tasks:</title>
          <p>Ensuring fairness and explainability As mentioned
in section 2.1, data quality plays an essential role.
Combating bias and ensuring fairness are an important aspect
of preprocessing. This especially applies to automated
solutions [3]. Therefore, we would particularly like to
highlight the need for such a tool to help data scientists
identify biases and unfairness present in the data or
arising during preprocessing. For the former, there should
be tool support for extensive testing; for the latter, the
data changes need to be measured as described in [4]. To
date, there are no standard methods for measuring data
changes. Developing appropriate metrics and
descriptions are important research topics in future studies [4].</p>
          <p>
            The same applies to explainability, an issue often
associated with fairness. Even though many studies consider
the explainability of machine learning models themselves,
many decisions that afect the behavior of the models
are made in preprocessing [
            <xref ref-type="bibr" rid="ref6">30</xref>
            ]. Logging and
measuring data changes through benchmarks and analysis of
algorithm characteristics can help establish explainability
and examine preprocessing steps in terms of
introducing bias [4]. Moreover, consumer labels such as those
envisioned for machine learning models in [
            <xref ref-type="bibr" rid="ref7">31</xref>
            ] could
also support combating bias and ensuring fairness and
explainability [4].
          </p>
          <p>
            Nevertheless, the topic of fairness is a very complex
one and dificult to grasp. For example, in [
            <xref ref-type="bibr" rid="ref8">32</xref>
            ], almost
twenty diferent types of bias were presented. Moreover,
most fairness metrics do not consider the social
consequences of decisions based on predictions by machine
learning models [2]. In [
            <xref ref-type="bibr" rid="ref9">33</xref>
            ], it is argued that current
fairness metrics and research in ethical AI are not suficient.
Many practical issues need to be considered, several of
which have been presented here. Nevertheless, the
proposed approaches will provide initial help to address this
complex issue, since social-minded measures are crucial
for data quality.
can be made according to the arrival time and
intervals of the data.
• Usability: To ensure credibility, in addition to
verification by domain experts, automated checks
could be made according to the range of data or
accepted values.
• Reliability: At this point, an evaluation according
to precision and recall could take place as well
as the examination of integrity constraints and
functional dependencies and the adherence to
formats. For recall and precision ground truth is
required, but often not available.
• Relevance: Whether data is suitable for a use case
primarily depends on the goal of the analysis or
prediction and must be assessed by a domain
expert. This expert can be supported in the
decisionmaking process by providing automated
information about the distribution of the data, warning
of protected features and fairness metrics. As
explained above, an audit based on fairness metrics
is not suficient for an assessment according to
ethical AI, but it would be a starting point to
counteract unfairness and bias. Measuring changes
in data can further support this, as well as the
consumer labels mentioned.
• Presentation Quality: This also requires human
involvement and can be supported by automated
checks based of certain standards and
specifications.
          </p>
          <p>
            In order to take a first step in this direction and to
create a basis for working on the aspects mentioned above,
we are extending the data generator implemented as part
of the EvoBench project [
            <xref ref-type="bibr" rid="ref10">34</xref>
            ]. Since data sets are essential
for testing new approaches, we aim to generate test data
for evaluating data preparation pipelines. Therefore, we
generate targeted data sets with specific error types that
are as close to reality as possible. These can then be used
to run tests and evaluate approaches.
          </p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>4.2. Workflow Properties</title>
        <p>• Availability: Data accessibility comes into play
before preprocessing and is therefore not
considered here. With regard to timelessness, checks</p>
        <sec id="sec-3-3-1">
          <title>In section 2.2, we described the challenges currently facing the preparation process itself. In this context, we see</title>
          <p>Comprehensive integrated evaluation As de- the following tasks:
scribed in section 2, the evaluation of preprocessing is
usually dificult, especially due to its iterative nature.</p>
          <p>According to the before mentioned aspect to ensure
fairness and measure data changes, we propose a
comprehensive evaluation related to all dimensions of
data quality, described in [11] and presented in this
paper in section 2.1.</p>
          <p>High-level application and abstraction from
ITknowledge One of the most frequently cited
challenges relates to human involvement and required user
expertise. Human involvement by domain experts is
indispensable in data preprocessing. To make their work
as easy as possible and to overcome the aforementioned
discrepancy between data analysts and infrastructure
engineers, domain experts must be able to use tools without
in-depth IT-knowledge.
• Visual and interactive approaches
• Approaches based on natural language
process</p>
          <p>ing
• Example-oriented approaches</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>These are intended to help select the appropriate prepro</title>
          <p>
            cessing method according to the data and the objective
and provide an intelligent guidance for data cleaning. To
support users in choosing the appropriate preprocessing
methods, the aforementioned consumer labels [
            <xref ref-type="bibr" rid="ref7">31</xref>
            ] are
conceivable, for example.
          </p>
          <p>Support of semi-structured and unstructured data
As described in section 3, most tools are only suitable for
relational data. The data quality of semi-structured and
unstructured data needs to be better explored as well as
tools developed for such data formats.</p>
          <p>This will also be one of the research questions that we
want to investigate next.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusion and Future Work</title>
      <p>
        Hence, we suggest that the tool should have high-level used need to be scalable. As described in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], this is not
functions that can be applied without IT-knowledge. First the case with most current tools.
approaches in this direction could be:
      </p>
      <sec id="sec-4-1">
        <title>In this paper, we have identified the challenges that arise</title>
        <p>in data preprocessing and, depending on them, presented
Minimizing human involvement Even though hu- a series of tasks that need to be addressed in future work.
man involvement is indispensable, it is very expensive. For maximum benefit, the diferent aspects must be
comTherefore, the manual efort of data preprocessing must bined, which is why we have proposed the concept of a
be reduced as much as possible. This underlines the need holistic tool.
for intelligent and automated solutions that combine in- Furthermore, to produce a holistic data preparation
dividual tools in a pipeline and optimize it, as referred to tool, we intend to investigate the following research
quesin [15] and [16]. tions in future work:</p>
        <p>
          For this purpose, the usage of machine learning
techniques is conceivable, for example. This could be applied
to automatically predict the default configuration settings
of tools and pipelines, similar to the automated tuning of
database systems, e.g. in [
          <xref ref-type="bibr" rid="ref11">35</xref>
          ]. Raha [
          <xref ref-type="bibr" rid="ref2">26</xref>
          ] and Baran [
          <xref ref-type="bibr" rid="ref3">27</xref>
          ]
are, as described, first such approaches in this direction.
        </p>
        <p>Alternatively, the use of dictionaries or knowledge stores
could be a possible solution.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Data Lineage In the challenges mentioned in sec</title>
        <p>tion 2.2 the need for data lineage and the associated
reproducibility of data preprocessing was emphasized.</p>
        <p>For this reason, the proposed tool should provide data
lineage tracking capabilities, as well as rollback
capabilities. Despite the iterative nature of data preprocessing,
in addition to being measured, changes must be tracked
and, if necessary, reversed. Furthermore, results of the
preprocessing pipeline ought to be reproducible. This
once again highlights the need for tracking data lineage,
as the raw data and ready preprocessed data alone are
not suficient for reproducibility, as shown in [21].</p>
        <p>This means that in addition to these data, the
preprocessing process itself also needs to be documented and
a description must be provided which algorithms were
used to modify data and which parameters were applied.</p>
        <p>One of the next research questions we want to explore
will be how best to accomplish this.</p>
        <p>Scalability In section 2.2, the challenges posed by
growing data volumes have been described. To cope
with the increasing amount of data and the resulting
challenges mentioned, the proposed tool and algorithms
• Which preprocessing algorithms are suitable for</p>
        <p>which data?
• Which of these algorithms are covered by which</p>
        <p>existing tools?
• What is the most eficient way to combine
diferent data preprocessing tools in a pipeline in terms
of accuracy and to minimize false positives?
• How can the pipeline be evaluated automatically,
with respect to all the aspects of data quality
mentioned above?
• How can the data quality of semi-structured and</p>
        <p>unstructured data be ensured?
• How can data changes be reliably measured?
• What measures need to be taken to detect bias in</p>
        <p>the data and counteract unfairness?
• How can data lineage, especially in iterative
processes, be reliably tracked and reproducibility
ensured?
• How can machine learning be used to
automatically preprocess data or generate suggestions for
domain experts?
[2] E. Pitoura, Social-minded Measures of Data Quality: Scientific Data (2016) 160018. URL: https://doi.org/
Fairness, Diversity, and Lack of Bias, ACM J. Data 10.1038/sdata.2016.18.</p>
        <p>Inf. Qual. (2020) 12:1–12:8. URL: https://dl.acm.org/ [14] S. Krishnan, et al., Towards reliable interactive data
doi/10.1145/3404193. cleaning: a user survey and recommendations, in:
[3] S. Schelter, J. Stoyanovich, Taming technical bias Proceedings of the Workshop on
Human-In-thein machine learning pipelines, IEEE Data Engi- Loop Data Analytics, HILDA@SIGMOD 2016, San
neering Bulletin (Special Issue on Interdisciplinary Francisco, CA, USA, June 26 - July 01, 2016, ACM,
Perspectives on Fairness and Artificial Intelligence 2016, p. 9. URL: https://doi.org/10.1145/2939502.</p>
        <p>Systems) (2020) 39–50. 2939511.
[4] M. Klettke, A. Lutsch, U. Störl, Kurz erklärt: Mea- [15] Z. Abedjan, et al., Detecting Data Errors: Where are
suring data changes in data engineering and their we and what needs to be done?, Proc. VLDB Endow.
impact on explainability and algorithm fairness, (2016) 993–1004. URL: http://www.vldb.org/pvldb/
Datenbank-Spektrum (2021). URL: https://doi.org/ vol9/p993-abedjan.pdf .</p>
        <p>10.1007/s13222-021-00392-w. [16] M. Hameed, F. Naumann, Data Preparation: A
Sur[5] F. Endel, H. Piringer, Data Wrangling: Making vey of Commercial Tools, SIGMOD Rec. (2020) 18–
data useful again, IFAC-PapersOnLine (2015) 111– 29. URL: https://doi.org/10.1145/3444831.3444835.
112. URL: https://www.sciencedirect.com/science/ [17] C. A. Goble, et al., FAIR Computational Workflows,
article/pii/S2405896315001986, 8th Vienna Interna- Data Intell. (2020) 108–121. URL: https://doi.org/10.
tional Conferenceon Mathematical Modelling. 1162/dint_a_00033.
[6] W. Fan, F. Geerts, Foundations of Data Qual- [18] W. Mauerer, S. Scherzinger, Nullius in Verba:
Reity Management, Morgan &amp; Claypool Pub- producibility for Database Systems Research,
Relishers, 2012. URL: https://doi.org/10.2200/ visited, in: 37th IEEE International Conference
S00439ED1V01Y201207DTM030. on Data Engineering, ICDE 2021, Chania, Greece,
[7] M. Klettke, U. Störl, Four Generations in Data April 19-22, 2021, IEEE, 2021, pp. 2377–2380. URL:
Engineering for Data Science: The Past, Presence https://doi.org/10.1109/ICDE51399.2021.00270.
and Future of a Field of Science, Datenbank- [19] ACM, Artifact review and badging – version
Spektrum (2021). URL: https://doi.org/10.1007/ 2.0, 2021. URL: https://www.acm.org/publications/
s13222-021-00399-3. policies/artifact-review-badging.
[8] M. Mahdavi, et al., Towards Automated Data Clean- [20] L. Rupprecht, et al., Improving Reproducibility
ing Workflows, in: Proceedings of the Confer- of Data Science Pipelines through Transparent
ence on "Lernen, Wissen, Daten, Analysen", Berlin, Provenance Capture, Proc. VLDB Endow. (2020)
Germany, September 30 - October 2, 2019, CEUR- 3354–3368. URL: http://www.vldb.org/pvldb/vol13/
WS.org, 2019, pp. 10–19. URL: http://ceur-ws.org/ p3354-rupprecht.pdf .</p>
        <p>Vol-2454/paper_8.pdf . [21] M. Pawlik, et al., A Link is not Enough
[9] P. Li, et al., CleanML: A Study for Evaluating the - Reproducibility of Data,
DatenbankImpact of Data Cleaning on ML Classification Tasks, Spektrum (2019) 107–115. URL: https:
in: 37th IEEE International Conference on Data //doi.org/10.1007/s13222-019-00317-8.
Engineering, ICDE 2021, Chania, Greece, April 19- [22] M. Boehm, A. Kumar, J. Yang, Data Management
22, 2021, IEEE, 2021, pp. 13–24. URL: https://doi. in Machine Learning Systems, Morgan &amp;
Clayorg/10.1109/ICDE51399.2021.00009. pool Publishers, 2019. URL: https://doi.org/10.2200/
[10] I. F. Ilyas, Efective Data Cleaning with Continuous S00895ED1V01Y201901DTM057.</p>
        <p>Evaluation, IEEE Data Eng. Bull. (2016) 38–46. URL: [23] S. Krishnan, et al., BoostClean: Automated
Erhttp://sites.computer.org/debull/A16june/p38.pdf . ror Detection and Repair for Machine Learning,
[11] L. Cai, Y. Zhu, The Challenges of Data Quality CoRR (2017). URL: http://arxiv.org/abs/1711.01299.
and Data Quality Assessment in the Big Data Era, arXiv:1711.01299.</p>
        <p>Data Sci. J. (2015) 2. URL: https://doi.org/10.5334/ [24] T. Rekatsinas, et al., HoloClean:
Holisdsj-2015-002. tic Data Repairs with Probabilistic Inference,
[12] F. Sidi, et al., Data quality: A survey of data qual- CoRR (2017). URL: http://arxiv.org/abs/1702.00820.
ity dimensions, in: 2012 International Conference arXiv:1702.00820.
on Information Retrieval &amp; Knowledge Manage- [25] A. Heidari, et al., HoloDetect: Few-Shot
Learnment, Kuala Lumpur, Malaysia, March 13-15, 2012, ing for Error Detection, in: Proceedings of the
IEEE, 2012, pp. 300–304. URL: https://doi.org/10. 2019 International Conference on Management of
1109/InfRKM.2012.6204995. Data, SIGMOD Conference 2019, Amsterdam, The
[13] M. D. Wilkinson, et al., The FAIR Guiding Principles Netherlands, June 30 - July 5, 2019, ACM, 2019,
for scientific data management and stewardship, pp. 829–846. URL: https://doi.org/10.1145/3299869.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Ridzuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. M. N. Wan</given-names>
            <surname>Zainon</surname>
          </string-name>
          ,
          <article-title>A Review on Data Cleansing Methods for Big Data, Procedia Computer Science (</article-title>
          <year>2019</year>
          )
          <fpage>731</fpage>
          -
          <lpage>738</lpage>
          . URL: https://www.sciencedirect.com/science/ article/pii/S1877050919318885, the Fifth Information Systems International Conference,
          <volume>23</volume>
          -
          <issue>24</issue>
          <year>July 2019</year>
          , Surabaya, Indonesia.
          <volume>3319888</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mahdavi</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Raha</surname>
            :
            <given-names>A</given-names>
          </string-name>
          <string-name>
            <surname>Configuration-Free Error</surname>
          </string-name>
          Detection System,
          <source>in: Proceedings of the 2019 International Conference on Management of Data, SIGMOD Conference</source>
          <year>2019</year>
          , Amsterdam, The Netherlands, June 30 - July 5,
          <year>2019</year>
          , ACM,
          <year>2019</year>
          , pp.
          <fpage>865</fpage>
          -
          <lpage>882</lpage>
          . URL: https://doi.org/10.1145/3299869. 3324956.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mahdavi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Abedjan</surname>
          </string-name>
          ,
          <article-title>Baran: Efective Error Correction via a Unified Context Representation and Transfer Learning</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <article-title>(</article-title>
          <year>2020</year>
          )
          <fpage>1948</fpage>
          -
          <lpage>1961</lpage>
          . URL: http://www.vldb.org/pvldb/vol13/ p1948-mahdavi.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>D.</given-names>
            <surname>Baylor</surname>
          </string-name>
          , et al.,
          <article-title>TFX: A TensorFlow-Based Production-Scale Machine Learning Platform</article-title>
          ,
          <source>in: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , Halifax,
          <string-name>
            <surname>NS</surname>
          </string-name>
          , Canada,
          <source>August 13 - 17</source>
          ,
          <year>2017</year>
          , ACM,
          <year>2017</year>
          , pp.
          <fpage>1387</fpage>
          -
          <lpage>1395</lpage>
          . URL: https: //doi.org/10.1145/3097983.3098021.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chu</surname>
          </string-name>
          , et al.,
          <article-title>Data Cleaning: Overview and Emerging Challenges</article-title>
          ,
          <source>in: Proceedings of the 2016 International Conference on Management of Data, SIGMOD Conference</source>
          <year>2016</year>
          , San Francisco, CA, USA, June 26 - July 01,
          <year>2016</year>
          , ACM,
          <year>2016</year>
          , pp.
          <fpage>2201</fpage>
          -
          <lpage>2206</lpage>
          . URL: https://doi.org/10.1145/2882903.2912574.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>C. V. G.</given-names>
            <surname>Zelaya</surname>
          </string-name>
          ,
          <article-title>Towards Explaining the Efects of Data Preprocessing on Machine Learning</article-title>
          ,
          <source>in: 35th IEEE International Conference on Data Engineering, ICDE</source>
          <year>2019</year>
          , Macao, China, April 8-
          <issue>11</issue>
          ,
          <year>2019</year>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>2086</fpage>
          -
          <lpage>2090</lpage>
          . URL: https://doi.org/10. 1109/ICDE.
          <year>2019</year>
          .
          <volume>00245</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>C.</given-names>
            <surname>Seifert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Scherzinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wiese</surname>
          </string-name>
          ,
          <article-title>Towards Generating Consumer Labels for Machine Learning Models</article-title>
          ,
          <source>in: 2019 IEEE First International Conference on Cognitive Machine Intelligence (CogMI)</source>
          , Los Angeles, CA, USA, December
          <volume>12</volume>
          -
          <issue>14</issue>
          ,
          <year>2019</year>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>179</lpage>
          . URL: https://doi.org/10.1109/CogMI48466.
          <year>2019</year>
          .
          <volume>00033</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>N.</given-names>
            <surname>Mehrabi</surname>
          </string-name>
          , et al.,
          <source>A Survey on Bias and Fairness in Machine Learning, ACM Comput. Surv</source>
          . (
          <year>2021</year>
          )
          <volume>115</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>115</lpage>
          :
          <fpage>35</fpage>
          . URL: https://doi.org/10.1145/3457607.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Storchan</surname>
          </string-name>
          , E. Kurshan, Beyond Fairness Metrics:
          <article-title>Roadblocks and Challenges for Ethical AI in Practice</article-title>
          ,
          <source>CoRR</source>
          (
          <year>2021</year>
          ). URL: https://arxiv.org/ abs/2108.06217. arXiv:
          <volume>2108</volume>
          .
          <fpage>06217</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conrad</surname>
          </string-name>
          , et al.,
          <article-title>EvoBench: Benchmarking Schema Evolution in NoSQL, in: Performance Evaluation and Benchmarking -</article-title>
          13th
          <source>TPC Technology Conference, TPCTC 2021</source>
          , Copenhagen, Denmark,
          <year>August</year>
          ,
          <year>2021</year>
          , Springer,
          <year>2021</year>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>49</lpage>
          . URL: https: //doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -94437-
          <issue>7</issue>
          _
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Aken</surname>
          </string-name>
          , et al.,
          <source>An Inquiry into Machine Learning-based Automatic Configuration Tuning Services on Real-World Database Management Systems, Proc. VLDB Endow</source>
          .
          <article-title>(</article-title>
          <year>2021</year>
          )
          <fpage>1241</fpage>
          -
          <lpage>1253</lpage>
          . URL: http://www.vldb.org/pvldb/vol14/p1241-aken.pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>