<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Data-Driven Decision-Making with Incomplete Data: Optimisation of Product Ordering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tetyana Hess</string-name>
          <email>tetyana.hess@studium.fernuni-hagen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Uta Störl</string-name>
          <email>uta.stoerl@fernuni-hagen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>36th GI-Workshop on Foundations of Databases</institution>
          ,
          <addr-line>Grundlagen von Datenbanken</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>FernUniversität in Hagen</institution>
          ,
          <addr-line>Hagen</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>The paper analyses the concept of applying Data-Driven Decision Making (DDDM) in the retail sector with a particular focus on order planning. The objective is to systematically review the current state of research, summarise existing findings, and identify research gaps. Special attention is given to the question of which methods already exist to enable well-founded decisions in areas such as sales forecasting and product ordering, even when data is incomplete. The quality of the underlying data is also considered, which includes both transactional data and the maintenance of master data. The goal of this paper is to study the use of DDDM in situations where data is missing or faulty, with a practical focus on order planning. The main focus here is on developing a resilient forecasting system that ensures stable and transparent decisions for ordering goods, depending on stock levels, shelf life, and under conditions where the data is incomplete or missing. To address issues related to data quality, two strategies are used. The first strategy is to use synthetic datasets that simulate real retail scenarios [1] and the second is to use the application of the Multiple Imputation by Chained Equations (MICE) method, a statistically sound technique for multivariate imputation of missing values [2]. This paper contributes to the practical operationalisation of data-driven decision-making (DDDM) by integrating modern imputation methods and simulation-based data modelling. The relevance of this work is determined by conducting an analysis of the existing literature that deals with and investigates data-driven decision-making in retail under conditions of incomplete or missing data.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Data-Driven Decision Making</kwd>
        <kwd>Incomplete Data</kwd>
        <kwd>Data Quality</kwd>
        <kwd>Syntetic Dataset</kwd>
        <kwd>MICE</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Data-Driven Decision-Making (DDDM) refers to the systematic collection, analysis, examination,
and interpretation of data, usually through the application of analytics or machine learning methods
and techniques, to make informed decisions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In retail, data-driven decision-making has gained
significant importance, especially given its seasonality, the limited shelf life of products, complex
logistical dependencies and routes, and the high level of competition.
      </p>
      <p>Theoretically, other industries can also apply the basic methodological foundation of data-driven
decision-making. In this work, however, the research is focused specifically on the retail sector, because
one of the authors has significant practical experience in a leading German retail company.</p>
      <p>To illustrate this practical dependency, consider a store manager who can physically inspect the
shelves and immediately see what is missing and what needs to be ordered. An SCM manager responsible
for ordering products for multiple stores and not physically present faces a bigger challenge because
they have to rely on the data in the system. Sometimes there are discrepancies between the stock on the
shelf, meaning in the store, and the stock in the database. It happens due to theft, unrecorded discounts,
spoilage, or data entry errors. These diferences between actual and system-recorded data highlight the
problem of incomplete or inaccurate data.</p>
      <p>
        This work is based on practical experience in retail data analysis, where efective sales forecasting and
ordering decisions are crucial but complicated by the fact that the data is often incomplete, incorrect, or
outdated [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4, 5, 6</xref>
        ]. This is especially true for perishable goods and stock values that are only updated
periodically after inventory checks.
      </p>
      <p>
        The aim of this paper is to provide a systematic overview of scientific methods for decision support
under conditions of data uncertainty. The focus is on imputation methods, in particular Multiple
Imputation by Chained Equations (MICE) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], as well as the generation of synthetic datasets to simulate
realistic retail scenarios [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>This paper delivers a structured overview of data sources and forecasting models for ordering
decisions and identifying practically relevant variables. It builds a foundation for developing a robust
forecasting model, which will be validated in future work by using a Monte Carlo simulation.</p>
      <p>The structure of this paper is organized as follows: Background introduces the basic concepts of
data-driven decision-making. State of the Art provides a systematic review of the existing literature
and practice-oriented solution approaches. Research Questions formulates key questions, which are
answered in Concept through methodically sound concepts. Conclusion and Future Work closes with a
critical reflection and an outlook on further research.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>
        The concept of Data-Driven Decision-Making (DDDM) means systematically making decisions based on
the data we have in the system and have analyzed, including machine learning methods [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Especially
in the retail sector, data-driven decision-making (DDDM) helps with managing demand, inventory, and
pricing in such a dynamically evolving sector.
      </p>
      <p>In the academic literature, data-driven decision-making is presented as a combination of human
expert understanding and algorithmic machine learning support, making this approach integrated and
allowing dynamic response to market change.</p>
      <p>
        A suitable model for illustration is the DECAS framework by Elgendy et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>This model diferentiates between the main components: “Data-Driven” with data and analytics
and “Decision-Making” with decision-making process, decision makers and the decision itself. The
visualization (Figure 1) shows the importance of structured data preparation for well-founded decisions.</p>
      <p>Figure 1 shows that data-driven decision-making describes a process where decisions are not only
based on intuition but are systematically supported and validated by data.</p>
    </sec>
    <sec id="sec-3">
      <title>3. State of the Art</title>
      <p>Despite technological progress, there are many challenges in the practical implementation of DDDM.
These challenges must be solved to fully use the benefits of Big Data and modern analytics tools. One
of the biggest challenges is the data quality and reliability of data.</p>
      <sec id="sec-3-1">
        <title>3.1. Data Quality as a Success Factor</title>
        <p>
          A large number of studies confirm that data-driven decisions require high-quality data. McAfee and
Brynjolfsson [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] demonstrate that data-driven companies are more innovative and eficient because
they can respond more quickly to market changes. Janssen et al. [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] point out that low-quality data can
lead to bias and uncertainty in analysis results, which can reduce the quality of decisions.
        </p>
        <p>
          High-quality and reliable data are the foundation of data-driven decision-making. However, if data is
incomplete or missing, companies risk making decisions based on incorrect or insuficient information,
which can lead to incorrect decisions and serious consequences. According to Janssen et al. [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], the
quality of data is a key factor that strongly influences decision-making , because low-quality
data often causes bias and uncertainty in the results.
        </p>
        <p>The quality of decisions in retail is mainly determined by the availability, consistency, and accuracy
of the underlying data. In conclusion, data that is inconsistent, missing, or faulty represents a challenge
in the implementation of forecasting models or data-driven decision systems.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Forecasting Models with Incomplete Data</title>
        <p>
          A key challenge of the Forecasting Models is the integration of incomplete sales data into predictive
models. Pedregal et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] present the Tobit Exponential Smoothing method for censored time series,
which allows robust forecasts even when the available data is limited. Other approaches use model-based
classification frameworks to distinguish between zero values and classify demand into diferent types.
This improves the accuracy of forecasts [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
        <sec id="sec-3-2-1">
          <title>Basic Principle of Exponential Smoothing</title>
          <p>
            Exponential Smoothing (ETS) is one of the most widely used forecasting methods in both practice and
research. The method is based on the foundational work by Charles C. Holt in 1957 [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ].
          </p>
          <p>Holt describes an efective method for time series forecasting. This method is based on exponentially
weighted moving averages. The idea of exponential smoothing is that both seasonal fluctuations and
forecast trends are considered. At the same time, the requirements for using this algorithm do not
involve huge volumes of data or large computational resources.</p>
          <p>Therefore, among other things, in exponential smoothing, when we observe historical data - from
today backward, for example - the coeficient applied over time will be larger for more recent data than
for older and older observations.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Why do Zero Values happen?</title>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Imputation Techniques for Missing Values</title>
        <sec id="sec-3-3-1">
          <title>Multiple Imputation by Chained Equations (MICE)</title>
          <p>
            The method of Multiple Imputation by Chained Equations (MICE) has especially become widely used
for handling missing or inaccurate data. This approach takes into account multivariate relationships
between variables and has proven to be a dynamic and reliable tool, particularly when working with
complex datasets, such as those found in medical or psychological research [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ].
          </p>
          <p>
            It is important to know why the data is missing or incomplete. In statistics, there are three main
mechanisms of missing value: MCAR (Missing Completely At Random), MAR (Missing At Random),
and MNAR (Missing Not At Random) [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ].
          </p>
          <p>The MCAR mechanism means that the data is missing completely at random. MAR implies that the
probability of missingness depends on observed variables and not on the missing values themselves.
MNAR means that data are missing for reasons directly related to the missing values.</p>
          <p>
            To determine which missingness mechanism is present, specific statistical tests are used. Little’s test
[
            <xref ref-type="bibr" rid="ref14">14</xref>
            ] will be used for MCAR. For MAR and MNAR, pattern analysis or logistic regression models are
often applied.
          </p>
          <p>MICE is particularly useful under the Missing at Random (MAR)1mechanism. The MICE method
works on the principle of an iterative chain of regressions. In this case for each variable with missing
values, a separate prediction model is built using other known variables. The process is then repeated
several times until the results become consistent. Depending on the type of data, diferent models are
used:
• linear regression (for numerical data),
• logistic regression (for binary data),
• ordinal logistic regression (for ordinal data).</p>
          <p>Thanks to this flexibility, the MICE method is considered a powerful tool for working with datasets
that contain missing values.</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>Synthetic Datasets</title>
          <p>In addition to the MICE method described above, other approaches enable data imputation. One of
them, which is gaining increasing importance, is synthetic datasets. They are particularly relevant when
working with confidential data. The goal of such synthetic data is to create an artificial dataset that
preserves the statistical properties of the real data but contains no confidential information.</p>
          <p>
            What approaches exist? Three will be highlighted in this work:
• Generative Adversarial Networks (GANs) [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] – There are two neural networks in a GAN:
the generator, which aims to create realistic data, and the discriminator, which tries to tell the
diference between real data and data that was developed [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ]. This adversarial learning process
creates synthetic data that is very similar to the original data.
• Variational Autoencoders (VAEs) [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] use variational inference and the structure of
autoencoders to make new data. This is especially useful when the original data’s distribution is
hard to understand or unknown.
• Monte Carlo-based Methods – This method is one of the most commonly used tools of
stochastic modeling that is also employed to generate synthetic datasets without using any
personally identifiable information. An interesting study in this context is “Nested Stochastic
Valuation of Large Variable Annuity Portfolios: Monte Carlo Simulation and Synthetic Datasets”
by Guojun Gan and Emiliano A. Valdez (2018) [
            <xref ref-type="bibr" rid="ref19">19</xref>
            ]. A synthetic dataset is generated to realistically
1Missing at Random (MAR)[
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]: For any individual, the probability that the value of a variable  is missing does not
depend on the (unobserved) value of  itself, but only on the observed values of the other variables. Formally, this means:
 ( = mis | 1, . . . , ) =
=  ( = mis | 1, . . . , − 1, +1, . . . , ).
simulate complex financial processes, such as the assessment of important insurance portfolios,
without using sensitive data. The data is generated through Monte Carlo simulations combined
with a nested stochastic simulation framework. These models show how markets change over
time [
            <xref ref-type="bibr" rid="ref19">19</xref>
            ].
          </p>
          <p>
            A key advantage of synthetic data is its compliance with data protection [
            <xref ref-type="bibr" rid="ref20 ref21">20, 21</xref>
            ]. This allows
simulation and data analysis without exposing confidential information – an important factor in
data-driven fields such as retail.
          </p>
          <p>
            Actual research suggests combining MICE and synthetic data generation. MICE is used to impute
real data gaps, while synthetic data improves diversity and robustness in forecasting models [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ]. In
retail this hybrid strategy is useful.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Research Questions</title>
      <p>The quality of data-based forecasts in the retail sector plays a big role in decision-making and strongly
depends on the completeness, consistency, and relevance of the underlying data.</p>
      <p>The authors present the following research questions:
• RQ1. Which methods are suitable for systematically identifying data gaps and anomalies, and
how can data quality be assessed beyond traditional metrics?
• RQ2. Which methods are appropriate for analyzing and improving the semantic, syntactic, and
pragmatic quality of master data?
• RQ3. What criteria determine whether data is suitable for accurate forecasting?
• RQ4. Which imputation techniques are particularly suitable for use in the retail sector?</p>
    </sec>
    <sec id="sec-5">
      <title>5. Concept</title>
      <p>The quality, completeness, and consistency of the underlying data have a significant impact on
data-driven decisions in retail. For this reason, the focus of this chapter is on the methodological
framework for answering the defined research questions. Additionally, ideas and perspectives are
discussed regarding which methodological approaches can be followed and which research areas should
be explored in future work.</p>
      <sec id="sec-5-1">
        <title>RQ1. Which methods are suitable for systematically identifying data gaps and anomalies, and how can data quality be assessed beyond traditional metrics?</title>
        <p>
          To systematically identify data gaps and anomalies, classical methods of Exploratory Data Analysis
(EDA) [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] are applied. These help to detect structural patterns, outliers, and distributions both visually
and statistically. Diferent visualization techniques such as boxplots, heatmaps, and time series plots
support intuitive detection of missing values, extreme values, and unusual distribution patterns.
        </p>
        <p>In addition, specialized anomaly detection methods [23, 24] are used to identify unusual,
implausible, or faulty data points automatically. There are diferent statistical methods (e.g., Z-scores),
machine learning algorithms (e.g., Isolation Forest), and time series-based approaches (e.g., seasonal
decomposition).</p>
        <p>The combination of exploratory analysis and algorithmic anomaly detection allows for robust data
validation, covering both obvious and subtle anomalies. This makes quality control more efective and
lfexible in fields like retail.</p>
        <p>Beyond traditional metrics, a context-based assessment of data quality is recommended to reflect
practical retail requirements. In addition to the well-known 15 dimensions of Wang and Strong (1997)
[25], which include aspects such as accuracy, timeliness, and completeness, the analysis should be
extended to include domain-specific quality indicators.</p>
        <p>Therefore, the concept also integrates retail-specific Key Performance Indicators (KPIs), such as stock
accuracy and sales volume.</p>
        <p>A key element of the concept is the use of exploratory data analysis (EDA), particularly univariate
and bivariate methods. The study begins with the assumption of a normal distribution, because it is a
simple and understandable way to identify outliers using sigma intervals (e.g., ±1σ, ±2σ, ±3σ). KPIs
such as sales volume, order quantity, and stock are checked to determine whether they lie within these
intervals. Values outside these intervals are marked as potential outliers that require further review.</p>
        <p>At the same time, the assumption of normality must be critically questioned. In retail, sales and
stock data often show right-skewed distributions, for example due to the Pareto efect (20% of products
generate 80% of sales). As a result the methodological study also includes testing alternative statistical
models, such as log-normal distributions, to ensure meaningful outlier detection.</p>
        <p>To illustrate this, a sample dataset is used with KPIs such as daily sales, order quantity, and stock.
This dataset with inserted faulty values was manually created for demonstration purposes to simulate
typical retail data problems.</p>
        <p>In the first step, an univariate analysis is performed (Figure 2a). The stock data shows a mean of
approximately 1610 units under the assumption of a normal distribution. Sigma intervals show the
separation of normal values and outliers. Stock values below zero are implausible and immediately
lfagged as erroneous. Stock values below zero are clearly implausible and immediately flagged as
erroneous. The values beyond the ±2σ range are marked as potential outliers for further inspection.
(a) Stock
(b) Sales Quantity</p>
        <p>A similar procedure is applied to sales volumes (Figure 2b). Again, assuming a normal distribution,
clear upper and lower boundaries can be established.</p>
        <p>By combining classical dimensions, domain-specific KPIs, and well-considered analytical techniques,
a more nuanced evaluation of data quality is achieved.</p>
      </sec>
      <sec id="sec-5-2">
        <title>RQ2. Which methods are appropriate for analyzing and improving the semantic, syntactic, and pragmatic quality of master data?</title>
        <p>This paper presents a practical approach for analyzing and improving master data quality, based on
classical methods and exploratory data analysis (EDA).</p>
        <p>
          While rule-based checks, similarity analysis, and ontologies are established methods [
          <xref ref-type="bibr" rid="ref4">4, 26</xref>
          ], the
integration of anomaly detection and EDA can reveal semantic and pragmatic quality issues at an early
stage.
        </p>
        <p>• Syntactic quality</p>
        <p>In addition to standard checks (e.g., format and value ranges), EDA methods can reveal frequent
deviations from syntactic standards. For example, length analysis may identify invalid IDs, or
analysis of BBD (best-before date) clusters may reveal unrealistic values (e.g., very high or low
years) or placeholders such as “01.01.1900”.
• Semantic quality</p>
        <p>Beyond classical ontology use, EDA can help identify semantic inconsistencies. For example,
grouped bar plots or heatmaps can be used to check whether product assignments to categories
or product IDs are consistent (e.g., revenue distribution by product group). Another example is
whether specific product groups have unusually short or long shelf lives compared to industry
averages.
• Pragmatic quality</p>
        <p>This is especially important because it helps identify inconsistencies in master data.. For example,
comparing sales velocity with the remaining shelf life may show that products with high BBD
values but low sales are either slow-moving goods or contain incorrect shelf-life data.</p>
        <p>Extending the analysis of these three aspects of data quality (semantic, pragmatic, and syntactic)
with EDA provides greater flexibility and helps identify not only syntactic errors but also semantic
and pragmatic inconsistencies. Especially in the retail sector, where master data plays a major role in
analysing a wide product range, this ofers valuable opportunities to ensure high data quality.</p>
      </sec>
      <sec id="sec-5-3">
        <title>RQ3. What criteria determine whether data is suitable for accurate forecasting?</title>
        <p>For accurate forecasting, it is essential to check the data for suitability, that is, for quality. Firstly, the
data completeness is very important, because missing values can distort time patterns. It’s also crucial
that historical data goes back long enough to accurately identify recurring cycles and trends.</p>
        <p>Another key criterion is autocorrelation, which measures the dependency of observations over time.
In this context, autocorrelation should be analysed per product group. It is important to distinguish
between short-term autocorrelation and seasonal autocorrelation, as both imply diferent modelling
approaches [27]. Stable and significant autocorrelation may indicate high predictability, while strongly
lfuctuating patterns suggest irregular behaviour.</p>
        <p>Variance stability is also important, as the heteroskedastic time series makes the estimation of forecast
intervals more dificult [27].</p>
        <p>Additionally, the distribution of the data matters for forecasts—special attention should be paid to
asymmetry and sensitivity to outliers.</p>
        <p>In conclusion, the data should be relevant and have appropriate granularity. For example, in the
context of forecasting and sales prediction, the data should be available at least on a daily basis.</p>
      </sec>
      <sec id="sec-5-4">
        <title>RQ4. Which imputation techniques are particularly suitable for use in the retail sector?</title>
        <p>
          The study “Inventory record inaccuracy in grocery retailing: Impact of promotions and product
perishability, and targeted efect of audits” [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] shows that accurate inventory records in retail can
increase sales by up to 11%. Why is the stock data not always accurate? This is because stock levels are
typically a calculated value and inventory checks are only done once a month or even less frequently.
As a result, incorrect inventory data can lead to wrong or missed product orders. Errors in delivery
recording may also cause inflated inventory levels. Accurate inventory is essential for sales forecasting,
since forecasts are based on historical values. If products are not available on the shelf due to incorrect
stock data, they are also missing from sales data.
        </p>
        <p>
          As described, data quality in retail is a critical success factor [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Faulty or incomplete data can
cause major ineficiencies across stock and sales processes. A frequently neglected aspect is the proper
handling of missing data and distortions, which directly afect stock and sales forecasts. This section
therefore, focuses on the suitability of imputation methods, especially the Monte Carlo [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] method,
and illustrates them with practical examples from retail.
        </p>
        <p>A key advantage of Monte Carlo imputation is that missing values are not replaced by deterministic
statistics such as mean values but by stochastic sampling from estimated probability distributions. This
is especially relevant in retail, where both sales quantity and inventory data often show strong skewness
and outliers.</p>
        <p>Choosing the correct underlying distribution is critical. One central criterion is based on the Pareto
Principle, which states that a small number of products in many categories generate the majority of
the revenue. For such heterogeneous data sets, assuming a normal distribution is not suficient. Instead,
specialized distributions are needed to reflect the real structure of the data.</p>
        <p>For items with highly uneven sales patterns – especially slow-moving non-food products or
promotional items with rare sales peaks – the Pareto distribution2 is a suitable choice.</p>
        <p>It allows modeling of rare but high sales values and addresses the phenomenon of "few products,
high sales share." Careful attention must be paid to the choice of the minimum value  and the shape
parameter  .</p>
        <p>For regularly sold products, such as fast-moving consumer goods or fresh products with recurring
demand, the log-normal distribution3 is more appropriate.</p>
        <p>Since sales are always positive and often right-skewed, the log-normal distribution is well-suited.
Its logarithmic values follow a normal distribution. The parameters  and  can be estimated from
the log-transformed data, providing a realistic representation of mean and variance while preserving
skewness in the original data.</p>
        <p>Finally, the Monte Carlo imputation ofers a well-founded method to fill data gaps in the retail sector
while considering the sales structure. Selecting the correct distribution is very important to avoid bias
in forecasting models and to support commercially reasonable decisions.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Future Work</title>
      <p>
        This paper is based on the scientifically supported assumption that data-driven decision-making in retail
– especially in forecasting and sales prediction – is strongly influenced by the quality and completeness
of the underlying data [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This assumption is confirmed by both academic literature and practical
experience in the retail industry. The aim of this exploratory study was to identify and discuss initial
approaches for addressing data quality challenges. The focus was on methods for identifying data gaps
and anomalies, as well as on techniques for analysing and improving data quality.
      </p>
      <p>Within the scope of this work, exploratory data analysis, synthetic datasets, and Monte Carlo
simulations were considered as potential tools to improve the quality of the available data and to close
missing values realistically.</p>
      <p>Future research should further develop the methods discussed here and systematically compare them
with established approaches from the current scientific literature. Particular focus should be given to
the selection and evaluation of suitable imputation methods in order to assess their efectiveness and
applicability in the specific context of retail. Moreover, it would be of particular interest to validate the
developed methods using real data from the retail sector, thereby formulating well-founded practical
recommendations for the application of data-driven decision-making based on empirically verified
evidence.</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the author used GPT-4 in order to: Grammar and spelling check.
2A continuous random variable  is called Pareto distributed Par(, min) with parameters  &gt; 0 and min &gt; 0 if it has the
probability density function
 () =
{︃   min
 +1
0
for  ≥ min,
for  &lt; min
density function
3A continuous random variable  is called log-normally distributed with parameters  ∈ R and  &gt; 0, if it has the probability
 () =</p>
      <p>1
 √2</p>
      <p>︂(
exp −
(ln  −  )2 )︂
[23] G. M. Tavares, V. G. T. da Costa, V. E. Martins, P. Ceravolo, S. B. Jr., Leveraging anomaly detection
in business process with data stream mining, Braz. J. Inf. Syst. 12 (2019) 54–75.
[24] V. Chandola, A. Banerjee, V. Kumar, Anomaly detection: A survey, ACM Comput. Surv. 41 (2009)
15:1–15:58.
[25] D. M. Strong, Y. W. Lee, R. Y. Wang, Data quality in context, Commun. ACM 40 (1997) 103–110.
[26] M. Melkonian, C. Juigné, O. Dameron, G. Rabut, E. Becker, Towards a reproducible interactome:
semantic-based detection of redundancies to unify protein interaction databases, Bioinform. 38
(2022) 1685–1691.
[27] R. J. Hyndman, G. Athanasopoulos, Forecasting: Principles and practice | 2.8 autocorrelation and
3.4 evaluating forecast accuracy (2018). URL: https://www.otexts.com/fpp3/acf.html.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mabry</surname>
          </string-name>
          , G. Cheng,
          <article-title>Advancing retail data science: Comprehensive evaluation of synthetic data</article-title>
          ,
          <source>CoRR abs/2406</source>
          .13130 (
          <year>2024</year>
          ). arXiv:
          <volume>2406</volume>
          .
          <fpage>13130</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>S. van Buuren</surname>
          </string-name>
          ,
          <article-title>Flexible imputation of missing data</article-title>
          , CRC press (
          <year>2018</year>
          )
          <fpage>8</fpage>
          -
          <lpage>10</lpage>
          ,
          <fpage>120</fpage>
          -
          <lpage>129</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Elgendy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Elragal</surname>
          </string-name>
          , T. Päivärinta,
          <article-title>Decas: A modern data-driven decision theory for big data and analytics</article-title>
          ,
          <source>Journal of Decision Systems</source>
          <volume>31</volume>
          (
          <year>2022</year>
          )
          <fpage>337</fpage>
          -
          <lpage>373</lpage>
          . URL: https://doi.org/10.1080/12460125.
          <year>2021</year>
          .
          <volume>1894674</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Becker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matzner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Mueller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Winkelman</surname>
          </string-name>
          ,
          <article-title>Towards a semantic data quality management using ontologies to assess master data quality in retailing</article-title>
          , AIS eLibrary (
          <year>2008</year>
          )
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          . URL: https: //aisel.aisnet.org/amcis2008/129.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Winkelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. B. und Christian</given-names>
            <surname>Janiesch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Becker</surname>
          </string-name>
          ,
          <article-title>Improving the quality of article master data: Specification of an integrated master data platform for promotions in retail</article-title>
          , AIS eLibrary (
          <year>2008</year>
          ). URL: https://aisel.aisnet.org/cgi/viewcontent.cgi?article=1138&amp;context=ecis2008.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Rekik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Oliva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Glock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Syntetos</surname>
          </string-name>
          ,
          <article-title>Inventory record inaccuracy in grocery retailing: Impact of promotions and product perishability, and targeted efect of audits (</article-title>
          <year>2025</year>
          ). arXiv:
          <volume>2506</volume>
          .
          <fpage>05357</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>McAfee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Brynjolfsson</surname>
          </string-name>
          ,
          <article-title>Big data: The management revolution</article-title>
          ,
          <source>Harvard Business Review</source>
          (
          <year>2012</year>
          ). URL: https://hbr.org/
          <year>2012</year>
          /10/big
          <article-title>-data-the-management-revolution.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Janssen</surname>
          </string-name>
          , H. van der Voort, A. Wahyudi,
          <article-title>Factors influencing big data decision-making quality</article-title>
          ,
          <source>Journal of Business Research</source>
          (
          <year>2017</year>
          )
          <fpage>338</fpage>
          -
          <lpage>345</lpage>
          . URL: https://doi.org/10.1016/j.jbusres.
          <year>2016</year>
          .
          <volume>08</volume>
          .007.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Pedregal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Trapero</surname>
          </string-name>
          , E. Holgado,
          <article-title>Tobit exponential smoothing, towards an enhanced demand planning in the presence of censored data (</article-title>
          <year>2024</year>
          ). arXiv:
          <volume>2407</volume>
          .
          <fpage>17920</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>I.</given-names>
            <surname>Svetunkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sroginis</surname>
          </string-name>
          ,
          <article-title>Why do zeroes happen? A model-based approach for demand classification</article-title>
          ,
          <source>CoRR</source>
          (
          <year>2025</year>
          ). arXiv:
          <volume>2504</volume>
          .
          <fpage>05894</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Holt</surname>
          </string-name>
          ,
          <article-title>Forecasting seasonals and trends by exponentially weighted moving averages</article-title>
          ,
          <string-name>
            <given-names>O.N.R.</given-names>
            <surname>Research Memorandum</surname>
          </string-name>
          (
          <year>1957</year>
          )
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>M. J. Azur</surname>
            ,
            <given-names>E. A.</given-names>
          </string-name>
          <string-name>
            <surname>Stuart</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Frangakis</surname>
            ,
            <given-names>P. J.</given-names>
          </string-name>
          <string-name>
            <surname>Leaf</surname>
          </string-name>
          ,
          <article-title>Multiple imputation by chained equations: what is it and how does it work?</article-title>
          ,
          <source>National Library of Medicine</source>
          (
          <year>2011</year>
          )
          <fpage>40</fpage>
          -
          <lpage>49</lpage>
          . URL: https://pubmed.ncbi. nlm.nih.gov/21499542/.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>I.</given-names>
            <surname>Jansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hens</surname>
          </string-name>
          , G. Molenberghs,
          <string-name>
            <given-names>M.</given-names>
            <surname>Aerts</surname>
          </string-name>
          , G. Verbeke,
          <string-name>
            <surname>M. G. Kenward,</surname>
          </string-name>
          <article-title>The nature of sensitivity in monotone missing not at random models</article-title>
          ,
          <source>Comput. Stat. Data Anal</source>
          .
          <volume>50</volume>
          (
          <year>2006</year>
          )
          <fpage>830</fpage>
          -
          <lpage>858</lpage>
          . doi:
          <volume>10</volume>
          . 1016/J.CSDA.
          <year>2004</year>
          .
          <volume>10</volume>
          .009.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R. J. A.</given-names>
            <surname>Little</surname>
          </string-name>
          ,
          <article-title>Inference about means from incomplete multivariate data</article-title>
          ,
          <source>Biometrika</source>
          (
          <year>1976</year>
          )
          <fpage>593</fpage>
          -
          <lpage>604</lpage>
          . URL: https://academic.oup.com/biomet/article/63/3/593/270937.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Silva-Ramírez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pino-Mejías</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>López-Coello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cubiles-de-</surname>
          </string-name>
          la-Vega,
          <article-title>Missing value imputation on missing completely at random data using multilayer perceptrons</article-title>
          ,
          <source>Neural Networks</source>
          <volume>24</volume>
          (
          <year>2011</year>
          )
          <fpage>121</fpage>
          -
          <lpage>129</lpage>
          . URL: https://doi.org/10.1016/j.neunet.
          <year>2010</year>
          .
          <volume>09</volume>
          .008.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>X.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , L. Yuan,
          <article-title>Improved generative adversarial imputation networks for missing data</article-title>
          ,
          <source>Appl. Intell</source>
          .
          <volume>54</volume>
          (
          <year>2024</year>
          )
          <fpage>11068</fpage>
          -
          <lpage>11082</lpage>
          . URL: https://doi.org/10.1007/s10489-024-05814-2.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sinha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Wellman</surname>
          </string-name>
          ,
          <article-title>Generating realistic stock market order streams</article-title>
          , AAAI Press,
          <year>2020</year>
          , pp.
          <fpage>727</fpage>
          -
          <lpage>734</lpage>
          . URL: https://doi.org/10.1609/aaai.v34i01.
          <fpage>5415</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>B.</given-names>
            <surname>Roskams-Hieter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wade</surname>
          </string-name>
          ,
          <article-title>Leveraging variational autoencoders for multiple data imputation</article-title>
          ,
          <source>in: Machine Learning and Knowledge Discovery in Databases: ECML PKDD</source>
          , Springer,
          <year>2023</year>
          , pp.
          <fpage>491</fpage>
          -
          <lpage>506</lpage>
          . URL: https://doi.org/10.1007/978-3-
          <fpage>031</fpage>
          -43412-9_
          <fpage>29</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>G.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Valdez</surname>
          </string-name>
          ,
          <article-title>Nested stochastic valuation of large variable annuity portfolios: Monte carlo simulation and synthetic datasets</article-title>
          ,
          <source>Data</source>
          <volume>3</volume>
          (
          <year>2018</year>
          )
          <article-title>31</article-title>
          . URL: https://doi.org/10.3390/data3030031.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>L.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Skoularidou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cuesta-Infante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Veeramachaneni</surname>
          </string-name>
          ,
          <article-title>Modeling tabular data using conditional gans</article-title>
          ,
          <source>NeurIPS</source>
          (
          <year>2019</year>
          ). URL: https://dl.acm.org/doi/10.5555/3454287.3454946.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mohapatra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Kerschbaum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <article-title>Diferentially private data generation with missing data</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>17</volume>
          (
          <year>2024</year>
          )
          <fpage>2022</fpage>
          -
          <lpage>2035</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>W.</given-names>
            <surname>Polasek</surname>
          </string-name>
          ,
          <article-title>Explorative daten-analyse: Eda; einführung in die deskriptive statistik</article-title>
          , Springer-Verlag (
          <year>1988</year>
          )
          <fpage>3</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>