<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Benchmarking Tabular Data Synthesis for User Guidance</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maria F. Davila</string-name>
          <email>maria.davila@ofis.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Supervised By: Wolfram Wingerath</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabian Panse</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CEUR Workshop Proceedings</institution>
          ,
          <addr-line>CEUR-WS.org</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Carl von Ossietzky University of Oldenburg</institution>
          ,
          <addr-line>Ammerländer Heerstraße 114-118, 26129, Oldenburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Hasso Plattner Institute</institution>
          ,
          <addr-line>Prof.-Dr.-Helmert-Straße 2-3, 14482 Potsdam</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>OFFIS - Institute for Informatics</institution>
          ,
          <addr-line>Escherweg 2, Oldenburg, 26121</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <fpage>25</fpage>
      <lpage>28</lpage>
      <abstract>
        <p>This paper presents our research on evaluating the suitability of a Tabular Data Synthesis (TDS) tool for use-case specific requirements. The main goal is to develop a platform that allows users, for example researchers from other fields, to select a suitable TDS tool for their real-world application. In the course of developing such a platform, three contributions are currently planned: Firstly, a decision guide for users formulated by compiling the reported performance of leading tools against a set of functional and non-functional requirements. Secondly, a benchmarking framework for TDS tools based on these identified requirements. Lastly, a customizable tool selection platform, developed through extensive benchmarking of predominant TDS tools. This platform must provide a number of possible tools based on specific use case constraints and allow for community-based expansion, thereby ofering a dynamic and adaptable solution for TDS tool selection.</p>
      </abstract>
      <kwd-group>
        <kwd>Tabular data synthesis</kwd>
        <kwd>Deep generative models</kwd>
        <kwd>Artificial data generation</kwd>
        <kwd>Benchmarking</kwd>
        <kwd>Customized tool selection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-3">
      <title>2. Related</title>
    </sec>
    <sec id="sec-4">
      <title>Work</title>
      <sec id="sec-4-1">
        <title>Tabular data synthesis (TDS) is a method which creates</title>
        <p>Our focus is tabular data, which is structured data
orgarealistic artificial data that mirror the distribution and
nized into rows representing individual data points, and
structure of real relational datasets, and it provides a
columns representing diferent features. Our work can
solution [1] for data scarcity in data-driven applications.
be classified as a recommendation system for TDS tools.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Data synthesis is mainly popular for images [2] and</title>
      </sec>
      <sec id="sec-4-3">
        <title>However, to the best of our knowledge, there are no TDS</title>
        <p>text [3], however TDS has demonstrated impressive out- tools’ recommender systems.
comes in the generation of highly realistic artificial tables</p>
      </sec>
      <sec id="sec-4-4">
        <title>Platforms such as Synthetic Data Vault (SDV) [7] and</title>
        <p>[4, 5, 6]. The goal of this research is to determine how the
its enterprise version DataCebo, Gretel AI [8] and Mostly
iftness-for-use of a TDS tool can be assessed, given
con</p>
      </sec>
      <sec id="sec-4-5">
        <title>AI [9] ofer the possibility to generate tabular data. They</title>
        <p>crete real-world applications. The main contributions of
implement some leading TDS models to fit the widest
this research are: 1) A decision guide to select a suitable
range of applications possible, yet we find there is
curTDS tool for an application, developed by compiling the
rently no universal TDS tool which works well for all
reported performance of the predominant tools on our
datasets. These platforms do not report on the specific
identified functional and non-functional requirements.
limitations of their models for each application, partly</p>
        <sec id="sec-4-5-1">
          <title>2) A benchmarking framework for TDS tools, using</title>
          <p>our previously identified functional and non-functional
because there is no standard framework.</p>
          <p>Relevant surveys for our work include Hernandez’s
requirements as performance indicators. 3) A customiz- [10] review for health records, Fan’s [11] analysis of
Genable tool selection platform, developed by
benchmarkerative Adversarial Networks (GAN) across diferent data
ing predominant TDS tools (cf. Contribution 2). The
types, Figueira’s [12] survey on evaluation methods,
Broplatform is customizable because it outputs a number
phy’s [13] exploration on time series generation, Koo and
of suitable tools. The suitability is estimated based on</p>
        </sec>
      </sec>
      <sec id="sec-4-6">
        <title>Kim’s [14] review on generative difusion models, and</title>
        <p>use-case specific constraints. The platform is expandable</p>
      </sec>
      <sec id="sec-4-7">
        <title>Lin’s [15] review with focus on time-series difusion. because it allows community-based updates.</title>
        <p>nEvelop-O
LGOBE
(M. F. Davila)
CEUR
htp:/ceur-ws.org</p>
        <p>ISN1613-073
Published in the Proceedings of the Workshops of the EDBT/ICDT 2024</p>
        <p>(M. F. Davila)
www.offis.de/offis/person/maria-fernanda-davila-restrepo.html
© 2024 Copyright © 2024 for this paper by its authors. Use permitted under</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>3. Purposes of TDS</title>
      <p>Reviewing the diferent surveys and TDS tool papers, we
identified five reasons for the synthesis of tabular data.
• Missing value imputation: Incomplete entries
often occur in real-world datasets, potentially distorting
analysis. TDS is employed to fill these gaps with
plausible values.
• Dataset balancing: Some classes having significantly
more instances than others can cause bias in
datadriven models. TDS could balance the dataset by
generating records for the underrepresented classes.
• Dataset augmentation: Increasing the dataset size
by creating new artificial records help, for example, to
increase robustness in machine learning models.
• Privacy protection: Generating artificial data that
adheres to privacy regulations allows secure data sharing
while safeguarding sensitive information.
• Customized data generation: Generating data with
specific external constraints allows the creation of
scenario-based data, particularly valuable when
original data is unavailable. For example, generating
environmental datasets for diferent future scenarios is
essential for data-driven meteorological models [16].
The protection of privacy can be the sole reason for the
synthesis (e.g., if sensitive data need to be shared), but it
can also be combined with any of the other purposes.</p>
    </sec>
    <sec id="sec-6">
      <title>4. Challenges of TDS</title>
      <sec id="sec-6-1">
        <title>For all domains of data synthesis where privacy protec</title>
        <p>tion is of interest, one challenge is the privacy vs. utility
trade-of [17]. Data utility describes the data’s
efectiveness in fulfilling its intended purpose, besides privacy.
Balancing data privacy and utility is a fundamental
challenge in data synthesis, because enhancing privacy often
diminishes data utility and vice versa [18].</p>
        <p>The challenge in TDS is to accurately capture the main
information and structure of the input dataset, to be able
to replicate it in a synthetic dataset that is useful for
the desired purpose. We summarize the challenges of
accurately capturing and replicating this information as
follows:
• Handling Missing Values: Capturing the real
column distribution and column correlations in a dataset
with missing values is challenging, because the gaps
distort statistical properties.
• Addressing Class Imbalance: In a dataset with class
imbalance the model could over fit, or sufer mode
collapse. However, it is crucial to efectively learn from
such imbalances to identify and understand outliers,
which are often indicative of anomalies [ 19].
• Complex Column Relations: Real-world datasets
include correlations between columns, and sometimes
these columns belong to diferent tables within the
dataset. Capturing and preserving these relations is
particularly challenging, for example, when there are
multiple tables interconnected through foreign key
references [20].
• Temporal dependencies: Temporal columns add
complexity to the synthesis process. This is
particularly dificult for long-term relations because it requires
the model to retain information over long periods of
time.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>5. TDS Tool Requirements</title>
      <sec id="sec-7-1">
        <title>Our focus are deep generative models, a subcategory</title>
        <p>of data-driven TDS tools, which leverages deep
learning techniques to model joint probability distributions
of the dataset. Compiling existing surveys shows that
the currently predominant tools belong to the following
categories: Variational Autoencoders (VAE), Generative
Adversarial Networks (GAN), Normalizing Flows (NF),
Graph Neural Networks (GNN), Difusion Probabilistic
Models, and Transformers (often LLMs). The sampling
method SMOTE is also often included because it achieves
good performance for its simplicity [5].</p>
        <p>We compiled a list of requirements reported for the
diferent TDS tools, as shown in Table 1.
• Diversity of Column Types: Diferent from images
composed by pixels and text composed by words and
phrases, tabular data often contains various column Combining the purposes described in Section 3, the
types, such as numerical, categorical, text, temporal, requirements listed in Table 1, and a compilation of TDS
and mixed. tools (Table 3), we mark what requirements each TDS
• Complex Distributions: Real-world columns can tool was reported to fulfill. As a result, we created a
have complex distributions and capturing the real dis- first attempt for a decision guide for users who wants to
tribution of a column is crucial to generate realistic synthesize tabular data for a real-world application. The
data. result is shown in Table 2.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>6. Research Question</title>
      <p>In the course of creating the guide, we identified some
research gaps: 1) there are no reported tools for
customized generation of time series datasets, 2) many ad- The main goal is to develop a platform that allows users
vanced tools are reported to violate integrity constraints to select a suitable TDS tool for their use-case specific
[18], 3) transformer models capture correlations well requirements. This divides into the research questions:
with significantly less pre-processing but the results seem RQ1 What are the use-case specific functional and
nonto not fully capture the complex distributions of single functional requirements that drive the selection
column [6], 4) there are no tools to efectively gener- process between TDS tools?
ate multiple related tables, preserving the column cor- RQ2 How can the suitability of a specific tool be
evalrelations and integrity constraints, 6) there is no trans- uated based on the identified requirements
usparency in reporting the computational costs or resources ing a standardized benchmarking framework?
required. Which metrics can be used to assess the
performance of the tools?</p>
      <sec id="sec-8-1">
        <title>C3 A customizable tool selection platform, devel</title>
        <p>oped by benchmarking predominant TDS tools
(cf. Contribution 2).</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>7. Conclusions</title>
      <p>of the VLDB Endowment 13 (2020) 1962–1975. doi:10.14778/
3407790.3407802.</p>
      <p>Synthetic tabular data is a solution for the scarcity and [12] A. Figueira, B. Vaz, Survey on synthetic data generation,
lack of diversity of real-world datasets for data-driven evaluation methods and gans, Mathematics 10 (2022) 2733.
doi:10.3390/math10152733.
applications. From the research gaps we identified, our [13] E. Brophy, Z. Wang, Q. She, T. Ward, Generative adversarial
focus is addressing how to evaluate the suitability of TDS networks in time series, ACM Comput. Surv. 55 (2023). doi:10.
tools for use-case specific requirements. 1145/3559540.</p>
      <p>We successfully identified the main purposes, chal- [14] H. Koo, T. E. Kim, A comprehensive survey on generative
lenges and some of the requirements for TDS tools, how- difusion models for structured data, arXiv:2306.04139 (2023).
doi:10.48550/arXiv:2306.04139.
ever the decision-making process has so many dimen- [15] L. Lin, Z. Li, R. Li, X. Li, J. Gao, A comprehensive survey on
gensions that it cannot be adequately presented as a table. erative difusion models for structured data, arXiv:2305.00624
For that reason, we aim to break down the decision- (2023). doi:10.48550/arXiv:2305.00624.
making process involved in selecting a TDS tool for real- [16] A. Nandy, C. Duan, H. J. Kulik, Data-driven scenarios of climate
world applications into functional and non-functional change, Resilient Urban Futures. The Urban Book Series (2021).
URL: https://doi.org/10.1007/978-3-030-63131-4_13.
requirements. Examples of functional requirements is [17] N. Park, M. Mohammadi, K. Gorde, S. Jajodia, H. Park, Y. Kim,
the ability to handle multiple column types and distribu- Data synthesis based on generative adversarial networks,
Protions. Examples of non-functional requirements is the ceedings of the VLDB Endowment 14 (2018). doi:10.48550/
resource eficiency of the tool (time and memory), or its arXiv:1806.03384.
scalability. Afterwards, those requirements will base a [18] C. Ge, S. Mohapatra, X. He, I. F. Ilyas, Kamino, Proceedings of
the VLDB Endowment 14 (2020). doi:10.48550/arXiv.2012.
benchmarking framework for TDS tools, used to
evalu15713.
ate the tools’ suitability for specific use cases. Finally, a [19] R. C. Ripan, I. H. Sarker, Outlier detection approach for
efeccustomizable tool selection platform will be developed tively classifying cyber anomalies., Hybrid Intelligent Systems
to capture the full complexity of the decision-making (2021). URL: https://doi.org/10.1007/978-3-030-73050-5_27.
process and truly guide users in selection a suitable tool [20] P. Han, W. Xu, W. Lin, J. Cao, C. Liu, S. Duan, H. Zhu, C3-tgan,
TechRxiv (2023). doi:10.36227/techrxiv.24249643.v1.
for their real-world application. [21] N. V. Chawla, K. W. Bowyer, L. O. Hall, W. P. Kegelmeyer,</p>
      <p>The planned contributions add to the research field Smote, Journal of Artificial Intelligence Research 16 (2002)
by providing a baseline to benchmark TDS tools, allow- 321–357. doi:10.48550/arXiv.1106.1813.
ing researchers to easily identify gaps. This broadens [22] J. Yoon, J. Jordon, M. van der Schaar, Generating multilabel
the research space from machine learning researchers to discrete patient records using generative adversarial networks
(2017). URL: https://doi.org/10.48550/arXiv.1703.06490.
other data-centric fields, such as data management. The [23] J. Yoon, J. Jordon, M. van der Schaar, Pate-gan, International
contributions also bring TDS closer to users outside this Conference on Learning Representations (2019). URL: https:
research field, by simplifying the tool selection process. //openreview.net/forum?id=S1zk9iRqF7.
[24] L. Xie, K. Lin, S. Wang, F. Wang, J. Zhou, Diferentially
private generative adversarial network, arXiv.1802.06739 (2018).</p>
      <p>References doi:10.48550/arXiv.1802.06739.
[25] L. Xu, M. Skoularidou, A. Cuesta-Infante, K. Veeramachaneni,
[1] K. H. L. Minh, K. H. Le, Airgen, RTSI (2021). doi:10.1109/ Modeling tabular data using conditional gan, Proceedings</p>
      <p>RTSI50628.2021.9597364. NEURIPS 32 (2019).
[2] OpenAI, Dall-e: Creating images from text (2021). [26] Z. Zhao, A. Kunar, R. Birke, H. V. der Scheer, L. Y. Chen,
Ctab[3] OpenAI-ChatGPT, Version 4 (2023). gan, arXiv:2102.08369 (2021). doi:10.48550/arXiv.2102.
[4] Z. Zhao, A. Kunar, R. Birke, L. Y. Chen, Ctab-gan+, 08369.</p>
      <p>arXiv.2204.00401 (2022). doi:10.48550/arXiv.2204.00401. [27] Y. Zhang, N. A. Zaidi, J. Zhou, G. Li, Ganblr++, SDM (2022).
[5] A. Kotelnikov, D. Baranchuk, I. Rubachev, A. Babenko, Tab- [28] J. Yoon, D. Jarrett, M. van der Schaar, Time-series
generaddpm, arXiv:2209.15421 (2022). doi:10.48550/arXiv.2209. tive adversarial networks, Advances in Neural Information
15421. Processing Systems 32 (NeurIPS) (2019).
[6] V. Borisov, K. Seßler, T. Leemann, M. Pawelczyk, G. Kas- [29] Z. Lin, A. Jain, C. Wang, G. Fanti, V. Sekar, Using gans for
neci, Language models are realistic tabular data generators, sharing networked time series data, arXiv:1909.13403 (2019).
arXiv.2210.06280 (2023). doi:10.48550/arXiv.2210.06280. doi:10.48550/arXiv.1909.13403.
[7] N. Patki, R. Wedge, K. Veeramachaneni, The synthetic data [30] A. Desai, C. Freeman, Z. Wang, I. Beaver, Timevae,
vault, DSAA (2016) 399–410. doi:10.1109/DSAA.2016.49. arXiv:2111.08095 (2021). doi:10.48550/arXiv.2111.08095.
[8] K. Boyd, Create synthetic time-series data with doppelganger [31] H. Lim, M. Kim, S. Park, N. Park, Regular time-series
generaand pytorch, 2022. tion using sgm, arXiv.2301.08518 (2023). doi:10.48550/arXiv.
[9] Mostly ai, 2023. URL: https://mostly.ai/. 2301.08518.
[10] M. Hernandez, G. Epelde, A. Alberdi, R. Cilla, D. Rankin, [32] T. Liu, Z. Qian, J. Berrevoets, M. van der Schaar,
GogSynthetic data generation for tabular health records, Neu- gle, The Eleventh International Conference on Learning
Reprocomputing 493 (2022) 28–45. doi:10.1016/j.neucom.2022. resentations (2023). URL: https://openreview.net/forum?id=
04.053. fPVRcJqspu.
[11] J. Fan, T. Liu, G. Li, J. Chen, Y. Shen, X. Du, Relational data [33] A. V. Solatorio, O. Dupriez, Realtabformer, arXiv.2302.02041
synthesis using generative adversarial networks, Proceedings (2023). doi:10.48550/arXiv.2302.02041.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>