<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>G. Recupito);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Transformer-Based Approach to Data Quality</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gilberto Recupito</string-name>
          <email>grecupito@unisa.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giammaria Giordano</string-name>
          <email>giagiordano@unisa.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dario Di Nucci</string-name>
          <email>ddinucci@unisa.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Palomba</string-name>
          <email>fpalomba@unisa.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Data Smells, Data Quality, Software Engineering for AI, MLOps, Empirical Software Engineering</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>MLOps25: Workshop on Machine Learning Operations</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SeSa Lab - University of Salerno</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>In recent years, the integrity of data used in machine learning pipelines has become increasingly critical, as even state-of-the-art models are constrained by the quality of their input. Among the various threats to data reliability, data smells-subtle and often semantic anomalies-pose unique challenges due to their nuanced and context-sensitive nature. This paper presents an innovative approach to detect semantic data smells using BERT, a transformer-based language model originally designed to understand natural language. We focus specifically on two underexplored categories: Split Value Smells and Multiple Value Smells, which respectively reflect improper data fragmentation and conflation. To facilitate model training, we constructed a large and heterogeneous corpus comprising synthetic and real datasets, utilizing data augmentation techniques via Faker and SDV. We trained two binary classifiers using fine-tuned BERT models, achieving high performance (F1 scores of 0.88 and 0.98) in detecting these smells. Our findings demonstrate that transformer-based models are not only efective in capturing structural patterns in tabular data but also capable of generalizing across diverse semantic anomalies. This work establishes a foundation for the broader application of language models in data quality assurance, opening new avenues for semantic-level data cleaning automation.</p>
      </abstract>
      <kwd-group>
        <kwd>Quality</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        As machine learning (ML)-enabled systems evolve from experimental prototypes to deployed
applications, ensuring their reliability requires attention to all stages of the ML lifecycle, from data preparation
to model deployment and monitoring [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. To support this shift, MLOps has emerged as a collection of
engineering practices that extend the DevOps principles to ML systems [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A key aspect of MLOps is
continuous quality assurance, which involves maintaining the performance, integrity, and stability of
ML systems throughout their operation. Although current MLOps eforts often emphasize monitoring
model performance and retraining strategies, data quality remains a critical and challenging factor [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Low-quality data can silently introduce errors, degrade model performance, and increase the risk of
biased or misleading predictions. Among the most subtle and impactful threats to data quality are data
smells, data value-based indications of latent data quality issues caused by poor practices that may
lead to future failures [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Although certain types of data smell—particularly those related to syntax
or formatting—can be addressed using existing validation tools, current approaches fall short when it
comes to semantic data smells. These involve deeper contextual and structural inconsistencies that are
dificult to detect with rule-based techniques [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
      </p>
      <p>In this preliminary work, our goal is to investigate initial opportunities for modeling and identifying
subtle semantic anomalies in tabular data. To start exploring these capabilities, we focused on two
simple data quality issues related to the distribution of information among the columns of a dataset:
Split Value Smells and Multiple Value Smells. By transforming tabular data into textual representations,</p>
      <p>CEUR
Workshop</p>
      <p>ISSN1613-0073
we leveraged the ability of BERT to capture semantic patterns and irregularities. Our approach supports
automated semantic data quality checks that can be integrated into MLOps pipelines, reducing manual
overhead, and improving the robustness of ML-enabled systems over time.</p>
      <p>
        To train and evaluate our models, we created a synthetic dataset using a combination of the Faker
library [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and the Synthetic Data Vault (SDV) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We fine-tuned two separate BERT-based
classifiers and evaluated their performance using standard classification metrics. The model trained for
Split Value Smells achieved an accuracy of 88.99% with perfect precision and an F1 score of 0.8763. The
model trained for Multiple Value Smells achieved an even higher accuracy of 98% and an F1 score of
0.9795. We provide the following contributions:
• We introduce a novel BERT-based approach to detect semantic data smells in tabular datasets.
• We define and formalize two complex smell types— Split Value Smells and Multiple Value Smells—and
propose an automated way to inject them into synthetic data.
      </p>
      <p>• We build two meta-datasets to explore further research on semantic data smells.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background and Related Work</title>
      <p>In this section, we outline key data quality issues relevant to ML-enabled systems, establishing the
broader context in which these issues arise. We then review related work on the detection of such
issues, with a particular focus on recent advances in data smell identification.</p>
      <sec id="sec-2-1">
        <title>2.1. Data Quality Assurance in MLOps Environments</title>
        <p>
          In MLOps pipelines, data quality plays a central role in ensuring long-term performance, reliability,
and maintainability of ML-enabled systems. Unlike traditional software systems, where source code is
the primary artifact, ML-enabled systems are highly dependent on data, an inherently dynamic and
evolving asset. Renggli et al. [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] introduced a data quality-driven perspective of MLOps, arguing that
many MLOps challenges are fundamentally rooted in data management. They emphasized how key
data quality dimensions are distributed across diferent stages of the ML pipeline. Degraded data quality
directly impacts model performance. Mohammed et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] investigated this relationship, demonstrating
that data completeness and consistency issues have a significant impact on both classification and
regression tasks. This finding supports earlier work by the data management community, which focused
on automating data quality assurance practices and tools [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Therefore, current data validation
techniques are typically integrated into CI/CD pipelines and focus on syntactic validations, such as
type checking, missing values, or schema conformity.
        </p>
        <p>
          In addition to classical data quality issues, researchers have introduced the concept of data smells,
recurring patterns in data that suggest latent quality problems. Foidl et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] proposed a taxonomy of
these smells, distinguishing between syntactic and semantic issues. Building on this work, Recupito
et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] expanded the taxonomy by introducing two new types of semantic data smells relevant to
tabular datasets: Split Value Smells and Multiple Value Smells.
        </p>
        <p>The former refers to cases where a single piece of information is unnecessarily distributed across
multiple columns, for example, splitting a date of birth into separate month, day and year components, as
illustrated in Table 1. Such fragmentation can hinder downstream processing and reduce data usability.
The latter involves instances where multiple distinct pieces of information are inappropriately merged
into a single column, for example, by storing both an email address and a phone number in the same field,
as shown in Table 2. This practice complicates data interpretation and parsing. Both smells represent
semantic anomalies that go beyond syntactic validation and require contextual understanding to detect
and resolve efectively. These semantic issues are particularly problematic in production-grade ML
systems, where they can silently propagate through automated workflows, resulting in weak models and
unreliable predictions. However, such techniques often fail to capture more complex violations of the
data’s semantic intent, such as improperly merged or split values. Overcoming these limitations requires
more adaptive and context-aware strategies that go beyond rule-based checks and can understand the
structure and meaning of the data within its application context.
…
→
→
→
→
→
→</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Related Work</title>
        <p>
          Recent studies have increasingly focused on formalizing and detecting data smells, recurring patterns
in data that may indicate latent quality issues. Foidl et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] introduced a taxonomy of data smells
and proposed DSD, a tool to detect a subset of syntactic data smells. They also developed a machine
learning–based approach for identifying format inconsistency smells. However, most existing detection
methods remain limited to easily measurable features, such as null values or outliers, and overlook
semantic-level anomalies that span multiple fields or require contextual interpretation. Li et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]
presented CleanML, a tool that automatically detects and cleans data quality issues, including
missing, extreme, and duplicate values. Their empirical study demonstrated that applying data cleaning
techniques can significantly improve the performance of downstream ML models. However, these
approaches largely focus on surface-level issues and fall short of addressing complex data semantics,
particularly in structured data used in ML pipelines.
        </p>
        <p>
          In parallel, recent advances in transformer-based models—particularly BERT and its derivatives—have
shown strong potential for processing tabular and hybrid (text-tabular) data. These models have been
successfully adapted to capture the structure and semantics of tables in diverse tasks. For example,
TaBERT [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and TAPAS [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] jointly model natural language and tabular data for Table Question
Answering, incorporating table-specific attention mechanisms and positional encodings. For semantic
table understanding, models such as TURL [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] and TUTA [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] leverage structure-aware transformers
to perform entity linking, relation extraction, and hierarchical parsing. In self-supervised settings,
TABBIE [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] and TABNER [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] pretrain BERT-based architectures for tasks like corrupted cell detection
and named entity recognition in spreadsheets, demonstrating the models’ ability to learn meaningful
structural patterns. These works illustrate BERT’s strong contextual embedding capabilities, making it a
compelling foundation for detecting semantic anomalies in structured data. Building on this foundation,
our work bridges the gap between data quality and research communities by applying BERT to semantic
data validation. Specifically, we reinterpret rows in tabular datasets as natural language sequences and
ifne-tune BERT to detect data smells that manifest semantic and contextual irregularities. Our approach
        </p>
        <sec id="sec-2-2-1">
          <title>Contribution to the State of the Art.</title>
          <p>While existing work on data quality predominantly addresses structural and syntactic issues, this
work explores the capabilities of transformer-based models of detecting semantic data smells that
require semantic understanding.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Research Method</title>
      <p>The goal of this study is to explore whether BERT can support automated detection of semantic smells
in tabular data sets. We focus our investigation on two key categories: Split Value Smells, which occur
when a single logical entity (e.g., a person’s full name or an address) is improperly fragmented across
multiple columns; and Multiple Value Smells, which result from the incorrect merging of distinct values,
such as a phone number and an email, into a single cell.</p>
      <p>
        This study is carried out from a dual perspective. On the one hand, our aim is to contribute to the
body of knowledge by understanding the applicability of language models in semantic data quality
tasks, thereby addressing the interests of researchers in empirical data-centric ML. On the other hand,
we respond to the practical needs of developers and data engineers, who routinely face the challenge of
detecting and correcting these types of anomaly at scale. In terms of reporting, we follow the guidelines
of Wohlin et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. We formulated two research questions to guide our investigation:
contributes to data quality assurance by ofering a flexible, learning-based solution to identify complex
data quality issues.
pendent values are improperly merged into a single cell (i.e., Multiple Value Smells). For this task,
we fine-tuned BERT to process individual cells as input, leveraging its token-level attention to learn
whether a given entry contains signs of conflation—e.g., delimiters, unexpected semantic shifts, or
irregular formatting. As with RQ1, we used a combination of accuracy, precision, recall, and F1 score to
assess performance and complemented this with confusion matrix analysis to uncover boundary cases
and failure modes.
      </p>
      <p>To conduct our study, we constructed a synthetic dataset of labeled tabular entries. These were
designed to represent realistic examples of both clean and anomalous data.</p>
      <p>Where possible, we
introduced smells based on known real-world issues observed in open datasets, supplemented by
manually engineered cases to ensure balanced coverage across scenarios.</p>
      <p>All records were serialized in a column-based format suitable for the input structure of BERT, and
each instance was annotated with a binary label corresponding to the presence of a smell, allowing us
to train and evaluate the models under consistent conditions.
in more detail in the following sections.</p>
      <p>Model Fine-Tuning</p>
      <p>Kaggle
Data Augmentation
Faker SDV
20 Datasets for Smells
Split Value Multiple</p>
      <p>Smell Value Smell
4040000,0k0 0
rroowwssppeerr  
dsamtaesellt
Split Value</p>
      <p>Smell
Multiple
Value Smell</p>
      <p>BERT</p>
      <p>Analysis of results</p>
      <p>Bert Fine-Tuning
Split Value Multiple</p>
      <p>Smell Value Smell</p>
      <p>Research Questions
SDpelSittmeVcetalillnuge VDaMleuuteeltcSitpmilneegll</p>
      <sec id="sec-3-1">
        <title>3.1. Dataset Construction</title>
        <p>
          Due to the limited availability of well-defined semantic data smells datasets, we constructed two
comprehensive meta-datasets1 for detecting Split Value Smells and Multiple Value Smells, respectively [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
Source and Synthetic Data Generation. We collected real-world datasets from public repositories
such as Kaggle.2 However, since real datasets rarely include explicit annotations of data smells, we
employed synthetic data generation to create controlled and diverse examples. Specifically, we used the
Python Faker library [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] to generate human-readable tabular data that can be afected by these data
smells, including names, addresses, and dates of birth, enabling us to simulate plausible examples of
well-formed and smelly data structures.
        </p>
        <p>
          To enhance diversity and quantity, we expanded the dataset using Synthetic Data Vault (SDV) [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ],
which enables the generation of synthetic tabular data while preserving the statistical properties of the
original data. We trained SDV synthesizers in our initial datasets to generate additional records, thus
balancing and scaling the data. The validity of the generated data was confirmed using the built-in
diagnostic tools of SDV3. All data sets achieved 100% synthetic data validation scores, indicating high
ifdelity in capturing the original data distributions and relational structures.
        </p>
        <p>Data Smell Injection. We injected two data smells: Split Value Smells and Multiple Value Smells. On
the one hand, we programmatically fragmented single logical values (e.g., full addresses or dates) into
multiple columns. On the other hand, we concatenated logically distinct values into single fields, such
as merging “email” and “phone number” into a single string-separated column.</p>
        <p>
          Tables 1 and 2 show examples of these transformations. In total, we generated 40 base datasets (20
per smell type), each consisting of 5,000 rows. These were then aggregated into two meta-datasets
with approximately 400,000 labeled records. Due to the injection strategy, the class distribution in each
raw meta-dataset was naturally imbalanced. For example, in the case of Split Value Smells, the smelly
examples were overrepresented because each split introduced multiple rows. Conversely, Multiple Value
Smells tended to be underrepresented due to the merging of values into fewer fields. In addition, row
deduplication is performed across all data sets to ensure diversity. Finally, an undersampling strategy
was applied to mitigate learning bias. The final version used for training and evaluation contained a
balanced 50/50 distribution of positive and negative labels. Specifically, we applied stratified sampling
to preserve variation between diferent base datasets and structural patterns. After this step, two
meta-datasets, each containing 5,000 rows, are created.
1A meta-dataset refers to a collection of datasets, where each entry captures general information about an individual dataset.
2https://www.kaggle.com/
3SDV diagnostic tool: https://docs.sdv.dev/sdv/single-table-data/evaluation/diagnostic
Preprocessing and Labeling. Each row in the meta-dataset was labeled as “smelly” or “not smelly”
based on whether it contained a known transformation. Additionally, to prepare the data for model
training with BERT, we transformed each row into a structured text format (e.g., Feature: Value),
allowing the model to treat tabular records as natural language sequences, as recommended in the recent
literature [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. After preprocessing, we serialized each data row into a text sequence and analyzed the
sequence length of each row to conform to BERT input requirements. The 94% of the sequences were
less than 50 tokens, with a maximum sequence length of 110 tokens.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Model Fine-Tuning</title>
        <p>
          To efectively detect semantic data smells within tabular data, we adopted a fine-tuning approach on
a pretrained BERT model. This decision leverages the strong contextualization capabilities of BERT,
which, although initially designed for natural language processing tasks, has recently demonstrated
high efectiveness in structured data domains [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>
          Preprocessing and Input Representation. The fine-tuning process began transforming tabular
records into text sequences. Each row was serialized into a single textual string by concatenating
feature names with their corresponding values (e.g., Feature: Value). This format ensures that the
BERT tokenizer can efectively segment and encode the data while preserving the structure of the
original features. All inputs were tokenized using the BERT base uncased tokenizer and padded or
truncated to a maximum sequence length of 128 tokens. Given the binary nature of our task (smelly
vs. not smelly), each data sample was associated with a binary label. The input sequences were then
divided into training (80%) and validation (20%) sets using stratified sampling to maintain label balance.
Model Architecture and Training Setup. We used the bert-base-uncased model from the
Hugging Face Transformers library.4 A classification head—consisting of a dropout layer followed by a
dense layer—was added on top of the final [CLS] token representation. This setup enables binary
classification by mapping the contextualized embedding to a single probability score. The training was
carried out using AdamW optimizer [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], with a learning rate of 2 × 10−5, the Binary Cross-Entropy
loss function, and a batch size of eight, three epochs, and a maximum sequence length of 128 tokens.
Each model was trained independently to avoid confusion caused by overlapping structural patterns
between smell types. This phase was performed on a NVIDIA Tesla T4 through Google Colab.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Public Data Availability</title>
        <p>
          To support replicability and facilitate future research, we have made all materials, including scripts and
datasets, publicly available in a permanent online appendix [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results of our Work</title>
      <p>This section reports on the findings of the study. To improve clarity, we discuss each RQ separately.</p>
      <sec id="sec-4-1">
        <title>4.1. On the Use of BERT in Detecting Split Value Smell</title>
        <p>As shown in Table 3a, the model achieved an accuracy of 88.99%, with perfect precision (1.000) and an
F1 score of 0.8763. Although these results confirm the ability of the model to detect Split Value Smells
with high reliability, the recall was 0.7799, indicating that some true positives were missed.</p>
        <p>Table 3b shows the confusion matrix for the Split Value Smells model. In particular, the model did
not produce false positives, resulting in perfect precision. However, it did not detect 11% of smelly
instances, which were mistakenly classified as clean (false negatives), highlighting a common trade-of
in binary classification: the model favors precision over recall. Such behavior is particularly desirable
4Available at: https://huggingface.co/google-bert/bert-base-uncased
detrimental if undetected anomalies significantly impact downstream tasks.
when trained on relatively limited data.</p>
        <p>One notable challenge lies in the semantic variability of Split Value Smells. Unlike syntactic anomalies,
these smells are often context-dependent. For example, splitting an address into “Street”, “City”, and
“Zip Code” might be acceptable in some domains but smelly in others if those components are only
meaningful when considered as a whole. This ambiguity limits the model’s ability to generalize across
domains and may explain the relatively lower recall. Moreover, diferent instances of splitting (e.g., full
name into first, middle, and last names) lack a consistent pattern, which increases intra-class variability
and training noise. The zero false-positive rate of the model demonstrates that it serves as a reliable
iflter for high-confidence cases. In practice, this makes it well suited for human-in-the-loop data quality
workflows, where automatic detection aids but does not replace manual review. A flagged instance can
be trusted to be problematic, while borderline or context-sensitive cases may still require user judgment.</p>
        <p>In summary, BERT can efectively detect Split Value Smell, ofering high precision and generalizability
despite the challenges posed by semantic ambiguity. Future improvements could involve integrating
auxiliary metadata to capture implicit semantic groupings better.</p>
        <p>ILST</p>
        <p>RQ1 — Summary of the Results.</p>
        <p>BERT efectively detects Split Value Smells with high precision and generalizability. Although
the recall is modest due to the semantic variability of smells, the model serves as a strong,
highconfidence filter in human-in-the-loop data cleaning workflows.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. On the Use of BERT in Detecting Multiple Value Smells</title>
        <p>As shown in Table 4a, the BERT-based model for detecting Multiple Value Smells achieved exceptional
results. The model reported an overall accuracy of 98.00%, with perfect precision (1.000) and a remarkably
high recall of 0.9599, which led to an F1 score of 0.9795. These metrics indicate that the model not
only correctly classifies true positives but also maintains a near-zero false negative rate, substantially
outperforming the counterpart model developed for Split Value Smells.
precision makes the model both sensitive and specific, a highly desirable trait in anomaly detection.
from the training accuracy, implying that the model avoids overfitting and maintains robustness.</p>
        <p>The superior performance of the Multiple Value model can be attributed to the more uniform
representation of the smells. Unlike Split Value Smells, which vary significantly in their structural form,
Multiple Value Smells often exhibit consistent patterns that are more easily captured by the token-based
architecture of BERT. These regularities provide stronger learning signals, enabling the model to
distinguish between smelly instances and clean data with high reliability. From a practical perspective, the
high recall achieved by the model ensures that very few smells are overlooked, making it suitable for
autonomous data cleaning tools or alert systems in large-scale data pipelines. Furthermore, perfect
precision allows its output to be used confidently without requiring downstream verification, minimizing
the risk of incorrect flagging.</p>
        <p>In conclusion, BERT is not only capable of detecting Multiple Value Smells but does so with near-optimal
performance. The consistent structure of these smells aligns well with BERT’s contextual encoding
capabilities, enhancing the utility of the model in detecting semantic level anomalies in tabular data.</p>
        <p>ILST</p>
        <p>RQ2 — Summary of the Results.</p>
        <p>BERT achieves near-perfect detection of Multiple Value Smells, with 98% accuracy, perfect precision,
and very high recall. The model is efectively generalized and benefits from the uniformity of the
smell structure, making it suitable for reliable automated anomaly detection in tabular data.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Implications</title>
      <p>The results obtained in this study have important implications for both the development of data quality
systems in practice and future research on semantic anomaly detection. By demonstrating that BERT
can be fine-tuned to detect data anomalies with high performance, this work opens new perspectives
on how language models can be leveraged to address long-standing issues in structured data validation.
lFask</p>
      <sec id="sec-5-1">
        <title>Implications for Practitioners</title>
        <p>From an applied standpoint, our findings suggest that transformer-based models are not only
theoretically interesting but also practically helpful in improving data quality workflows. The model trained to
detect Multiple Value Smells achieved near-perfect performance, with high recall and perfect precision,
meaning that data engineers and analysts can rely on it to flag problematic entries with minimal risk of
false positives automatically.</p>
        <p>Equally important, the Split Value Smells detector, despite a lower recall, also showed perfect precision,
indicating that its predictions are highly reliable when a smell is detected. The model could thus serve
as a high-confidence filter embedded in human-in-the-loop workflows, where semantic anomalies are
surfaced to data curators or engineers for review.</p>
        <p>Furthermore, the text-based input representation—where each row is serialized into a “natural
language-like” form—lowers the barrier to integration. It allows this kind of model to be embedded in
tools without deep coupling to schema logic or data format assumptions, making deployment across
diverse data platforms more feasible.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Implications for Researchers</title>
        <p>For researchers, this study highlights the underexplored potential of using language models for
structured data quality tasks. Although most prior work in NLP and tabular learning focuses on downstream
prediction or representation learning, our results demonstrate that pre-trained models, such as BERT,
can be adapted to detect subtle structural inconsistencies typically overlooked by both traditional
rule-based systems and data-centric AI methods.</p>
        <p>In addition, the successful use of synthetic smells, systematically injected into otherwise clean data,
demonstrates a promising way to create large-scale training datasets when labeled examples are scarce.
This result could be a useful methodological direction for further research on other types of semantic
anomalies, including those related to believability, completeness, or cross-column dependencies.
iLghtbul</p>
        <sec id="sec-5-2-1">
          <title>Implications — Summary</title>
          <p>RWENCH
RWENCH</p>
          <p>BERT can act as a high-precision semantic smell detector, enabling safe automation in data
quality workflows.</p>
          <p>The approach fits well within human-in-the-loop pipelines due to its low false-positive rate.
This work supports a novel line of research using language models to detect contextual
inconsistencies in structured data.</p>
          <p>Synthetic smell injection is a scalable solution to train data in semantic validation tasks.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Threats to Validity</title>
      <p>Although the results presented are encouraging, several threats to validity must be acknowledged.
Internal Validity. A potential threat to internal validity concerns the synthetic nature of the datasets
used. Although we used data generation techniques that simulate realistic tabular structures and
semantics, the smells were injected programmatically using deterministic rules. This injection may
not fully capture the complexity and variability of smells that arise in organically grown data sets.
Furthermore, the use of undersampling to achieve class balance may afect the natural distribution of
“smelly” vs. “clean” instances.</p>
      <p>External Validity. The external validity of our findings may be limited due to the reliance on synthetic
and semi-synthetic datasets. Although the data sets were designed to reflect common patterns in
realworld data, we did not perform experiments on large-scale industry data sets or diverse domains such as
ifnance, healthcare, or e-commerce. As such, the generalizability of the trained models across diferent
contexts remains to be evaluated.</p>
      <p>Construct Validity. We defined Split and Multiple Value Smells based on logical criteria and
formalized these into transformation rules. However, these definitions may not encompass all the nuances
recognized by domain experts. Additionally, converting tabular rows into text sequences for BERT
input involves abstraction choices that may omit structural cues present in multi-column formats.
Conclusion Validity. The conclusions drawn from this study are based on standard classification
metrics, including accuracy, precision, recall, and F1 score. Although these metrics provide a reliable
indication of model performance, a further evaluation of operational pipelines or user-facing tools is
necessary to assess practical efectiveness. Moreover, the limited number of training epochs and the
small amount of available GPU resources could constrain the models’ full learning potential.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions</title>
      <p>This paper investigated the applicability of transformer-based language models, specifically BERT,
for the automated detection of semantic data smells in tabular datasets. Our goal was to assess
the extent to which BERT can be fine-tuned to identify structural anomalies that require contextual
understanding—namely, Split Value Smells and Multiple Value Smells.</p>
      <p>Our results revealed that BERT performs well in both scenarios. The model trained for Split Value
Smells achieved high precision and generalizability, despite the intrinsic variability of how such smells
manifest in real-world schemas. In contrast, the model trained to detect Multiple Value Smells achieved
near-perfect performance, supported by the regularity of the merged value patterns. These outcomes
suggest that pre-trained language models can capture subtle semantic cues within structured data when
appropriately adapted.</p>
      <p>At the same time, our findings highlight that semantic variability and feature inconsistency can impact
model recall, particularly with fragmented values, underscoring the need for more comprehensive data
representations and possibly multi-column modeling strategies to leverage the potential of BERT in
this domain fully.</p>
      <p>Our future research agenda includes extending this work to additional categories of semantic smells
and exploring cross-dataset generalization. We also plan to integrate metadata-aware modeling and
assess the efectiveness of alternative transformer architectures optimized for tabular data. Ultimately,
our goal is to support the development of intelligent data validation tools that can operate at the
semantic level across diverse domains and formats.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This work has been partially supported by (1) the QUAL-AI national research projects funded by the
EU - NGEU and the MUR under the PRIN 2022 program (Contracts 2022B3BP5S), and (2) the project
“FAIR” (PE0000013). Additionally, the work is an extended version of Nicolò Gallotta’s Bachelor’s thesis,
developed at the University of Salerno in 2024. We gratefully acknowledge his valuable contribution to
the early stages of this research.</p>
    </sec>
    <sec id="sec-9">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used ChatGPT-4o and Grammarly to check grammar
and spelling.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Martínez-Fernández</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bogner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Franch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Oriol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Siebert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Trendowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Vollmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wagner</surname>
          </string-name>
          ,
          <article-title>Software engineering for ai-based systems: a survey</article-title>
          ,
          <source>ACM Transactions on Software Engineering and Methodology (TOSEM) 31</source>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>59</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kreuzberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kühl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hirschl</surname>
          </string-name>
          ,
          <article-title>Machine learning operations (mlops): Overview, definition, and architecture</article-title>
          ,
          <source>IEEE access 11</source>
          (
          <year>2023</year>
          )
          <fpage>31866</fpage>
          -
          <lpage>31879</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Baier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Jöhren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Seebacher</surname>
          </string-name>
          ,
          <article-title>Challenges in the deployment and operation of machine learning in practice</article-title>
          .,
          <source>in: ECIS</source>
          , volume
          <volume>1</volume>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Foidl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Felderer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ramler</surname>
          </string-name>
          ,
          <article-title>Data smells: Categories, causes and consequences, and detection of suspicious data in ai-based systems</article-title>
          ,
          <source>arXiv preprint arXiv:2203.10384</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Recupito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rapacciuolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Di</given-names>
            <surname>Nucci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Palomba</surname>
          </string-name>
          ,
          <article-title>Unmasking data secrets: An empirical investigation into data smells and their impact on data quality</article-title>
          ,
          <source>in: Proceedings of the IEEE/ACM 3rd International Conference on AI Engineering-Software Engineering for AI</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>53</fpage>
          -
          <lpage>63</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>G. F.</surname>
          </string-name>
          et al.,
          <article-title>Faker: Python package for generating fake data</article-title>
          ,
          <year>2024</year>
          . https://faker.readthedocs.io/.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Patki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wedge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Veeramachaneni</surname>
          </string-name>
          ,
          <article-title>The synthetic data vault</article-title>
          ,
          <source>in: IEEE International Conference on Data Science and Advanced Analytics (DSAA)</source>
          , IEEE,
          <year>2016</year>
          , pp.
          <fpage>399</fpage>
          -
          <lpage>410</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Renggli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rimanic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. M.</given-names>
            <surname>Gürel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Karlaš</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <article-title>A data quality-driven view of mlops</article-title>
          ,
          <source>IEEE Data Engineering Bulletin</source>
          <volume>44</volume>
          (
          <year>2021</year>
          )
          <fpage>11</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mohammed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Budach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Feuerpfeil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ihde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nathansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Noack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Patzlaf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Harmouch</surname>
          </string-name>
          ,
          <article-title>The efects of data quality on machine learning performance on tabular data</article-title>
          ,
          <source>Information Systems</source>
          <volume>132</volume>
          (
          <year>2025</year>
          )
          <fpage>102549</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G.</given-names>
            <surname>Recupito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pecorelli</surname>
          </string-name>
          , G. Catolino,
          <string-name>
            <given-names>S.</given-names>
            <surname>Moreschini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Di</given-names>
            <surname>Nucci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Palomba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Tamburri</surname>
          </string-name>
          ,
          <article-title>A multivocal literature review of mlops tools and features</article-title>
          ,
          <source>in: 2022 48th Euromicro Conference on Software Engineering and Advanced Applications (SEAA)</source>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>84</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Blase</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Cleanml:
          <article-title>A study for evaluating the impact of data cleaning on ml classification tasks</article-title>
          ,
          <source>in: 2021 IEEE 37th International Conference on Data Engineering (ICDE)</source>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.</given-names>
            <surname>Yin</surname>
          </string-name>
          , G. Neubig, W.-t. Yih,
          <string-name>
            <given-names>S.</given-names>
            <surname>Riedel</surname>
          </string-name>
          ,
          <article-title>Tabert: Pretraining for joint understanding of textual and tabular data</article-title>
          ,
          <source>in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>8413</fpage>
          -
          <lpage>8426</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Herzig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Nowak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mueller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Piccinno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Eisenschlos</surname>
          </string-name>
          , Tapas:
          <article-title>Weakly supervised table parsing via pre-training</article-title>
          ,
          <source>in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>4320</fpage>
          -
          <lpage>4333</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>X.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lees</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>Turl: Table understanding through representation learning</article-title>
          ,
          <source>ACM SIGMOD Record</source>
          <volume>51</volume>
          (
          <year>2022</year>
          )
          <fpage>33</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Fu</surname>
          </string-name>
          , S. Han,
          <string-name>
            <surname>D</surname>
          </string-name>
          . Zhang, Tuta:
          <article-title>Tree-based transformers for generally structured table pre-training</article-title>
          ,
          <source>in: Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery &amp; Data Mining</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1780</fpage>
          -
          <lpage>1790</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>H.</given-names>
            <surname>Iida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Thai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Manjunatha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Iyyer</surname>
          </string-name>
          , Tabbie:
          <article-title>Pretrained representations of tabular data</article-title>
          ,
          <source>in: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Koleva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ringsquandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Buckley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Tresp</surname>
          </string-name>
          ,
          <article-title>Named entity recognition in industrial tables using tabular language models</article-title>
          ,
          <source>in: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>348</fpage>
          -
          <lpage>356</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>C.</given-names>
            <surname>Wohlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Runeson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Höst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Ohlsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Regnell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wesslén</surname>
          </string-name>
          , Experimentation in software engineering, Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>I.</given-names>
            <surname>Loshchilov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hutter</surname>
          </string-name>
          ,
          <article-title>Decoupled weight decay regularization</article-title>
          ,
          <source>in: International Conference on Learning Representations (ICLR)</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>G.</given-names>
            <surname>Recupito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Giordano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Di</given-names>
            <surname>Nucci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Palomba</surname>
          </string-name>
          ,
          <article-title>Detecting semantic data smells with bert: Replication package</article-title>
          , https://doi.org/10.6084/m9.figshare.
          <volume>29328182</volume>
          .
          <year>v1</year>
          ,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>