<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Heterogeneity Detection and Comparison for Cost-Aware Graph Schema Evolution and Transformation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dominique Hausler</string-name>
          <email>dominique.hausler@ur.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Meike Klettke</string-name>
          <email>meike.klettke@ur.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Regensburg, Data Engineering Group</institution>
          ,
          <addr-line>Regensburg, Bavaria</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <abstract>
        <p>Graph databases allow direct data insertion without predefining a schema. While flexible, this often leads to a relaxed schema, especially without the usage of schema checking component, such as constraints. Consequently, heterogeneous structures emerge, requiring identification during data exploration. Furthermore, evolving such a relaxed schema can increase this heterogeneity even more. In this paper, we present Graph Data Profiles (GDPs) which capture initial heterogeneity by extracting statistical metadata (e.g., the number of labels and property keys for each node and edge). Additionally, these include uniqueness constraints (i.e., unique column combinations), existence constraints (mandatory or optional properties), missing edges, type constraints and valid inclusion dependencies. To support schema evolution and transformation, we introduce the GDP Dif, which underlines changes in heterogeneity. Thereby, inclusion dependencies help to detect in-version redundancies and evolution operations such as renaming or moving between versions. In contrast to other works, we address both schema evolution - essential for aligning a schema with changing requirements - and schema transformation, an frequent task in database monitoring. Our approach includes a cost calculation based on (1) a formalization of evolution or transformation operations and (2) heterogeneity metrics derived from the GDPs. Subsequently, GDPs ensure the detection of heterogeneity during data exploration, while the GDP Dif highlights the impact of an operation on the schema. This ensures user-informed decision making, supporting users to efectively clean unintended heterogeneity, such as typos.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Schema Evolution</kwd>
        <kwd>Schema Transformation</kwd>
        <kwd>Graph Data Profiling</kwd>
        <kwd>Heterogeneity</kwd>
        <kwd>Inclusion Dependencies</kwd>
        <kwd>Constraints</kwd>
        <kwd>Cost Functions</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Schema evolution is essential to retain the functionality of a system with changing external and internal
requirements. Thereby, even simple evolution operations can cause complex underlying schema changes,
especially when multiple software engineers operate on a single graph database. Furthermore, users
cannot estimate whether an evolution operation (add, rename, delete, copy, move, split, merge)
afects all selected data, a subset of the data, or none at all. Since schema-less databases allow direct
population without a predefined schema, issues regarding heterogeneous structures such as typos
frequently arise.</p>
      <p>Both schema evolution and transformation share identical components but are distinguished by which
component needs to be derived. Evolutionary tasks take a source schema and evolution operations to
generate a target schema. In contrast, schema transformation compares a source and a target schema to
identify schema modifications. We envision a modular framework that integrates both scenarios, aiming
to detect, visualize, and estimate the impact of schema evolution and transformation on heterogeneity.</p>
      <p>To address the issue of relaxed schemata, for instance, in collaborative environments, we present
Graph Data Profiles (GDPs) to capture heterogeneity in an initial data exploration phase by extracting
statistical metadata (e.g. number of occurrences for labels) and diferent constraints (e.g., uniqueness,
inclusion dependencies). Besides illustrating and quantifying heterogeneity, the GDPs enable
userinformed decisions during schema evolution or version comparison (schema transformation). More
precisely, they underline potential syntactic and semantic errors. These insights allow users to align the
schema with their intentions by cleaning accidental errors with our tool Nautilus. Thus, our approach
ofers a modular solution for data exploration, schema evolution, and transformation.</p>
      <p>To compare changes in irregularity a Graph Data Profile Dif (GDP Dif) is created (see Table 1),
enabling users to track how the degree of heterogeneity changed through a preview option during
schema evolution. In terms of transformation, the Difs give insights into how heterogeneity was
modified between versions. Additionally, a formalization for each evolution operation (see Table 2) is
provided. These formal definitions are combined with the heterogeneity metrics stored in the GDPs to
estimate an operation’s complexity. This is essential, as heterogeneity substantially impacts the costs of
schema changes.</p>
      <p>Contribution We propose GDPs to display heterogeneity in schema-less graph databases by
extracting statistical metadata to make implicit structures explicit during data exploration. Moreover,
the GDPs incorporate constraints, namely uniqueness, existence, type constraints, and inclusion
dependencies. Users benefit from GDP Difs when monitoring pipelines, as they depict changes in the
degree of heterogeneity between versions. During schema evolution, a preview of the emerging GDP
Dif alongside the operation’s cost is displayed, illustrating its impact. Inclusion dependencies, hereby,
support an automatized detection of adjustments, caused by rename or move operations, and in-version
redundancies emerging when copying. Due to the modularity of our framework, it is applicable for
both schema evolution and transformation.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Schema Evolution Schema Evolution is defined as the process of modifying a schema to align
with new requirements. Due to its importance, it is subject in a variety of works [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4">1, 2, 3, 4</xref>
        ]. Even
though graph databases are schema-less, PG-Schema resembles the state-of-the-art schema description
language [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ] for which constraints are defined in [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] present approaches on schema
extraction, considering challenges emerging in schema-flexible property graphs. A model-driven
approach for schema evolution is presented in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Extending these, we present a modular framework
to detect initial irregularities through GDPs alongside GDP Difs, aiding the estimation of an operation’s
impact and allowing a heterogeneity comparison between versions. Depicting heterogeneity supports
user-informed decisions, e.g. to clean relaxed structures.
      </p>
      <p>
        Schema Transformation Schema transformation is a generalization of schema evolution, and can
take place between two arbitrary schemata or databases. It requires schema matching (to identify
diferences between a source and target schema) and mapping strategies (to transform a source database
into a target databases). Various matching techniques are categorized in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] schemata are
converted into graphs, followed by graph matching methods. Pattern matching approaches are shown
in [
        <xref ref-type="bibr" rid="ref14 ref15 ref16">14, 15, 16</xref>
        ]. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] analyzes transformation between RDF and property graphs. Rafe et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] present
a graph meta model to represent semantics and estimate changes via graph transformation rules.
Andersen et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] automatically detect a minimal rule set, describing graph modifications. Our
framework compares the GDPs extracted from source and target database and generates GDP Difs,
emphasizing the aspect of heterogeneity. We formalize evolution and transformation operations, to be
later incorporated in our cost model.
      </p>
      <p>
        Graph Data Profiling Data profiling aims to gain insights into a dataset. In [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] renown profiling
strategies are adapted to graph data and categorized systematically. In alignment with this work, we
extract structural metadata, such as the number of occurrences, and constraints, such as uniqueness
and existence, to be integrated in our GDPs during data exploration. Work on graph dependencies
is presented in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] and [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. We, instead, integrate distance and similarity metrics often used to
detect inclusion dependency candidates. Moreover, we plan on integrating our bottom-up inclusion
dependency search [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] to (1) detect identical in-version data, resulting from copy or move operations
[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], and (2) to automatically identify rename or move operations during schema transformation.
Cost Functions Cost functions are used to quantify and estimate the complexity of a task or
operation. de Oliveira Werneck et al. [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] present a learning-approach for graph matching tasks. [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]
calculates costs for sequential path optimization, while Chen et al. [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] aim to improve multi-level logic
optimization algorithms. To determine similarities between graph patterns, Neuhaus and Bunke [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]
show how cost functions can be learned automatically. The metrics stored in the GDPs as well as the
formalized operations are integrated in our cost model, which aim to foresee the complexity and efect
of an operation.
      </p>
      <p>
        Data Quality Marchesin et al. [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] highlight the importance of data quality in Knowledge Graphs
(KGs). This results from the integration of LLMs to enrich KG environments, making an automatized
data quality evaluation necessary to ensure reliability. To increase the acceptance of AI systems used for
classification tasks, insights into data quality are essential [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] conducts a review on data quality
dimensions, whereas [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] summarizes available tools. We include the concept of GDPs to emphasize
heterogeneity, aiming for user-informed decisions to clean unintentionally emerged, relaxed structures
by executing evolution operations. These are defined in our evolution language Geo [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ], intuitively
describing an operation’s efect, thus, widening the range of users to non-domain experts.
      </p>
      <p>Source
Graph
Target
Graph</p>
      <p>Execute the</p>
      <p>Defined
BE Operation
1</p>
      <p>Profile
Extraction
Algorithm</p>
      <p>User</p>
      <p>Cost-Aware Schema Evolution &amp; Transformation
2
Source Graph
Data Profiles</p>
      <p>AT
Target Graph
Data Profiles</p>
      <p>Define</p>
      <p>Evolution
AE Operation</p>
      <p>Profile</p>
      <p>Comparison
3 Algorithm</p>
      <p>Identify</p>
      <p>Evolution</p>
      <p>BT Operations
4
Graph Data
Profile Diffs
Visualized for User-informed decision making
Automatized Processes
Mandaory Interaction</p>
      <p>Calculate</p>
      <p>Costs per
5 Operation</p>
      <p>Evolution</p>
      <p>Operations
6b
Target Graph</p>
      <p>Schema &amp;
6aProfile Diffs</p>
    </sec>
    <sec id="sec-3">
      <title>3. Use Cases</title>
      <sec id="sec-3-1">
        <title>Typical use cases in which our framework is beneficial:</title>
        <p>Use Case 1: In-Version Heterogeneity In a medical company, multiple software engineers
simultaneously manage a graph database without schema constraints, leading to irregularities. To detect these
variations, a data analyst uses our GDPs to clean unintentionally created variations by defining evolution
operations in our evolution language Geo. Subsequently, data quality is improved by homogenizing the
schema in alignment with user needs.</p>
        <p>Use Case 2: Heterogeneity Between Versions In the second scenario, a data analyst compares
diferent database versions. Our framework thereby, automatically generates GDP Difs to identify
transformation operations including their costs. This allows the analyst to quantify whether
heterogeneity increased e.g. by adding new data variants or decreased through homogenization. Understanding
schema changes when monitoring or optimizing pipelines is essential, for instance to draw conclusions
for similar pipelines in the manner of a knowledge base.</p>
        <p>Interim Conclusion Both scenarios outline how error-prone schemata in collaborative environments
are. Thus, illustrating the need to detect heterogeneous structures through our GDPs. Moreover, GDP
Difs ensure the estimation of an evolution operation’s impact and costs. During monitoring tasks,
the framework assists users by automatically detecting transformation operations, reducing manual
adjustments and overcoming a missing documentation of the evolutionary process.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. System Overview</title>
      <p>Figure 1 shows the workflow of our envisioned, modular framework. Blue elements indexed with E are
schema evolution specific, turquoise elements marked with T illustrate schema transformation, whereby
gray components are part of both tasks. In Step 1, a source graph database serves as input, followed by the
extraction of the schema components required to generate GDPs in Step 2. For schema transformation,
the GDPs are extracted and generated for both the source and target graph data (Step A ). In contrast, for
schema evolution, the GDPs are visualized to the user – indicated by the eye icon – to depict in-version
heterogeneity, i.e., synonyms or typos. Based on the information on heterogeneous structures in the
GDPs, the user can define an evolution operation (Step A  ), for instance, to decrease heterogeneity by
renaming labels with typos. Thereafter, the GDP Difs are generated in Step 3 and displayed in Step 4.
In Step 5, the metrics stored in the GDPs and the formalized operations (see Table 2) are incorporated in
the cost functions to illustrate the complexity of evolution or transformation operations. The output
consists of an interactive schema visualization and the GDP Difs for each entity, i.e., nodes and edges,
(Step 6a) together with an evolution or transformation operation log-file (Step 6b). During schema
evolution, the user then can decide whether the operation is executed (Step B ) or not. Moreover, the
GDP Difs incorporate inclusion dependencies to automatically resolve ambiguities when comparing
versions in Step B .</p>
    </sec>
    <sec id="sec-5">
      <title>5. Framework</title>
      <p>We present a modular framework that extends the graph evolution tool Nautilus – implementing our
evolution language Geo (Graph Evolution Operation) [34] – with Graph Data Profiles (GDPs). These
depict heterogeneous graph data and form the basis for predicting an operation’s efort.
Heterogeneity Presumptions As a result of the modeling freedom ofered by schema-less databases,
we distinguish between three ways of storing graph data. (1) Analogous to relational databases, users
can predefine a non-relaxed schema in a schema-first approach, utilizing constraints. (2) Just like in Use
Case 1, the data can be directly inserted without using schema constraints. (3) The third scenario is
a partially schema controlled usage, where homogeneous structures are forced by constraints, while
heterogeneous parts are kept schema-free. If schema-first approaches allow relaxed structures like
optional properties, they also belong to this class. Our approach focuses on the third case, as it is
often unclear whether observed heterogeneity is intentional. Hereby, our GDPs support users in
homogenization – for instance, by enforcing existence on formerly optional properties – and increasing
heterogeneity – e.g., by adding variants or evolving subsets. This allows users to choose between
homogeneity of data and structural variations in the data when evolving graph data. Additionally, the
GDP Difs give insight into heterogeneity modifications during version comparison like in Use Case 2.</p>
      <p>Graph Data Profiles (GDPs) To make heterogeneous structures explicit as needed in Use Case 1,
we present GDPs incorporating statistical and structural metadata – extracted by the algorithm used
in Step 1 of Figure 1 – for each label of a given graph database. The term label is used for both node
labels and edge types. Since relaxed schemata are common in collaborative environments, detecting
irregularities during data exploration provides deep insights into the dataset’s structure. Emphasizing
heterogeneity additionally depicts candidates for data cleaning, thereby enabling the analyst to optimize
the schema in alignment with data quality aspects. Distance metrics display syntactic and semantic
in-version candidates, whereby valid inclusion dependencies aid identifying identical information.
Furthermore, constraints highlight homogeneous structures that could be schema-controlled.
GDP Dif Table 1 depicts an exemplary GDP Dif, serving as a preview before executing an evolution
operation. Consequently, supporting user-driven decision making by depicting schema changes in the
context of heterogeneity. Such a preview is also beneficial to perform data cleaning (Use Case 1). When
monitoring pipelines, the GDP Difs facilitate a deeper understanding of schema modification and each
operation’s efect on the dataset’s heterogeneity, as described in Use Case 2.</p>
      <p>The GDP Dif in Table 1 illustratively shows Client nodes, formerly labeled Customer. The p
in the header shows that the label Customer was deleted. Changes in structure metadata such as
occurrence are icon-coded by ( for increasing and " for decreasing numbers. 25.56% of the Client
nodes are additionally labeled KeyAccount, i.e., depicting multi-labeling. The Dif shows a decreasing
number of multi-labeled Clients, caused by increasing their number from 65 to 90. Furthermore,
syntactic and semantic heterogeneity metrics – namely Levenshtein to detect typos and Wu-Palmer
for synonyms – are shown to detect structural irregularities. Precisely, the variation customer was
identified for the Client label.</p>
      <p>In order to analyze properties, constraints need to be made explicit. This includes deriving uniqueness
constraints, resembling unique column combinations, together with existence and type constraints
from the data. Moreover, valid inclusion dependencies serve an in-version detection of identical data,
caused by copy operations, as it is the case for . _ ⊆ . _ .
Here, the GDP Dif illustrates that this error was later cleaned by p under Levenshtein. A relaxed
candidate generation gives insights into renamed or moved elements (syntactic and semantic metrics). In
terms of inclusion dependencies, false positives might occur, e.g., ._ ⊆  . .
Nevertheless, user-centric cleaning decisions ensure robustness, while depicting potential overlaps.
Another key aspect is illustrated in the super-type cell. Here the associated label for each property is
depicted. Solely for node entities all associated edges are displayed, with the objective of identifying
syntactic and semantic errors, as it is the case for p BOUGHT, as well as to display missing and required
edges (= existence constraint).</p>
      <p>̂︀</p>
      <p>Based on the GDP Dif, the following representative conclusions can be drawn: The deletion of
the associated edge labeled BOUGHT, formerly connecting Customer and produtc, a variant of the
(Customer)-[BOUGHT]→(Product) pattern, represents a homogenization action. Consequently, a
GDP Dif-driven analysis reveals a renaming process. In the properties section the key firRst_name
from Cusotmer was added, resembling a move operation.</p>
      <p>Cost Functions Besides GDPs, we envision a cost model to capture the impact of each operation
executed during schema evolution or transformation. In this context, heterogeneity can cause severe
structural changes; consequently, the cost model accounts for both an operation’s complexity and its
efect on data irregularities.</p>
      <p>For the complexity estimation, each evolution operation is formalized in an implementation near
manner, as exemplified in Table 2 for add. This allows users to predict the complexity of the Cypher
code. In alignment with our evolution language Geo, every operation is partitioned by entity types
(nodes and edge) and features (label, property). In our notation,  represents the aligning Cypher
command, in this context CREATE or MERGE.  is the finite set of available labels, subdivided into
node labels  and edge labels  , whereby  and  resemble a precise instance. Properties are
defined as key-value pairs  ×  , whereby  is a Cypher variable to refer to a defined pattern, like in
CREATE(n:Client). Since Geo translates directly into Graph Manipulation Operations (GMOs), the
formalization is based on the aligning Geo statement.</p>
      <p>
        To ensure a cost-aware schema evolution and schema transformation, the cost function are further
refined by the following aspects derived from the metadata stored in the GDPs and categorized according
to [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>• Structure metadata: The number of afected entities to estimate pattern complexity.
• Cardinalities and value distribution: Analysis of whether constraints become invalid as
heterogeneity increases or whether new constraints emerge by homogenizing former irregularities.
• Paths metadata: Identification of important nodes with numerous edges, prone to cause a structural
break upon their modification.
• Patterns, data types &amp; domains: Use of syntactic distance and semantic similarity to detect potential
heterogeniouse structures.
• Functional: Deriving graph inclusion dependencies to detect renamed, copied or moved structures
in and between versions or databases.</p>
      <p>For each aspect, a weight is defined in association of its efect on heterogeneous structures, whereby
a value of one resembles homogeneity. Costs are sensitive to irregularities, thus, requiring the
composition of GDP metadata and operation complexity. Our approach considers cost-estimations for
monitoring data engineering pipelines and supports user decisions via a GDP Dif preview together
with the emerging target schema in terms of schema evolution. The tool displays the resulting costs
approximation to the user.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Data Quality Aspects</title>
      <sec id="sec-6-1">
        <title>The framework considers two data quality dimensions:</title>
        <p>Syntactic Accuracy The integration of GDPs as well as the GDP Dif preview aim to support data
cleansing by improving syntactic accuracy. This is accomplished by depicting candidates for typical
errors occurring in relaxed schemata, such as typos, by measuring the syntactic distance. Subsequently,
users can utilize the evolution form to clean the data accordingly (see Use Case 1). Since users cannot
estimate the impact of an evolution operation on the dataset, we propose a GDP Dif preview serving
as control mechanism before actually conducting the operation on the original dataset. Thus, the GDPs
depict heterogeneity during data exploration caused by the schema-less environment. In addition, the
GDP Dif allows estimating the impact of schema evolution and gives insight into irregularity changes
during schema transformation (see Use Case 2).</p>
        <p>Semantic Consistency The GDPs specialize in detecting semantic errors to facilitate user-driven
decision making, thereby, increasing data consistency. To achieve this, the semantic similarity between
entity types and properties is extracted, allowing refinements of structures that emerged unintentionally
during schema evolution. On top of that, constraints are included to illustrate homogeneous parts
suitable to be put under schema control. Regarding data dependencies, GDPs store information on
valid inclusion dependencies, which provide insights into (1) identical in-version information as well as
(2) evolution or transformation operations like rename, move or copy. Each GDP represents the left
side of an inclusion dependency. Moreover, the GDP Difs make increases and decreases in semantic
irregularities during monitoring tasks visible.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Implementation</title>
      <p>Nautilus implements our evolution language Geo, closing the gap for an evolution language for
property graphs. Geo aims to widen the range of users by intuitively defining each operation, requiring
no prior knowledge of Cypher. The implementation is available on GitHub1. The tool includes a form
to define the renown evolution operations add, rename, delete, copy, move, spit, merge and the
graph-specific transform operation (i.e., converting a node into an edge and vice versa). Moreover, a
ifrst version of the GDPs is integrated in a data exploration option. This includes information on label
occurrences; partitioned by nodes and edges. For each label all associated properties of the initial and
latest schema versions are depicted, allowing a manual comparison.</p>
      <p>We plan to extend the current data profiles in Nautilus by integrating the GDPs. This will highlight
irregularities in the source graph and enable users to clean variations by evolving the schema with
our evolution form. After defining the evolution operation, the resulting GDP Dif (see Table 2) is
displayed in a preview along with the estimated costs. This supports users in estimating the impact of
an operation on the schema and its underlying data before executing it on the original data. To evaluate
the accuracy of the presented GDPs, we will generate a dataset with heterogeneous structures, such
as optional properties, missing edges, and variations, serving as gold standard. Initial tests revealed
a high accuracy of the current data profile generation. In the future, we will extend the evaluation
by determining precision, recall and F1 score derived from true positives, false positives, and false
negatives.</p>
      <sec id="sec-7-1">
        <title>1Nautilus: Implementation</title>
        <p>Nautilus-Graph-Schema-Evolution
of a graph evolution language:
https://github.com/DominiqueHausler/</p>
        <sec id="sec-7-1-1">
          <title>Operation</title>
          <p>node
add
edge
label
property</p>
          <p>Formalized Graph Manipulation Operation
( :  { })
with  ∈ {CREATE, MERGE},
 ⊆  × ,   ∈ 
() − [ :   { }] → ()
with  ∈ {CREATE, MERGE},
 ⊆  × ,   ∈ 
SET  :</p>
          <p>outputting ′ () = () ∪ {}
SET . = 
with  ′() =  () ∪ {(, )}</p>
        </sec>
        <sec id="sec-7-1-2">
          <title>Evolution language – Geo</title>
          <p>"add node with label" label
"add relationship with type" label
"add label" label "to node with label" labels
"add" ("unique")? ("mandatory")? "property" property
"with datatype" datatype "to node with label" label</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>8. Conclusion</title>
      <p>Our presented framework facilitates user-informed decision making in the terms of schema evolution
and transformation – two closely intertwined approaches. For schema transformation, the framework
outlines heterogeneity changes in form of Graph Data Profile Difs (GDP Difs), additionally calculating
the costs for each transformation operation. Thus, the framework can be seamlessly integrated into
monitoring components of data engineering pipelines. Evolutionary tasks benefit from a GDP Dif
preview, allowing users to estimate an operation’s impact and costs before executing it on the
original data. Additionally, GDPs highlight potential errors through heterogeneity metrics, enabling the
user to improve data quality by defining evolution operations to clean them. To determine whether
heterogeneity in or between versions increased, decreased or remained consistent, the cost functions
incorporate the information stored in the GDPs alongside the formalized operation. Users, therefore,
are provided a precise estimation of an operation’s impact together with a comprehensive analysis of
relaxed schemata in schema-less graph databases.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>The work of Dominique Hausler has been funded by Deutsche Forschungsgemeinschaft (German
Research Foundation) – 552691270.</p>
    </sec>
    <sec id="sec-10">
      <title>Declaration on Generative AI</title>
      <p>The logo was generated with DALL-E 3. ChatGPT and AI Studio were used for rephrasing and to find
synonyms.
in Neo4j, in: ER (Companion), volume 3618 of CEUR Workshop Proceedings, CEUR-WS.org, 2023.
[34] D. Hausler, M. Klettke, Nautilus: Implementation of an evolution approach for graph databases,
in: MoDELS (Companion), ACM, 2024, pp. 11–15.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Brahmia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Grandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Oliboni</surname>
          </string-name>
          ,
          <article-title>A literature review on schema evolution in databases</article-title>
          ,
          <source>Computing Open</source>
          <volume>02</volume>
          (
          <year>2024</year>
          )
          <fpage>2430001</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Eckwert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guckert</surname>
          </string-name>
          , G. Taentzer,
          <article-title>EvolveDB: Evolving relational database schemas in a modeldriven way</article-title>
          ,
          <source>Software and Systems Modeling</source>
          (
          <year>2025</year>
          )
          <fpage>1</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Cleve</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gobert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Meurice</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Maes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Weber</surname>
          </string-name>
          ,
          <article-title>Understanding database schema evolution: A case study</article-title>
          ,
          <source>Sci. Comput</source>
          . Program.
          <volume>97</volume>
          (
          <year>2015</year>
          )
          <fpage>113</fpage>
          -
          <lpage>121</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>U.</given-names>
            <surname>Störl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Klettke</surname>
          </string-name>
          , S. Scherzinger,
          <article-title>NoSQL schema evolution and data migration: State-of-the-art and opportunities</article-title>
          , in: EDBT, OpenProceedings.org,
          <year>2020</year>
          , pp.
          <fpage>655</fpage>
          -
          <lpage>658</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bonifati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Furniss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Green</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Harmer</surname>
          </string-name>
          , E. Oshurko,
          <string-name>
            <given-names>H.</given-names>
            <surname>Voigt</surname>
          </string-name>
          ,
          <article-title>Schema validation and evolution for graph databases</article-title>
          ,
          <source>in: ER</source>
          , volume
          <volume>11788</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2019</year>
          , pp.
          <fpage>448</fpage>
          -
          <lpage>456</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Angles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bonifati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dumbrava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Fletcher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Green</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hidders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Libkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Marsault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Martens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Murlak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Plantikow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Savkovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sequeda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Staworko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tomaszuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Voigt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Vrgoc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zivkovic</surname>
          </string-name>
          , PG-Schema:
          <article-title>Schemas for Property Graphs</article-title>
          ,
          <source>Proc. ACM Manag. Data</source>
          <volume>1</volume>
          (
          <year>2023</year>
          )
          <volume>198</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>198</lpage>
          :
          <fpage>25</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Tomaszuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E. L.</given-names>
            <surname>Gayo</surname>
          </string-name>
          ,
          <article-title>On Property Constraints in PG-Schema, in: K-CAP</article-title>
          , ACM,
          <year>2025</year>
          , pp.
          <fpage>69</fpage>
          -
          <lpage>73</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Angles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bonifati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dumbrava</surname>
          </string-name>
          , G. Fletcher,
          <string-name>
            <given-names>K. W.</given-names>
            <surname>Hare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hidders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. E.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Libkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Martens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Murlak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Perryman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Savkovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Sequeda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Staworko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tomaszuk</surname>
          </string-name>
          , PG-Keys:
          <article-title>keys for Property Graphs</article-title>
          , in: SIGMOD Conference, ACM,
          <year>2021</year>
          , pp.
          <fpage>2423</fpage>
          -
          <lpage>2436</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bonifati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dumbrava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Martinez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ghasemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jafré</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Luton</surname>
          </string-name>
          , T. Pickles,
          <article-title>DiscoPG: Property Graph schema discovery and exploration</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>15</volume>
          (
          <year>2022</year>
          )
          <fpage>3654</fpage>
          -
          <lpage>3657</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sideri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Troullinou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Ymeralli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Efthymiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Plexousakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kondylakis</surname>
          </string-name>
          ,
          <article-title>PG-HIVE: hybrid incremental schema discovery for Property Graphs</article-title>
          ,
          <source>CoRR abs/2512</source>
          .01092 (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Terwilliger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Bernstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Unnithan</surname>
          </string-name>
          ,
          <article-title>Worry-free database upgrades: automated modeldriven evolution of schemas and complex mappings</article-title>
          , in: SIGMOD Conference, ACM,
          <year>2010</year>
          , pp.
          <fpage>1191</fpage>
          -
          <lpage>1194</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E.</given-names>
            <surname>Rahm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Bernstein</surname>
          </string-name>
          ,
          <article-title>A Survey of Approaches to Automatic Schema Matching</article-title>
          , VLDB J.
          <volume>10</volume>
          (
          <year>2001</year>
          )
          <fpage>334</fpage>
          -
          <lpage>350</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sun</surname>
          </string-name>
          , G. Wang,
          <article-title>Schema matching based on SQL statements</article-title>
          ,
          <source>Distributed Parallel Databases</source>
          <volume>38</volume>
          (
          <year>2020</year>
          )
          <fpage>193</fpage>
          -
          <lpage>226</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>W.</given-names>
            <surname>Martens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Niewerth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Popp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rojas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vansummeren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Vrgoc</surname>
          </string-name>
          ,
          <article-title>Representing paths in graph database pattern matching</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>16</volume>
          (
          <year>2023</year>
          )
          <fpage>1790</fpage>
          -
          <lpage>1803</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>P.</given-names>
            <surname>Barceló</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Reutter</surname>
          </string-name>
          ,
          <article-title>Schema mappings and data exchange for graph databases</article-title>
          , in: ICDT, ACM,
          <year>2013</year>
          , pp.
          <fpage>189</fpage>
          -
          <lpage>200</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>N.</given-names>
            <surname>Francis</surname>
          </string-name>
          , L. Libkin,
          <article-title>Schema mappings for data graphs</article-title>
          , in: PODS, ACM,
          <year>2017</year>
          , pp.
          <fpage>389</fpage>
          -
          <lpage>401</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bruyat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Champin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Médini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Laforest</surname>
          </string-name>
          ,
          <article-title>PRSC: from PG to RDF and back, using schemas</article-title>
          ,
          <source>Semantic Web</source>
          <volume>15</volume>
          (
          <year>2024</year>
          )
          <fpage>2555</fpage>
          -
          <lpage>2595</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>V.</given-names>
            <surname>Rafe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Golparian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rasoolzadeh</surname>
          </string-name>
          ,
          <article-title>Using graph transformation systems to formalize tropos diagrams</article-title>
          ,
          <source>J. Vis. Lang. Comput</source>
          .
          <volume>30</volume>
          (
          <year>2015</year>
          )
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Andersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Davoodi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fagerberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Flamm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Fontana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kolcák</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. V. F. P.</given-names>
            <surname>Laurent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Merkle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nøjgaard</surname>
          </string-name>
          ,
          <source>Automated inference of graph transformation rules</source>
          ,
          <source>CoRR abs/2404</source>
          .02692 (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Maiolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Etcheverry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Marotta</surname>
          </string-name>
          ,
          <article-title>Data profiling in Property Graph databases</article-title>
          ,
          <source>ACM J. Data Inf. Qual</source>
          .
          <volume>12</volume>
          (
          <year>2020</year>
          )
          <volume>20</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          :
          <fpage>27</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Shimomura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. H. L.</given-names>
            <surname>Fletcher</surname>
          </string-name>
          , N. Yakovets, ProGGD
          <article-title>- data profiling on Knowledge Graphs using graph generating dependencies</article-title>
          , in: ISWC (Posters/Demos/Industry), volume
          <volume>3632</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>M.</given-names>
            <surname>Manouvrier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Belhajjame</surname>
          </string-name>
          ,
          <article-title>Graph functional dependencies: Analysis and translation to PG-Schema, Inf</article-title>
          . Syst.
          <volume>136</volume>
          (
          <year>2026</year>
          )
          <fpage>102633</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hausler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Conrad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sperling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Störl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Klettke</surname>
          </string-name>
          ,
          <article-title>Discovering inclusion dependencies in a multi-model scenario</article-title>
          ,
          <source>in: ER</source>
          , volume
          <volume>16189</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2025</year>
          , pp.
          <fpage>242</fpage>
          -
          <lpage>260</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Klettke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Awolin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Störl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Scherzinger</surname>
          </string-name>
          ,
          <article-title>Uncovering the evolution history of data lakes</article-title>
          , in: IEEE BigData, IEEE Computer Society,
          <year>2017</year>
          , pp.
          <fpage>2462</fpage>
          -
          <lpage>2471</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>R. de Oliveira</surname>
            <given-names>Werneck</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Raveaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tabbone</surname>
          </string-name>
          , R. da Silva Torres,
          <article-title>Learning cost functions for graph matching</article-title>
          , in: S+SSPR, volume
          <volume>11004</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2018</year>
          , pp.
          <fpage>345</fpage>
          -
          <lpage>354</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>J. AbuBekr</surname>
            , I. Chikalov,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Hussain</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Moshkov</surname>
          </string-name>
          ,
          <article-title>Sequential optimization of paths in directed graphs relative to diferent cost functions</article-title>
          ,
          <source>in: ICCS</source>
          , volume
          <volume>4</volume>
          of Procedia Computer Science, Elsevier,
          <year>2011</year>
          , pp.
          <fpage>1272</fpage>
          -
          <lpage>1277</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zuo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , E-Syn:
          <article-title>E-graph rewriting with technology-aware cost functions for logic synthesis</article-title>
          , in: DAC, ACM,
          <year>2024</year>
          , pp.
          <volume>124</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>124</lpage>
          :
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>M.</given-names>
            <surname>Neuhaus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bunke</surname>
          </string-name>
          ,
          <article-title>Automatic learning of cost functions for graph edit distance</article-title>
          ,
          <source>Inf. Sci</source>
          .
          <volume>177</volume>
          (
          <year>2007</year>
          )
          <fpage>239</fpage>
          -
          <lpage>247</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>S.</given-names>
            <surname>Marchesin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Alonso</surname>
          </string-name>
          ,
          <article-title>Large language models and data quality for Knowledge Graphs, Inf</article-title>
          . Process. Manag.
          <volume>62</volume>
          (
          <year>2025</year>
          )
          <fpage>104281</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sadhukhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <article-title>A graph theoretic approach to assess quality of data for classification task</article-title>
          ,
          <source>Data Knowl. Eng</source>
          .
          <volume>158</volume>
          (
          <year>2025</year>
          )
          <fpage>102421</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>F.</given-names>
            <surname>Sidi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. H. S.</given-names>
            <surname>Panah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Afendey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Jabar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ibrahim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mustapha</surname>
          </string-name>
          ,
          <article-title>Data quality: A survey of data quality dimensions</article-title>
          , in: CAMP, IEEE,
          <year>2012</year>
          , pp.
          <fpage>300</fpage>
          -
          <lpage>304</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ehrlinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wöß</surname>
          </string-name>
          ,
          <article-title>A survey of data quality measurement and monitoring tools</article-title>
          ,
          <source>Frontiers Big Data</source>
          <volume>5</volume>
          (
          <year>2022</year>
          )
          <fpage>850611</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hausler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Klettke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Störl</surname>
          </string-name>
          ,
          <article-title>A language for graph database evolution and its implementation</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>