<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards New Data Quality Rules for Modeling Data Change</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nishttha Sharma</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Supervised by: Dr. Fei Chiang McMaster University</institution>
          ,
          <addr-line>Hamilton ON</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Data is not static, and attribute value changes often trigger changes in another set of attributes. Traditional methods for analyzing data changes often treat these changes in isolation, failing to consider the broader context in which they occur. This lack of contextual awareness limits the ability to capture relationships between attributes or interpret their significance, especially when distinguishing between normal variations and potential anomalies. In this paper, we discuss the importance of context-awareness and the need to identify normal change behaviour. To achieve this, we introduce a new data quality rule, called change rule, capable of capturing changes in both antecedent and consequent attributes within ordered tuples of a relational instance.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Data Dependencies</kwd>
        <kwd>Dynamic Data Dependencies</kwd>
        <kwd>Change Exploration</kwd>
        <kwd>Change Dependency</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>In real-world datasets, values rarely remain static as data
continuously changes over time. These changes often carry
critical information, revealing patterns, trends, and triggers
that are essential for understanding environmental
conditions, system, user behaviour and trends. Existing database
systems have limited functionality to manage changes, and
to identify abnormal changes, often relying on triggers to
recognize out-of-bound changes. In this work, we consider
changes to relational attributes for an entity. To simplify
our setting, attribute changes are modeled as a sequence
of ordered tuples, implicitly with respect to time. Hence, a
tuple represents the value each attribute holds for an entity
at a specific point in time.</p>
      <p>
        Data changes occur in numeric and non-numeric
attributes. Changes to numeric attributes are often measured
using absolute diference, percentage change, rate of change,
rolling average [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. While these metrics are easy to compute,
they fail to capture the broader context of the change, such
as the influence of related attributes or the significance of
the change.
      </p>
      <p>
        For non-numeric attributes, changes are often measured
using edit distances (Levenshtein, Jaro-Winkler, Hamming)
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] or set-based coeficients (Overlap, Jaccard, Dice) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
However, these metrics are insuficient because they ignore
the semantic meaning of the changes and the context in
which they occur. Context is critical because it provides
the necessary information to interpret the significance of a
change. Without context, changes are reduced to isolated
events, which can lead to misleading interpretations of the
data change.
      </p>
      <p>Example 1. Table 1 shows two employees (Emp) E1 and
E2 and their Position, Salary and number of employees
managed (EmpMng) as of a specific Year. Consider the
following changes and the need for greater context:
Numeric attribute value changes: As observed in tuples
1 − 3 of Table 1, after only two years as a Software
Developer, E1 was promoted to the position of Senior Software
Developer accompanied by a significant increase in salary
($68,400 to $82,000). Whereas tuples 9 − 13 show that E2
spent four years as a Software Developer before being
pro</p>
      <p>Year</p>
      <p>Emp
EmpMng

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17</p>
      <p>Position</p>
      <p>Salary
moted to Senior Software Developer with a similar salary
increase as E1’s (from $71,500 to $80,000). Changes in salary
are typically quantified using percentage change (+19.9%
for E1 and +11.9% for E2). While this provides a numerical
summary of the change, it fails to account for the broader
context. For instance, E1 received a larger raise after a
shorter tenure and took on the responsibility of managing
four employees, whereas E2 had to wait twice as long for a
similar promotion and gained the responsibility of
managing two fewer employees compared to E1.</p>
      <p>Non-numeric changes within and between classes:
Traditional edit distance metrics such as Levenshtein distance
(LD) quantify changes based on character modifications.
The transition from Software Developer to Senior Software
Developer has an LD = 7, whereas for Senior Software
Developer to Lead Developer, LD = 13. These values suggest that
the latter change is almost twice as significant as the former
despite both changes being promotions to the next
position within the same class (development roles), as shown in
Figure 1.</p>
      <p>The implications of a change can be much greater
between diferent classes. For instance, the LD between Lead
Developer and Manager is 11 which suggests that this
transition is smaller than the transition from Senior Software
Developer to Lead Developer (LD = 13). However, this
interpretation is misleading. The change from Lead Developer
to Manager represents a more significant career shift as
compare to the change from Senior Software Developer
to Lead Developer (where both positions are in the same
class), as it involves moving from a development role to
a managerial position (2 levels up as per Figure 1) which
is accompanied by a significant increase in the number of
people managed. Existing distance measures fail to capture
semantic interpretations of the data.</p>
      <p>Problem 1: The need for context. The example
highlights that not all changes are equally significant. Context
is often needed to interpret data change, and there is a need
to augment existing distance metrics with context.</p>
      <p>
        While identifying (significant) changes is important, it is
equally critical to diferentiate between normal changes vs
abnormal changes. Traditional methods have used
declarative methods such as data dependencies of the form  →  ,
where ,  are attribute sets, representing antecedent and
consequent attributes. Order Dependencies (OD) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
Sequential Dependencies (SD) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and Diferential
Dependencies (DD) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] specify expected relationships between
attribute sets. ODs introduce ordering relationships but
do not explicitly quantify changes in attribute values. SDs
model consequent attribute changes but do not account for
variations in the antecedent attributes. DDs, while
addressing changes in both antecedent and consequent attributes,
apply to unordered data.
      </p>
      <p>Example 2. Consider a sequential dependency (SD)
stating that when ordered by Position, the change in Salary
between consecutive tuples should be between 5% and 20%.
For E1, the SD is violated between 3 and 4 with a salary
change of 2.6% falling below the range. It is also violated
between 7 and 8 , where the salary change (23.8%) exceeds
the upper bound. These violations help in identifying
abnormal changes. However, we also want to identify patterns
where diferent changes in the antecedent attributes, such
as changes within Position will elicit diferent changes in
the consequent (Salary). For instance, with no change in
position, salary still changes annually by 2% to 10%.
Whenever there is a promotion (change in position &gt; 0), the salary
always changes by 10% to 25%. The existing dependencies
do not capture relationships of this form.</p>
      <p>To address this, we define a data quality rule called
change rule. The change rule captures relationships
between changes in attribute values of an ordered relational
instance.</p>
      <p>Problem 2: Diferentiating normal vs. abnormal
data change. Existing dependencies do not capture the
dependence between changes from antecedent attributes
to consequent attributes on ordered tuples. A declarative
specification is needed that models the expected range of
value change between attribute sets. We propose change
rules to address this problem.
1.1. Challenges
• Context representation: Context helps to interpret the
significance of data changes. Which attributes, and which
subset of values are used to provide this context? Is this
context time-dependent? How are existing distance
measures augmented to consider this context?
• Eficient rule mining: Manual specification of change
rules is not practically feasible, and automated solutions
are needed. Determining dependent sets of attributes is
important towards identifying meaningful data change.
Exhaustive enumeration of all attribute sets and their
values is not feasible, and eficient methods to evaluate
the large space of attribute sets are needed.
• Filtering spurious changes. Rule mining is known to
produce spurious rules. Determining which changes are
most relevant and defining (support) measures that filter
less meaningful changes is necessary.</p>
      <sec id="sec-1-1">
        <title>1.2. Contributions</title>
        <p>
          We expect to make the following contributions.
• Context-aware change metric: A metric for
quantifying changes in both numeric and non-numeric attributes,
augmenting them with contextual information from
related attributes.
• Change rules: A new rule that captures the relationship
of changes from one attribute set  to another attribute
set  across ordered tuples.
• Change rule discovery algorithm: An eficient
discovery algorithm for change rules over ordered datasets. The
algorithm adapts the FastDD method to handle ordered
data [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], and identifies changes in sequential attribute
values using context-aware metrics.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>We discuss the relationship of our work to existing
metrics, data dependencies, association rules and statistical/ML
approaches.</p>
      <sec id="sec-2-1">
        <title>2.1. Similarity, Distance Metrics</title>
        <p>
          Traditional numeric metrics analyze individual attributes in
isolation, missing contextual relationships between changes
in diferent attributes. Measures of central tendency (mean,
median, mode) summarize values but can be skewed by
outliers. Dispersion metrics (variance, standard deviation, IQR)
capture data spread, but also prone to outlier sensitivity.
Shape distribution measures (skewness, kurtosis, CV)
describe asymmetry and variability but can be biased when
data is highly skewed or sparse [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>
          For non-numeric (categorical, text) data, cosine similarity
is commonly used. Cosine similarity measures the cosine
of the angle between vectors [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and is often used with
embeddings to capture semantic similarity. Overlap, Jaccard,
and Dice Coeficients [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] are used to quantify the similarity
and diversity of sets. Edit distance such as Levenshtein,
Jaro-Winkler, Hamming quantifies the number of operations
needed to transform one string into another [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>While these metrics are widely used, they do not capture
semantic distances. For example, "Software Developer" and
"Senior Software Developer" have a high edit distance
despite being closely related in meaning. Embedding-based
approaches (e.g., BERT) address this by capturing contextual
meaning but require pre-trained models and domain-specific
tuning. An efective approach for measuring semantic
similarity between non-numeric values is to compute cosine
similarity on BERT embeddings, which allows for a
contextaware representation of the data.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Data Dependencies</title>
        <p>
          Order Dependencies (ODs) extend functional dependencies
by enforcing ordering relationships [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. They ensure that a
positive change in the antecedent corresponds to a positive
change in the consequent. However, the semantics of ODs
do not declaratively capture the change in any attribute
values.
        </p>
        <p>
          Sequential Dependencies (SDs) declaratively specify the
change in consequent attributes [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. They enforce
constraints on how the consequent changes in response to an
instance ordered on the antecedent, i.e. when the instance
is ordered on , the changes in the consecutive  -values
will be within a range . However, they fail to capture the
change in the antecedent. Conditional SDs (CSDs) focus on
identifying intervals within ordered data that satisfy a given
SD. They prefer larger, contiguous intervals that capture
a substantial portion of the data satisfying the embedded
SD. However, the continuity of these intervals requires a
trade-of with the specificity of the bound , which is not
addressed in the paper.
        </p>
        <p>
          Diferential Dependencies (DDs) model diferences
between any two tuples in a relation independent of the
tuple ordering, i.e., if the antecedent attribute diferences lie
within a range , then the consequence attribute value
diferences must lie within a range  [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. By not capturing
order, DDs miss critical contextual information like trends
or patterns across consecutive tuples.
        </p>
        <p>
          TSDDs [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] designed for time-series data capture temporal
relationships by treating data within a given time window
as an ordered set and supporting real-valued function
operations. However, similar to SDs, they do not account for
changes in the antecedent attributes over time. Additionally,
selecting an optimal time window remains a challenge, as
an overly narrow window may overlook significant trends,
while a broader one risks diluting the relevance of
dependencies.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Association Rules</title>
        <p>
          Association rules identify co-occurrences of items within
a dataset, typically expressed in the form of {, } → ,
stating that if items  and  appear together, then  is
likely to appear as well [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Unlike data dependencies,
which enforce constraints that all instances must satisfy,
association rules identify probabilistic relationships without
guaranteeing consistency. Dependencies ensure structural
integrity, while association rules uncover patterns that may
not hold universally.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Statistical and Machine Learning</title>
      </sec>
      <sec id="sec-2-5">
        <title>Approaches</title>
        <p>Statistical and machine learning approaches leverage
patterns in historical data to identify deviations that fall
outside expected behavior. Statistical methods rely on
predeifned thresholds and assumptions about data distribution,
while machine learning approaches adapt to complex,
highdimensional datasets.</p>
        <p>
          Statistical and machine learning approaches ofer
complementary techniques for identifying and diferentiating
normal and abnormal changes in data. Statistical techniques
include rule-based thresholds and hypothesis testing. For
example, Z-scores and modified Z-scores are commonly
used to detect anomalies by measuring how far a data point
deviates from the mean, relative to the standard deviation
[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. For example, if a data point’s z-score exceeds a certain
threshold (e.g., 3), it may be flagged as abnormal. Similarly,
control charts and statistical process control (SPC) methods
monitor data streams over time, flagging points that fall
outside control limits as potential anomalies [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>
          Machine learning provides various techniques for
distinguishing normal from abnormal changes in data,
particularly through anomaly detection algorithms. Isolation
Forest [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] isolates anomalies by partitioning the dataset into
smaller subsets. Points that require fewer partitions to be
isolated are identified as anomalies. This method works well
in high-dimensional data but may struggle with datasets
containing overlapping clusters or anomalies that are close
to the decision boundary.
        </p>
        <p>While numerous anomaly detection methods exist, our
approach specifically targets anomalies in the change of
attribute values. We achieve this by defining a change rule that
not only identifies abnormal behavior but also captures the
relationships between changes across multiple attributes.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Preliminaries</title>
      <p>Let  be a relational schema on attributes 1, 2, ...,  ,
and  and  be sets of attributes such that  ⊆ 
and  ⊆ . Let  = {1, 2, ...,  } be a relational
instance of  with  tuples, ordered on X (implicitly
ordered on time). The distance between consecutive
tuples in  for an attribute  is given via a context-aware
distance measure: ([], +1[]). We define a
permissible range for  as  = (, ), where ,  are
real values, i.e., if ([], +1[]) ∈ , then  ≤
([], +1[]) ≤ .</p>
      <p>We define a support function (,  )) that
measures the relative strength of a change rule  in . Naturally,
we seek high-support rules to ensure that they have
suficient evidence in the instance. We introduce change rules
in the next section, and focus on their discovery (as part of
Problem 2).</p>
      <p>Problem Definition: Given a minimum support threshold
 , find all change rules Σ such that  satisfies Σ (  |= Σ) ,
such that for all  ∈ Σ , (,  ) ≥  .
A change rule is a novel data quality rule which describes
a relationship between the changes in attributes within .
It states that when the change in the antecedent is within
some range  = (, ), then the corresponding change
in the consequent will also be within a defined range  =
(, ).</p>
      <p>DEFINITION 1. Let  be the permutation of tuples of
 increasing on  (that is,  (1)[] &lt;  (2)[] &lt;
. . . &lt;  ()[]). Change rule  :  →
 holds over  if for all  such that 1 ≤  ≤
 − 1, when ( ()[],  (+1)[]) ∈  then
( ()[ ],  (+1)[ ]) ∈ .</p>
      <p>When ordered on X, if the  between any two
consecutive -values is within the range  then the  between
the corresponding  -values must be within . A change
rule with a minimum support threshold  holds when at
least  % pairs of consecutive tuples in the instance satisfy
the conditions of the change rule.</p>
      <p>Example 3. Consider the change rule over Table 1:
 :  (5,15) → (0.1,0.25)</p>
      <p>This rule states that if the change in Position is between
5 and 15, then the change in Salary will be between 10%
to 25%. This holds true for most of the table except when
E2 is promoted from Senior Software Developer to Lead
Developer. In this case, the salary increase is only 8.7%,
which is below the expected 10% to 25% increase. This
deviation from the rule highlights that the employee received a
smaller-than-normal raise with their promotion.</p>
      <sec id="sec-3-1">
        <title>4.1. Discovery of Change Rules</title>
        <p>
          We build upon the Diferential Dependency discovery
algorithm, FastDD [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] over unordered data.
• Dif-Set Construction: Encodes pairwise diferences
between all tuples into a dif-set, where each element
represents a diferential constraint violation (e.g., [] −
 [] &gt; ), where  and  are any two tuples in a
relational instance and  is a numerical value. For change
rules, we modify this step by using a sorted instance  on
the antecedent attributes  to compute the  between
consecutive pairs of tuples. This eliminates the
redundant comparisons by restricting dif-set construction to
adjacent tuple pairs in the sorted instance .
• Set Cover Enumeration: Finds minimal subsets of
differential functions (antecedent) that cover all violations
of the consequent. For change rules, instead of fixed
thresholds, use intervals  and  for antecedent and
consequent gaps. That is, we find the minimal subsets of
( ()[],  (+1)[]) ∈  that cover all violations
of ( ()[ ],  (+1)[ ]) ∈ .
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusion and Next Steps</title>
      <p>Data changes over time, however, we want to capture
relationships between these changes. In this paper, we discussed
the importance of context-awareness when capturing these
changes and the relevance of identifying normal change
behaviour. We introduced a new data rule, called change
rule, that captures the relationship between the changes in
antecedent and the changes in the consequent.</p>
      <p>As next steps, we plan to address the aforementioned
problems and challenges:
• Exploring transformer-based embeddings (e.g.,
BERT), to quantify and accurately capture
contextaware changes in numeric and non-numeric data
without compromising semantic information.
• Optimize Set Cover Enumeration by developing an
eficient method to minimize the search space when
identifying minimal subsets of antecedent changes
that explain consequent violations.
• Consider the lagged efects of earlier changes on
subsequent changes, i.e., the change in an attribute
at one time step influences changes at a later time
step.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Heckert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Filliben</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Croarkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hembree</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. F.</given-names>
            <surname>Guthrie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tobias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Prinz</surname>
          </string-name>
          , Handbook 151:
          <article-title>Nist/sematech e-handbook of statistical methods (</article-title>
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G.</given-names>
            <surname>Navarro</surname>
          </string-name>
          ,
          <article-title>A guided tour to approximate string matching, ACM computing surveys (CSUR) 33 (</article-title>
          <year>2001</year>
          )
          <fpage>31</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Cardinal</surname>
          </string-name>
          ,
          <article-title>Similarity measures and graph adjacency with sets (</article-title>
          <year>2022</year>
          ).
          <article-title>URL: towardsdatascience.com/ similarity-measures-and-graph-adjacency-with-</article-title>
          <string-name>
            <surname>sets</surname>
          </string-name>
          , [Online; posted 28-Oct-2022].
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Szlichta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Godfrey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Golab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kargar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <article-title>Efective and complete discovery of order dependencies via set-based axiomatization</article-title>
          ,
          <source>arXiv preprint arXiv:1608.06169</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Golab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Karlof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Korn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Saha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <article-title>Sequential dependencies</article-title>
          ,
          <source>Proceedings of the VLDB Endowment</source>
          <volume>2</volume>
          (
          <year>2009</year>
          )
          <fpage>574</fpage>
          -
          <lpage>585</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Song</surname>
          </string-name>
          , L. Chen,
          <article-title>Diferential dependencies: Reasoning and discovery</article-title>
          ,
          <source>ACM Transactions on Database Systems (TODS) 36</source>
          (
          <year>2011</year>
          )
          <fpage>1</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tan</surname>
          </string-name>
          , S. Ma,
          <article-title>Eficient diferential dependency discovery</article-title>
          ,
          <source>Proceedings of the VLDB Endowment</source>
          <volume>17</volume>
          (
          <year>2024</year>
          )
          <fpage>1552</fpage>
          -
          <lpage>1564</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>W. H.</given-names>
            <surname>Gomaa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Fahmy</surname>
          </string-name>
          ,
          <article-title>A survey of text similarity approaches</article-title>
          ,
          <source>international journal of Computer Applications</source>
          <volume>68</volume>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>X.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          , Tsddiscover:
          <article-title>Discovering data dependency for time series data</article-title>
          ,
          <source>in: 2024 IEEE 40th International Conference on Data Engineering (ICDE)</source>
          , IEEE,
          <year>2024</year>
          , pp.
          <fpage>3668</fpage>
          -
          <lpage>3681</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Srikant</surname>
          </string-name>
          , et al.,
          <article-title>Fast algorithms for mining association rules</article-title>
          ,
          <source>in: Proc. 20th int. conf. very large data bases, VLDB</source>
          , volume
          <volume>1215</volume>
          ,
          <string-name>
            <surname>Santiago</surname>
          </string-name>
          ,
          <year>1994</year>
          , pp.
          <fpage>487</fpage>
          -
          <lpage>499</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <article-title>Statistical process control charts as a tool for analyzing big data, Big and Complex Data Analysis: Methodologies and Applications (</article-title>
          <year>2017</year>
          )
          <fpage>123</fpage>
          -
          <lpage>138</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>F. T.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. M.</given-names>
            <surname>Ting</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.-H.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <article-title>Isolation forest</article-title>
          , in: 2008 eighth ieee international
          <source>conference on data mining, IEEE</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>413</fpage>
          -
          <lpage>422</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>