<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overlap-Based Duplicate Table Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>(Discussion Paper)</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Zecchini</string-name>
          <email>luca.zecchini@unimore.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tobias Bleifuß</string-name>
          <email>tobias.bleifuss@hpi.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Simonini</string-name>
          <email>giovanni.simonini@unimore.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sonia Bergamaschi</string-name>
          <email>sonia.bergamaschi@unimore.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felix Naumann</string-name>
          <email>felix.naumann@hpi.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hasso Plattner Institute, University of Potsdam</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Modena and Reggio Emilia</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Both the Web and data lakes contain much redundant data in the form of largely overlapping pairs of tables. In many cases, this overlap is not accidental and provides meaningful information about the relatedness of the tables. In particular, we focus on the largest overlap between two tables, i.e., their largest common subtable. The largest overlap can help us discover multiple coexisting versions of the same table, which possibly difer in the completeness and correctness of the conveyed information. Automatically detecting these highly similar, duplicate tables would allow us to guarantee their consistency through data cleaning or change propagation, but also to eliminate redundancy to free up storage space or to save additional work for the editors. Unfortunately, detecting the largest overlap is a computationally challenging problem, requiring to carefully permute columns and rows. We introduce therefore Sloth, our solution to eficiently detect the largest overlap between two tables. As we experimentally demonstrate on real-world datasets, Sloth is not only efective in solving this task, but can impact on multiple additional use cases, such as detecting potential copying across sources or automatically discovering candidate multi-column joins.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Table Overlap</kwd>
        <kwd>Table Matching</kwd>
        <kwd>Related Tables</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Overlapping Tables</title>
      <p>
        The Web contains a huge amount of structured data in tabular form. In 2008, it was already
possible to retrieve more than 14 billion tables [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], and nowadays more than 2 million tables
can coexist in the English version of Wikipedia alone [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Web tables often have a very dynamic existence. A particularly representative case is that
of Wikipedia [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], where tables are frequently edited or updated, moved within their page or
to another page, copied to related pages or elsewhere, with frequent episodes of carelessness,
conflicts among editors [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and even vandalism [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Because of this dynamism and the
heterogeneity of the community of Wikipedia editors, it can be very dificult to guarantee data
quality, which is fundamental for an encyclopedia, whose content should always be correct,
complete, and updated.
(a) A pair of tables about football teams and stadiums.
      </p>
      <p>(b) The largest overlap between the two tables.</p>
      <p>
        Among the 2.13 million tables existing in Wikipedia at the time of our latest snapshot, we
surprisingly discovered that about 6.5 million pairs of tables present an overlap equal to at
least half of the area of the smaller table, for an estimated redundancy of 63.49 MB. Even
more surprisingly, we detected 5.9 million pairs of coexisting tables with identical content,
highlighting the massive difusion of copy-and-paste practices in Wikipedia [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        In particular, we focus on the largest overlap between the two tables [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], i.e., their largest
common rectangular subtable, as depicted in Figure 1. In many cases, this overlap is not
accidental, but gives meaningful insights about the relatedness of the tables and the quality
of their content. The nature of tabular data allows changing the order of columns and rows
(Figure 1b), making the detection of the largest overlap computationally challenging.
      </p>
      <p>
        The ability to detect the largest overlap between two tables, and in particular to retrieve pairs
of highly similar tables, defined as duplicate or matching tables, can lead to several benefits,
such as verifying the consistency of the information conveyed by the tables, pointing out cases
of incompleteness or inconsistencies. This is not only relevant for Web tables: every scenario
where a table can be duplicated at a certain point in time, with an independent development
for the diferent copies, is prone to the insurgence of inconsistencies. For instance, when data
scientists retrieve datasets from the enterprise’s data lake, perform transformations (e.g., join,
wrangling, etc.) for their analysis, then store back the new datasets into the data lake [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Depending on the context, a user might desire to ensure the consistency of the information
present in duplicate tables through operations of data cleaning [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and change propagation [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ],
or it might be more convenient to directly prevent the rise of inconsistencies by eliminating
this redundancy. In fact, avoiding redundancy not only allows to save disk space, but also to
lighten the workload for website editors, who would just have to focus on a single consistent
table instead of performing every editing multiple times, exposing to the risk of inconsistencies
or missed updates. Whatever the purpose, one must first detect such duplicate tables.
      </p>
      <p>
        While the existing literature widely recognizes the importance of discovering related tables [
        <xref ref-type="bibr" rid="ref10 ref11 ref9">9,
10, 11</xref>
        ] on the Web or in data lakes for enriching the conveyed information, proposing many
approaches for the eficient detection of unionable tables [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12, 13, 14</xref>
        ] or joinable tables [
        <xref ref-type="bibr" rid="ref15">15, 16,
17, 18, 19</xref>
        ], the task of detecting duplicate tables is only investigated in some specific or limited
scenarios (e.g., to detect the subsequent versions of a table throughout the history of a Wikipedia
page [20] or restricted to the basic cases of perfect duplicates and row containment [21]).
      </p>
      <p>
        To fill this gap, we recently proposed Sloth [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], a novel solution to determine the largest
overlap between a given pair of tables, i.e., the maximal contiguous rectangular area of identical
cells that can be achieved by reordering columns and rows of both tables.
      </p>
      <p>
        More formally, given a bijective attribute mapping  :  ⊆  →  ⊆  defined
between two tables () and ( ), we refer to the table overlap  = [ ] ∩+ [ ] as
the intersection under the bag semantics (i.e., which allows duplicates) between the bags of
tuples obtained through the projection of () on  and ( ) on  . Defined  as the
set of all possible overlaps between the two tables and the overlap area  = | | · |  | as
the number of cells contained in the overlap  , the set of the largest overlaps * = {* ∈
 | * ≥  , ∀ ∈ } is composed of the overlaps with the maximum area, in most
cases just one. All details of our formalization can be found in the full research paper [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>At the core of Sloth lies the first algorithm designed to detect the largest overlap between
two tables. First, our algorithm detects the pairs of attributes across the two tables that share
some cell values. By combining these pairs of attributes, it is possible to obtain a complete
overview of all potentially non-empty overlaps existing between the tables (i.e., the candidates
to be the largest one) in the form of a lattice [22]. The combined pairs determine an upper bound
for the area of the candidates. Thus, our algorithm aims to exploit this bounding mechanism to
prioritize candidates and detect the largest overlap as soon as possible, minimizing the number
of candidates for which we need to compute the actual area.</p>
      <p>
        Since this task is computationally challenging, Sloth also relies on a greedy variant for the
algorithm based on beam search [23, 24] to deal with critical pairs for which the exact technique
struggles to produce a result in a reasonable time for the user. Section 2 provides an overview
of both algorithms, which are described in detail in the full research paper [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Beyond the eficiency of Sloth and the quality of the results produced by the greedy algorithm,
through our experimental evaluation we were able to highlight multiple relevant real-world
use cases, such as the detection of highly overlapping tables in Wikipedia, the recognition of
potential copying across tables from diferent sources [ 25], and the automated discovery of
candidate multi-column joins in a corpus of relational tables. An excerpt from the full research
paper [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is reported in Section 3, while Section 4 briefly compares to the existing literature.
Finally, Section 5 presents the future directions of our research and concludes the paper.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Largest Overlap Detection</title>
      <sec id="sec-2-1">
        <title>1https://github.com/dbmodena/sloth</title>
        <p>defined to spot critical cases that would not provide a result in a reasonable time, activating its
greedy variant (Figure 2c). Both the timeout and the parameter for the greedy algorithm (i.e.,
the beam width  ) can be edited by the users according to their needs (e.g., a faster computation
or a better accuracy). Additional parameters can define the minimum area Δ to consider the
largest overlap as relevant and even minimum/maximum width/height thresholds.</p>
        <p>
          Our exact algorithm (Algorithm 1) needs to consider all mappings that can potentially
determine the largest overlap, denoted as candidates. To identify the candidates, our algorithm
ifrst considers all possible single-attribute mappings, i.e., those mappings for which  is
represented by a single attribute , checking the area of the overlap between [] and [ ()].
We call seeds those single-attribute mappings whose area is greater than zero, and we collect
them in a dedicated list (Line 2) sorted by descending area, as depicted in Figure 3b. All details
about the invoked functions can be found in the full research paper [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
        <p>Since every multi-attribute mapping is a combination of some single-attribute mappings, the
candidates can be considered as the combinations of the seeds and modeled as the nodes of a
lattice, as depicted in Figure 3b, where every level  contains the combinations of  seeds. Due
to the bijectivity of the mapping, some seeds cannot be combined (e.g., 0 and 4 in Figure 3b,
both considering the same attribute City from ), hence possibly producing a semilattice.</p>
        <p>Moving up within the lattice increases the width of the overlap, but not necessarily its area,
as its height may decrease as new columns are added. In particular, the seed with the minimum
area (equal to its height) in the combination defines an upper bound for the height of the
candidate, and therefore for its area. This bounding mechanism can be exploited both to prune
the lattice and to prioritize the candidates based on their potential area. We define therefore the
pruning threshold  (Line 3), which always contains the maximum between Δ and the maximum
actual area of a candidate that we know so far (initially the area of the first seed in the list, then
possibly updated every time we verify a new candidate).</p>
        <p>Our algorithm exploits the upper bound defined by the seeds to manage two priority queues
(i.e., max heap structures), aiming to minimize both the number of candidates that need to be
materialized and those among them whose actual area needs to be computed: (i) Levels (Line 4),
containing the representations of the levels of the lattice, used to generate the candidates
incrementally by decreasing potential area; (ii) Candidates (Line 5), containing the generated
candidates, used to progressively verify their actual area and detect the largest overlap.</p>
        <p>In particular, we iterate on the priority queues until both of them are emptied (Line 6),
terminating early as soon as all largest overlaps are detected. At each iteration, first we need to
ensure that at least one of the candidates with the potential largest overlap has been generated
and inserted into Candidates (Lines 7-8), then we can check the candidate at the top of Candidates
(Line 10). If it has already been verified (i.e., its actual overlap has already been computed), none
of the other candidates (among both the ones already generated and the ones yet to generate)
can present a greater area, hence it is one of the largest overlaps, and we can add it to the result
set (Line 12); otherwise, we need to compute its overlap (and therefore its actual area) and
reinsert it into the priority queue if it can still be part of the result set (Line 14).</p>
        <p>Since our exact algorithm generates the candidates by combining the seeds, the detection of
a very large number of seeds (e.g., for a pair of wide tables with some values repeated across
several columns in both) may produce a huge lattice, making it sometimes impossible to generate
the candidates in a reasonable amount of time. For a result as close as possible to the exact
largest overlap, we designed therefore a greedy variant inspired by beam search, a heuristic
search algorithm that performs a breadth-first search in a tree by only expanding the  most
promising nodes at each level, where the parameter  is denoted as beam width.</p>
        <p>Our greedy algorithm is designed to bottom-up traverse the lattice generated from the seeds.
At the beginning, we consider a maximum of  seeds with the greatest area. To find candidates
for the second level, we combine each of them with every other seed and drop repeated and
invalid combinations. After this generating step, we verify all new candidates and again select
only the  candidates with the greatest actual area among them. For the third and every further
level, we combine the selected candidates of the previous level with every seed that is not
already part of the candidate and then again verify their area and limit their number to  , while
the pruning threshold  is updated and can determine an early stopping.</p>
        <p>(a) Two input tables () and ( ) about football teams.
(b) The sorted list of detected seeds and the valid candidates in the semilattice generated from the seeds.</p>
        <p>For every node we report the upper bounds for its area and its height (between round brackets).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental Evaluation</title>
      <p>
        Our experiments, whose configurations and results are reported and discussed in detail in the
full research paper [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], aim to evaluate the performance of Sloth (considering both the exact
algorithm and its greedy variant) and its application to three real-world use cases.
      </p>
      <p>In Table 1, we present the datasets used in our evaluation. We denote as wiki_history the
Wikipedia table matching dataset from the IANVS project2, which captures the evolution of
all 3.5M tables present in the English Wikipedia throughout its entire history (until
September 1, 2019). We separately consider the most recent snapshot from this dataset, denoted as
wiki_latest. Beyond the Wikipedia scenario, we employ uni_dwh [16], a real-world university
data warehouse, and two datasets3 reporting the information about stock symbols and flights
captured from diferent sources across multiple days [ 25], both in their original version (i.e.,
raw) and in the one obtained through schema alignment (i.e., clean).</p>
      <p>
        First, we evaluate the performance of Sloth on a representative subset of the wiki_history
dataset and the uni_dwh dataset, chosen to cover both the case of Web tables and the one of
relational database tables in our analysis. Our experiments [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] confirm that the main factor
      </p>
      <sec id="sec-3-1">
        <title>2https://hpi.de/naumann/projects/data-profiling-and-analytics/change-exploration.html 3https://lunadong.com/fusionDataSets.htm</title>
        <p>leading to timeouts is the number of seeds, which directly afects the size of the lattice, hence the
number of candidates that potentially need to be generated. This aspect is strictly correlated to
the width of the tables and the amount of repeated cell values across them. Further, we analyze
the impact on the runtime of the diferent tasks (i.e., seed detection, candidate generation, and
candidate verification) for both the exact and the greedy algorithm, and also of the beam width
for the latter. We also show the impact of the table height on the seed detection runtime and
how our similarity estimation significantly difers from two widely adopted similarity metrics
based on set semantics, such as Jaccard similarity and overlap set similarity [17, 26]. Finally, we
evaluate the accuracy of the greedy algorithm, showing that, even in absence of approximation
guarantees on the quality of its result, it is generally able to detect largest overlaps of the same
area as those discovered by the exact algorithm.</p>
        <p>
          Moving to the real-world use cases, first we use Sloth to quantify the amount of largely
overlapping pairs of tables coexisting in Wikipedia, using wiki_latest. The results, reported in
Table 2, highlight the redundancy of the information conveyed by Wikipedia tables and the
difusion of copy-and-paste practices in the encyclopedia. Then, we show that Sloth can be
used to detect potential copying across multiple sources, allowing us to retrieve all clusters of
sources with declared copying dependencies in the stock and flight datasets [ 25] with a minimum
efort, without the need for schema alignment, also leading to the discovery of meaningful
additional sources. Finally, we use Sloth to automatically detect candidate multi-column joins
in uni_dwh, a task not supported by any of the existing solutions [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], as highlighted by showing
the limitations of a simple adaptation of Josie [17] in such a scenario.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Related Work</title>
      <p>
        Even if a plethora of algorithms have been designed to eficiently discover related tables [
        <xref ref-type="bibr" rid="ref10 ref11 ref9">9, 10,
11</xref>
        ], especially unionable [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12, 13, 14</xref>
        ] and joinable ones [
        <xref ref-type="bibr" rid="ref15">15, 16, 17, 18, 19</xref>
        ], none of them can
exactly compute the largest overlap between two tables.
      </p>
      <p>For instance, Josie [17] considers a join column from a query table and finds the top- k columns
whose set of cell values presents the largest intersection with the one of the join column. Since
it operates on single columns using the set semantics, it is impossible to use it for detecting
the largest overlap. At most, an adaptation considering entire tables under the bag semantics
would produce an upper bound for the largest overlap, so it might be used to enhance scalability
by passing only the most promising pairs to Sloth. Similarly, Mate [19] is the only system
supporting the discovery of multi-column joins (using a dedicated hashing function named
XASH), but it requires to provide a set of columns as input. Since the set that yields the largest
overlap is not known up-front, it cannot be easily employed as an alternative to Sloth.</p>
      <p>Moving to the previous approaches to duplicate table detection, Bleifuß et al. [20] aim to
ifnd matching tables across subsequent versions of a Wikipedia page. Their solution (i.e., a
multi-stage matching process based on Jaccard similarity) is designed for one-to-one matches
among a limited number of tables and exploits specific aspects such as the position of the tables
inside the page, hence it is not generalizable. Koch et al. [21] use instead XASH and define two
tables as duplicates if they contain the same set of tuples, only tackling the cases of perfect
duplicates or row containment.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and Future Work</title>
      <p>We presented Sloth, a method to eficiently determine the largest overlap between two tables,
allowing to detect duplicate tables. This leads to several benefits both on the Web and in
data lakes. For instance, it allows spotting and solving common data quality issues, such as
inconsistent or incomplete information. Also, it helps eliminate redundancy to free up storage
space or to save additional work for the editors, preventing the insurgence of data quality
problems. Through our experimental evaluation we assessed the performance of Sloth in
real-world scenarios, considering Web tables from Wikipedia and relational tables from a data
warehouse, up to use cases such as the detection of potential copying across multiple sources
and the discovery of candidate multi-column joins in a corpus of relational tables.</p>
      <p>Moving beyond the results presented in this paper, we plan to broaden our research in multiple
directions. First, we want to design updatable indexes to enable overlap-based duplicate table
detection at scale, allowing users to provide a table as a query and retrieve the top-k tables
presenting the largest overlap with it. Then, we want to enable Sloth to detect not only the
largest, but also the best overlap between two tables, defining quality metrics to capture how
meaningful an overlap is. Finally, we plan to evaluate the use of table embeddings [27] to
estimate the largest overlap and to analyze the impact of table deduplication on the performance
of tabular language models [28], similarly to what has already been demonstrated for their
textual counterparts [29].
14778/2994509.2994534.
[16] R. Castro Fernandez, Z. Abedjan, F. Koko, G. Yuan, S. Madden, M. Stonebraker, Aurum: A
Data Discovery System, in: Proceedings of the IEEE International Conference on Data
Engineering (ICDE), 2018, pp. 1001–1012. doi:10.1109/ICDE.2018.00094.
[17] E. Zhu, D. Deng, F. Nargesian, R. J. Miller, JOSIE: Overlap Set Similarity Search for Finding
Joinable Tables in Data Lakes, in: Proceedings of the ACM International Conference on
Management of Data (SIGMOD), 2019, pp. 847–864. doi:10.1145/3299869.3300065.
[18] Y. Dong, K. Takeoka, C. Xiao, M. Oyamada, Eficient Joinable Table Discovery in Data
Lakes: A High-Dimensional Similarity-Based Approach, in: Proceedings of the IEEE
International Conference on Data Engineering (ICDE), 2021, pp. 456–467. doi:10.1109/
ICDE51399.2021.00046.
[19] M. Esmailoghli, J. Quiané-Ruiz, Z. Abedjan, MATE: Multi-Attribute Table Extraction,
Proceedings of the VLDB Endowment (PVLDB) 15 (2022) 1684–1696. doi:10.14778/
3529337.3529353.
[20] T. Bleifuß, L. Bornemann, D. V. Kalashnikov, F. Naumann, D. Srivastava, Structured
Object Matching across Web Page Revisions, in: Proceedings of the IEEE International
Conference on Data Engineering (ICDE), 2021, pp. 1284–1295. doi:10.1109/ICDE51399.
2021.00115.
[21] M. Koch, M. Esmailoghli, S. Auer, Z. Abedjan, Duplicate Table Detection with Xash, in:
Proceedings of the Conference on Database Systems for Business, Technology and Web
(BTW), 2023, pp. 367–390. doi:10.18420/BTW2023-18.
[22] G. Birkhof, Lattice Theory, 1940. doi: 10.1090/coll/025.
[23] B. T. Lowerre, The HARPY Speech Recognition System, Ph.D. thesis, Carnegie Mellon</p>
      <p>University, 1976.
[24] X. Huang, J. Baker, R. Reddy, A Historical Perspective of Speech Recognition,
Communications of the ACM (CACM) 57 (2014) 94–103. doi:10.1145/2500887.
[25] X. Li, X. L. Dong, K. Lyons, W. Meng, D. Srivastava, Truth Finding on the Deep Web: Is
the Problem Solved?, Proceedings of the VLDB Endowment (PVLDB) 6 (2012) 97–108.
doi:10.14778/2535568.2448943.
[26] D. Deng, Y. Tao, G. Li, Overlap Set Similarity Joins with Theoretical Guarantees, in:
Proceedings of the ACM International Conference on Management of Data (SIGMOD),
2018, pp. 905–920. doi:10.1145/3183713.3183748.
[27] R. Cappuzzo, P. Papotti, S. Thirumuruganathan, Creating Embeddings of
Heterogeneous Relational Datasets for Data Integration Tasks, in: Proceedings of the ACM
International Conference on Management of Data (SIGMOD), 2020, pp. 1335–1349.
doi:10.1145/3318464.3389742.
[28] G. Badaro, M. Saeed, P. Papotti, Transformers for Tabular Data Representation: A Survey
of Models and Applications, Transactions of the Association for Computational Linguistics
(TACL) 11 (2023) 227–249. doi:10.1162/tacl_a_00544.
[29] K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, N. Carlini,
Deduplicating Training Data Makes Language Models Better, in: Proceedings of the Annual
Meeting of the Association for Computational Linguistics (ACL), 2022, pp. 8424–8445.
doi:10.18653/v1/2022.acl-long.577.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Cafarella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Halevy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          , E. Wu,
          <string-name>
            <surname>Y. Zhang,</surname>
          </string-name>
          <article-title>WebTables: Exploring the Power of Tables on the Web, Proceedings of the VLDB Endowment (PVLDB) 1 (</article-title>
          <year>2008</year>
          )
          <fpage>538</fpage>
          -
          <lpage>549</lpage>
          . doi:
          <volume>10</volume>
          .14778/1453856.1453916.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bleifuß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bornemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          , The Secret Life of Wikipedia Tables,
          <source>in: Proceedings of the Workshop on Search, Exploration, and Analysis in Heterogeneous Datastores (SEA Data @ VLDB)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>20</fpage>
          -
          <lpage>26</lpage>
          . URL: https: //ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2929</volume>
          /paper4.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bykau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Korn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Velegrakis</surname>
          </string-name>
          ,
          <string-name>
            <surname>Fine-Grained Controversy</surname>
          </string-name>
          Detection in Wikipedia,
          <source>in: Proceedings of the IEEE International Conference on Data Engineering (ICDE)</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>1573</fpage>
          -
          <lpage>1584</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICDE.
          <year>2015</year>
          .
          <volume>7113426</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gerling</surname>
          </string-name>
          , Automatic Vandalism Detection in Wikipedia,
          <source>in: Proceedings of the European Conference on Information Retrieval (ECIR)</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>663</fpage>
          -
          <lpage>668</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>540</fpage>
          -78646-7_
          <fpage>75</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zecchini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bleifuß</surname>
          </string-name>
          , G. Simonini,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <article-title>Determining the Largest Overlap between Tables, Proceedings of the ACM on Management of Data (PACMMOD) 2 (</article-title>
          <year>2024</year>
          )
          <volume>48</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>48</lpage>
          :
          <fpage>26</fpage>
          . doi:
          <volume>10</volume>
          .1145/3639303.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>F.</given-names>
            <surname>Nargesian</surname>
          </string-name>
          , E. Zhu,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Pu</surname>
          </string-name>
          , P. C.
          <article-title>Arocena, Data Lake Management: Challenges and Opportunities</article-title>
          ,
          <source>Proceedings of the VLDB Endowment (PVLDB) 12</source>
          (
          <year>2019</year>
          )
          <fpage>1986</fpage>
          -
          <lpage>1989</lpage>
          . doi:
          <volume>10</volume>
          .14778/3352063.3352116.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>I. F.</given-names>
            <surname>Ilyas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chu</surname>
          </string-name>
          , Data Cleaning,
          <year>2019</year>
          . doi:
          <volume>10</volume>
          .1145/3310205.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bleifuß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bornemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          , Exploring Change:
          <article-title>A New Dimension of Data Analytics, Proceedings of the VLDB Endowment (PVLDB) 12 (</article-title>
          <year>2018</year>
          )
          <fpage>85</fpage>
          -
          <lpage>98</lpage>
          . doi:
          <volume>10</volume>
          .14778/3282495.3282496.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>A. Das Sarma</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Fang</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Halevy</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Xin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
          </string-name>
          , Finding Related Tables,
          <source>in: Proceedings of the ACM International Conference on Management of Data (SIGMOD)</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>817</fpage>
          -
          <lpage>828</lpage>
          . doi:
          <volume>10</volume>
          .1145/2213836.2213962.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bogatu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A. A.</given-names>
            <surname>Fernandes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. W.</given-names>
            <surname>Paton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Konstantinou</surname>
          </string-name>
          ,
          <article-title>Dataset Discovery in Data Lakes</article-title>
          ,
          <source>in: Proceedings of the IEEE International Conference on Data Engineering (ICDE)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>709</fpage>
          -
          <lpage>720</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICDE48307.
          <year>2020</year>
          .
          <volume>00067</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z. G.</given-names>
            <surname>Ives</surname>
          </string-name>
          ,
          <article-title>Finding Related Tables in Data Lakes for Interactive Data Science</article-title>
          ,
          <source>in: Proceedings of the ACM International Conference on Management of Data (SIGMOD)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1951</fpage>
          -
          <lpage>1966</lpage>
          . doi:
          <volume>10</volume>
          .1145/3318464.3389726.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>M. J. Cafarella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Halevy</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Khoussainova</surname>
          </string-name>
          ,
          <article-title>Data Integration for the Relational Web, Proceedings of the VLDB Endowment (PVLDB) 2 (</article-title>
          <year>2009</year>
          )
          <fpage>1090</fpage>
          -
          <lpage>1101</lpage>
          . doi:
          <volume>10</volume>
          .14778/1687627. 1687750.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>O.</given-names>
            <surname>Lehmberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <article-title>Stitching Web Tables for Improving Matching Quality</article-title>
          ,
          <source>Proceedings of the VLDB Endowment (PVLDB) 10</source>
          (
          <year>2017</year>
          )
          <fpage>1502</fpage>
          -
          <lpage>1513</lpage>
          . doi:
          <volume>10</volume>
          .14778/3137628. 3137657.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>F.</given-names>
            <surname>Nargesian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <source>Table Union Search on Open Data, Proceedings of the VLDB Endowment (PVLDB) 11</source>
          (
          <year>2018</year>
          )
          <fpage>813</fpage>
          -
          <lpage>825</lpage>
          . doi:
          <volume>10</volume>
          .14778/3192965.3192973.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nargesian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          , LSH Ensemble:
          <article-title>Internet-Scale Domain Search, Proceedings of the VLDB Endowment (PVLDB) 9 (</article-title>
          <year>2016</year>
          )
          <fpage>1185</fpage>
          -
          <lpage>1196</lpage>
          . doi:10.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>