<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Every Data Lake Has a Past: Analytical Exploration of Wikipedia History as a Temporal Data Lake</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mahdi Esmailoghli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Steven Purtzel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roee Shraga</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Renée J. Miller</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthias Weidlich</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Humboldt-Universität zu Berlin</institution>
          ,
          <addr-line>Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Waterloo</institution>
          ,
          <addr-line>Waterloo</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Worcester Polytechnic Institute (WPI)</institution>
          ,
          <addr-line>Worcester</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <abstract>
        <p>The rise of data lakes has created complex and temporally rich repositories in which tabular data exists in multiple versions. Current data discovery methods fail to utilize this crucial temporal dimension, treating each version individually. This limits the efectiveness of data discovery and integration for downstream tasks, e.g., machine learning model training. To address this, we conduct an analytical study focusing on the Wikipedia table history data lake to characterize its temporal dimensions. Our work provides essential statistics on table evolution and revision types. This foundational understanding of table evolution can help guide the development of future data lake and data management systems capable of leveraging temporal table properties. Our exploratory analyses reveal distinct patterns in how table versions transform over time, identifying specific update frequencies, editor behaviors, and rollback prevalence that characterize the lake's lifecycle.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Temporal Data Lakes</kwd>
        <kwd>Data Discovery</kwd>
        <kwd>Data Analysis</kwd>
        <kwd>Wikipedia History</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The proliferation of data lakes over the past decade has established them as a primary resource for data
scientists. Data lakes are widely capitalized to enrich existing datasets for various downstream tasks,
including training data augmentation for machine learning (ML) models [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Data search or discovery
within this context typically uses several established operators: join discovery [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ], where a query
table is horizontally extended by joining with tables from the lake; union discovery [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which involves
the vertical aggregation of rows from tables sharing similar schemata; correlation discovery [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ] aims
at finding informative features to improve a model’s predictive accuracy; and integrated discovery that
combines a variety of these operators [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ].
      </p>
      <p>
        The evolution of data lakes has inherently resulted in temporally-rich corpora of data, particularly
tabular data, where tables often exist in multiple versions [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. These table corpora are referred to as
temporal data lakes [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Examples of such data lakes include: tables encapsulated within Wikipedia
pages, where each change to the page generates a new version of its corresponding tables, GitHub
repositories, where updates to the codebase lead to new versions of accompanying data, and open data
portals, where organizations and governments frequently provide updated reports [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        However, current data discovery approaches are rendered impractical when applied to temporal
data lakes because they treat each table version as an isolated entity. However, the state of the art has
shown that utilizing the temporal dimension of data lakes can significantly improve the eficacy of data
integration solutions. As one example, Bornemann et al. [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ] showed that inclusion dependencies
that are consistently valid across diferent versions of a Wikipedia table tend to be more semantically
correct than dependencies observed in a single instance of a table.
      </p>
      <p>
        Wikipedia history is a publicly available source of versioned data. Tables within Wikipedia pages have
been widely used before to enrich data at hand [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], such as in the WikiTables [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] project. However,
the evolution of these tables is not well studied, which is a vital prerequisite for efectively leveraging
such rich corpora. For instance, the bottom layer in Figure 1 demonstrates an example of a simple
table evolution, which goes through schema modification, row addition, content edits, and structure
re-ordering. Without comprehending and being able to systematically track these evolutions, one
cannot leverage temporal data lakes to build a reliable data discovery system.
      </p>
      <p>
        In this paper, we take an exploratory approach to analyzing the temporal dimensions of Wikipedia
tables (Wiki Lake), validated and curated in previous research [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which represent a prominent and
accessible example of a large-scale temporal data lake. This corpus provides necessary labels, allowing
us to accurately track the lineage of individual table versions.
      </p>
      <p>We aim to move beyond the analysis of static data snapshots by characterizing the fundamental
mechanics of table evolution across the entire data lake. Specifically, we focus on three interrelated
temporal dimensions: the frequency and distribution of versions, the semantic nature of changes, e.g.,
schema modifications vs. content updates, and the lifecycle of updates, e.g., persistent changes vs. the
ones rolled back. By quantifying these dynamics, we seek to identify evolutionary patterns and editor
behaviors that distinguish stable data from volatile revisions.</p>
      <p>We acknowledge that this work represents only the initial step toward understanding the complex
concept of temporally-rich tabular data lakes. Private data lakes may have diferent temporal patterns
than the Wiki Lake. Nonetheless, our study provides a first step in characterizing the temporal evolution
of tables. Further research and exploration are essential to develop robust data management systems
tailored for modern and temporal data lakes.</p>
      <p>In the remainder of this paper, we first review related work and then provide a brief overview of
the Wiki Lake and how it organizes table versions. Subsequently, we define the temporal dimensions
we aim to understand and outline the exploratory actions taken to investigate them. Following each
dimension, we conduct an analysis, extracting statistics from the data to characterize the defined data
lake attributes. Finally, we discuss the lessons learned as well as future directions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>A large body of prior work has analyzed Wikipedia through its revision histories, treating article
evolution as temporal events, from which high-level semantic behaviors can be inferred. For instance, a
body of research has explored conflicts such as edit wars through analyzing revision metadata [ 16, 17].
Another line of research focuses on rapid rollbacks, whether they help discover vandalism [18, 19] or
recurring maintenance [20]. In contrast, we do not target behaviors at the article level, but study table
evolution at scale, treating tables as first-class temporal citizens in the data lake.</p>
      <p>
        While the domain of temporal data lakes remains significantly under-explored, recent work has
begun to address the analysis and utilization of these repositories, with a particular emphasis on the
evolving histories of the Wiki Lake. Notably, Bleifuß et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] conducted an analytical characterization
of the Wiki Lake. Although their work acknowledges the temporal aspect of the data lake, their primary
investigation targets a broader spectrum of properties, such as the lifecycle of tables, including creation
and deletion, and provenance regarding contributor activity. In contrast, this paper is dedicated to a
more fine-granular examination of a data lake’s temporal dimensions. We prioritize a detailed analysis
of schema evolution dynamics, the correlation between distinct modification types, and the semantic
significance of change intervals within this specific context. Moreover, Bleifuß et al. [ 21, 22] and
Esmailoghli and Weidlich [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] introduce specific systems and corresponding data structures, including
Change-cube and a hierarchical index. In this paper, we merely focus on an analytical approach to
explore the temporal aspect of the data lake to understand the data before building such data structures.
      </p>
      <p>
        In data management, the eficient storage of multiple table versions has received considerable attention
in the literature [23, 24, 25, 26, 27, 28]. However, characterizations of the evolution itself have only
recently been studied [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and this work is limited to two versions, not the broad characterization of
evolution in a data lake that we are undertaking.
Extracted Tables
      </p>
      <p>Table A1</p>
      <p>Table A2</p>
      <p>Table B2</p>
    </sec>
    <sec id="sec-3">
      <title>3. Wikipedia, a Temporal Data Lake</title>
      <p>
        We study the Wiki Lake, which is derived from Wikipedia revision history dumps that track all
modifications to Wikipedia pages. This corpus encapsulates the tables embedded within these pages
and stores the discrete changes applied to each.1 We utilize the corpus prepared by Bornemann et
al. [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ], in which columns across versions are aligned,2 enabling the analysis of schema evolution.
      </p>
      <p>The data lake comprises 512 JSON files with a collective size of 809GB, containing 2.8M tables. The
data is structured hierarchically around pages, including PageID, PageTitle, and their corresponding
tables. Each table entity contains a TableId and a chronological sequence of table versions. Note that
versions of a table are contained in the same page’s history. Each version includes its unique identifier, a
timestamp, contributor metadata, e.g., username or IP address, column identifiers, and the table content
represented as a list of row values. Figure 1 illustrates an abstract representation of this data model.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Dimensions and Analysis</title>
      <p>
        Traditional data lake research has focused primarily on structural and relational characteristics. These
include volumetric measures, such as the number of tables [
        <xref ref-type="bibr" rid="ref3">3, 29</xref>
        ], dimensionality of tables, including
the average number of rows and columns, and how tables are relevant to either each other [30] or to a
given input table [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Similarly, Wikipedia has been examined through various lenses, ranging from
the topological connectivity of its pages and entities [31] to the semantic relevance of WikiTables [32].
Extensive research has also leveraged the Wikipedia revision history as a corpus for supervised learning,
utilizing human-annotated updates to improve data integration tasks [
        <xref ref-type="bibr" rid="ref16 ref17">33, 34, 35</xref>
        ].
      </p>
      <p>
        Despite these advancements, the temporal dynamics of Wiki Lake remain largely underexplored.
While current literature acknowledges Wikipedia as a repository of interconnected tables, there is a lack
of comprehensive analysis regarding the evolution of these tables throughout their lifecycle [
        <xref ref-type="bibr" rid="ref15 ref18">36, 15</xref>
        ].
      </p>
      <p>To address this, we categorize and analyze the temporal dimensions of the Wiki Lake by focusing on
Lake Overview, Change Interval, and Rollback.</p>
      <sec id="sec-4-1">
        <title>4.1. Lake Overview</title>
        <p>First, we characterize the aggregate statistics of the Wiki Lake. It comprises 2.8M tables. Each table
contains 14 versions, averaging 5 columns and 11 rows per version. The distribution of version counts
conforms to a power-law distribution illustrated in Figure 2, characterized by a substantial number of
sparsely updated tables and a small subset of highly volatile ones. Notably, the tables exhibiting the
highest revision frequency are found within the pages List of social networking websites, America’s Next
Top Model, and List of the verified oldest people, recording 10k, 7k, and 6k changes, respectively.
1https://dumps.wikimedia.org/
2https://github.com/HPI-Information-Systems/tindResources</p>
        <p>
          To further investigate these changes, we track schema and row-level evolution across 37M consecutive
table versions. To maintain column provenance across revisions, we utilize generated unique column
identifiers [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. For row tracking, we employ a heuristic to identify a surrogate Primary Key (PK).
Specifically, we find the first attribute where the uniqueness ratio, i.e., the cardinality of unique values
relative to the total row count, exceeds 90%. We begin with the leftmost column, based on our
observation that the first column is typically the most identifying one. Figure 3 presents a Venn diagram
characterizing four distinct classes of structural and content modifications: column addition, column
deletion, row insertion, and row deletion. The distribution indicates that the most frequent transformation
between successive versions is the concurrent insertion and deletion of rows. This pattern can also be
interpreted as modifying existing row values, including updates to the identified PK.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Change Interval</title>
        <p>
          Change interval is defined as the time diference between two consecutive revisions. Analyzing this
provides insights into the content-agnostic behavioral dynamics of table evolution. Specifically, change
intervals can reveal the underlying intent of modifications, such as anomalous or fraudulent edits,
and unveil the inherent semantics of the data. For instance, tables representing national GDP are
not expected to be updated daily. We explore the distribution of change intervals across the lake to
establish expected temporal baselines within various contexts. For instance, in the context of edit wars,3
continuous edits are performed by diferent editors due to the inherent controversy of the Wikipedia
article and conflicting views. One expects shorter change intervals on such tables [
          <xref ref-type="bibr" rid="ref19">37</xref>
          ] as opposed to
more fact-driven content, e.g., historical records, which typically remain static for extended periods and
undergo updates only upon the release of oficial reports.
        </p>
        <p>Figure 4 illustrates the global distribution of change intervals, revealing a bimodal distribution: a
short-term peak concentrated around one minute and a long-term peak spanning weekly to monthly
intervals. This characterizes two distinct behavioral patterns within the Wiki Lake. The first represents
high-frequency revisions, often attributable to iterative editing by a single contributor. In practice,
substantial modifications to Wikipedia pages are non-atomic, requiring a sequence of successive
commits. Additionally, short-term intervals often result from the correction of erroneous or malicious
modifications by other users, moderators, or automated bots. These entities identify inappropriate edits
and execute rollbacks to restore the table to its prior state.</p>
        <p>Figure 5 depicts the distribution of change intervals aggregated at the table level. In this analysis, we
utilize the median change interval per table to mitigate the influence of temporal outliers. The figure
demonstrates that the high frequency of near-instantaneous changes diminishes when aggregated,
suggesting that while rapid edits are numerous in total volume, they do not constitute the primary
mode of long-term temporal evolution for the majority of tables.</p>
        <p>To further understand the characteristics of temporal intervals, we analyze their variance using
the Coeficient of Variation (  ). The  is defined as the ratio of the standard deviation of change</p>
        <sec id="sec-4-2-1">
          <title>3https://en.wikipedia.org/wiki/List_of_edit_wars_on_Wikipedia</title>
          <p>1sec
1sec
1min</p>
          <p>1hour 1day1we1ekmonth 1year
Change Interval
1min</p>
          <p>1hour 1day1we1ekmonth 1year
Change Interval
intervals ( ) to their mean ( ),  =  . It serves as a normalized measure of interval dispersion.
Intuitively, tables updated at regular frequencies, such as those documenting the Grammy Award for Song
of the Year, should exhibit a low  , with a perfectly periodic interval yielding  = 0. Conversely,
tables subject to random update patterns result in higher  values, typically  &gt; 1.</p>
          <p>Figure 6 illustrates the distribution of  across the Wiki Lake. Tables with  &lt; 1 are characterized
as having relatively deterministic update schedules, whereas those with  &gt; 1 exhibit higher temporal
dynamicity. Our analysis reveals a significant population of tables with near-static update patterns,
followed by a broad distribution of tables displaying varying degrees of temporal irregularity across a
wide spectrum of  values.</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Rollback</title>
        <p>A rollback is the restoration of a table,
column, or row to a previous state. Within the
Wikipedia ecosystem, rollbacks constitute a
critical mechanism for maintaining data
integrity and trustworthiness, typically executed
by moderators or automated systems4.
Rolltbiaocnkssoanretarbelceos,rtdheedreafsorine,diedpeenntidficeantitonmiosdnificeac-- 0 0 C2oefficient4of Variatio6n (CV) 8 10
essary to detect incorrect information and also Figure 6: Coeficient of Variation (CV) distribution.
avoid duplicate data. We analyze rollbacks to
distinguish between reversionary actions and evolutionary content updates. Specifically, we investigate
the relevant dimensions to temporal rollbacks, the profiles of the contributors involved, and the extent
to which a table’s revision history can serve as a proxy for its immunity to fraudulent data.</p>
        <p>Figure 7 (left) illustrates the distribution of temporal intervals between consecutive revisions across
the lake, categorized by contributor consistency, i.e., whether or not the same user initiated the successive
changes. The experiment reveals a significant correlation between temporal proximity and contributor
identity. High-frequency revisions, i.e., those occurring within a very short temporal window, are
predominantly executed by the same user. This behavior is largely attributed to the non-atomic nature
of complex updates and subsequent self-corrections.</p>
        <p>In contrast, limiting the analysis to rollbacks  →  → , where  and  represent distinct
table versions, results in a diferent distribution. Figure 7 (right) presents the temporal distribution
restricted exclusively to these reversionary cycles. Several key observations emerge: (i) rollbacks
account for about 7% of the 37M consecutive changes, (ii) the distribution deviates significantly from
the global baseline, showing a disproportionately high density of near-instantaneous changes, and (iii)
unlike general edits, rollbacks are significantly more likely to be performed by a diferent user.</p>
        <sec id="sec-4-3-1">
          <title>4https://en.wikipedia.org/wiki/Wikipedia:Bots/Requests_for_approval/ClueBot_NG</title>
          <p>Static/Dynamic threshold at 1.0
Distribution of ALL Changes
(Total: 37,077,741)
7 1eU6ser Relationship</p>
          <p>Different User
6 Same User
Distribution of ROLLBACKS Only
(Total: 2,519,119)
User Relationship</p>
          <p>Different User
Same User
8
7
t 6
n
ou5
C
e
t 4
u
l
o
s3
b
A
2
1
0
1e5
t 5
n
u
o
C4
e
t
u
l 3
o
s
b
A2
1
0</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Future Work</title>
      <p>
        The exploratory analysis of the Wiki Lake yields several key insights and future directions:
Not all change is evolution. We observe that raw revision counts significantly overestimate
meaningful table evolution. While tables exhibit a high volume of revisions, a substantial fraction of these
changes occur within very short intervals and often reflect non-atomic edits, self-corrections, vandalism,
or moderation-driven rollbacks rather than genuine data evolution. In contrast, long-term intervals
reflect the natural evolution of the data. Future discovery systems must diferentiate between these
“maintenance” intermediate table revisions and “evolutionary” stable ones to avoid being overwhelmed
by them. It will be interesting to see how these insights carry over to other temporal data lakes where
changes are not made through Wiki pages, but rather through data cleaning or integration operations.
The persistence of structural integrity. Despite the high volume of content updates, the fundamental
structure of tables remains stable. Our experiments show that in 80% of the changes the schema of
tables remains unchanged. For data lakes that may not be as entity-based as the Wikipedia tables, we
hypothesize that there will nonetheless be some structural integrity as tables evolve.
Future systems. Our findings suggest that temporal information is not merely supplementary metadata
but a foundational signal for reasoning about data lakes. Metrics such as change interval regularity and
rollback frequency provide strong indicators of table reliability, maturity, and semantic intent. These
signals are not entirely orthogonal to traditional discovery features such as schema similarity or value
overlap, and hence should be considered for downstream tasks of temporally-valid data discovery.
Data discovery over temporal data lakes. One future direction is the development of a data discovery
architecture that can reason over the semantics of evolving tables during the retrieval process. By
incorporating this reasoning, a data discovery system could move beyond value and structure matching
to understand the temporal validity of the retrieved results according to the user-provided query.
Semantic evolution analysis. Another angle for future research is to characterize the semantic
trajectories of changes, determining whether tables evolve according to predictable patterns. Current
advancements in Tabular Foundation Models (TFM) [
        <xref ref-type="bibr" rid="ref20">38</xref>
        ] and table embeddings [
        <xref ref-type="bibr" rid="ref21">39</xref>
        ] provide the tools
to project tables into high-dimensional semantic spaces. This allows for a robust analysis that moves
beyond simple syntax or schema changes to capture the underlying evolution of temporal semantics.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used Grammarly and Gemini3 to: Grammar and spelling
check. Furthermore, the authors used Gemini3 for Figures to: Generate images. After using these tools
and services, the authors reviewed and edited the content as needed and take full responsibility for the
publication’s content.
[16] A. Kittur, B. Suh, B. A. Pendleton, E. H. Chi, He says, she says: conflict and coordination in wikipedia, in:
Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI07, ACM, 2007, p.
453–462. URL: http://dx.doi.org/10.1145/1240624.1240698. doi:10.1145/1240624.1240698.
[17] T. Yasseri, R. Sumi, A. Rung, A. Kornai, J. Kertész, Dynamics of conflicts in wikipedia, PLoS ONE 7 (2012)
e38869. URL: http://dx.doi.org/10.1371/journal.pone.0038869. doi:10.1371/journal.pone.0038869.
[18] A. G. West, S. Kannan, I. Lee, Detecting wikipedia vandalism via spatio-temporal analysis of revision
metadata?, in: Proceedings of the Third European Workshop on System Security, EuroSys ’10, ACM, 2010, p.
22–28. URL: http://dx.doi.org/10.1145/1752046.1752050. doi:10.1145/1752046.1752050.
[19] B. T. Adler, L. de Alfaro, S. M. Mola-Velasco, P. Rosso, A. G. West, Wikipedia Vandalism Detection: Combining
Natural Language, Metadata, and Reputation Features, Springer Berlin Heidelberg, 2011, p. 277–288. URL:
http://dx.doi.org/10.1007/978-3-642-19437-5_23. doi:10.1007/978-3-642-19437-5_23.
[20] J. Daxenberger, I. Gurevych, Automatically classifying edit categories in Wikipedia revisions, in: D. Yarowsky,
T. Baldwin, A. Korhonen, K. Livescu, S. Bethard (Eds.), Proceedings of the 2013 Conference on Empirical
Methods in Natural Language Processing, Association for Computational Linguistics, Seattle, Washington,
USA, 2013, pp. 578–589. URL: https://aclanthology.org/D13-1055/.
[21] T. Bleifuß, T. Johnson, D. V. Kalashnikov, F. Naumann, V. Shkapenyuk, D. Srivastava, Enabling change
exploration: Vision paper, in: M. A. Sharaf, Y. Velegrakis (Eds.), Proceedings of the ExploreDB’17, Chicago,
IL, USA, May 19, 2017, ACM, 2017, pp. 1:1–1:3. URL: https://doi.org/10.1145/3077331.3077340. doi:10.1145/
3077331.3077340.
[22] T. Bleifuß, L. Bornemann, D. V. Kalashnikov, F. Naumann, D. Srivastava, Dbchex: Interactive exploration
of data and schema change, in: 9th Biennial Conference on Innovative Data Systems Research, CIDR
2019, Asilomar, CA, USA, January 13-16, 2019, Online Proceedings, www.cidrdb.org, 2019. URL: http:
//cidrdb.org/cidr2019/papers/p65-bleifuss-cidr19.pdf.
[23] A. P. Bhardwaj, S. Bhattacherjee, A. Chavan, A. Deshpande, A. J. Elmore, S. Madden, A. G. Parameswaran,
Datahub: Collaborative data science &amp; dataset version management at scale, in: Seventh Biennial Conference
on Innovative Data Systems Research, CIDR 2015, Asilomar, CA, USA, January 4-7, 2015, Online Proceedings,
www.cidrdb.org, 2015. URL: http://cidrdb.org/cidr2015/Papers/CIDR15_Paper18.pdf.
[24] S. Bhattacherjee, A. Chavan, S. Huang, A. Deshpande, A. G. Parameswaran, Principles of dataset versioning:
Exploring the recreation/storage tradeof, Proc. VLDB Endow. 8 (2015) 1346–1357. URL: http://www.vldb.
org/pvldb/vol8/p1346-bhattacherjee.pdf. doi:10.14778/2824032.2824035.
[25] S. Huang, L. Xu, J. Liu, A. J. Elmore, A. G. Parameswaran, Orpheusdb: Bolt-on versioning for relational
databases, Proc. VLDB Endow. 10 (2017) 1130–1141. URL: http://www.vldb.org/pvldb/vol10/p1130-huang.pdf.
doi:10.14778/3115404.3115417.
[26] M. E. Schüle, J. Schmeißer, T. Blum, A. Kemper, T. Neumann, Tardisdb: Extending sql to support versioning,
in: Proceedings of the 2021 International Conference on Management of Data, SIGMOD/PODS ’21, ACM,
2021, p. 2775–2778. URL: http://dx.doi.org/10.1145/3448016.3452767. doi:10.1145/3448016.3452767.
[27] G. S. Yilmaz, T. Wattanawaroon, L. Xu, A. Nigam, A. J. Elmore, A. Parameswaran, Datadif: User-interpretable
data transformation summaries for collaborative data analysis, in: Proceedings of the 2018 International
Conference on Management of Data, SIGMOD/PODS ’18, ACM, 2018, p. 1769–1772. URL: http://dx.doi.org/
10.1145/3183713.3193564. doi:10.1145/3183713.3193564.
[28] M. Armbrust, T. Das, L. Sun, B. Yavuz, S. Zhu, M. Murthy, J. Torres, H. van Hovell, A. Ionescu, A. Łuszczak,
M. Świtakowski, M. Szafrański, X. Li, T. Ueshin, M. Mokhtar, P. Boncz, A. Ghodsi, S. Paranjpye, P. Senster,
R. Xin, M. Zaharia, Delta lake: high-performance acid table storage over cloud object stores, Proceedings
of the VLDB Endowment 13 (2020) 3411–3424. URL: http://dx.doi.org/10.14778/3415478.3415560. doi:10.
14778/3415478.3415560.
[29] Z. Abedjan, M. Esmailoghli, S. Galhorta, Data discovery in data lakes: Operations, indexes, systems, Proc.</p>
      <p>VLDB Endow. 18 (2025) 5455–5459. URL: https://www.vldb.org/pvldb/vol18/p5455-abedjan.pdf.
[30] R. C. Fernandez, Z. Abedjan, F. Koko, G. Yuan, S. Madden, M. Stonebraker, Aurum: A data discovery
system, in: 34th IEEE International Conference on Data Engineering, ICDE 2018, Paris, France, April
16-19, 2018, IEEE Computer Society, 2018, pp. 1001–1012. URL: https://doi.org/10.1109/ICDE.2018.00094.
doi:10.1109/ICDE.2018.00094.
[31] N. Heist, H. Paulheim, Caligraph: A knowledge graph from wikipedia categories and lists, Semantic Web 16
(2025). URL: https://doi.org/10.1177/22104968251361349. doi:10.1177/22104968251361349.
[32] E. Muñoz, A. Hogan, A. Mileo, Dreta: Extracting RDF from wikitables, in: E. Blomqvist, T. Groza (Eds.),
Proceedings of the ISWC 2013 Posters &amp; Demonstrations Track, Sydney, Australia, October 23, 2013, volume
1035 of CEUR Workshop Proceedings, CEUR-WS.org, 2013, pp. 89–92. URL: https://ceur-ws.org/Vol-1035/
iswc2013_demo_23.pdf.
[33] G. Karagiannis, I. Trummer, S. Jo, S. Khandelwal, X. Wang, C. Yu, Mining an "anti-knowledge base" from</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Chepurko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Marcus</surname>
          </string-name>
          , E. Zgraggen,
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Fernandez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kraska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Karger</surname>
          </string-name>
          ,
          <article-title>ARDA: automatic relational data augmentation for machine learning 13 (</article-title>
          <year>2020</year>
          )
          <fpage>1373</fpage>
          -
          <lpage>1387</lpage>
          . URL: http://www.vldb.org/pvldb/vol13/ p1373-chepurko.pdf.
          <source>doi:10.14778/3397230</source>
          .3397235.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Esmailoghli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Quiané-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Abedjan</surname>
          </string-name>
          ,
          <article-title>MATE: multi-attribute table extraction</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>15</volume>
          (
          <year>2022</year>
          )
          <fpage>1684</fpage>
          -
          <lpage>1696</lpage>
          . URL: https://www.vldb.org/pvldb/vol15/p1684-esmailoghli.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nargesian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <article-title>JOSIE: overlap set similarity search for finding joinable tables in data lakes</article-title>
          , in: SIGMOD, ACM,
          <year>2019</year>
          , pp.
          <fpage>847</fpage>
          -
          <lpage>864</lpage>
          . URL: https://doi.org/10.1145/3299869.3300065. doi:
          <volume>10</volume>
          . 1145/3299869.3300065.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , R. J.
          <string-name>
            <surname>Miller</surname>
          </string-name>
          ,
          <article-title>Semantics-aware dataset discovery from data lakes with contextualized column-based representation learning</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>16</volume>
          (
          <year>2023</year>
          )
          <fpage>1726</fpage>
          -
          <lpage>1739</lpage>
          . URL: https://www.vldb.org/pvldb/vol16/p1726-fan.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Esmailoghli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Quiané-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Abedjan</surname>
          </string-name>
          ,
          <article-title>COCOA: correlation coeficient-aware data augmentation, in: EDBT, OpenProceedings</article-title>
          .org,
          <year>2021</year>
          , pp.
          <fpage>331</fpage>
          -
          <lpage>336</lpage>
          . URL: https://doi.org/10.5441/002/edbt.
          <year>2021</year>
          .
          <volume>30</volume>
          . doi:
          <volume>10</volume>
          . 5441/002/edbt.
          <year>2021</year>
          .
          <volume>30</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A. S. R.</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bessa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Musco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Freire</surname>
          </string-name>
          ,
          <article-title>A sketch-based index for correlated dataset search</article-title>
          ,
          <source>in: 38th IEEE International Conference on Data Engineering, ICDE</source>
          <year>2022</year>
          ,
          <string-name>
            <given-names>Kuala</given-names>
            <surname>Lumpur</surname>
          </string-name>
          , Malaysia, May 9-
          <issue>12</issue>
          ,
          <year>2022</year>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>2928</fpage>
          -
          <lpage>2941</lpage>
          . URL: https://doi.org/10.1109/ICDE53745.
          <year>2022</year>
          .
          <volume>00264</volume>
          . doi:
          <volume>10</volume>
          .1109/ICDE53745.
          <year>2022</year>
          .
          <volume>00264</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Becktepe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Esmailoghli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Koch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Abedjan</surname>
          </string-name>
          ,
          <string-name>
            <surname>Demonstrating</surname>
            <given-names>MATE</given-names>
          </string-name>
          and
          <article-title>COCOA for data discovery</article-title>
          , in: S.
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Pandis</surname>
            ,
            <given-names>K. S.</given-names>
          </string-name>
          <string-name>
            <surname>Candan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Amer-Yahia</surname>
          </string-name>
          (Eds.),
          <source>Companion of the 2023 International Conference on Management of Data</source>
          , SIGMOD/PODS 2023, Seattle, WA, USA, June 18-23,
          <year>2023</year>
          , ACM,
          <year>2023</year>
          , pp.
          <fpage>119</fpage>
          -
          <lpage>122</lpage>
          . URL: https://doi.org/10.1145/3555041.3589716. doi:
          <volume>10</volume>
          .1145/3555041.3589716.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Esmailoghli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schnell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Abedjan</surname>
          </string-name>
          ,
          <article-title>BLEND: A unified data discovery system</article-title>
          ,
          <source>in: 41st IEEE International Conference on Data Engineering, ICDE</source>
          <year>2025</year>
          ,
          <string-name>
            <given-names>Hong</given-names>
            <surname>Kong</surname>
          </string-name>
          , May
          <volume>19</volume>
          -23,
          <year>2025</year>
          , IEEE,
          <year>2025</year>
          , pp.
          <fpage>737</fpage>
          -
          <lpage>750</lpage>
          . URL: https://doi.org/10.1109/ICDE65448.
          <year>2025</year>
          .
          <volume>00061</volume>
          . doi:
          <volume>10</volume>
          .1109/ICDE65448.
          <year>2025</year>
          .
          <volume>00061</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Shraga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <article-title>Explaining dataset changes for semantic data versioning with explain-da-v</article-title>
          ,
          <source>Proceedings of the VLDB Endowment</source>
          <volume>16</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Esmailoghli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Weidlich</surname>
          </string-name>
          ,
          <article-title>The past still matters: A temporally-valid data discovery system</article-title>
          ,
          <source>CoRR abs/2510</source>
          .13662 (
          <year>2025</year>
          ). URL: https://doi.org/10.48550/arXiv.2510.13662. doi:
          <volume>10</volume>
          .48550/ARXIV.2510.13662. arXiv:
          <volume>2510</volume>
          .
          <fpage>13662</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bornemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bleifuß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nargesian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <article-title>Eficient discovery of temporal inclusion dependencies in wikipedia tables</article-title>
          , in: EDBT, OpenProceedings.org,
          <year>2024</year>
          , pp.
          <fpage>399</fpage>
          -
          <lpage>411</lpage>
          . URL: https://doi.org/10.48786/edbt.
          <year>2024</year>
          .
          <volume>35</volume>
          . doi:
          <volume>10</volume>
          .48786/EDBT.
          <year>2024</year>
          .
          <volume>35</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bleifuß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bornemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <article-title>Structured object matching across web page revisions</article-title>
          ,
          <source>in: 37th IEEE International Conference on Data Engineering, ICDE</source>
          <year>2021</year>
          , Chania, Greece,
          <source>April 19-22</source>
          ,
          <year>2021</year>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>1284</fpage>
          -
          <lpage>1295</lpage>
          . URL: https://doi.org/10.1109/ICDE51399.
          <year>2021</year>
          .
          <volume>00115</volume>
          . doi:
          <volume>10</volume>
          .1109/ICDE51399.
          <year>2021</year>
          .
          <volume>00115</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Tschirschnitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Papenbrock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <article-title>Detecting inclusion dependencies on very many tables</article-title>
          ,
          <source>ACM Trans. Database Syst</source>
          .
          <volume>42</volume>
          (
          <year>2017</year>
          )
          <volume>18</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          :
          <fpage>29</fpage>
          . URL: https://doi.org/10.1145/3105959. doi:
          <volume>10</volume>
          .1145/3105959.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Bhagavatula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Noraset</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Downey</surname>
          </string-name>
          ,
          <article-title>Methods for exploring and mining tables on wikipedia</article-title>
          , in: D. H.
          <string-name>
            <surname>Chau</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Vreeken</surname>
          </string-name>
          , M. van
          <string-name>
            <surname>Leeuwen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          Faloutsos (Eds.),
          <source>Proceedings of the ACM SIGKDD Workshop on Interactive Data Exploration and Analytics</source>
          ,
          <source>IDEA@KDD</source>
          <year>2013</year>
          , Chicago, Illinois, USA,
          <year>August 11</year>
          ,
          <year>2013</year>
          , ACM,
          <year>2013</year>
          , pp.
          <fpage>18</fpage>
          -
          <lpage>26</lpage>
          . URL: https://doi.org/10.1145/2501511.2501516. doi:
          <volume>10</volume>
          .1145/2501511.2501516.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bleifuß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bornemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <article-title>The secret life of wikipedia tables</article-title>
          , in: D.
          <string-name>
            <surname>Mottin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lissandrini</surname>
            ,
            <given-names>S. B.</given-names>
          </string-name>
          <string-name>
            <surname>Roy</surname>
          </string-name>
          , Y. Velegrakis (Eds.),
          <source>Proceedings of the 2nd Workshop on Search</source>
          , Exploration, and
          <article-title>Analysis in Heterogeneous Datastores (SEA-Data 2021) co-located with 47th International Conference on Very Large Data Bases (VLDB</article-title>
          <year>2021</year>
          ), Copenhagen, Denmark,
          <year>August 20</year>
          ,
          <year>2021</year>
          , volume
          <volume>2929</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>20</fpage>
          -
          <lpage>26</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2929</volume>
          /paper4.pdf.
          <article-title>wikipedia updates with applications to fact checking and beyond</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>13</volume>
          (
          <year>2019</year>
          )
          <fpage>561</fpage>
          -
          <lpage>573</lpage>
          . URL: http://www.vldb.org/pvldb/vol13/p561-karagiannis.pdf.
          <source>doi:10.14778/3372716</source>
          .3372727.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mahdavi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Abedjan</surname>
          </string-name>
          ,
          <article-title>Baran: Efective error correction via a unified context representation and transfer learning</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>13</volume>
          (
          <year>2020</year>
          )
          <fpage>1948</fpage>
          -
          <lpage>1961</lpage>
          . URL: http://www.vldb.org/pvldb/vol13/p1948-mahdavi. pdf.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bleifuß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bornemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <article-title>Exploring change - A new dimension of data analytics</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>12</volume>
          (
          <year>2018</year>
          )
          <fpage>85</fpage>
          -
          <lpage>98</lpage>
          . URL: http://www.vldb.org/pvldb/ vol12/p85-bleifuss.pdf.
          <source>doi:10.14778/3282495</source>
          .3282496.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bornemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bleifuß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <article-title>Natural key discovery in wikipedia tables</article-title>
          , in: Y.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>King</surname>
          </string-name>
          , T. Liu, M. van Steen (Eds.),
          <source>WWW '20: The Web Conference</source>
          <year>2020</year>
          , Taipei, Taiwan,
          <source>April 20-24</source>
          ,
          <year>2020</year>
          , ACM / IW3C2,
          <year>2020</year>
          , pp.
          <fpage>2789</fpage>
          -
          <lpage>2795</lpage>
          . URL: https://doi.org/10.1145/3366423.3380039. doi:
          <volume>10</volume>
          .1145/3366423.3380039.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sumi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yasseri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kornai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kertész</surname>
          </string-name>
          ,
          <article-title>Edit wars in wikipedia</article-title>
          , in: PASSAT/SocialCom 2011, Privacy, Security,
          <source>Risk and Trust (PASSAT)</source>
          ,
          <source>2011 IEEE Third International Conference on and 2011 IEEE Third International Conference on Social Computing (SocialCom)</source>
          , Boston, MA, USA,
          <fpage>9</fpage>
          -
          <lpage>11</lpage>
          Oct.,
          <year>2011</year>
          , IEEE Computer Society,
          <year>2011</year>
          , pp.
          <fpage>724</fpage>
          -
          <lpage>727</lpage>
          . URL: https://doi.org/10.1109/PASSAT/SocialCom.
          <year>2011</year>
          .
          <volume>47</volume>
          . doi:
          <volume>10</volume>
          .1109/PASSAT/SOCIALCOM.
          <year>2011</year>
          .
          <volume>47</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>P.</given-names>
            <surname>Papotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Binnig</surname>
          </string-name>
          ,
          <article-title>Panel on neural relational data: Tabular foundation models, llms</article-title>
          ... or both?,
          <source>Proc. VLDB Endow</source>
          .
          <volume>18</volume>
          (
          <year>2025</year>
          )
          <fpage>5513</fpage>
          -
          <lpage>5515</lpage>
          . URL: https://www.vldb.org/pvldb/vol18/p5513-paolo.pdf.
          <source>doi:10.14778/ 3750601</source>
          .3760519.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>G.</given-names>
            <surname>Shrestha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Akula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Yannam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pyayt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. N.</given-names>
            <surname>Gubanov</surname>
          </string-name>
          ,
          <article-title>Tabular embeddings for tables with bidimensional hierarchical metadata and nesting</article-title>
          , in: A.
          <string-name>
            <surname>Simitsis</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Kemme</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Queralt</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Romero</surname>
          </string-name>
          , P. Jovanovic (Eds.),
          <source>Proceedings 28th International Conference on Extending Database Technology, EDBT</source>
          <year>2025</year>
          , Barcelona, Spain, March
          <volume>25</volume>
          -28,
          <year>2025</year>
          , OpenProceedings.org,
          <year>2025</year>
          , pp.
          <fpage>92</fpage>
          -
          <lpage>105</lpage>
          . URL: https://doi.org/10.48786/edbt.
          <year>2025</year>
          .
          <volume>08</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>