<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>SEBD</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Entity Resolution On-Demand for Querying Dirty Datasets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>(Discussion Paper)</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Simonini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Zecchini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felix Naumann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sonia Bergamaschi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hasso Plattner Institute, University of Potsdam</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Modena and Reggio Emilia</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>31</volume>
      <fpage>02</fpage>
      <lpage>05</lpage>
      <abstract>
        <p>Entity Resolution (ER) is the process of identifying and merging records that refer to the same real-world entity. ER is usually applied as an expensive cleaning step on the entire data before consuming it, yet the relevance of cleaned entities ultimately depends on the user's specific application, which may only require a small portion of the entities. We introduce BrewER, a framework designed to evaluate SQL SP queries on unclean data while progressively providing results as if they were obtained from cleaned data. BrewER aims at cleaning a single entity at a time, adhering to an ORDER BY predicate, thus it inherently supports top-k queries and stop-and-resume execution. This approach can save a significant amount of resources for various applications. BrewER has been implemented as an open-source Python library and can be seamlessly employed with existing ER tools and algorithms. We thoroughly demonstrated its eficiency through its evaluation on four real-world datasets.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Entity Resolution</kwd>
        <kwd>Data Integration</kwd>
        <kwd>ELT</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Entity Resolution On-Demand</title>
      <p>(a) Offline ER
(b) Query
Choose a matching function: μ()
- μ_customers_DL_Transfer_1
- μ_customers_DL_Transfer_2
...
- μ_electronics_DL_custom_n
...
a
t
a
d
y
itr
d
new data</p>
      <p>Choose resolution functions:
α1() = MIN(&lt;customer_id&gt;)
α2() = MAX(&lt;data_month_usage&gt;)
α3() = AVG(&lt;calls_month&gt;)</p>
      <sec id="sec-1-1">
        <title>Batch ER</title>
        <p>Q1
SELECT customer_id
FROM customer
WHERE data_month_usage &gt; ’10GB’</p>
        <p>ORDER BY calls_month DESC
Query
exec.</p>
        <p>ll
a
c
e
r
Qcomparisons
cleaned data
cleaned data clean results
(c) Query w/ BrewER
Q2
SELECT TOP 50</p>
        <p>MIN(customer_id),</p>
        <p>AVG(calls_month)
FROM customer
GROUP BY ENTITY WITH MATCHER μ</p>
        <p>HAVING MAX(data_month_usage)&gt;’10GB’
new data ORDER BY AVG(calls_month) DESC
a
t
a
d
y
itr
d</p>
      </sec>
      <sec id="sec-1-2">
        <title>BrewER</title>
        <sec id="sec-1-2-1">
          <title>Fusion</title>
        </sec>
        <sec id="sec-1-2-2">
          <title>Matching</title>
          <p>ll
a
c
e
r</p>
          <p>Qcomparisons
progressive
clean results
1.1. Executing queries on dirty data with existing solutions
Traditionally, ER is employed as a cleaning step before using the data. Yet, in many practical
scenarios this might not be convenient:
Example 1.1. Ellen is a data scientist building a machine learning model to predict customer
churn for a telecom company, with the following requirements: (i) she has limited time to add new
data to her dataset, which will contain duplicates; (ii) she has business priorities: it is better to have
clean data for high-value customers (i.e., those that make more phone calls) than for low-value
ones, and only customers with a certain data usage (e.g., those that have a monthly data usage
greater than 10 GB) should be considered—she can express this with Query 1 in Figure 1b.</p>
          <p>For ER, Ellen already has a matching function to choose (adapting some internal pre-trained
deep learning models) and she knows rules for resolving the conflicts in the attribute values of the
clusters of matching records (e.g., AVG(calls_month), MAX(data_month_usage), etc.).</p>
          <p>The example above depicts a common scenario for data practitioners (e.g., data scientists),
characterized by:
• An information need: only some entities are relevant and some entities are more relevant
than others;
• Time constraints: data might become outdated quickly and/or users want to do a fast
exploration of relevant portions of cleaned data.</p>
          <p>Example 1.2. To get correct results for the query (i.e., taking into account that some records
are duplicates) Ellen employs a traditional ER framework to clean the entire dataset (Figure 1a).
However, she soon realizes that ER is the bottleneck due to its inherently quadratic complexity
and the cost of the matching function, which involves expensive operations based on deep neural
network models. As a result, it takes a significant amount of time to clean the entire dataset using
ER. Furthermore, to build the right ER pipeline is not a trivial task: she would need to debug the ER
pipeline with the data at hand (e.g., to check if the matcher she is employing is performing well for
high-value customers), but she cannot stop the ER process after receiving a handful of the entities to
inspect—this is because those entities might not be relevant for the query or might be only partially
resolved, which could lead to incorrect results. Alternatively, she would have to manually select
records from the dataset to test the ER pipeline, which is also time-consuming.</p>
          <p>
            The motivating example highlights the need for a more eficient and targeted approach to ER
that prioritizes cleaning.
1.2. A novel approach to execute queries on dirty data
We propose BrewER1 [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ], an ER framework that aims to provide an eficient and targeted
approach to ER by evaluating SQL SP (Selection and Projection) queries on dirty data and
returning results as if they were issued on cleaned data. The key feature of BrewER is its ability
to perform ER progressively, guided by an ORDER BY clause, to incrementally return the most
relevant results to the data scientist. This approach avoids matching and resolving entities
that are not part of the final result, thereby saving time and resources. Additionally, BrewER
inherently supports top-k queries and allows for stop-and-resume execution. To enable this
progressive approach to ER, BrewER introduces a special “GROUP BY ENTITY WITH MATCHER
[matcher of choice]” operator, meaning that matching records should be grouped according
to the selected matcher—then, filtering conditions on the entities can be applied by means of the
HAVING clause. Overall, BrewER ofers a more eficient and targeted approach to ER, which can
save time and resources, especially for data scientists dealing with large and complex datasets.
Example 1.3. Using BrewER, Ellen can easily adapt her original SQL query to work with dirty
data by employing a special GROUP BY statement and moving the selection statements into the
HAVING clause, predicated on each group (i.e., each entity), as shown in Figure 1c. Additionally, she
specifies the resolution functions for ER within the SQL query as aggregate functions. Once Ellen
has specified her new, equivalent query, BrewER executes it directly on the dirty data, applying
1This paper is a revisited version of the one published at the 48ℎ International Conference on Very Large Databases (VLDB 2022).
μ:&lt;[match],
          </p>
          <p>[non-match]&gt;
matchDB NMaotncMhaLticshtLsi,sts</p>
          <p>BrewER</p>
          <p>Entity
Matching</p>
          <p>Blocking</p>
          <p>Query
progressive
cleaned results</p>
          <p>ER progressively on the right portion of the data to yield correct results incrementally. This allows
Ellen to receive the first entities in a fraction of the time required by existing ER frameworks. She
can explore new data without completely cleaning it and maximize the ER eforts on the entities
she actually needs for her task. Moreover, BrewER allows Ellen to stop the execution at any time
with the guarantee that the results produced so far are correct. She can then inspect the results of
the ER process for entities of interest and debug it if needed. This feature saves time and resources
compared to traditional ER frameworks, where debugging and testing the ER pipeline requires a
complete cleaning of the entire dataset. Finally, BrewER keeps track of both executed comparisons
and resolved entities to avoid repeating the same operations when multiple queries are issued on
the same data. This further improves the eficiency of the ER process, allowing Ellen to perform her
analysis more quickly and accurately.</p>
          <p>
            More generally, our proposed approach is well-suited for addressing one of the major
challenges in data lake management systems [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]: to support extraction and cleaning as part of the
integration pipeline on-demand. Similarly, on-demand data transformation that returns results
in a timely manner is a fundamental requirement of ELT (Extract-Load-Transform) pipelines,
especially when combined with top-k queries for debugging transformations [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ].
          </p>
          <p>The main contributions of BrewER can be summarized as follows. We formalize the concept
of ER-on-demand, which involves progressively cleaning and emitting entities that satisfy
queries issued directly on dirty datasets, and propose an algorithm for it. Finally, we implement
this algorithm in an open-source system and extensively evaluate it on four real-world datasets,
demonstrating its efectiveness.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. The BrewER Framework</title>
      <p>
        BrewER is designed as a flexible and adaptable framework, as illustrated in Figure 2. It is
implemented as a Python library, whose code is publicly available on GitHub2. This approach
allows the seamless integration of BrewER into Python workflows in Jupyter3 notebooks,
as we show below in Section 3. Being BrewER agnostic towards the selected blocking and
matching functions, users can integrate it with their preferred binary matching libraries (such
as DeepMatcher [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], Ditto [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], etc.) and blocking techniques (like Magellan [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], JedAI [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ],
2https://github.com/dbmodena/BrewER
3https://jupyter.org
      </p>
      <p>SELECT [TOP ] ⟨ ()⟩</p>
      <p>FROM 
[WHERE  ]
GROUP BY ENTITY WITH MATCHER</p>
      <p>
        [HAVING ⟨ () {LIKE|IN|&lt;|≤ |&gt;|≥ |=} ⟩]
[ORDER BY  () [ASC|DESC]]
or SparkER [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]). The system then performs ER in an on-demand fashion while executing the
user’s query with the algorithm presented in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>In particular, BrewER builds on the output of the blocking function and performs a preliminary
ifltering of the blocks, keeping only the ones containing records whose values might lead to
the generation of an entity appearing in the result of the query. The records of the blocks
that pass the filtering are then inserted in a priority queue, keeping for each one the list of its
candidate matches. The priority is defined according to the value of the attribute appearing
in the ORDER BY clause, in ascending or descending order. BrewER iterates on the priority
queue, considering at each iteration the head element: if it is a record, its candidate matches
are checked generating a completely resolved entity; otherwise (i.e., it is a completely resolved
entity), it is emitted or discarded based on whether or not it satisfies the query.</p>
      <p>To optimize the performance of the matching functions and avoid re-comparing candidate
pairs, BrewER maintains separate databases for the lists of matching and non-matching records
for each matching function adopted by the user. To save space, users may opt to store only the
ifnal resolved entities—the resolution functions cannot change across queries in this case.
2.1. Supported queries</p>
    </sec>
    <sec id="sec-3">
      <title>3. Experiments and Demonstration</title>
      <p>
        Through our experimental evaluation [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], we point out the benefits that BrewER can generate
in terms of elapsed time and saved resources. In Table 1, we report the characteristics of the
datasets used in our experiments, presenting significant diferences regarding the size (in terms
of both records and attributes) and the domain, covering commercial products (cameras in the
Dataset
SIGMOD20
SIGMOD21
Altosight
Funding
case of SIGMOD204, part of the Alaska benchmark [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], and USB sticks for SIGMOD215 and its
superset provided by Altosight6) and organizations (Funding7 [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]).
      </p>
      <p>
        In Figure 4, we show the results obtained by running batches of conjunctive (i.e., with the
HAVING conditions in AND) and disjunctive (i.e., with the HAVING conditions in OR) queries
with BrewER on the four datasets. The plot shows the average number of comparisons needed
to reach a certain query recall (i.e., the emission of a certain percentage of resulting entities).
BrewER is able to return the entities with a high priority in a small fraction of the time
required for performing the entire cleaning process, which has to be carried out by the batch
algorithms to get the query results (here we consider as a baseline QDA [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], a query-driven
4http://www.inf.uniroma3.it/db/sigmod2020contest
5https://dbgroup.ing.unimo.it/sigmod21contest
6https://altosight.com
7https://raw.githubusercontent.com/qcri/data_civilizer_system/master/grecord_service/gr/data/address/address.csv
batch approach and the closest prior work to BrewER). Figure 5, reporting a similar experiment
(considering in this case the queries in the batches yielding the largest and the smallest result
sets), allows us to highlight the significant diference in terms of elapsed time compared to the
traditional batch approach. In our full research paper [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], you can find out more details about
the described experiments and several additional experiments covering the impact of blocking,
of diferent aggregate function, and the shortcomings of the existing related approaches. In
our demonstration [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], we show the benefits of BrewER for data practitioners, addressing in
particular two scenarios, depicted in Figure 6 and described below.
3.1. Querying dirty datasets
BrewER makes it possible to run queries on dirty datasets obtaining the progressive emission of
the cleaned resulting entities (Figure 6b), as soon as they are obtained, avoiding the
inconsistencies that would be raised by running the query directly on the dirty dataset (Figure 6a). BrewER
inherently supports top-k queries, thus Ellen can run such a query for the quick emission of the
results with the highest priority; once inspected the returned entities, she can decide whether
to resume the query to get the complete result set.
3.2. ER pipeline debugging
BrewER makes it also possible to obtain early insights on the quality of the ongoing cleaning
process, assessing the goodness of the chosen combination of blocking and matching functions.
As depicted in Figure 6c, Ellen can exploit top-k queries to check the absence of inconsistencies
in the result set. If some issues are spotted (e.g., duplicate records not matched because of a too
aggressive blocking function or a weak matching function), she can intervene and redesign the
ER pipeline, saving a significant amount of time and resources compared to batch solutions,
which allow to perform such controls only after the completion of the cleaning process.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Related Work</title>
      <p>
        The shortcomings of the traditional batch approach to ER are pointed out in literature and
diferent solutions have been proposed to overcome its limitations in dynamic scenarios. In
particular, the related work can be grouped into two main research directions: progressive
approaches and query-driven approaches. These methods present several significant diferences
compared to BrewER [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], whose on-demand approach implies the co-existence of both aspects.
4.1. Progressive ER
Progressive approaches to ER [
        <xref ref-type="bibr" rid="ref17 ref18 ref19 ref20 ref21">17, 18, 19, 20, 21</xref>
        ] try to maximize the impact of ER on a dirty
dataset in a limited amount of time. The key idea of these methods is to prioritize the comparisons
for the candidate pairs of records for which the probability to match is higher. Thus, it is not
possible for the user to define a priority based on their interests, as done in BrewER through
the ORDER BY clause of the query. Furthermore, operating at match level and not at entity level,
these approaches do not guarantee to dispose of clean entities before completing the ER process,
while BrewER progressively returns the clean entities appearing in the result of the query,
according to the user-defined priority.
4.2. Query-driven ER
Query-driven approaches to ER [
        <xref ref-type="bibr" rid="ref15 ref22">15, 22</xref>
        ] aim at performing ER only on the portion of the dataset
which is needed to answer the query. These solutions are the closest prior works to BrewER,
operating at block level to detect the comparisons that are not relevant for the query at hand.
Nevertheless, query-driven approaches are not designed to support the progressive emission of
the entities; thus, it is needed to wait for the end of the cleaning process to be able to inspect
the results of the query. Moreover, they can support only a limited range of aggregate functions
(e.g., the average or the majority voting are not supported).
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>
        We presented BrewER [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], a framework for Entity Resolution (ER) that allows users to filter
entities of interest from dirty data without having to clean the entire dataset. BrewER achieves
this by evaluating SQL SP queries on dirty data and progressively returning results as if they
were issued on cleaned data. The system is flexible and adaptable, allowing users to integrate
their preferred binary matching and blocking techniques. BrewER has been implemented as an
open-source Python library, which can be seamlessly integrated in data science workflows (e.g.,
in Jupyter notebooks) We demonstrated the eficacy of BrewER on four real-world datasets
and have shown that its overhead is negligible in real-world use cases. Future work includes
exploring how to support SQL SPJ queries for multi-table dirty datasets and additional features
for ER pipeline debugging.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Christen</surname>
          </string-name>
          , Data Matching:
          <article-title>Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection, Data-Centric Systems and Applications</article-title>
          (DCSA), Springer,
          <year>2012</year>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>642</fpage>
          -31164-2.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>X. L.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <source>Big Data Integration, Synthesis Lectures on Data Management (SLDM)</source>
          , Morgan &amp; Claypool Publishers,
          <year>2015</year>
          . doi:
          <volume>10</volume>
          .2200/ S00578ED1V01Y201404DTM040.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bleiholder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <source>Data Fusion, ACM Computing Surveys (CSUR) 41</source>
          (
          <year>2008</year>
          ) 1:
          <fpage>1</fpage>
          -
          <lpage>1</lpage>
          :
          <fpage>41</fpage>
          . doi:
          <volume>10</volume>
          .1145/1456650.1456651.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Papadakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Skoutas</surname>
          </string-name>
          , E. Thanos, T. Palpanas,
          <article-title>Blocking and Filtering Techniques for Entity Resolution: A Survey, ACM Computing Surveys (CSUR) 53 (</article-title>
          <year>2021</year>
          )
          <volume>31</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>31</lpage>
          :
          <fpage>42</fpage>
          . doi:
          <volume>10</volume>
          .1145/3377455.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Simonini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zecchini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Entity</given-names>
            <surname>Resolution</surname>
          </string-name>
          On-Demand,
          <source>Proceedings of the VLDB Endowment (PVLDB) 15</source>
          (
          <year>2022</year>
          )
          <fpage>1506</fpage>
          -
          <lpage>1518</lpage>
          . doi:
          <volume>10</volume>
          .14778/ 3523210.3523226.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>F.</given-names>
            <surname>Nargesian</surname>
          </string-name>
          , E. Zhu,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Pu</surname>
          </string-name>
          , P. C.
          <article-title>Arocena, Data Lake Management: Challenges and Opportunities</article-title>
          ,
          <source>Proceedings of the VLDB Endowment (PVLDB) 12</source>
          (
          <year>2019</year>
          )
          <fpage>1986</fpage>
          -
          <lpage>1989</lpage>
          . doi:
          <volume>10</volume>
          .14778/3352063.3352116.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Dolan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dunlap</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Hellerstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Welton</surname>
          </string-name>
          , MAD Skills:
          <article-title>New Analysis Practices for Big Data</article-title>
          ,
          <source>Proceedings of the VLDB Endowment (PVLDB) 2</source>
          (
          <year>2009</year>
          )
          <fpage>1481</fpage>
          -
          <lpage>1492</lpage>
          . doi:
          <volume>10</volume>
          .14778/1687553.1687576.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mudgal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rekatsinas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Doan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Park</surname>
          </string-name>
          , G. Krishnan,
          <string-name>
            <given-names>R.</given-names>
            <surname>Deep</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Arcaute</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Raghavendra</surname>
          </string-name>
          ,
          <article-title>Deep Learning for Entity Matching: A Design Space Exploration</article-title>
          ,
          <source>in: Proceedings of the International Conference on Management of Data (SIGMOD)</source>
          , ACM,
          <year>2018</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>34</lpage>
          . doi:
          <volume>10</volume>
          .1145/3183713.3196926.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Suhara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Doan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <article-title>Deep Entity Matching with Pre-Trained Language Models</article-title>
          ,
          <source>Proceedings of the VLDB Endowment (PVLDB) 14</source>
          (
          <year>2020</year>
          )
          <fpage>50</fpage>
          -
          <lpage>60</lpage>
          . doi:
          <volume>10</volume>
          .14778/ 3421424.3421431.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>Konda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>P. Suganthan G. C.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Doan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ardalan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Ballard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Panahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Naughton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Prasad</surname>
          </string-name>
          , G. Krishnan,
          <string-name>
            <given-names>R.</given-names>
            <surname>Deep</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Raghavendra</surname>
          </string-name>
          ,
          <source>Magellan: Toward Building Entity Matching Management Systems, Proceedings of the VLDB Endowment (PVLDB) 9</source>
          (
          <year>2016</year>
          )
          <fpage>1197</fpage>
          -
          <lpage>1208</lpage>
          . doi:
          <volume>10</volume>
          .14778/2994509.2994535.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>G.</given-names>
            <surname>Papadakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Mandilaras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gagliardelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Simonini</surname>
          </string-name>
          , E. Thanos, G. Giannakopoulos,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Palpanas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Koubarakis</surname>
          </string-name>
          ,
          <article-title>Three-dimensional Entity Resolution with JedAI, Information Systems</article-title>
          (IS)
          <volume>93</volume>
          (
          <year>2020</year>
          )
          <volume>101565</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>101565</lpage>
          :
          <fpage>17</fpage>
          . doi:
          <volume>10</volume>
          .1016/j.is.
          <year>2020</year>
          .
          <volume>101565</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>L.</given-names>
            <surname>Gagliardelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Simonini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Beneventano</surname>
          </string-name>
          , S. Bergamaschi, SparkER: Scaling Entity Resolution in Spark,
          <source>in: Proceedings of the International Conference on Extending Database Technology (EDBT)</source>
          ,
          <source>OpenProceedings.org</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>602</fpage>
          -
          <lpage>605</lpage>
          . doi:
          <volume>10</volume>
          .5441/ 002/edbt.
          <year>2019</year>
          .
          <volume>66</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>V.</given-names>
            <surname>Crescenzi</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. De Angelis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Firmani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Mazzei</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Merialdo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Piai</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <article-title>Alaska: A Flexible Benchmark for Data Integration Tasks, arXiv preprint (</article-title>
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          . 48550/arXiv.2101.11259.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Abedjan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Elmagarmid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. F.</given-names>
            <surname>Ilyas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Madden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ouzzani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stonebraker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <article-title>Unsupervised String Transformation Learning for Entity Consolidation</article-title>
          ,
          <source>in: Proceedings of the International Conference on Data Engineering (ICDE)</source>
          ,
          <source>IEEE Computer Society</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>196</fpage>
          -
          <lpage>207</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICDE.
          <year>2019</year>
          .
          <volume>00026</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>H.</given-names>
            <surname>Altwaijry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mehrotra</surname>
          </string-name>
          ,
          <string-name>
            <surname>Query-Driven Approach</surname>
          </string-name>
          to Entity Resolution,
          <source>Proceedings of the VLDB Endowment (PVLDB) 6</source>
          (
          <year>2013</year>
          )
          <fpage>1846</fpage>
          -
          <lpage>1857</lpage>
          . doi:
          <volume>10</volume>
          .14778/ 2556549.2556567.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zecchini</surname>
          </string-name>
          , G. Simonini,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          , BrewER: Entity Resolution OnDemand,
          <source>Proceedings of the VLDB Endowment (PVLDB) 16</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Whang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Marmaros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Garcia-Molina</surname>
          </string-name>
          ,
          <article-title>Pay-As-You-Go Entity Resolution, IEEE Transactions on Knowledge and Data Engineering (TKDE) 25 (</article-title>
          <year>2013</year>
          )
          <fpage>1111</fpage>
          -
          <lpage>1124</lpage>
          . doi:
          <volume>10</volume>
          . 1109/TKDE.
          <year>2012</year>
          .
          <volume>43</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>T.</given-names>
            <surname>Papenbrock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Heise</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          , Progressive Duplicate Detection,
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          <article-title>Transactions on Knowledge and Data Engineering (TKDE) 27 (</article-title>
          <year>2015</year>
          )
          <fpage>1316</fpage>
          -
          <lpage>1329</lpage>
          . doi:
          <volume>10</volume>
          .1109/TKDE.
          <year>2014</year>
          .
          <volume>2359666</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>D.</given-names>
            <surname>Firmani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Saha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <article-title>Online Entity Resolution Using an Oracle, Proceedings of the VLDB Endowment (PVLDB) 9 (</article-title>
          <year>2016</year>
          )
          <fpage>384</fpage>
          -
          <lpage>395</lpage>
          . doi:
          <volume>10</volume>
          .14778/2876473.2876474.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>G.</given-names>
            <surname>Simonini</surname>
          </string-name>
          , G. Papadakis,
          <string-name>
            <given-names>T.</given-names>
            <surname>Palpanas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          ,
          <article-title>Schema-agnostic Progressive Entity Resolution</article-title>
          ,
          <source>in: Proceedings of the International Conference on Data Engineering (ICDE)</source>
          ,
          <source>IEEE Computer Society</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>53</fpage>
          -
          <lpage>64</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICDE.
          <year>2018</year>
          .
          <volume>00015</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>L.</given-names>
            <surname>Gazzarri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Herschel</surname>
          </string-name>
          ,
          <article-title>Progressive Entity Resolution over Incremental Data</article-title>
          ,
          <source>in: Proceedings of the International Conference on Extending Database Technology (EDBT)</source>
          ,
          <source>OpenProceedings.org</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>80</fpage>
          -
          <lpage>91</lpage>
          . doi:
          <volume>10</volume>
          .48786/edbt.
          <year>2023</year>
          .
          <volume>07</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>H.</given-names>
            <surname>Altwaijry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mehrotra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <article-title>QuERy: A Framework for Integrating Entity Resolution with Query Processing</article-title>
          ,
          <source>Proceedings of the VLDB Endowment (PVLDB) 9</source>
          (
          <year>2015</year>
          )
          <fpage>120</fpage>
          -
          <lpage>131</lpage>
          . doi:
          <volume>10</volume>
          .14778/2850583.2850587.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>