<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Progressive Search-driven Entity Resolution</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alberto Pietrangelo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Simonini</string-name>
          <email>giovanni.simonini@unimore.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sonia Bergamaschi</string-name>
          <email>sonia.bergamaschi@unimore.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ioannis Koumarelas</string-name>
          <email>ioannis.koumarelas@hpi.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felix Naumann</string-name>
          <email>felix.naumann@hpi.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hasso Plattner Institute, University of Potsdam</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universita degli Studi di Modena e Reggio Emilia</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Keyword-search systems for databases aim to answer a user query composed of a few terms with a ranked list of records. They are powerful and easy-to-use data exploration tools for a wide range of contexts. For instance, given a product database gathered scraping e-commerce websites, these systems enable even non-technical users to explore the item set (e.g., to check whether it contains certain products or not, or to discover the price of an item). However, if the database contains dirty records (i.e., incomplete and duplicated records), a preprocessing step to clean the data is required. One fundamental data cleaning step is Entity Resolution, i.e., the task of identifying and fusing together all the records that refer to the same real-word entity. This task is typically executed on the whole data, independently of: (i) the portion of the entities that a user may indicate through keywords, and (ii) the order priority that a user might express through an order by clause. This paper describes a rst step to solve the problem of progressive search-driven Entity Resolution: resolving all the entities described by a user through a handful of keywords, progressively (according to an order by clause). We discuss the features of our method, named SearchER and showcase some examples of keyword queries on two real-world datasets obtained with a demonstrative prototype that we have built.</p>
      </abstract>
      <kwd-group>
        <kwd>Keyword search</kwd>
        <kwd>Entity Resolution</kwd>
        <kwd>Data Cleaning</kwd>
        <kwd>Pay- as-you-go</kwd>
        <kwd>Query-driven Data Integration</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Entity Resolution (ER) is a fundamental task for data cleaning and
integration [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ][
        <xref ref-type="bibr" rid="ref7">7</xref>
        ][
        <xref ref-type="bibr" rid="ref11">11</xref>
        ][
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]: it aims to identify di erent representations of the same
realworld entity in a given dataset. Typically, an ER work ow is composed of three
main sub-tasks: blocking, matching, and resolution. The listed steps may depend
on other tasks themselves; e.g.: blocking may require schema-alignment for the
blocking rules de nition; the match function may require the generation of a
labeled training set to properly train a classi er, etc. We adopt this simpli
cation for the sake of the presentation. Blocking is typically employed to avoid the
quadratic complexity of the nave solution (which compares all possible pairs
of records) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Basically, blocking provides the set of candidate pairs that are
actually compared through a match function, i.e., a binary function that takes
as input two records r1 and r2 and answers the question \do r1 and r2 refer to
the same real-world entity?' '|we say that r1 and r2 are matching (r1 r2) if the
answer is yes, non-matching (r16 r2) otherwise. Finally, a resolve function takes
as input all the records referring to a single entity (i.e., a set of matching records)
and returns a single representative record, resolving con icts of matches (e.g.,
it may result that r1 r2, r2 r3, but r16 r3), and inconsistent attribute values
(e.g., r1 r2, but some attributes have di erent values). This task is also known
as Data Fusion.
      </p>
    </sec>
    <sec id="sec-2">
      <title>The Challenge</title>
      <p>
        When the computational resources and/or the time are critical components for
ER [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], progressive ER aims to yield record pairs progressively, trying to
maximize the recall in case of early termination [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ][
        <xref ref-type="bibr" rid="ref14">14</xref>
        ][
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Yet, these techniques are
speci cally designed to yield any match as soon as possible, without considering
any indication of the user (i.e., the query); adapting them to the progressive
search-driven problem, in a keyword search system [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], is not trivial. This can
be particularly useful for data exploration [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Consider the following example.
Example: Given a dirty dataset of e-commerce products gathered from several
sources on the Web, say that a user wants to nd all the entities (i.e., resolved
records) referring to Apple iPhone 8 smartphones, ordered by decreasing price.
(This is represented in the keyword-based query in Query 1.) Furthermore, say
that the user has a limited time budget for such a task and wants to retrieve as
many entities as possible within that budget.
Query. 1. Keyword query issued to retrieve all the \Apple iPhone 8" in the database,
ordered by decreasing price.
      </p>
      <p>Employing a traditional batch approach to ER, i.e., resolving the whole
dataset and then executing the query, might be quite expensive on real datasets
composed of millions of entities. Even existing progressive ER approaches are
useless with such a problem, since they aim to approximate the optimal order
of the record comparisons on the basis of the matching likelihood of the record
pairs|the matching likelihood is assessed through heuristics, such as the
similarity of some attributes. Thus, to generate the nal result for Query 1, they
have to perform the whole ER process; this is because the most expensive Apple
iPhone 8 entity might correspond to records that have the lowest matching
likelihood in the dataset.</p>
    </sec>
    <sec id="sec-3">
      <title>Our approach</title>
      <p>In this paper, we investigate the problem of the progressive search-driven Entity
Resolution. We propose a rst attempt to address this problem by envisioning
an ER method, called SearchER, which enables users to express the relevance
of the entities of their interest by means of a keyword query combined with the
selection of an attribute that determines the ordering (descending or ascending).</p>
      <p>From the user point of view, SearchER takes as input: (i) a dirty dataset,
(ii) a blocking function, (iii) a user keyword query (which de nes the entities of
interest), (iv) an attribute for the ordering predicate, and (v) the ordering type
(i.e., ascending or descending ). Then, SearchER returns as output the solution
for the user query, progressively. Under the hood SearchER exploits the blocking
function to de ne the space of possible comparisons; then it identi es some
initial candidate records to be resolved (i.e., those containing the keywords) and
iteratively explores comparisons that involve these records. For ordering entities,
SearchER considers not only the ordering attribute, but also the number of
keywords that are associated to the entities. For example, considering Query 1,
it may happen that an entity with a low price containing all the keywords is
emitted before an entity with a high price, but that does not contain all the
keywords. In other words, we assume that the keywords are more important
than the ordering for the user. For this reason, we say that the nal solution is
approximate, since the nal ordering might not be completely respected.</p>
      <p>We also built a rst experimental prototype of SearchER and tested it with
some queries on two real-world datasets. This very preliminary result does not
intend to show the e cacy of our method in general, yet it allows us to showcase
that such an approach to ER is actually feasible and promising.</p>
      <p>The reminder of this paper is organized as follows: Section 2 introduces the
preliminaries; Section 3 describes the envisioned SearchER method; Section 4
reports our preliminary experiments on two real-world datasets; Section 5
describes main related work; nally, Section 6 concludes the paper and discusses
the ongoing and future work.
2</p>
      <sec id="sec-3-1">
        <title>Preliminaries</title>
        <p>
          To perform ER, the nave comparison of all possible pairs of records has a
quadratic complexity, thus for scaling to large datasets blocking techniques are
generally employed [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]: the records are indexed according to heuristics (called
blocking criteria) into clusters (possibly overlapping) and the all-pair
comparison is executed only within each cluster (a.k.a. block ). Thus, the e cacy of
blocking (i.e., how many matches are indexed in the blocks) strictly depends on
the de nition of the blocking criteria. Intuitively, large and overlapping blocks
capture more matches than small and non-overlapping ones, but at the expense
of e ciency.
        </p>
        <p>An approach that has been shown to achieve high accuracy is meta-blocking,
which operates on large and overlapping blocks, restructuring them to lter out
non promising comparisons.</p>
        <p>Meta-blocking relies on the assumption that the matching likelihood of any
two records is analogous to their degree of co-occurrence in a block collection.
This means that a block collection B has to be generated by a blocking method
that yields redundancy-positive blocks, where the similarity of two records is
proportional to the number of blocks they share.</p>
        <p>
          Based on redundancy, which is common for blocking methods [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ],
metablocking represents the block collection as a blocking graph. This is an undirected
weighted graph GB(VB; EB), where VB is the set of nodes, and EB is the set of
weighted edges. Every node ni 2 VB represents a record ri 2 R, while every edge
ei;j represents a comparison ci;j 2 B R R. A weighting function is employed
to weight the edges, leveraging the co-occurrence patterns of records in B: each
edge is assigned a weight that is derived exclusively from the (characteristics
of the) blocks its adjacent records have in common. For example, the ARCS
function sums the inverse cardinality of common blocks, assigning higher scores
to pairs of records sharing smaller (i.e., more distinctive) blocks.
        </p>
        <p>For the problem of progressive search-driven ER, we propose a revised and
extended version of the blocking graph model, as explained in Section 3.
3</p>
      </sec>
      <sec id="sec-3-2">
        <title>The SearchER</title>
      </sec>
      <sec id="sec-3-3">
        <title>Method</title>
        <p>At its core, SearchER employs a blocking function to generate the block collection
that is exploited for building blocking graph. In the following we describe the
novel node-weighting and edge-weighting strategies employed by SearchER for
building the blocking graph, and nally we describe how the blocking graph is
employed by SearchER for the query evaluation.</p>
        <p>Revised node weighting: The nodes are weighted according to the
likelihood of appearing in the nal solution. This introduces the concept of
RecordRelevance; the intuition is explained with the following example:
Example: Consider Query 1, issued by a user that is looking for iPhone entities
and requires the results to be generated progressively, from the most expensive to
the cheapest ones. Say that a record r1 refers to entity "1 and has a high price.
Say also that r1 has many edges connecting it to many records with low prices,
and that these edges have a high matching likelihood. So, it is likely that the nal
(resolved) price of "1 will not be high. This means that it is likely that "1 will not
belong to the nal solution. Hence, the Record-Relevance (i.e., the node-weight)
of r1 should be low.</p>
        <p>
          Record-Relevance computation: The Record-Relevance is assessed by
adapting tf-idf [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], a widely employed information retrieval measure, in the following
way: rstly, a weight is assigned to each node proportionally to the number of
In our preliminary experiment we employ Token Blocking [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], which considers each
token as a blocking key, regardless of the attribute in which it appears.
        </p>
        <p>
          Revised edge weighting: The weight of the edges captures the likelihood of
a ecting the nal solution. The weighting schema considers both the matching
likelihood of an edge (as in meta-blocking [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]), and its Edge-Relevance.
Example: Consider Query 1. Say that a record r1 is connected to only one record
r2 with its same price, and both belong to the real-world entity "1 (i.e., r1:year =
r2:year and r1 r2). Then, performing the comparison of r1 and r2 will not
change the rank of "1. Hence, the Edge-Relevance (i.e., the edge-weight) of the
edge connecting r1 and r2 should be low.
        </p>
        <p>The Edge-Relevance also depends on the resolve function employed by the user.
For the intuition consider the following example:
Example: Consider Query 1. Say that r1 (with a high price r1:year) is connected
to another record r2 (with r2:year r1:year). Then, the Record-Relevance of r1
(and r2) should change on the basis of the resolve function: if the resolve function
assigns MAX(r1:year; r2:year) as nal price of the entity, it is more likely for r1
(and r2) to be part of the answer for Query 1; which is not true if the resolve
function assigns MIN(r1:year; r2:year) as nal price.</p>
        <p>
          Edge-Relevance computation: The Edge-Relevance between two nodes is
computed as the weight in a traditional blocking graph, normalized by the relative
variation (e.g., price = ri:year rj:year) of the attribute value that is
employed for the ordering clause. The edge weight has to be stored also with a
sign that indicates the direction of the variation: a positive value means that
the value of the adjacent node ri with the smaller id is greater than the other
node rj with a higher id (i.e., i &lt; j)|this is just a convention, it could be the
other way around. Thus, the edge weight can be interpreted di erently on the
basis of the resolve function (e.g., MIN/MAX price of the matching records).
Query evaluation: Processing a user's query, the blocking graph is not built
in its entirety (i.e., considering all the records/nodes); instead, only the portion
of the graph that is valuable for the query is considered. In fact, given a block
collection, the node-centric subgraph of the blocking graph for any node ri (i.e.,
the subgraph involving ri and its neighbours only) can be e ciently built using
algorithms described in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Thus, SearchER considers only the
nodecentric subgraphs of the nodes that correspond to the records containing at least
one of the given keywords, and merges subgraphs that share nodes. Notice that
the resulting blocking graph can be disconnected.
        </p>
        <p>When looking at the obtained blocking graph, we observe in preliminary
experiments that for queries with a limited number of keywords (ideally the vast
majority), many nodes tend to have the same weights. SearchER divides them
into levels and resolves the records starting from the highest levels and sorts
the comparisons within each level according to their edge-weights. As soon as a
node is evaluated (i.e., all the comparisons involving it have been performed),
the resulting entity is emitted.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Preliminary experimental results</title>
        <p>
          We devised a prototype of SearchER that takes advantage of the proposed data
model and performed preliminary experiments by issuing keyword queries on
top of two well-known, real-world datasets (for which the ground-truth of the
matching records is known). The rst dataset is CDDB [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], which contains CD
entities described with: name, artist, category, genre, and year elds along with
the track titles. The second dataset is CORA [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], which contains computer science
research articles described with: a liation, author, location, title, venue and
year.
        </p>
        <p>In the following we report the results obtained for the following queries:
q1: \Rock"
q2: \Rock and Roll"
q3: \Genetic Algorithms Applications"
q4: \Arti cial Intelligence"
q5: \Jazz Music"
Ordering: Year
Sort: DESCENDING
issued on CDDB
issued on CDDB
issued on CORA
issued on CORA
issued on CDDB</p>
        <p>
          For each query, the entities are sorted by descending values of the attribute year for
both the datasets. As resolve function we employed M AX(ri:year; rj:year), yet very
similar results have been obtained employing M IN (ri:year; rj:year). In our evaluation
we consider as baselines: Standard Blocking (SB) and Sorted Neighbour (SN), with the
same con gurations of [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]: the ER process is executed employing these methods, and
then we performed the queries on top of the resolved entity set (using tf-idf for ranking
the results). The results for q1, q2, q3 and q4 are shown in Figure 2: it reports the
number of emitted entities that satisfy the corresponding query (y-axis), as function
of the number of pairwise comparisons performed (x-axis).
        </p>
        <p>The results show that SearchER starts emitting entities much earlier (in terms of
compared pairs) than the other methods; this is because traditional blocking techniques
have to wait until the last comparison in order to evaluate the user query. Surprisingly,
we observe that for some queries (q3 and q4 in Figure 2c,d), the baselines cannot identify
as many matches as SearchER. This is due to the underlying blocking techniques that
fail to yield some of the candidates that are actually matches, and which are correctly
identi ed as such in the blocking graph.</p>
        <p>Similarly, for query q5, we report the number of entities found, in function of the
number of comparisons (Figure 3a) and in function of the time (Figure 3b).
Showing that the advantage of SearchER is evident also considering the execution w.r.t.
execution time.
5</p>
      </sec>
      <sec id="sec-3-5">
        <title>Related Work</title>
        <p>
          Query Driven Approach|To avoid the ER evaluation on the whole dataset
before answering any query, recent works proposed a Query Driven Approach (QDA) to
ER [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ][
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], which aims to resolve only the portion of the entities in a dataset that are
We report that the proposed method provides an approximate solution: some of the
emissions are not in the corrected order and a tolerance range has been considered.
actually relevant to the user, by exploiting SPJ SQL clauses. The ORDER BY clause is
out of the scope of QDA.
        </p>
        <p>Thus, Query 1, where the relevance ordering is expressed through an ORDER BY
clause, will force QDA to resolve all the entities in the dataset and then sort them by
price. Furthermore, QDA strictly relies on batch blocking, which is employed as initial
step of the ER. Hence, adapting QDA to work in a pay-as-you-go fashion is not trivial.
6</p>
      </sec>
      <sec id="sec-3-6">
        <title>Conclusion and Future Work</title>
        <p>In this paper, we have presented a preliminary study of the problem of progressive
search-driven Entity Resolution, namely the task of deduplicating records of a dirty
dataset by following a user query that expresses: (i) the keywords representing the
entities of interest (e.g., \Apple iPhone 8"); (ii) the ordering of interest (e.g., from
the most expensive to the cheapest). We have proposed a rst solution for solving this
problem, and proposed an approximate method, called SearchER.</p>
        <p>We believe that the proposed method paves the way for further investigation of
the problem. In particular, we are currently devising a method to provide an exact</p>
        <p>
          Pietrangelo et al.
solution for a wide range of resolution function; e.g., w.r.t. Query 1:
minimum=maximum=user def ined f unction([price]) for determining the nal price of an
entity. We are also investigating how to exploit di erent blocking techniques (other than
Standard Blocking and meta-blocking) and advanced similarity functions for the
matching phase [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Finally, we will compare approximate solutions (such as SearchER) and
exact solutions, to thoroughly study their characteristics and de ning the trade-o s to
guide practitioners.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Altwaijry</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalashnikov</surname>
            ,
            <given-names>D.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mehrotra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Query-driven approach to entity resolution</article-title>
          .
          <source>PVLDB</source>
          <volume>6</volume>
          (
          <issue>14</issue>
          ),
          <year>1846</year>
          {
          <year>1857</year>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Altwaijry</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalashnikov</surname>
            ,
            <given-names>D.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mehrotra</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>QDA: A query-driven approach to entity resolution</article-title>
          .
          <source>IEEE TKDE 29(2)</source>
          ,
          <volume>402</volume>
          {
          <fpage>417</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Benedetti</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beneventano</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergamaschi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simonini</surname>
          </string-name>
          , G.:
          <article-title>Computing interdocument similarity with context semantic analysis</article-title>
          .
          <source>Information Systems</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bergamaschi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferrari</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guerra</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simonini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velegrakis</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Providing insight into data source topics</article-title>
          .
          <source>J. Data Semantics</source>
          <volume>5</volume>
          (
          <issue>4</issue>
          ),
          <volume>211</volume>
          {
          <fpage>228</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bergamaschi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guerra</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simonini</surname>
          </string-name>
          , G.:
          <article-title>Keyword search over relational databases: Issues, approaches and open challenges</article-title>
          .
          <source>In: Bridging Between Information Retrieval and Databases</source>
          . pp.
          <volume>54</volume>
          {
          <issue>73</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Christen</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A survey of indexing techniques for scalable record linkage and deduplication</article-title>
          .
          <source>IEEE TKDE 24(9)</source>
          ,
          <volume>1537</volume>
          {
          <fpage>1555</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>X.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Big Data Integration</article-title>
          . Morgan &amp;
          <string-name>
            <surname>Claypool</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Getoor</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Machanavajjhala</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Entity resolution: Theory, practice &amp; open challenges</article-title>
          .
          <source>PVLDB</source>
          <volume>5</volume>
          (
          <issue>12</issue>
          ),
          <year>2018</year>
          {
          <year>2019</year>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Madhavan</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>X.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Halevy</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Je ery</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ko</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Web-scale data integration: You can a ord to pay as you go</article-title>
          .
          <source>In: CIDR</source>
          . pp.
          <volume>342</volume>
          {
          <issue>350</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Schutze, H.:
          <article-title>Introduction to information retrieval</article-title>
          . Cambridge University Press (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Naumann</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herschel</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>An Introduction to Duplicate Detection</article-title>
          .
          <source>Synthesis Lectures on Data Management</source>
          , Morgan &amp; Claypool Publishers (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Papadakis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papastefanatos</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palpanas</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koubarakis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Scaling entity resolution to large, heterogeneous data with enhanced meta-blocking</article-title>
          .
          <source>In: EDBT</source>
          . pp.
          <volume>221</volume>
          {
          <issue>232</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Papadakis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Svirsky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palpanas</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Comparative analysis of approximate blocking techniques for entity resolution</article-title>
          .
          <source>PVLDB</source>
          <volume>9</volume>
          (
          <issue>9</issue>
          ),
          <volume>684</volume>
          {
          <fpage>695</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Papenbrock</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heise</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naumann</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Progressive duplicate detection</article-title>
          .
          <source>IEEE TKDE 27(5)</source>
          ,
          <volume>1316</volume>
          {
          <fpage>1329</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Simonini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergamaschi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jagadish</surname>
            ,
            <given-names>H.V.</given-names>
          </string-name>
          :
          <article-title>BLAST: a loosely schema-aware meta-blocking approach for entity resolution</article-title>
          .
          <source>PVLDB</source>
          <volume>9</volume>
          (
          <issue>12</issue>
          ),
          <volume>1173</volume>
          {
          <fpage>1184</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Simonini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papadakis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palpanas</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergamaschi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Schema-agnostic progressive entity resolution</article-title>
          .
          <source>In: IEEE ICDE</source>
          . pp.
          <volume>53</volume>
          {
          <issue>64</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Simonini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Big data exploration with faceted browsing</article-title>
          .
          <source>In: HPCS</source>
          . pp.
          <volume>541</volume>
          {
          <issue>544</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Whang</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marmaros</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Molina</surname>
          </string-name>
          , H.:
          <article-title>Pay-as-you-go entity resolution</article-title>
          .
          <source>IEEE TKDE 25(5)</source>
          ,
          <volume>1111</volume>
          {
          <fpage>1124</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fiameni</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simonini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergamaschi</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>SOPJ: A scalable online provenance join for data integration</article-title>
          .
          <source>In: HPCS</source>
          . pp.
          <volume>79</volume>
          {
          <issue>85</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>