<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Unsupervised Product Entity Resolution using Graph Representation Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mozhdeh Gheini</string-name>
          <email>gheini@isi.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mayank Kejriwal</string-name>
          <email>kejriwal@isi.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>USC Information Sciences Institute</institution>
          ,
          <addr-line>Marina del Rey, California</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <abstract>
        <p>Entity Resolution (ER) is defined as the algorithmic problem of determining when two or more entities refer to the same underlying entity. In the e-commerce domain, the problem tends to arise when the same product is advertised on multiple platforms, but with slightly (or even very) diferent descriptions, prices and other attributes. While ER has been well-explored for domains like bibliographic citations, biomedicine, patient records and even restaurants, work on product ER is not as prominent. In this paper, we report preliminary results on an unsupervised product ER system that is simple and extremely lightweight. The system is able to reduce mean rank reductions on some challenging product ER benchmarks by 50-70% compared to a text-only benchmark by leveraging a combination of text and neural graph embeddings.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Information systems → Clustering and classification.</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        With the proliferation of datasets on the Web, it is not
uncommon to find the same entity in multiple datasets, all referred to in
slightly different ways [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Entity Resolution (ER) is the
algorithmic problem of determining when two or more entities refer to
the same underlying entity [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. ER is a difficult problem that
has been studied for more than 50 years, in fields as wide-ranging
as biomedicine, movies, bibliographic citations and census records
[
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Although human level performance has not been achieved,
considerable progress has been made, across research areas as
diverse as Semantic and World Wide Web, knowledge discovery and
databases.
      </p>
      <p>A particular domain-specific kind of ER that is extremely
relevant to e-commerce is product ER. In the e-commerce domain, the
ER problem tends to arise when the same product is advertised on
multiple platforms, but with slightly (or even very) diferent
descriptions, prices and other attributes. Product ER can be dificult both
because of the wide variety of available products and e-commerce
platforms, but also because of artifacts like missing and noisy data,
especially when the data has been acquired in the first place by
crawling and scraping webpages, followed by semi-automatic
information extraction techniques like wrapper induction. The need
is for an unsupervised, lightweight solution that, given a query
entity, can rank candidate entities in other datasets and platforms
in descending order of match probability. For an unsupervised ER
to be truly viable, a matching entity, if one exists, generally needs
to be in the top 5, on average. Text-based methods like tf-idf are not
currently able to achieve this, however, on popular benchmarks, as
we illustrate subsequently.</p>
      <p>
        The core contribution in this paper is a Product Entity Resolution
system that is unsupervised, lightweight and that uses a
combination of text and graph-theoretic techniques to leverage not just a
description of the product, but also the context of the dataset in
which it occurs, to make an intelligent matching decision. The core
intuition is to model each product entity as a node in a graph, with
other nodes (e.g., prices, manufacturer) representing the
information set of the entity. Since more than one entity can have the same
price, manufacturer etc., these entities end up sharing context. Next,
once the entities and their information sets have been represented
in an appropriate graph-theoretic way, we embed the nodes in the
graph using a well-known neural embedding algorithm like
DeepWalk [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. An embedding in this context is a dense, real-valued,
low-dimensional vector. The goal is to embed the product entity
nodes in such a way, without using any training data, as to ensure
that nodes with similar embeddings (in a cosine similarity space)
will have high likelihood of being matches.
      </p>
      <p>Our guiding hypothesis in this paper is that neither a
puretext approach nor graph embedding, by itself, will yield optimal
performance on the product ER problem. Rather, they will have to
be combined in some way to yield a low mean rank. We conduct a
set of experiments to show that this is indeed the case. By using two
well-known and challenging benchmarks against various baselines,
we demonstrate that both text and graph-theoretic approaches can
together contribute to an average mean rank of less than 5, making
such a system closer to viable deployment.</p>
      <p>The rest of this paper is structured as follows. Section 2 covers
some relevant related work, while Section 3 outlines our approach.
We describe empirical results in Section 4, with Section 5 concluding
the paper.
2</p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>This paper draws on work in two broad areas of research, namely
entity resolution and graph embeddings. Below, we individually
cover pertinent aspects of these fields.
2.1</p>
    </sec>
    <sec id="sec-4">
      <title>Entity Resolution</title>
      <p>
        Entity Resolution (ER) is a problem that has been around for more
than 50 years in the AI literature [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], with the earliest versions of the
problem concerning the linking of patient records. More recently,
both link prediction and entity resolution (ER) were both recognized
as important steps in the overall link mining community about a
decade ago [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In the Semantic Web community, instance matching
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], link discovery [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] and class matching [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] are specific
examples of such sparse edge-discovery tasks. Other applications
include protein structure prediction (bioinformatics) [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ],
clickthrough rate prediction (advertising) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], social media and network
science [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Good overviews of ER were provided both by
Getoor and Machanavajjhala [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and in the book by Christen [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        A variety of recent papers have started to consider product
datasets in the suite of benchmarks that they evaluate. For
example, [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] considers both Amazon-Google Products and Abt-Buy,
the two publicly available benchmarks that we also consider in
this paper, in their evaluation. Performance on these datasets is
considerably lower, even for supervised systems, compared to
bibliographic datasets like DBLP and ACM, demonstrating that the
problem deserves to be looked at in a domain-specific way for
higher performance. Recently, Zhu et al. [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] was able to achieve
better performance using a high-complexity k-partite graph
clustering algorithm that was also unsupervised. However, they did
not consider the problem from the purview of ranking or IR and
their method had considerable runtime and modeling complexity.
In contrast, our method is lightweight and directly uses IR metrics
to evaluate.
      </p>
      <p>
        Broadly speaking, a typical ER pipeline consists of two phases:
blocking and matching [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Blocking is motivated by the fact that,
in the worst case, one would have to do a pairwise comparison
between every pair of entities to determine the matching pairs.
In this vein, blocking refers to a set of ‘divide-and-conquer’-style
techniques for approximately grouping a set of entities so that
pairwise comparisons (matching), which are more expensive, are only
conducted on pairs of entities that share a block. Using an indexing
function known as a blocking scheme, a blocking algorithm clusters
approximately similarly entities into (possibly overlapping)
clusters known as blocks. Only entities sharing a block are candidates
for further analysis in the second similarity step. State-of-the-art
similarity algorithms in various communities are now framed in
terms of machine learning, typically as binary classification [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>An alternative to a batch blocking-matching workflow is to adopt
an IR-centric workflow whereby a list of candidate entities needs
to be ranked when given a query entity. We adopt this IR-centric
workflow, since we recognize that, for an enterprise-grade system,
manual perusal and tuning will be necessary unless the accuracy of
the approach is very high. Hence, blocking does not apply. However,
as we show in the experimental section, a bag-of-words model can
be used for blocking-like pruning of entities that are not likely
to be matches, leading to significantly higher performance when
combined with the graph-theoretic approach.
2.2</p>
    </sec>
    <sec id="sec-5">
      <title>Graph Embeddings</title>
      <p>With the advent of neural networks, representation learning has
emerged as an influential area of research in modern times. As input,
neural networks need distributed representations of raw input that
are able to capture the similarity between diferent data points.</p>
      <p>
        For instance, in the Natural Language Processing (NLP)
community, the use of word representations goes back to as early as
1986, as detailed in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. More recently, the benefits of
using distributed representations for words to statistical language
modeling, a fundamental task in NLP, was shown in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The need
for these representations have inspired and given rise to diferent
word embeddings, such as word2vec [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], GloVe [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], and FastText
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], without which many current NLP systems would not have been
able to perform as well.
      </p>
      <p>
        Other forms of inputs, besides textual inputs, are no exception
when it comes to the need for rich distributed vector space
representations. Hence in the graph community, similar ideas have
been adopted to embed graph nodes and edges. In this work, we
use Deep Walk [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Like with word embeddings, numerous other
graph embedding algorithms have been proposed including, but not
limited to node2vec [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], GraphEmbed [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ], and LINE [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
Conceptually, these could be substituted for DeepWalk in this paper, an
option we are exploring in future research.
3
      </p>
    </sec>
    <sec id="sec-6">
      <title>APPROACH</title>
      <p>Graphs are an important representation and modeling tool in many
domains and for many problems, ranging from social media and
networks applications to information science. With the advent of
machine learning approaches and their conformity to vector inputs,
we need to transform graphs into vector spaces to be able to have
the best of both worlds: the richness of graph representations and
efectiveness of machine learning approaches.</p>
      <p>
        Graph embeddings are low-dimensional vector representations
of graphs that try to represent the structure of the graph to guide
analytical problem solving but without sacrificing eficiency [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. As
outlined and discussed in detail in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], there exist diferent graph
embedding methods and techniques depending on the information
set of a node that needs to be preserved. For the product ER
problem, we hypothesize that the second-order proximity of the nodes
(which is the similarity between nodes’ neighborhoods) is an
important information set, considering that the nodes themselves (the
product names) only contain limited information and are inherently
ambiguous. In this paper, we adopt a random walk-based approach,
where the goal is to encode a node’s information set by collecting
a set of random walks starting from that node.
      </p>
      <p>
        More specifically, we use DeepWalk for this purpose [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], which
has been made available on GitHub1. DeepWalk is heavily inspired
by sequence embedding architectures like word2vec in the NLP
community. In DeepWalk, each node of the graph roughly
corresponds to a ‘word’, and hence, each random walk can be considered
as a sentence in a language. Then, a neural model from the NLP
ifeld (in this case Skip-gram [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]) is used to capture second-order
1https://github.com/phanein/deepwalk
proximity information by ensuring that nodes that share a
similar context (in this case, neighborhoods) would achieve vector
embeddings that are close together in a cosine similarity space.
      </p>
      <p>Before DeepWalk (or any graph embedding for that matter) can
be employed, a product dataset, generally represented as key-value
or semi-structured table with missing values, must be represented
as a graph. To come up with a graph representation for a given
dataset, we investigate two diferent strategies:
• Simple Nodes: In this method, we assign a node to each
product entity (‘row’ in the table) as well as to each attribute
(a ‘cell’ of that row e.g., the price of the product). Entity nodes
are then linked to their corresponding to attribute nodes. So
for instance, if ‘Apple iPhone 6’ is $600, the entity node
corresponding to Apple iPhone 6 is linked to the attribute
node representing the price value of $600. Note that if some
other product similarly has a price of $600, it would also be
linked to that node. For text attributes (such as description),
we model a separate node for each unique token in the text
value and link the entity node to all tokens occurring in its
text attribute value (modeled as a bag of words). In essence,
this setting can be seen as a collection of star graphs where
entity nodes are the centers of local star graphs.
• Aggregated Nodes: This method is very similar to the
previous one except that we treat prices specially by combining
similar prices together and representing them with a single
node. Specifically, in our experiments, we divide prices into
bins of width $5 to achieve such a grouping. This is an
example of domain-specific graph representation, since prices
are clearly important when deciding whether to link entities.
Although we have only considered one such grouping in this
paper, other such groupings are also possible (and not just
for prices), a possibility we are exploring in future research.
By way of example, consider the two product entities in Table 1. We
now consider the graph representations for these sample records
using the two strategies above. Figure 1 shows the Simple Nodes
representation, and Figure 2 the Aggregated Nodes representation2.
In both figures, entity nodes have dashed borders and attribute
nodes have solid borders.</p>
      <p>Note that, because we are only considering unsupervised Entity
Resolution in this paper, there are no direct edges linking entities
together (only second-order edges, where entities are linked via a
token or price node etc.).</p>
      <p>Once the product datasets have been modeled in this way, they
are embedded using DeepWalk. The result is an embedding (a
continuous, real-valued vector) for each node in the graph i.e. we get
a vector for attributes, tokens and entities. Rather than directly
compute matches, we take an IR-centric approach whereby, for
each entity in the test set for which a match exists, we generate a
ranking between that source entity and all other entities (candidates)
by computing the cosine similarities between the vectors of the
source entity and each candidate entity, followed by the ranking
of the candidate entities by using the scores in descending order.
Using the withheld gold standard, we know the rank of the ‘true’
entity matching the source entity, which is used to compute metrics
2In the figure, we assume a bin width of $10 for illustration purposes compared to the
actual experimental bin width of $5.
We conducted a preliminary set of experiments to explore answers
to two research questions:</p>
      <p>First, How well does the graph embedding (whether on the
simple or aggregated nodes representation) do on the product ER
problem compared to traditional text-based baselines? Is the graph
embedding adequate by itself?</p>
      <p>Second, Does the aggregated nodes representation help
compared to the simple nodes representation?</p>
      <p>
        We conduct a preliminary set of experiments to examine our
approach on two benchmark datasets from the e-commerce domain:
Amazon-Google Products and Abt-Buy Products [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
AmazonGoogle Products, as the name suggests, contains product entities
from the online retailers Amazon and Google; similarly, Abt-Buy
features products from two diferent sources. Besides the record
‘ID’, each record in both datasets has ‘title’, ‘description’,
‘manufacturer’, and ‘price’ fields. In the Amazon-Google dataset 1,363
Amazon records need to compared against 3,226 Google records,
with the ground truth containing 1,300 matching pairs. In the
AbtBuy dataset, 1,081 Abt records need to be compared against 1,092
Buy records with the ground truth containing 1,097 matching pairs.
tf-idf+Cosine Similarity. In countless prior empirical studies
across IR, tf-idf has continued to work well as a baseline. We use it
as a feature-engineered, word-centric baseline that we can compare
our graph embedding based method to. Specifically, we tokenize the
values of the ‘description’ fields, since the product descriptions are
highly indicative features. Next, for each record, we obtain a tf-idf
representation based on the tokens we obtained from the product
descriptions. For each record in Dataset 1 (which could be Amazon
or Google in Amazon-Google Products, and similarly, Abt or Buy in
Abt-Buy), we rank all the records in Dataset 2 by using the cosine
similarity between the records’ tf-idf representations.
      </p>
      <p>Graph Embedding+Cosine Similarity. This method is similar
to the above, except that we use the cosine similarity between the
graph embeddings of the entity nodes.</p>
      <p>tf-idf Re-ranking. For the previous two baselines, we got the
ranks from tf-idf and graph embedding methods separately.
However, we consider a ‘two-level’ approach also, whereby we re-rank
the top 100 entries in the original ranked list obtained by the tf-idf
method using the graph embedding+cosine similarity scores. In
this setting, the tf-idf serves as a ‘pruning’ mechanism (weeding
out everything except the top 100 entries), similar to blocking
algorithms in the ER literature, with graph embeddings having the
ifnal say.
4.2</p>
    </sec>
    <sec id="sec-7">
      <title>Metrics</title>
      <p>We use mean rank to evaluate our methods. For a given query record
(an ‘entity’), we consider the penalty to be the rank of the match to
that query based on the gold truth in the ranking of the
similaritybased method (whether based on graph embeddings, tf-idf or their
combination). For the Abt-Buy benchmark, query records can be
from either Abt or Buy (conversely, candidate entities that are
ranked for the query would be from Buy or Abt), while for
AmazonGoogle Products we only tested using Google records as queries for
the Simple Node Representation. For a given dataset, we average
this penalty for all records to get the mean rank. For example, if A
perfect entity resolver would therefore achieve a mean rank of 1.
4.3</p>
    </sec>
    <sec id="sec-8">
      <title>Results and Discussion</title>
      <p>In response to the first question, we conducted a preliminary
experiment whereby on the Amazon-Google Products dataset, using
Amazon entities as queries (and the simple nodes representation),
we computed the mean rank for all three methods mentioned earlier.
The tf-idf baseline was found to achieve average mean rank of 11.01,
while the graph embedding achieved mean rank of 1719.49. Thus,
by itself, the graph embedding was not a lot better than random,
and clearly a lot less viable than the tf-idf based method. Next, we
computed the mean rank for the re-ranking method, and found the
average mean rank to be 3.51, a significant reduction from both of
the other two methods. Hence, when used for re-ranking, the graph
embeddings significantly improve the tf-idf ranking. Although
simple, we believe that this is the first time a purely text-based method
and a graph embedding have been used in conjunction for the
product ER problem, and outperformed both individually.</p>
      <p>Second, Table 2 tabulates the results for all four benchmark
settings (both datasets, using queries from Abt/Buy and Google/Amazon
respectively) with graph embeddings using the simple nodes
representation. These results show that DeepWalk, although very poor by
itself, can significantly improve results when applied over TFIDF to
rerank its ranking. This improvement is consistent across datasets
as Table 2 shows.</p>
      <p>As mentioned earlier, we study two graph construction methods
(simple nodes and aggregated nodes). Our second research question
specifically asked whether the aggregated nodes representation
can help improve performance compared to the simple nodes
representation. As there are many empty price fields in the Abt-Buy
dataset, and we focus on the price field for node aggregation, we
only use the Amazon-Google dataset in this part. The results of
reranking TFIDF when we use aggregated nodes graph representation
is shown in Table 3.</p>
      <p>Our results show that, compared to simple node representation,
there is no significant or consistent gain when we aggregate price
nodes. While results do improve (by 0.04) compared to the
simple nodes representation (when using Google product entities as
queries), there is a reversal when using Amazon product entities as
queries. One reason for this may be due to how we build the graph
with respect to text attributes. Recall that we tokenize the text in
the description field and create a node for each token in the graph.
An entity node is then linked to a token node if and only if that
token appears in one of its text attributes. With this setting, for
an entity node at the center of a star, most outgoing edges will be
focused on text-based features and the price attribute only accounts
for one outgoing edge.</p>
      <p>One option that we’re exploring in future work is to upsample
the price nodes so that they have a stronger presence (via more
occurrences) in the random walks. Another option that we’re
considering is using other aggregation strategies.
5</p>
    </sec>
    <sec id="sec-9">
      <title>CONCLUSION</title>
      <p>Product Entity Resolution is a dificult problem with the potential
for high commercial impact, even with modest increases in metrics</p>
      <p>Dataset</p>
      <p>Simple Nodes Representation</p>
      <p>Query tf-idf Mean tf-idf Re-ranked
Record Rank Mean Rank
Google 8.27 4.18
Amazon 11.01 3.51
Abt 9.39 2.66</p>
      <p>Buy 11.88 2.76
like Mean Rank. In this paper, we illustrated a simple ‘two-level’
scheme that leverages both text and graph information to reduce
the mean rank on some competitive benchmarks from more than
11 to less than 5. Although the aggregated nodes representation
was found not to have an impact, this is likely due to the
preliminary nature of our experiments, since we did not try many such
aggregated representations. In the future, we will explore more
options for aggregation, and will explore variants of the re-ranking
scheme, along with exploring more options for graph embeddings
(e.g., LINE, node2vec). We believe that the best approach will be a
graph embedding specifically optimized for products rather than a
generic approach like LINE or DeepWalk.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Mohammad</given-names>
            <surname>Al</surname>
          </string-name>
          <string-name>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Vineet</given-names>
            <surname>Chaoji</surname>
          </string-name>
          , Saeed Salem, and
          <string-name>
            <given-names>Mohammed</given-names>
            <surname>Zaki</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Link prediction using supervised learning</article-title>
          .
          <source>In SDMÃŋ06: Workshop on Link Analysis, Counter-terrorism and Security.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          , Réjean Ducharme, Pascal Vincent, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Jauvin</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>A neural probabilistic language model</article-title>
          .
          <source>Journal of machine learning research 3</source>
          ,
          <string-name>
            <surname>Feb</surname>
          </string-name>
          (
          <year>2003</year>
          ),
          <fpage>1137</fpage>
          -
          <lpage>1155</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Indrajit</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lise</given-names>
            <surname>Getoor</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Collective entity resolution in relational data</article-title>
          .
          <source>ACM Transactions on Knowledge Discovery from Data (TKDD) 1</source>
          ,
          <issue>1</issue>
          (
          <year>2007</year>
          ),
          <fpage>5</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>5</volume>
          (
          <year>2017</year>
          ),
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Hongyun</given-names>
            <surname>Cai</surname>
          </string-name>
          , Vincent W Zheng, and
          <string-name>
            <surname>Kevin</surname>
            <given-names>Chen-Chuan</given-names>
          </string-name>
          <string-name>
            <surname>Chang</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A comprehensive survey of graph embedding: Problems, techniques, and applications</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>30</volume>
          ,
          <issue>9</issue>
          (
          <year>2018</year>
          ),
          <fpage>1616</fpage>
          -
          <lpage>1637</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Christen</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Data matching: concepts and techniques for record linkage, entity resolution, and duplicate detection</article-title>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Ivan</surname>
            <given-names>P</given-names>
          </string-name>
          <string-name>
            <surname>Fellegi and Alan B Sunter</surname>
          </string-name>
          .
          <year>1969</year>
          .
          <article-title>A theory for record linkage</article-title>
          .
          <source>J. Amer. Statist. Assoc</source>
          .
          <volume>64</volume>
          ,
          <issue>328</issue>
          (
          <year>1969</year>
          ),
          <fpage>1183</fpage>
          -
          <lpage>1210</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Lise</given-names>
            <surname>Getoor and Christopher P Diehl</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Link mining: a survey</article-title>
          .
          <source>ACM SIGKDD Explorations Newsletter</source>
          <volume>7</volume>
          ,
          <issue>2</issue>
          (
          <year>2005</year>
          ),
          <fpage>3</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Lise</given-names>
            <surname>Getoor</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ashwin</given-names>
            <surname>Machanavajjhala</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Entity resolution: theory, practice &amp; open challenges</article-title>
          .
          <source>Proceedings of the VLDB Endowment 5</source>
          ,
          <issue>12</issue>
          (
          <year>2012</year>
          ),
          <fpage>2018</fpage>
          -
          <lpage>2019</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Thore</surname>
            <given-names>Graepel</given-names>
          </string-name>
          , Joaquin Q Candela, Thomas Borchert, and
          <string-name>
            <given-names>Ralf</given-names>
            <surname>Herbrich</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Web-scale bayesian click-through rate prediction for sponsored search advertising in microsoft's bing search engine</article-title>
          .
          <source>In Proceedings of the 27th International Conference on Machine Learning (ICML-10)</source>
          .
          <fpage>13</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Aditya</given-names>
            <surname>Grover</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jure</given-names>
            <surname>Leskovec</surname>
          </string-name>
          .
          <year>2016</year>
          . node2vec:
          <article-title>Scalable feature learning for networks</article-title>
          .
          <source>In Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining. ACM</source>
          ,
          <volume>855</volume>
          -
          <fpage>864</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Geofrey</surname>
            <given-names>E</given-names>
          </string-name>
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          et al .
          <year>1986</year>
          .
          <article-title>Learning distributed representations of concepts</article-title>
          .
          <source>In Proceedings of the eighth annual conference of the cognitive science society</source>
          , Vol.
          <volume>1</volume>
          . Amherst, MA,
          <volume>12</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Mayank</given-names>
            <surname>Kejriwal</surname>
          </string-name>
          and Daniel P Miranker.
          <year>2013</year>
          .
          <article-title>An unsupervised algorithm for learning blocking schemes</article-title>
          .
          <source>In Data Mining (ICDM)</source>
          ,
          <source>2013 IEEE 13th International Conference on. IEEE</source>
          ,
          <fpage>340</fpage>
          -
          <lpage>349</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Lawrence</surname>
            <given-names>A Kelley</given-names>
          </string-name>
          and
          <source>Michael JE Sternberg</source>
          .
          <year>2009</year>
          .
          <article-title>Protein structure prediction on the Web: a case study using the Phyre server</article-title>
          .
          <source>Nature protocols 4</source>
          ,
          <issue>3</issue>
          (
          <year>2009</year>
          ),
          <fpage>363</fpage>
          -
          <lpage>371</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Hanna</given-names>
            <surname>Köpcke</surname>
          </string-name>
          and
          <string-name>
            <given-names>Erhard</given-names>
            <surname>Rahm</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Frameworks for entity matching: A comparison</article-title>
          .
          <source>Data &amp; Knowledge Engineering</source>
          <volume>69</volume>
          ,
          <issue>2</issue>
          (
          <year>2010</year>
          ),
          <fpage>197</fpage>
          -
          <lpage>210</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Hanna</surname>
            <given-names>Köpcke</given-names>
          </string-name>
          , Andreas Thor, and
          <string-name>
            <given-names>Erhard</given-names>
            <surname>Rahm</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Evaluation of entity resolution approaches on real-world match problems</article-title>
          .
          <source>Proceedings of the VLDB Endowment 3</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          (
          <year>2010</year>
          ),
          <fpage>484</fpage>
          -
          <lpage>493</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Linyuan</given-names>
            <surname>Lü</surname>
          </string-name>
          and
          <string-name>
            <given-names>Tao</given-names>
            <surname>Zhou</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Link prediction in complex networks: A survey</article-title>
          .
          <source>Physica A: Statistical Mechanics and its Applications</source>
          <volume>390</volume>
          ,
          <issue>6</issue>
          (
          <year>2011</year>
          ),
          <fpage>1150</fpage>
          -
          <lpage>1170</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Tomas</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jefrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Eficient estimation of word representations in vector space</article-title>
          .
          <source>arXiv preprint arXiv:1301.3781</source>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Tomas</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg S Corrado, and
          <string-name>
            <given-names>Jef</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>3111</volume>
          -
          <fpage>3119</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Axel-Cyrille Ngonga Ngomo</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>A time-eficient hybrid approach to link discovery</article-title>
          .
          <source>Ontology Matching</source>
          (
          <year>2011</year>
          ),
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Jefrey</surname>
            <given-names>Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Glove: Global vectors for word representation</article-title>
          .
          <source>In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP)</source>
          .
          <volume>1532</volume>
          -
          <fpage>1543</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Bryan</surname>
            <given-names>Perozzi</given-names>
          </string-name>
          , Rami Al-Rfou, and
          <string-name>
            <given-names>Steven</given-names>
            <surname>Skiena</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>DeepWalk: Online Learning of Social Representations</article-title>
          .
          <source>In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '14)</source>
          . ACM, New York, NY, USA,
          <fpage>701</fpage>
          -
          <lpage>710</lpage>
          . https://doi.org/10.1145/2623330.2623732
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>David</surname>
            <given-names>E Rumelhart</given-names>
          </string-name>
          , Geofrey E Hinton,
          <string-name>
            <surname>Ronald J Williams</surname>
          </string-name>
          , et al .
          <year>1988</year>
          .
          <article-title>Learning representations by back-propagating errors</article-title>
          .
          <source>Cognitive modeling 5</source>
          ,
          <issue>3</issue>
          (
          <year>1988</year>
          ),
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Salvatore</surname>
            <given-names>Scellato</given-names>
          </string-name>
          , Anastasios Noulas, and
          <string-name>
            <given-names>Cecilia</given-names>
            <surname>Mascolo</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Exploiting place features in link prediction on location-based social networks</article-title>
          .
          <source>In Proceedings of the 17th ACM SIGKDD international conference on Knowledge discovery and data mining. ACM</source>
          ,
          <volume>1046</volume>
          -
          <fpage>1054</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Shvaiko</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jérôme</given-names>
            <surname>Euzenat</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Ontology matching: state of the art and future challenges. Knowledge and Data Engineering</article-title>
          , IEEE Transactions on
          <volume>25</volume>
          ,
          <issue>1</issue>
          (
          <year>2013</year>
          ),
          <fpage>158</fpage>
          -
          <lpage>176</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Jian</surname>
            <given-names>Tang</given-names>
          </string-name>
          , Meng Qu, Mingzhe Wang, Ming Zhang, Jun Yan, and
          <string-name>
            <given-names>Qiaozhu</given-names>
            <surname>Mei</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Line: Large-scale information network embedding</article-title>
          .
          <source>In Proceedings of the 24th international conference on world wide web. International World Wide Web Conferences Steering Committee</source>
          ,
          <fpage>1067</fpage>
          -
          <lpage>1077</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Julius</surname>
            <given-names>Volz</given-names>
          </string-name>
          , Christian Bizer,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Gaedke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Georgi</given-names>
            <surname>Kobilarov</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Discovering and maintaining links on the web of data</article-title>
          .
          <source>In The Semantic Web-ISWC 2009</source>
          . Springer,
          <fpage>650</fpage>
          -
          <lpage>665</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>William</surname>
            <given-names>E</given-names>
          </string-name>
          <string-name>
            <surname>Winkler</surname>
            and
            <given-names>Yves</given-names>
          </string-name>
          <string-name>
            <surname>Thibaudeau</surname>
          </string-name>
          .
          <year>1991</year>
          .
          <article-title>An application of the Fellegi-Sunter model of record linkage to the 1990 US decennial census</article-title>
          .
          <source>Citeseer.</source>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Chao</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Keyang Zhang, Quan Yuan, Haoruo Peng, Yu Zheng, Tim Hanratty,
          <string-name>
            <given-names>Shaowen</given-names>
            <surname>Wang</surname>
          </string-name>
          , and Jiawei Han.
          <year>2017</year>
          . Regions, Periods, Activities:
          <article-title>Uncovering Urban Dynamics via Cross-Modal Representation Learning</article-title>
          .
          <source>In Proceedings of the 26th International Conference on World Wide Web (WWW '17)</source>
          .
          <source>International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, Switzerland</source>
          ,
          <fpage>361</fpage>
          -
          <lpage>370</lpage>
          . https://doi.org/10.1145/3038912.3052601
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Linhong</surname>
            <given-names>Zhu</given-names>
          </string-name>
          , Majid Ghasemi-Gol, Pedro Szekely,
          <source>Aram Galstyan, and Craig A Knoblock</source>
          .
          <year>2016</year>
          .
          <article-title>Unsupervised entity resolution on multi-type graphs</article-title>
          .
          <source>In International semantic web conference</source>
          . Springer,
          <fpage>649</fpage>
          -
          <lpage>667</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>