<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards explainable entity matching via comparison queries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alina Petrova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Egor V. Kostylev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bernardo Cuenca Grau</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ian Horrocks</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Oxford</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Nowadays there exists an abundance of heterogeneous Semantic Web data coming from multiple sources. As a result, matching Linked Data has become a tedious and non-transparent task. One way to facilitate entity matching across datasets is to provide human-readable explanations that highlight what the two entities have in common, as well as what differentiates the two entities. Entity comparison is an important information exploration problem that has recently gained considerable research attention [1, 2, 4]. In this paper we propose a solution towards explainable entity matching in Linked Data where entity comparison is used as a subroutine that assists in debugging and validation of matchings. To this end, we adopt the entity comparison framework in which explanations are modelled as unary conjunctive queries of restricted form [3, 4]. We concentrate on the data model where a dataset is an RDF graph-that is, a set of triples of IRIs and literals, jointly called entities. The basic building block of a query is a triple pattern, which is a triple of entities and variables. Then, a query is a non-empty finite set of triple patterns in which one variable, usually denoted by X, is an answer variable. The set Q(D) of answer entities to a query Q on a dataset D is defined as usual in databases. The two main notions of the framework are the similarity and difference queries for pairs of entities, which are defined as follows: a similarity query for entities a and b in a dataset D is a query Q satisfying fa; bg Q(D); a difference query for a relative to b is a query Q satisfying a 2 Q(D) and b 62 Q(D). In our prior work we proposed an algorithm for computing comparison queries that can be repurposed for similarity and difference queries [3]. The algorithm is based on the computation of a similarity tree-a data structure that represents commonalities and discrepancies in data for input entities a and b. It is a directed rooted tree with nodes and edges labelled by pairs of sets of entities such that the root is labelled by (fag; fbg) and every edge labelled (E1; E2) between nodes labelled (N1; N2) and (N10 ; N20 ) is justified in the sense that for every entity n in Ni, i 2 f1; 2g, there is a triple (n; e; n0) in the dataset with e 2 Ei and n0 2 N 0. i For instance, suppose there are 3 entities, Emma_Watson, Emily_Watson and E_Watson, that need to be either matched or disambiguated, and a data fragment given in Figure 1. Then the similarity trees for Emma_Watson and E_Watson, and for Emily_Watson and E_Watson are depicted in Figure 2 (where singleton sets f`g and pairs (f`g; f`g) are both written as ` for readability).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>A. Petrova et al.</p>
      <p>Each branch in a similarity tree can be treated as a separate similarity query,
in which each edge is encoded as a triple pattern, and each label (L1; L2) is
encoded as either an entity `, if L1 = L2 = f g
` ) or a fresh variable otherwise.</p>
      <p>For example, a query Q1 = (X; actedIn; Ballet_Shoes) is a similarity query for
Emma_Watson and E_Watson, while a query Q2 = (X; actedIn; Y ); (Y; year; Z)
is a similarity query for Emily_Watson and E_Watson. Moreover, each branch
involving non-entity labels can also be treated as a difference query, if instead
of some variables we take entities from one of the label sets. For example, query
Q3 = (X; actedIn; Little_Women); (Little_Women; year; 2019) is a difference query
for E_Watson relative to Emily_Watson.</p>
      <p>Both types of queries can assist in explaining why two entities should or
should not be merged: Q1 gives a good reason to match Emma_Watson and
E_Watson into one entity, Q2 is not specific enough to match the other pair, and
Q3 can act as an indicator that the two movies named Little_Women are indeed
two different movies, and Emily_Watson and E_Watson are different people.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Colucci</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giannini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Donini</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Di</given-names>
            <surname>Sciascio</surname>
          </string-name>
          , E.:
          <article-title>Finding commonalities in Linked Open Data</article-title>
          .
          <source>In: Proc. of CILC</source>
          . pp.
          <fpage>324</fpage>
          -
          <lpage>329</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>El</given-names>
            <surname>Hassad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Goasdoué</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Jaudoin</surname>
          </string-name>
          , H.:
          <article-title>Learning commonalities in SPARQL</article-title>
          .
          <source>In: Proc. of ISWC</source>
          . pp.
          <fpage>278</fpage>
          -
          <lpage>295</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Petrova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kostylev</surname>
            ,
            <given-names>E.V.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Cuenca</given-names>
            <surname>Grau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Horrocks</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          :
          <article-title>Query-based entity comparison in knowledge graphs revisited</article-title>
          .
          <source>In: Proc. of ISWC</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Petrova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherkhonov</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cuenca Grau</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horrocks</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Entity comparison in RDF graphs</article-title>
          .
          <source>In: Proc. of ISWC</source>
          . pp.
          <fpage>526</fpage>
          -
          <lpage>541</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>