<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Gold Standard Dataset for Large Knowledge Graphs Matching</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ziqi Zh</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>nk Hop</string-name>
          <email>f.hopfgartnerg@sheffield.ac.uk</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Systems, Umm Al Qura University</institution>
          ,
          <country country="SA">Saudi Arabia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Information School, The University of She eld</institution>
          ,
          <addr-line>She eld</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the last decade, a remarkable number of Knowledge Graphs (KGs) were developed, such as DBpedia, NELL and Google knowledge graph. These KGs are the core of many web-based applications such as query answering and semantic web navigation. The majority of these KGs are semi-automatically constructed, which has resulted in a signi cant degree of heterogeneity. KGs are highly complementary; thus, mapping them can bene t intelligent applications that require integrating di erent KGs such as recommendation systems and search engines. Although the problem of ontology matching has been investigated and a signi cant number of systems have been developed, the challenges of mapping large-scale KGs remain signi cant. In 2018, OAEI has introduced a speci c track for KG matching systems. Nonetheless, a major limitation of the current benchmark is their lack of representation of real-world KGs. In this work we introduce a gold standard dataset for matching the schema of large, automatically constructed, less-well structured KGs based on DBpedia and NELL. We evaluate OAEI's various participating systems on this dataset, and show that matching large-scale and domain independent KGs is a more challenging task. We believe that the dataset which we make public in this work makes the largest domain-independent gold standard dataset for matching KG classes.</p>
      </abstract>
      <kwd-group>
        <kwd>Knowledge Graphs Schema Matching Evaluation Dataset</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In the last decade, di erent KGs have been created as a result of years of
information extraction practices and crowdsourcing. DBpedia [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], YAGO [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ],
and NELL [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] are examples of large domain-independent KGs. Such KGs cover
multiple domains of knowledge such as medical, music, and publications. KGs
play a signi cant role in many applications such as reasoning, search engines
and e-commerce, while also being part of the linked open data domain [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>Copyright c 2020 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>Moreover, due to their automatically-constructed and independently-designed
nature, such KGs contain overlapping and complementary facts. For instance,
Bone and Artery are classi ed under BodyPart in NELL while being classi ed
as AnatomicalStructure in DBpedia.</p>
      <p>
        This problem of semantic heterogeneity has been thoroughly studied in the
Semantic Web community, with many ontology matching systems being
developed and surveyed [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Matching systems are annually evaluated through the
Ontology Alignment Evaluation Initiative (OAEI)3. While a new track for KG
matching has been introduced to OAEI's annual campaign since 2018, the
challenges of aligning large-scale KGs remain signi cant [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Currently, existing gold
standards are not well representative of real-world KGs. Such KGs are known
for sharing complementary facts about real-world entities such as people and
places, while current datasets are predominantly domain-dependent [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Further, the size of the existing gold standard does not accurately represent the
complexity of matching large-scale KGs that imply a signi cantly larger search
space due to orders of magnitude larger number of classes.
      </p>
      <p>
        This work proposes a gold standard dataset for matching the classes of large,
automatically constructed, inadequately structured, and domain-independent
KGs. The introduced benchmark is based on DBpedia and NELL. Although
both KGs are widely used in semantic web researches and can be considered
highly in uential, they are yet to be consolidated, even though the majority of
LOD cross-domain datasets, including KGs, are interlinked to DBpedia4 which
serves as a central link to many LOD datasets. According to [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], NELL is
considered as the most complementary KG to other larger KGs such as DBpedia
with an average of 10% gain of instances, while merging other large KGs can
only lead to a 5% gain. Therefore, we believe they are the best candidates for a
gold standard dataset for aligning large cross-domain KGs. We conduct an
experiment to evaluate the performance of OAEI's di erent participating systems
on this dataset, and show that mapping the classes of open KG is a much more
challenging task than the existing OAEI KG matching benchmark.
      </p>
      <p>The rest of this paper is structured as follows. We start by reviewing the
problem, and the current gold standard datasets for matching KGs in Section 2.
Then, we describe the process of building the proposed dataset in Section 3. In
Section 4, we present the results of evaluating current matching systems on the
proposed gold standard . We close with a discussion and a conclusion in Sections
5 and 6 respectively.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Ontology matching has been a well-studied problem which centers on discovering
corresponding entities across two distinct ontologies [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In the last decade, many
matching systems were developed and evaluated annually at the OAEI event.
      </p>
      <sec id="sec-2-1">
        <title>3 http://oaei.ontologymatching.org/ 4 https://lod-cloud.net</title>
        <p>The initiative provides over ten benchmark datasets in di erent tracks for
various matching systems to be evaluated. Examples of main tracks are Anatomy,
Conference, Complex Matching, Large Biomedical, and Interactive Matching.</p>
        <p>
          KGs are often compared to ontologies since both are used for data
representation purposes. Di erent from former ontology, open KGs are large-scale,
multi-domain and less well-formatted compared to ontologies [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. Similar to
ontologies, KGs entities also su er from semantic heterogeneity where the same
real-world entities are described using di erent terminologies.
        </p>
        <p>
          While there have been many well established matching systems for OAEI's
di erent matching tracks, the need for KG matchers remains an open area of
research [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Research in this domain has only been established since 2018, when
OAEI introduced a new track dedicated to KG matching5. Since then, ontology
matching tools have been evaluated on the provided benchmark, and multiple
KG matchers have participated in the latest version in 2019 [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Although
matching KGs has been a growing area of research recently, there is still a lack of gold
standard datasets that represent diverse KGs.
        </p>
        <p>
          The benchmark dataset currently used to evaluate systems in OAEI's KG
track is constructed from DBkWik [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], which is a KG created from wikis shared
on a wiki hosting platform. The individual KGs from the DBkWik project were
used to create the ground truth datasets for this track. The track consists of
ve test cases where each test case is aimed at matching both the schema,
including classes and properties, and the instance level of two KGs. The schema
level correspondences were built by ontology experts while the instance level
correspondences were automatically extracted [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. To the best of our knowledge
this gold standard is the only benchmark available to evaluate KG matching
systems. However, the number of mapped classes is considerably small, i.e., less than
50 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Therefore, this dataset does not represent the complexity of matching
real-world KGs where hundreds of classes can be matched.
        </p>
        <p>
          In terms of large domain-independent KGs, there are many published
according to the Semantic Web standards. Some of them are based on Wikipedia,
such as YAGO and DBpedia. Originally, DBpedia is a knowledge base
constructed from structured data embedded on Wikipedia [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. DBpedia also
involves crowdsourcing communities to maintain the quality of the mapping
between Wikipedia's articles and the structured knowledge in their KG. In
contrast, NELL is a fully automated learned KG under the Never-Ending Language
Learner project, which uses machine learning to read and extract knowledge
from free text on the web. It started with a seed KG that continuously evolves
by learning patterns from text to extract facts that are used to constantly grow
and update the seed KG [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Since its launch in 2010, NELL has grown to a
KG containing 50 million facts6. While the schema of the majority of Wikipedia
based KGs cover multiple types of properties, NELL graph schema is very
basic. It does not contain as many relations between instances [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Another KG
5 http://oaei.ontologymatching.org/2019/knowledgegraph/index.html
6 http://rtw.ml.cmu.edu/rtw/
of a taxonomy structure is WebIsALOD [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. However, the latter only covers
hypernymy relations and does not distinguish classes from instances.
3
3.1
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Approach</title>
      <sec id="sec-3-1">
        <title>Overview</title>
        <p>As mentioned earlier in Section 1, we use NELL and DBpedia as both KGs share
a signi cant amount of complementary facts. In this work, we deploy DBpedia
2016-10 version 7 using SPARQL query endpoint to return schema information.
As for NELL, since a query end point is not available we obtained schema
information by parsing a NELL dump le8 which contains every fact learned by the
project so far. As a result, DBpedia has over 750 classes while NELL has around
290 classes. Let P be the set of pair-wise classes across the two KGs, then the
number of all possible pairs is 218,660. Since our goal is to use human
annotators to identify all mappable pairs of classes, a greedy approach will lead to a
dataset that is expensive to annotate and likely to be overwhelmed with negative
pairs. Instead, we rst apply a Blocking Strategy to manually generate a set
of candidate pairs C which is a subset of P with signi cantly reduced number
of negative class pairs. Next, we perform a Candidate Filtering Strategy by
applying two similarity measures to each pair in C to further reduce the search
space for human annotators. Another screening was done after the ltering stage
to ensure that none of the discarded classes had a potential match in the
corresponding graph. Finally, for Dataset Annotation, we asked human annotators
to determine alignment of the resulting class pairs to construct the gold standard
dataset.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Generating Candidate Pairs</title>
        <p>Given the two KGs, we set one as source and one as target. Details about our
source and target choices will be explained later. Moreover, given P , the set
of all possible class pairs from the two KGs, we apply a Blocking Strategy
which requires manually screening the two KG class structures. The result of this
process is a set C which should eliminate as many true negatives as possible while
maintaining as many as (if not all) true positives. To illustrate the complexity of
the task, the classes named School in both KGs refer to di erent types of schools.
For instance, it is categorized as a subclass of EducationalInstitutions in
DBpedia while being a super class of HighSchool and University classes in
NELL. Given this structural inconsistency issue, a preliminary study aimed at
aligning the higher level of concepts across the two KGs was necessary.
7 https://wiki.dbpedia.org/develop/datasets/dbpedia-version-2016-10,visited
on 14-2-2020
8 http://rtw.ml.cmu.edu/rtw/resources, iteration number 1115, visited on
22-22020</p>
        <p>We manually created two subsets A and B , where the rst is a set of NELL
classes that have a possible corresponding class in DBpedia, and the second is a
set of DBpedia classes that have a possible corresponding class in NELL. These
two sets were created in two phases. First, we started by comparing the common
root classes across the two KGs, e.g., Person or Place. Then, all of their
nonroot (descendant) classes were added to A and B respectively. For instance, all
the descendant classes of P ersonnell and P ersondbp were added to each of A
and B respectively. Second, we examined other possible classes in which their
root classes do not share an overlap of words, i.e., they were not selected in
the rst step. A valid example are the two classes AcdemicSubjectdbp and its
possible equivalent class AcademicF ieldnell. While the former is a subclass of
TopicalConcept, the second is a subclass of everypromotedthing. The latter
is the root class of the KG taxonomic tree, i.e., the equivalent of OWL:Thing
in DBpedia. Therefore, our second screening phase was aimed at all descendant
classes in both KGs whose name values share overlapping words while their super
classes do not share overlapping words.</p>
        <p>As a result of this blocking strategy, a total of 18,492 candidate pairs were
generated in C as the product of A and B . We believe this blocking strategy
will not incorrectly discard any true positives because we have examined all
discarded classes to identify any possible match in the opposite KG. As shown
in Table 1, the number of distinct classes from DBpedia and NELL is 138 and 134
respectively. Nonetheless, this number of candidate pairs remains expensive for
an annotation task. Therefore, we proceed by the Candidate Filtering Strategy
to further reduce the numbers of pairs that need to be annotated by human
annotators while maintaining pair completeness.
In this section, we introduce the similarity measures applied to the candidate
pairs resulting from the prior phase. We apply a string-based and an
instancebased similarity measure combined with a low threshold to maximise the chance
to retain all true positives. We apply a String-based Similarity measure to
class names only since NELL does not o er other metadata descriptions of
classes. However, using only a string-based matcher can not guarantee a high
recall as both KGs use di erent names to describe the same classes. We then apply
an Instance-based Similarity measure to capture any possible true positive
pairs where string similarity could have failed to recognize them. To the best
of our knowledge, a matching approach that can handle a substantial number
of instances, such as in the case of KGs , is yet to be established. Therefore,</p>
        <p>
          Section 3.3 discusses the implementation of our preliminary instance-based
approach. We believe that combining both measures can ensure high (if not full)
recall of true positive pairs. It is also worth mentioning that due to the structural
irregularity in both KGs, structural-based similarity measures were excluded.
String-based Similarity Measure We apply the Levenshtein [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] edit
distance approach. This method has shown improvements over alternative
stringbased measures, particularly for matching classes [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Here the similarity between
class names in each candidate pair is measured. This value is then normalized by
dividing the value by the length of the longer string, i.e., class name, to produces
a value between [0.0, 1.0]. For this task we only retain a pair if the similarity
score of the two class names exceeds 0.4. State-of-the-art matching systems that
utilize an edit distance approach often apply a higher threshold, which can be
up to 0.8, to eliminate the number of false positive alignments [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Nonetheless,
in order to capture as many true positive pairs as possible, we use a threshold
that is twice lower than the state-of-the-art methods.
        </p>
        <p>Instance-based Similarity Measure This method casts the matching process
based on the principle of free-text index and search, which scales to very large
datasets. On a typical index/search scenario a collection of resources (i.e. web
documents) is indexed in a vector space where documents are represented with
weighted vectors of their text content. Weighting approaches, such as TF/IDF,
are used to weight term occurrences in the documents. A query given to a search
engine will also be converted into a vector representation and then matched
against all the vectors stored in the index. The matching is done by similarity
measures such as the cosine function where a ranked list of top K documents
related to the query is retrieved. Similarly, we propose to treat both KGs as a
collection of documents where each document corresponds to a class in a KG
and each term corresponds to the name of an instance. To map similar classes,
a query is built by sampling instance names from a source KG's class, and
matching against the index of the target KG. The equivalent class is determined
based on the search result, which is a ranked list of classes whose instance names
overlap with those in the query. We exploit Apache Solr 9, a state-of-the-art free
text index and search engine. The pseudocode for the entire similarity measure
is illustrated in Algorithm1.</p>
        <p>During the Indexing process, a separate index is created for the source and
target KGs. Classes from each KG are represented in documents that contain
the concatenation of the class's instance names. The documents' contents are
indexed using the standard Solr indexing process, including tokenisation,
stemming, lemmatization, lower casing, and term-weighing. For our particular task
an index is needed for NELL and DBpedia to perform the matching task. Thus,
we run the following query to obtain all instance names for each DBpedia class:</p>
        <p>SELECT ?name</p>
        <sec id="sec-3-2-1">
          <title>9 https://lucene.apache.org/solr/</title>
          <p>WHERE{ ?entity a &lt;http://dbpedia.org/ontology/%ClassName&gt;.
?entity rdfs:label ?name.</p>
          <p>Filter (lang(?name)="en")}
After each query, a new document representing a DBpedia class is added and
indexed in the designated DBpedia index. Similarly, an index was created for
NELL which contained indexed documents of instance names parsed from the
NELL facts dump.</p>
          <p>Algorithm 1 Instance-Based Similarity Measure
Require:
1: source a list of classes in Source KG
2: target a list of classes in Target KG
3: for Class an in source do
4: count = 1
5: candidate = [ ]
6: while count 30 do
7: query a concatenation of 20 instance names of class an
8: results search(query,target ) in the target index
9: for bn in results do
10: candidate.append(bn)
11: end for
12: count++
13: end while
14: candidate pairs
15: end for</p>
          <p>Top three frequent classes in candidate paired with an</p>
          <p>To perform the matching process, NELL and DBpedia were treated as source
and target respectively. Consequently, queries are generated by sampling
instances names from NELL's classes. This process can be performed in the
opposite direction; however, some of DBpedia's classes have missing instances. This
implies that a query cannot be created from such empty classes. For example,
classes such as State, Zoo, Profession are all leaf classes and supposed to be
populated with individuals but the links between class's name and its instances
are missing in the KG. A case in point is California10 and Florida11: both are
de ned in the data with classes (i.e., rdf:type) other than State. This problem
was encountered in 20 classes from the 138 classes selected from DBpedia. With
DBpedia being the center of the LOD datasets in mind, many options can be
explored in order to ful ll this gap. This includes using instances from SKOS
concepts or another KG that already has an established mapping with DBpedia,
such as WikiData or WebIsALOD. Nonetheless, we believe that performing a
one-way search is su cient for capturing all positive pairs for the annotation
task.</p>
          <p>In terms of the Search process, we aim to discover class pairs that share
a signi cant number of overlapping instance names across two KGs. Our
em10 http://dbpedia.org/page/California
11 http://dbpedia.org/page/Florida
pirical test on a smaller sample of the dataset showed that two key factors can
directly impact the search (matching) result. The rst one is the number of
instance names to be used in the query string. Due to these KGs' instances being
automatically extracted, and the large number of instances per class, using
either a too-large or too-small number of instance names to create queries will
result in no similar documents (classes) being retrieved or false positive pairs.
The second factor impacting the search result is the number of searches
(iterations) performed on each class to determine its equivalent class. Because of the
restriction of the query length, concatenating the names of all class instances
is not feasible. Moreover, by using a sample of instance names, di erent results
can be retrieved depending on the sample. Our experiment has shown that we
can obtain the maximum number of true positive pairs when concatenating 20
instances per query and performing 30 iterations per class.</p>
          <p>To demonstrate, for a class an in NELL, a random 20 instances of that
class are obtained and concatenated to form a query string. That query is then
matched against all documents (classes) in the target index, i.e., DBpedia.
Consequently, a list of classes whose instances overlap with those in the query are
retrieved. For example, if the following results were retrieved when sampling
instances from class Airportnell in the source KG:</p>
          <p>I t e r a t i o n 1 -&gt; {Airportdbp , Citydbp , P ortdbp }
I t e r a t i o n 2 -&gt; {Citydbp , P ortdbp , Airportdbp }
I t e r a t i o n 3 -&gt; {Airportdbp , P ortdbp }
I t e r a t i o n n -1 -&gt; {Airportdbp , Citydbp }</p>
          <p>I t e r a t i o n n -&gt; {Airportdbp , Citydbp , Streetdbp }</p>
          <p>By the end of the 30th iteration, we add three pairs of candidate alignments
for class Airportnell. Only the three most frequently retrieved classes among all
iterations are added as positive pairs with a non-zero as similarity score. For
the above example, the following pairs will be added: (Airportnell,Airportdbp),
(Airportnell,Citydbp), and (Airportnell, P ortdbp). Notice that Airportnell is not
matched to Streetdbp as the latter only appeared once during the search process.
Combining Similarity Measures As our goal for this particular task is to
discover potentially matching pairs to be annotated by human annotators, our
aim is to ensure a high (if not full) recall, which was achieved by combining the
two similarity measures. We applied the above mentioned similarity measures to
the 18,492 class pairs obtained in the prior phase. Only pairs that obtained a
similarity score higher than 0.4 by the String-based method or a non-zero value
by the Instance-based method were considered for the annotation task.
Following the above automated approach, we performed another manual screening to
discover remaining equivalent classes from NELL and DBpedia that were not
included in the potential pairs. By inspecting all pairs discarded by the ltering
process we were able to identify and recover 8 pairs. A total of 596 pairs were
created for the human annotation task.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Dataset Annotation</title>
        <p>In order to create a gold-standard dataset of matching classes, we asked human
annotators to determine the alignment for the previously discovered pairs, and
then aggregate their interpretations by the majority votes, as human annotators
can have di erent interpretation of correspondence. We have also performed a
study of the inter-annotator agreement (IAA). The dataset was annotated by
twenty research students and validated by two computer scientists. The
participants were provided with guiding instructions to complete the task. Several
labels were allowed to annotate pairs which are a match, not a match, more
general, and more specific. The latter two options are often used in the
ontology domain to label subsumption relation in ontologies. The reason we gave
the annotators this option is that it can be possible in a few cases. For example,
while DBpedia has two separate class for State and Province, NELL has one
class named StateOrProvince which combines both.</p>
        <p>Each participant annotated around 50 pairs on average. In order to observe
(IAA), 400 random pairs are duplicated among 12 annotators such that each
pair is annotated by 3 di erent annotators. The average IAA for this task was
measured using Cohen's kappa based on a sample of the dataset and it was 0.83.
The dataset was then validated by two experts. This was mainly to ensure that
the subsumption relations were used properly. Therefore, a subsumption relation
was only added to the dataset if there was an agreement by the experts. The gold
standard mapping resulting from this annotation task is publicly available as two
test cases12. The small test case includes a few instances per class, while the full
test case contains the full A-box information for the included classes. The latter
can be used to benchmark instance-based matching systems. The size of the gold
standard is 129 equivalent class pairs with 24 non-trivial matches, i.e., not an
exact matching string of class labels. Currently, the larger dataset in OAEI's KG
track carries only 15 class matches, while the maximum number of non-trivial
matches is 10. This makes the proposed dataset the largest domain-independent
gold standard for matching KG classes. This gold standard is considered as a
partial gold standard since some classes in both KGs have no equivalent class in
the corresponding KG.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>
        We evaluated the performance of the matching systems that participated in
the KG track in OAEI 2019 event on the proposed gold standard. The
Matching Evaluation Toolkit MELT [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] was used to perform this evaluation along
with the SEALS client. The following systems were evaluated: POMAP++ [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ],
AML [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], FCAMap-KG [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], LogMap [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], LogMapLt, LogMapKG, LogMapBio,
DOME [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], Wiktionary [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], and the string matcher used as a baseline for the
KG track.
12 https://github.com/OmaimaFallatah/KG_GoldeStandard
      </p>
      <p>We evaluated the class alignments resulting from each matcher based on
precision, recall, and f-measure. Results of the evaluation are shown in Table 4.
Since the proposed gold standard is only a partial gold standard, and to avoid
over-penalising systems that may discover reasonable matches that are not coded
in our gold standard, we ignore any predicted matches if neither of the classes
in that pair is present as a true positive pair with another class in our gold
standard. As an example, for a class an we only consider the alignment (an, bn)
as a false positive, if the gold standard has a true positive pair containing either
an or bn but not both in the same pair.</p>
      <p>As Table 4 shows, we have also evaluated the matchers on the
starwarsswtor test case, which is the largest dataset in the track in terms of the size of
class correspondences (which is 15). The best performing systems on the OAEI
dataset in terms of recall are DOME, Wiktionary and AML; however, DOME
and AML have obtained a lower recall (0.6) in our dataset, while Wiktionary is
one of the best performing systems on our dataset. In contrast, the second to best
performing matchers on the OAEI dataset, i.e., the LogMap family, obtained a
recall of 0.79, which is the best recall on our gold standard. Nonetheless, 27 out
of the 129 true positive pairs were not discovered by any matcher in the LogMap
family. Among the evaluated matchers, LogMApKG and FCAMap-KG are the
only systems that are particularly designed to match KGs. While the latter is
the second best performing system in OAEI's 2019 KG track, particularly in
matching classes, it has obtained a recall of 0.62 on our dataset. In terms of the
precision on our dataset, the scores are fairly high since most systems were only
able to discover trivial matches. However, the recall ranges between 0.6 and 0.79,
which shows that the dataset contains class correspondences that are di cult to
nd. Hence, all systems need further improvements in order to map the classes
of large and domain-independent KGs.</p>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>
        From the results presented above, the following three patterns were observed.
First, while current tools are able to produce high-quality results for well-formed
ontologies, such techniques are not as well-performing when applied on KGs that
lack textual descriptions. For instance, DOME is a matcher that trains a doc2vec
model using all available metadata descriptions for ontologies. This can explain
the matcher's low performance on our dataset as it requires a large amount of
text. Second, many ontology matching systems utilizes structural knowledge
available in well-structured ontologies such as disjoint axioms to re ne their
alignments [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Examples of systems that follow such an approach are AML and
LogMap. However, as a result of the lack of schematic information in NELL,
structural-based techniques can be di cult to apply in this case. Third,
matching strategies used when two resources are from a speci c domain setting are
not applicable for domain-independent settings where classes contain
information about real-world entities described with di erent terminologies. Therefore,
in order to tackle the problem of KG matching, the need for specialized matching
tools remains signi cant. Recently, many matching tools based on entity
embedding are being proposed but only tested with domain-dependent datasets or in
task-oriented settings, (e.g., [
        <xref ref-type="bibr" rid="ref21 ref7">7,21</xref>
        ]). Tailoring such methods for multi-domain
KG matching and testing them on our gold standard can lead to deeper
understanding and discovery in this domain.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this paper, we developed the largest gold standard dataset for matching the
classes of large KGs. Our gold standard is based on two highly in uential KGs,
and one of them is yet to be linked to the LOD. We evaluated several
state-of-theart matching tools on this dataset and showed that the task of matching large,
domain-independent KGs remains very challenging. We argue that matching
large, domain-independent and automatically constructed KGs has signi cant
utility and therefore, future work should be devoted further into this area. We
believe that our dataset and ndings will foster research in this direction.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Algergawy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faria</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferrara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fundulaki</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harrow</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hertling</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimenez-Ruiz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karam</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khiat</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lambrix</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , et al.:
          <article-title>Results of the ontology alignment evaluation initiative 2019</article-title>
          .
          <source>In: CEUR Workshop Proceedings</source>
          . pp.
          <volume>46</volume>
          {
          <issue>85</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Anam</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>Y.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>B.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          :
          <article-title>Review of ontology matching approaches and challenges</article-title>
          .
          <source>International Journal of Computer Science and Network</source>
          Solutions pp.
          <volume>1</volume>
          {
          <issue>27</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isele</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakob</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jentzsch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kontokostas</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mende</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morsey</surname>
            , M., van Kleef,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>DBpedia { A Large-scale, Multilingual Knowledge Base Extracted from Wikipedia</article-title>
          . Semantic Web pp.
          <volume>1</volume>
          {
          <issue>5</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Carlson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Betteridge</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kisiel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Settles</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hruschka</surname>
          </string-name>
          , E.R., Mitchell, T.M.:
          <article-title>Toward an architecture for never-ending language learning</article-title>
          .
          <source>In: Twenty-Fourth AAAI Conference on AI</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cheatham</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hitzler</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>String similarity metrics for ontology alignment</article-title>
          . In: International semantic web conference. pp.
          <volume>294</volume>
          {
          <issue>309</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , Zhang, S.:
          <article-title>Identifying mappings among knowledge graphs by formal concept analysis</article-title>
          .
          <source>In: OM@ ISWC</source>
          . pp.
          <volume>25</volume>
          {
          <issue>35</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dhouib</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zucker</surname>
            ,
            <given-names>C.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tettamanzi</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          :
          <article-title>An ontology alignment approach combining word embedding and the radius measure</article-title>
          .
          <source>In: International Conference on Semantic Systems</source>
          . pp.
          <volume>191</volume>
          {
          <issue>197</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Euzenat</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shvaiko</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Ontology matching: state of the art and future challenges</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Faria</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pesquita</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Santos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palmonari</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cruz</surname>
            ,
            <given-names>I.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Couto</surname>
            ,
            <given-names>F.M.:</given-names>
          </string-name>
          <article-title>The agreementmakerlight ontology matching system</article-title>
          .
          <source>In: OTM Confederated International Conferences "On the Move to Meaningful Internet Systems"</source>
          . pp.
          <volume>527</volume>
          {
          <fpage>541</fpage>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hertling</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>Webisalod: providing hypernymy relations extracted from the web as linked open data</article-title>
          .
          <source>In: International Semantic Web Conference</source>
          . pp.
          <volume>111</volume>
          {
          <issue>119</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Hertling</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>DOME results for OAEI 2018</article-title>
          . CEUR Workshop Proceedings pp.
          <volume>144</volume>
          {
          <issue>151</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Hertling</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>The knowledge graph track at oaei</article-title>
          .
          <source>In: European Semantic Web Conference</source>
          . pp.
          <volume>343</volume>
          {
          <issue>359</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hertling</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Portisch</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>Melt-matching evaluation toolkit</article-title>
          .
          <source>In: International Conference on Semantic Systems</source>
          . pp.
          <volume>231</volume>
          {
          <issue>245</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Jimenez-Ruiz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Logmap family participation in the OAEI 2019</article-title>
          . In: CEUR Workshop Proceedings. pp.
          <volume>160</volume>
          {
          <issue>163</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Laadhar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghozzi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Megdiche</surname>
            <given-names>Bousarsar</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Ravat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Teste</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ,
            <surname>Gargouri</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          :
          <article-title>Pomap: An e ective pairwise ontology matching system</article-title>
          .
          <source>In: 9th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management</source>
          . pp.
          <volume>161</volume>
          {
          <issue>168</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Levenshtein</surname>
            ,
            <given-names>V.I.</given-names>
          </string-name>
          :
          <article-title>Binary codes capable of correcting deletions, insertions, and reversals</article-title>
          .
          <source>In: Soviet Physics Doklady</source>
          . pp.
          <volume>707</volume>
          {
          <issue>710</issue>
          (
          <year>1966</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>Machine learning with and for semantic web knowledge graphs</article-title>
          .
          <source>In: Reasoning Web International Summer School</source>
          . pp.
          <volume>110</volume>
          {
          <fpage>141</fpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Portisch</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hladik</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>Wiktionary matcher</article-title>
          .
          <source>In: CEUR Workshop Proceedings</source>
          . pp.
          <volume>181</volume>
          {
          <issue>188</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Ringler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>One knowledge graph to rule them all?</article-title>
          <source>In: Joint German/Austrian Conference on Arti cial Intelligence</source>
          . pp.
          <volume>366</volume>
          {
          <issue>372</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasneci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Yago: a core of semantic knowledge</article-title>
          .
          <source>In: Proceedings of the 16th International Conference on World Wide Web</source>
          . pp.
          <volume>697</volume>
          {
          <issue>706</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Trisedya</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Zhang, R.:
          <article-title>Entity alignment between knowledge graphs using attribute embeddings</article-title>
          .
          <source>In: Proceedings of the AAAI Conference on AI</source>
          . pp.
          <volume>297</volume>
          {
          <issue>304</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gentile</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blomqvist</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Augenstein</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciravegna</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>An unsupervised data-driven method to discover equivalent relations in large Linked Datasets</article-title>
          . Semantic Web pp.
          <volume>197</volume>
          {
          <issue>223</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>