<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cleaning Noisy Knowledge Graphs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ankur Padia</string-name>
          <email>ankurpadia@umbc.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science and Electrical Engineering University of Maryland</institution>
          ,
          <addr-line>Baltimore County Baltimore. MD 21250</addr-line>
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>My dissertation research is developing an approach to identify and explain errors in a knowledge graph constructed by extracting entities and relations from text. Information extraction systems can automatically construct knowledge graphs from a large collection of documents, which might be drawn from news articles, Web pages, social media posts or discussion forums. The language understanding task is challenging and current extraction systems introduce many kinds of errors. Previous work on improving the quality of knowledge graphs uses additional evidence from background knowledge bases or Web searches. Such approaches are di cult to apply when emerging entities are present and/or only one knowledge graph is available. In order to address the problem I am using multiple complementary techniques including entitylinking, common sense reasoning, and linguistic analysis.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Problem Statement: Given a knowledge graph of entities and relations
extracted from text along with optional source documents, provenance
information, and con dence scores, how can we narrow down, identify
and explain likely errors.</p>
      <p>
        Many information extraction (IE) systems have been developed to extract
information from multiple sources like spreadsheets, Wikipedia Infoboxes,
discussion forum, and news articles [
        <xref ref-type="bibr" rid="ref10 ref17 ref4 ref5">5, 4, 17, 10</xref>
        ]. Fact are extracted following an IE
paradigm, either Open IE or model IE. Model IE based system populate an
ontology while Open IE extract possible facts without a target ontology. Extracted
ontology is serialized in a standard representation like RDF. However, the task to
extract information is challenging and current IE systems make many mistakes.
Given the importance of knowledge graphs (KG) for downstream applications
like Named Entity Recognizer/Typing, Semantic Role Labeling noise present in
the KG is sipped in during distant supervision hurting the performance of the
system.
      </p>
      <p>
        Errors are caused by many factors, including ambiguous, con icting,
erroneous and redundant information [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], or can be due to reasons like schema
violations, presence of outliers, and misuse of datatype properties [
        <xref ref-type="bibr" rid="ref12 ref18">12, 18</xref>
        ]. If the
validity of an atomic fact is suspect, we call the fact dubious. Aim of my thesis is
to build a pipeline to identify and explain such dubious facts present in a
knowledge graph. In this thesis, I plan to consider large general purpose knowledge
graphs which are generated an IE system.
      </p>
      <p>
        One of the solution is to use a con dence score to lter out correct candidates
facts. However, such a technique can assign high con dence to incorrect fact and
vice versa. For example, consider the fact from Never Ending Language Learner
(NELL) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] with great con dence aaron-mckie is an actor 1; in reality he
is a coach2.
      </p>
      <p>In this dissertation, I plan to exploring following aspects of the problem:
(1) How can a subset of dubious facts be e ectively narrowed down; (2) How
accurately can we con rm that a candidate is, in fact, incorrect; and (3) explain
why the fact is incorrect. I plan to address and answer these questions with
several complementary approaches, including linguistic analysis, common-sense
reasoning, and entity linking.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Relevancy</title>
      <p>
        Considering the importance of KG in real world applications, in this section
we perform quality assessment of two IE system: (1) Never Ending Language
Learner (NELL)[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], and (2) Kelvin[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. An automatic evaluation of the system
can be a research problem in itself, we evaluate the quality using con dence score
and expert engineered queries, respectively. Aim of the assessment is to highlight
that irrespective of the corpus size, targeted domains, and training methodology,
systems make mistakes jeopardizing its utility in downstream applications.
      </p>
      <p>
        Web scale KG: NELL [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is a semi-supervised, ontology driven, iterative
system that extracts facts from a Web corpus of more than one billion documents.
In each iteration, NELL learns new facts and assigns con dence score to it using
previous facts and available evidence. For example, as of the 990th iteration,
there were approximately 97 million facts. with ten million had con dence of
0.8 and above and nearly 77 million facts were in low range of 0.5 and 0.6.
Con dence scores are used to identify correct facts. However, for an IE system its
possible to have high con dence for incorrect fact and visa versa. Based solely on
con dence score, it's evident that NELL contains many dubious facts. Evaluating
the correctness of such a large knowledge graph is a research problem. To better
evaluate learning capability of an IE system, we additionally considered a more
controlled system { Kelvin.
      </p>
      <p>
        ColdStart based KG: Kelvin [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is an unsupervised information
extraction system that extracts facts from text to populate an ontology. It was initially
developed to take part in the NIST TAC ColdStart Knowledge Base Population
(CS-KBP) task3. The intention of TAC Coldstart KBP was to encourage
build
      </p>
      <sec id="sec-2-1">
        <title>1 http://rtw.ml.cmu.edu/rtw/kbbrowser/actor:aaron mckie</title>
        <p>2 https://en.wikipedia.org/wiki/Aaron McKie. A Google search query of
'AaronMcKie actor' mentions his name in a number of IMDB pages since he was the
subject of a documentary lm, which may have led to his being classi ed as an actor
3 http://tac.nist.gov/
ing knowledge base from scratch using set of text document without accessing
external resources like Web search engines or Wikipedia. The choice to analyze
Kelvin, beside other seven participants (Stanford, UMass, and others), is
motivated by easy access to internal HLTCOE4 resources and its relatively better
extraction performance. For TAC 2015, Kelvin learned 3.2 million facts from
total of 50K documents related to local news articles and web documents. Queries
with given subject and relation and missing object were engineered by expert
from Linguistic Data Consortium (LDC) were used to evaluate the accuracy of
information extraction. On an average, due to mistake in understanding of
relation among entities, Kelvin scored less then 30 F1 points making Kelvin a 2nd
ranked system with a small di erence from rst rank.</p>
        <p>
          Reasons for errors: An extracted fact can be incorrect due to multiple
reasons in addition to those discussed in [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Additional factors include the
following: (1) the choice of natural language processing libraries; (2) design
consideration of internal system components like classi ers and cross-document
co-reference resolution systems; (3) bad entity type assignments; (4) poor
indocument co-reference resolution; (5) missing mentions or extracting incorrect
mentions from the text; (6) learning techniques, semi-supervised (as in NELL) or
unsupervised (as in Kelvin); (7) heuristics employed; (8) choice of IE paradigm,
model based or open; (9) quality of inferencing either using rules or statistical
techniques; and (10) the nature of the text used to extract information.
        </p>
        <p>With the success of the proposed approach, I plan to contribute algorithms to
improve the quality of the KG increasing its utility in downstream applications.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Research Question(s)</title>
      <p>Narrow Can potential incorrect fact candidates be e ectively identi ed
Con rm How well can we con rm that a candidate is, in fact, incorrect
Explain How to rank or identify reasons why a fact might be wrong
In order to manage incorrect facts in a knowledge graph I am using a three stage
pipeline: Narrow Down, Con rm, and Explain. The Narrow Down is a triage
stage that selects a subset of all facts that are likely to be wrong, an important
step for very large KGs. Con rm is more granular then Narrow Down and aims to
classify the suspected facts as correct or incorrect. For each of the latter, Explain
provides a human-understandable explanation including the sources from which
the information was extracted and the likely cause of the error.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Hypotheses</title>
      <p>H1: Narrow Down: Metric like con dence or frequency count are insu cient to
identify likely possible facts
H2: Con rm: Better classi cation can be made when information at multiple
levels, i.e type and instance level, is taken into consideration (as in Sec. 6)</p>
      <sec id="sec-4-1">
        <title>4 http://hltcoe.jhu.edu/</title>
        <p>H3: Explain: Structure-based techniques are better candidates to locate
provenance information in web pages</p>
        <p>Narrow Down is an important phase in cleaning knowledge bases. By default,
all the facts could be considered dubious, but processing large KGs with millions
of facts will be demanding and redundant. Techniques like frequency count of
a predicate in the text and/or other KBs is of limited help, as it does not
capture interactions between di erent predicates e.g co-occurrence or three-way
interaction.</p>
        <p>
          Con rm helps to determine the credibility of the given facts. Previous
approaches like [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] have addressed this question to some extent with assumptions
that are di cult to apply in more common settings with only one KG. Such
assumption holds for entities and relations which are popular[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. For emerging
entities better approach that uses information at multiple level, i.e type and
instance, can perform better than previous method as shown in Sec. 5.
        </p>
        <p>
          Explain allows a person to interact with the system to understand a
classier's decision. For a given set of facts, it tries to identify appropriate provenance
information, which can be spread across the documents. Approaches likes [
          <xref ref-type="bibr" rid="ref7 ref9">9, 7</xref>
          ]
have been developed for multiple languages but are limited to relations from
DBpedia. Hence an approach that takes provenance information from multiple
sentences and documents is required.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Approach</title>
      <p>
        In order to address the problem, we use complementary approaches, Linguistic
Analytic (LA) and Entity Linking (EL), as shown in Figure 1. EL can provide
access to structured information from publicly available knowledge graphs like
DBPedia [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Yago [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], and Freebase [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. LA can be useful when EL fails to nd
corresponding entities among other knowledge graphs, especially for emerging
entities. Consider Figure 1, where an IE system makes a mistake to classify a
partial extracted entity \Broken Glass" as an organization. To determine the
credibility of the fact, information from multiple sources with an optional
ontology schema can provide complementary information for the classi er.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Preliminary Results</title>
      <p>In general, facts can be divided into popular and not-so-popular facts. Popular
facts involve entities, types, and relations for which relatively more information
is available, compared to not-so-popular facts, where there may only be one or
a few relevant documents.</p>
      <p>As an initial step, I have answered a part of second question which is to
identify if the type assertion of emerging entities is correct or incorrect
without explanation. For simplicity, we are assuming that all the fact present in the
knowledge graph are dubious and need to be veri ed. A signi cant number of
type error are made by current entity-extraction systems, even with ontologies
with limited number of type systems, as in TAC. My current approach takes
facts and natural language corpus as input and trains classi ers for each type,
combining features surrounding the entities (e.g., Mike Tyson) and
representative words from the concept (e.g., Boxer). The following steps are used to learn
signal words.
1. For each instance in the training set, retrieve the top-k documents that
contains the entity mention and extract surrounding context.
2. Combine surrounding context for all the entities to create a vocabulary set.
3. Entity Features (EF): A classi er is trained for each type t, with each word
in the vocabulary as a feature and each context is treated as instance. An
instance is considered positive if the context is extracted from an entity (e.g.,
Mike Tyson) belonging to the class and negative when the entity (e.g., Bill
Gates) belong to a disjoint class. Perform feature reduction (FR) to select
top-n high weighted words as features.
4. Concept Features (CF): Combine the documents for each entities of the class.</p>
      <p>Process document text using LSA to nd importance of each words. Perform
feature reduction to select the top-n words as features.
5. Combine CF and EF to train a classi er with documents as instances. Assign
positive and negative labels following disjoint class axioms. Values for each
feature are frequency counts of word present in the document.
6. Test it on the test data.</p>
      <p>Dataset: We evaluated the approach on expert engineered gold standard
which contained label of the entity, e.g Mike Tyson, followed by the expected
type and set of documents with named entity o set. Gold standard contains
2,350 entities of type person, organization and geo-political entity. We used
Approach AUC</p>
      <p>
        OpenEval 68.90
(EF + FR) + (CF + FR) 78.15
text corpus of roughly two million LDC documents [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] that included about one
million newswire documents, one million Web documents and 99K discussion
forum posts, hosted on an elasticsearch [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] instance using a single node and
default search settings. To select appropriate number of representative words for
each category we experimented with di erent values of the hyper-parameters
of the approach. To avoid biased result we partition the gold standard in size
60%n20%n20% and used 20% to decide the hyper-parameter.
      </p>
      <p>
        Baseline: We compared our algorithm with a more general state-of-the-art
baseline approach, OpenEval [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. OpenEval is an iterative approach to training
classi er for each relation and type. For a given fact, it iteratively uses web search
engine to obtain a set of document for training of the classi ers. Poor performing
classi ers are improved in next iteration by enhancing the queries. This process
is repeated for xed number of iterations. A similar process is conducted during
testing to conclude the credibility of unseen facts.
      </p>
      <p>Discussion: Table 1 shows improvement over state-of-the-art baseline. The
improvement achieved by our approach may be explained due to a combination of
surrounding words features along with class based keywords features. In order to
better understand this, consider our performance at multiple con guration
mentioned in Table 2. A minor performance gain is obtained when only surrounding
words are considered as features. Relatively better performance is achieved when
class level keywords are used as features without feature reduction. However, the
best performance is achieved when entity features and class features combined
after feature reduction. Hence supporting our second hypothesis.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Evaluation</title>
      <p>
        In order to measure if the questions asked in Section 3 are answered, I propose to
use evaluation approaches mentioned in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Manually engineered gold standard
would su ce to evaluate the performance of con rm phase. However, such an
approach might not be good candidate for Narrow Down as it aims to identify a
poor quality section of the graph and plan to use ranking. Each domain could be
ranked using expert perception and experience to later compare with the ranking
produced by the system. In case of explain I plan to construct a crowd source
gold standard expecting the human to assign a score to each of the justi cation
and use it to determine performance of the proposed algorithm.
8
      </p>
    </sec>
    <sec id="sec-8">
      <title>Lessons learned, Open Issues, and Future Work</title>
      <p>
        The main contributions of my PhD dissertation will be to answer the questions
in Section 3. To realize the framework, we performed experiments to determine
credibility of entity types for emerging entities. Tables 1 and 2 show that
combining information at multiple level, i.e type and instance level, yields better
classication for errors. One issue is the lack of gold standards on multiple datasets,
especially for the Explain. As of now, we have used gold standards available
from LDC [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] which follows well-de ned annotation guidelines for each entity
type and relations. However, creating such guidelines takes time, e ort and are
expensive. For future work, we plan to extend the existing approach outlined in
Section 5 for relations beyond entity types, like `hasSpouse' or `holdsPosition'.
We plan to develop a methodology to explain errors that are encountered in a
knowledge graphs and consider uents where a fact is correct for a time frame.
9
      </p>
    </sec>
    <sec id="sec-9">
      <title>Related Work</title>
      <p>
        There are some previous work on cleaning a KG but is often limited to
numerical values, or access ontology schema, or focus on popular entities and leaves a
gap to consider a single KG without schema information for emerging entities.
Existing approaches to deal with dubious facts can be divided into two
categories: internal and external [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Internal approaches use the facts mentioned
in the knowledge graph. Internal approach like [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] uses information about
distributions to identify outliers as incorrect facts. However, the approach is limited
to numerical literal values. More general approach [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] models ontological
constrains as rst order rules to use Probabilistic Soft Logic to infer correct KG
from a given KG. However, such approaches are limited to the relations with
prede ned semantics, e.g., label and mutualExclusion and are not applicable to
custom or natural relations like spouseOf.
      </p>
      <p>
        On the other hand, external approaches use resources beside the knowledge
graph, such as a Web search engine or access to background knowledge. External
approach like [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] uses a popular search engine to help user quickly select correct
provenance information for a given fact but covers relations limited to DBpedia.
Relatively broader approach is shown in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] which uses an ensemble of
knowledge graphs for the same set of documents with multiple extraction systems to
assess the correctness of the facts. However it's unclear how such methods can
be applied to single knowledge graphs like NELL and/or to emerging entities.
An approach similar to ours is demonstrated in OpenEval [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], which iteratively
gathers evidence via Web searches to train classi ers for each relation and type
and use them to determine the credibility of unseen facts. Since the approach
relies on Web searches, it works well on popular entities, types, and relations and
poorly on emerging entities. We overcome this limitation by using information at
multiple levels, i.e type and instance level, to determine of the fact for emerging
entities, as described in Sec. 5.
      </p>
      <p>Acknowledgments. I thank to my thesis advisor, Professor Tim Finin.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Auer</surname>
          </string-name>
          , S.e.a.:
          <article-title>Dbpedia: A nucleus for a web of open data</article-title>
          .
          <source>In: The Semantic Web</source>
          . Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bollacker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paritosh</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sturge</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , J.:
          <article-title>Freebase: a collaboratively created graph database for structuring human knowledge</article-title>
          .
          <source>In: SIGMOD. ACM</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Brank</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grobelnik</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mladenic</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>A survey of ontology evaluation techniques (</article-title>
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Carlson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Betteridge</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kisiel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Settles</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hruschka</surname>
            <given-names>Jr</given-names>
          </string-name>
          , E.R., Mitchell, T.M.:
          <article-title>Toward an architecture for never-ending language learning</article-title>
          .
          <source>In: AAAI</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gabrilovich</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heitz</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horn</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lao</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murphy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strohmann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Zhang, W.:
          <article-title>Knowledge vault: A web-scale approach to probabilistic knowledge fusion</article-title>
          .
          <source>In: SIGKDD. ACM</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ellis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Getman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fore</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuster</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bies</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strassel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Overview of linguistic resources for the TAC KBP 2015 evaluations: Methodologies and results</article-title>
          .
          <source>In: TAC KBP Workshop</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gerber</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Esteves</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Buhmann, L.,
          <string-name>
            <surname>Usbeck</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.C.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Speck</surname>
          </string-name>
          , R.:
          <article-title>Defactotemporal and multilingual deep fact validation</article-title>
          .
          <source>Journal of Web Semantics</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Gormley</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tong</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Elasticsearch: The De nitive Guide. " O'Reilly Media</surname>
          </string-name>
          ,
          <source>Inc."</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerber</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morsey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.C.N.</given-names>
          </string-name>
          :
          <article-title>Defacto-deep fact validation</article-title>
          .
          <source>In: ISWC</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. May eld, J.,
          <string-name>
            <surname>McNamee</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harmon</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Finin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawrie</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Kelvin:
          <article-title>Extracting knowledge from large text collections</article-title>
          .
          <source>In: AAAI Fall Symposium</source>
          . AAAI (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. Mitchell, T.,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hruschka</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , et. al, P.T.:
          <article-title>Never-ending learning</article-title>
          .
          <source>In: AAAI</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>Knowledge graph re nement: A survey of approaches and evaluation methods. Semantic web (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Pujara</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miao</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Getoor</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Knowledge graph identi cation</article-title>
          .
          <source>In: ISWC</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Samadi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veloso</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blum</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Openeval:
          <article-title>Web information query evaluation</article-title>
          .
          <source>In: AAAI</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasneci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Yago: A large ontology from wikipedia and wordnet</article-title>
          .
          <source>Journal of Web Semantics</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Wienand</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>Detecting incorrect numerical data in DBpedia</article-title>
          . In: ESWC (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>H.e.a.</given-names>
          </string-name>
          :
          <article-title>The wisdom of minority: Unsupervised slot lling validation based on multi-dimensional truth- nding</article-title>
          . In: COLING (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Zaveri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rula</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maurino</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pietrobon</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Quality assessment for linked data: A survey</article-title>
          .
          <source>Semantic Web</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>