<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>COPAAL - An Interface for Explaining Facts using Corroborative Paths</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Zafar Habeeb Syed</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nikit Srivastava</string-name>
          <email>nikit@mail.uni-paderborn.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Ro¨ der</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Axel-Cyrille Ngonga Ngomo</string-name>
          <email>axel.ngonga@upb.de</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data Science Group, Paderborn University</institution>
          ,
          <country>Germany zsyed</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Applied Informatics</institution>
          ,
          <addr-line>Leipzig</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>With the increasing uptake of knowledge graphs in domains as diverse as question answering, community-support systems and even personal assistants comes an increasing need for validated knowledge contained in these graphs. However, the sheer size and number of knowledge bases used in real-world applications makes manual fact checking impractical. Automated fact validation systems aim to compute the veracity of individual facts by evaluating the likelihood of these facts being true. In this demo, we present an interface for fact checking based on the COPAAL algorithm. Given triple whose veracity is to be evaluated, our interface provides (1) a score for the veracity of the triple, (2) evidence for the triple in the forms of paths, (3) explanation for the evidence in the form of verbalized RDF triples as well as (4) a graphical overview of the paths which support the input triple. We evaluate the performance of our fact checker, the quality of the verbalization we use and the usability of our user interface. The demo is available at http://copaal.dice-research.org/demo/.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The Web follows a participatory paradigm, which has led to more than 150 billion
RDF triples being published by thousands of independent data providers in more than
10,000 knowledge graphs (KGs).1 The largest open KGs contains billions of triples
pertaining to millions of entities. As open KGs are used in an increasing number of
applications, developers and end users have an increasing need to check the veracity
of triple before using them in applications, especially if these applications are
missioncritical. We developed the COPAAL approach [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to fact checking, which evaluates the
veracity of RDF triples by combining RDFS semantics with path search in knowledge
graphs. The approach was accepted as a full research paper at ISWC 2019. In this
corresponding demo paper2, we present (1) the user interface and REST service for fact
checking based on COPAAL as well as (2) supplementary evaluation results pertaining
to the verbalization of evidence and the system usability of the user interface. During the
1 http://lodstats.aksw.org/
2 Copyright c 2019 for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0).
demo at the conference, we will present the strengths and weaknesses of the approach
using selected examples as well as allow end users to check facts pertaining to DBpedia
resources for which they would like to see evidence.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Summary of the Approach</title>
      <p>
        With COPAAL, we address the following problem: Given an RDF knowledge graph G
and a triple (s; p; o), compute the likelihood that (s; p; o) is true. E.g., BarackObama
is a clearly a citizen of the USA, amongst others by virtue of having been born in
Hawaii, which is located in the USA. However, this fact is not available in
DBpedia 2016-10.3 Given this particular version of DBpedia, our approach can compute
paths between the resource BarackObama and USA, which corroborate the fact that
BarackObama is a national of the USA. The intuition behind our approach is that
certain sequences of properties (e.g., x birthplac!e y countr!y z) have a high mutual
information (MI) with certain predicates (e.g., nationality). We developed an
efficient approach for computing this MI (score) and combining the MI of several paths to
evaluate the veracity of particular facts. The exact computation details are given in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>System Overview</title>
      <p>We developed a web service and a UI so that users can easily interact and perform fact
validation using COPAAL. Figure 1 gives an overview of the three main components
of our user interface – input, path presentation and verbalization.</p>
      <p>(a) Input</p>
      <p>(b) Corroborative Paths and their explanations</p>
      <p>The input to COPAAL is a triple (s, p, o) whose correctness is to checked and for
which evidence is to be provided. A user can enter the subject, property and the object
directly into the interface (see Figure 1a). Note that the property must be an object
property in our demo. In addition, users can choose to have the evidence for their input
triple verbalized by selecting the ”verbalize” option. Finally, the user can forward the
input triple to the COPAAL service by clicking on the submit button.
3 http://downloads.dbpedia.org/2016-10/</p>
      <p>
        The COPAAL service computes corroborative paths for the input data and returns
a set of paths and their scores, i.e., paths through the input graph which have a high MI
with the input triple. In COPAAL, we visualize these paths using graphs4 (see Figure
1b). Note that the dotted path indicates the triple to be checked and the solid paths
represents the corroborative paths. One can navigate to view path explanations by clicking
on a path. The explanations are either sequences of triples or (if the user chose to have
verbalized evidence) sequences of sentences, which states the content of the
corroborative paths in simple English sentences. We used the the rule-based LD2NL framework,5
which is based on [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], to verbalize the triples in the paths.
      </p>
      <p>HawAasiiexpceocutnetdr,!youUrSaAp)praosacahmraeitnurnevsi dtheencpeafthor(Be.agr.a,cBkaOrbaacmkaO’bsanmataionabilirttyhp(slcaoc!ere
= 0.705, see Figure 1). Other paths pertaining to his alma mater, his presidency and his
political affiliation further corroborate that Barack Obama is a US citizen. COPAAL
computes a combined score which is also displayed to the end user. A binary (i.e.,
true/false) suggestion as to the truthfulness of the fact is also displayed to the user.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>
        Corroborative paths. We evaluated the performance of COPAAL on 17 datasets.
Details pertaining to the characteristics of all datasets as well as detailed insights derived
from the evaluation are given in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Our results on the four real-world datasets from [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
shown in Table 1 show that our approach clearly outperforms the state of the art. While
our approach can perform poorly on rare predicates, the AUC-ROC results suggest that
our approach is able to compute an appropriate score for most triples.
Verbalization. We evaluated the verbalization underlying our demo with two groups—
domain experts (66 persons) and non-experts (20 linguists). A set of triples and their
verbalization were shown to the volunteers. The experts were asked to rate the
verbalization regarding adequacy, fluency and completeness, i.e., whether all triples have
been covered. The non-experts were only asked to rate the fluency. The experiment was
carried out using 6 DBpedia resources.
4 We use d3 js – A JavaScript library for generating interactive and dynamic graphs
5 https://github.com/dice-group/ld2nl
0
      </p>
      <p>Our results revealed that verbalizing RDF is a difficult task. While the adequacy
of the verbalization was assigned an average score of 3.92 by experts (see Fig. 2), the
fluency was assigned a average score of 3.47 by experts and 3.0 by linguists (see Fig.
2). These results suggest is that (1) our framework generates sentences that are close
to that which a domain expert would also generate (adequacy). However (2) while the
sentences is grammatically sufficient for the experts, they are by linguists rated as being
grammatically passably good but still worthy of improvement.</p>
      <p>System usability We evaluated our user interface based on the System Usability Scale.
10 persons participated in the corresponding survey. Overall, we reached an SUS score
of 79.3 (school grade: A-), which means that most end users would be willing to use
the system and would recommend it to a friend if it were to be slightly improved.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>
        In this demo, we present a first interface for fact checking using COPAAL [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We
foresee a plethora of improvements in future works, including a natural-language
interface (both spoken and written) and a natural language output channel for the evidence.
Moreover, we will improve upon the verbalization of the paths.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements References</title>
      <p>This work has been supported by the German Federal Ministry of Transport and Digital
Infrastructure (BMVI) in the projects LIMBO (no. 19F2029I) and OPAL (no. 19F2028A).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Ngonga</given-names>
            <surname>Ngomo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.C.</given-names>
            , Bu¨hmann, L.,
            <surname>Unger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Lehmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Gerber</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>Sorry, i don't speak sparql: translating sparql queries into natural language</article-title>
          .
          <source>In: Proceedings of the 22nd international conference on World Wide Web</source>
          . pp.
          <fpage>977</fpage>
          -
          <lpage>988</lpage>
          . ACM (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Shiralkar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flammini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Menczer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciampaglia</surname>
            ,
            <given-names>G.L.</given-names>
          </string-name>
          :
          <article-title>Finding streams in knowledge graphs to support fact checking</article-title>
          .
          <source>In: 2017 IEEE International Conference on Data Mining (ICDM)</source>
          . pp.
          <fpage>859</fpage>
          -
          <lpage>864</lpage>
          . IEEE (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Syed</surname>
            ,
            <given-names>Z.H.</given-names>
          </string-name>
          , Ro¨der,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ngonga</given-names>
            <surname>Ngomo</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.C.</surname>
          </string-name>
          :
          <article-title>Unsupervised discovery of corroborative paths for fact validation</article-title>
          .
          <source>In: Proceedings of the International Semantic Web Conference</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>