<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Question Answering Benchmarks for Wikidata</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dennis Diefenbach</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Pellissier Tanon</string-name>
          <email>thomas.tanon@ens-lyon.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kamal Singh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierre Maret</string-name>
          <email>pierre.maretg@univ-st-etienne.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universite de Lyon, CNRS UMR 5516 Laboratoire Hubert Curien</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universite de Lyon, ENS de Lyon</institution>
          ,
          <addr-line>Inria, CNRS, Universite Claude-Bernard Lyon 1, LIP</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Wikidata is becoming an increasingly important knowledge base whose usage is spreading in the research community. However, most question answering systems evaluation datasets rely on Freebase or DBpedia. We present two new datasets in order to train and benchmark QA systems over Wikidata. The first is a translation of the popular SimpleQuestions dataset to Wikidata, the second is a dataset created by collecting user feedbacks.</p>
      </abstract>
      <kwd-group>
        <kwd>Wikidata</kwd>
        <kwd>Question Answering Datasets</kwd>
        <kwd>SimpleQuestions</kwd>
        <kwd>WebQuestions</kwd>
        <kwd>QALD</kwd>
        <kwd>SimpleQuestionsWikidata</kwd>
        <kwd>WDAquaCore0Questions</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        The most popular benchmarks for QA over Knowledge Bases are WebQuestions [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
SimpleQuestions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and QALD3 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ][
        <xref ref-type="bibr" rid="ref5">5</xref>
        ][
        <xref ref-type="bibr" rid="ref10">10</xref>
        ][
        <xref ref-type="bibr" rid="ref11">11</xref>
        ][
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Both WebQuestions and
SimpleQuestions were designed for Freebase. WebQuestions contains 5810 questions. They
can be answered using one reified statement with potentially some constraints like type
constraints or temporal constraints4. SimpleQuestions contains 108.442 questions which
can be answered using one triple pattern.
3 http://www.sc.cit-ec.uni-bielefeld.de/qald/
4 https://www.microsoft.com/en-us/download/details.aspx?id=52763
      </p>
      <p>Another popular benchmark is QALD. The number of questions and the datasets
used in the different QALD challenges are reported in Table 1. They generally can
be answered using up to 3 triple patterns. Sometimes modifiers like COUNT and
aggregation operators are needed. Note that only in the last QALD challenge a benchmark
for QA over Wikidata was presented. It contains 150 questions.</p>
      <p>
        Some other less known benchmarks on top of Freebase exists, like Free917 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] that
provides 917 questions annotated with lambda calculus forms.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3 SimpleQuestions to Wikidata</title>
      <p>
        In this section, we describe how we ported the SimpleQuestions dataset originally
designed for Freebase to Wikidata. The SimpleQuestions dataset [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] provides 108,442
questions, each annotated with a Freebase triple such that one of the acceptable answers
to the question is the subject of the triple.
      </p>
      <p>
        We mapped the Freebase triples to Wikidata using the same mapping process as [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]:
the subject and objects of triples, that are Freebase topics are mapped to Wikidata
items using automatically generated mappings and the properties are mapped using a
handmade mapping. When there is no equivalent property in Wikidata, but the Freebase
inverse property has an equivalent property PXX in Wikidata, we map the Freebase
property to a "fake" Wikidata property RXX ("R" indicating reverse). Note that not
every translated triple is required to exist in Wikidata. When migrating the data from
Freebase to Wikidata part of the information was lost. We therefore have created two
versions, the first containing questions that can be answered in Wikidata, the second
containing all questions. The QA community has concentrated so far on benchmarking
QA systems assuming that most of the questions in the benchmark are answerable.
But having many questions that are not answerable allows to tackle a new challenge,
i.e. let the QA system decide if it has the knowledge to answer the question or not.
      </p>
      <p>This new dataset contains 49.202 questions (21.957 of which are answerable over
Wikidata). One of the reason of the gap in size between the two datasets is that we
have only mapped 404 Freebase properties to Wikidata even if the dataset contains
1.837 properties. But the top 50 properties in the Freebase dataset provide 61% of the
triples and the top 100, 76%. It allowed to map 45% of the dataset with 108 properties.</p>
      <p>The datasets is available under the Creative Commons Attribution 3.0 licence at
https://github.com/askplatypus/wikidata-simplequestions. We offer it in the
same format as the original SimpleQuestions dataset and in QALD format. We call
it SimpleQuestionsWikidata.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Dataset using Logs and User feedback</title>
      <p>
        Creating large datasets for QA is a tedious and expensive task. For example, [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] report
that they spent several thousands dollars for the creation of WebQuestions (containing
5810 questions) using Amazon Mechanical Turk. In the following, we show how it is
possible to create a benchmark dataset for QA reducing the human effort and therefore
also the cost.
      </p>
      <p>
        The idea is to involve users of a QA system in the creation of the benchmark. In this
concrete case we collected the feedback given by users using WDAqua-core0 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], a QA
system available under www.wdaqua.eu/qa. The web-service is online since June 2016
and received 12302 requests (5231 from them are unique). The QA service is exposed
using Trill [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], a reusable front-end for QA systems. It contains a UI-component for
feedback (see figure 1). The UI-component can be used mainly in two situations. In the
first, end-users use the feedback component if they know that the answer is right or not.
The second situation involves expert users that are able to understand SPARQL queries.
If the answer is correct, the expert users can directly use the feedback component. If
not they can check the top-k SPARQL queries generated by the QA system, as shown
in Figure 2 and (if available) select the right one.
      </p>
      <p>The collected data contains 689 questions. Note that the questions are asked by real users
and are a mixture between keyword and full natural language questions. The questions
can be answered with maximal 2 triple patterns. Considering how the dataset was created
the benchmark can contain errors. The resulting benchmark is available in QALD format
at https://github.com/WDAqua/WDAquaCore0Questions under the Creative
Commons Attribution 3.0 licence. We call the generated dataset WDAquaCore0Questions.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We have presented two new datasets for the research community. The first dataset
is the translation of a popular benchmark over Freebase to Wikidata, namely
SimpleQuestions which contains 21.957 questions answerable over Wikidata. The
second is a dataset containing 689 questions generated using user feedback. By
offering these datasets we hope to move the QA community towards Wikidata, which</p>
      <p>Acknowledgments Parts of this work received funding from the European Union's
Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie grant
agreement No. 642795, project: Answering Questions using Web Data (WDAqua). It was also
supported by the LABEX MILYON (ANR-10-LABX-0070) of Universite de Lyon.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Berant</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frostig</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Semantic Parsing on Freebase from Question-Answer Pairs</article-title>
          . In: EMNLP (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bollacker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paritosh</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sturge</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , J.:
          <article-title>Freebase: a collaboratively created graph database for structuring human knowledge</article-title>
          .
          <source>In: Proceedings of the 2008 ACM SIGMOD international conference on Management of data</source>
          . pp.
          <volume>1247</volume>
          {
          <fpage>1250</fpage>
          .
          <string-name>
            <surname>AcM</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bordes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Usunier</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chopra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
          </string-name>
          , J.:
          <article-title>Large-scale simple question answering with memory networks</article-title>
          .
          <source>arXiv preprint arXiv:1506</source>
          .
          <year>02075</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cai</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yates</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Large-scale Semantic Parsing via Schema Matching and Lexicon Extension</article-title>
          .
          <source>In: ACL (1)</source>
          .
          <source>Citeseer</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Unger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cabrio</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.C.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Multilingual question answering over linked data (qald-3): Lab overview</article-title>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Diefenbach</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amjad</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Both</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maret</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Trill: A reusable front-end for qa systems</article-title>
          . In: ESWC P&amp;D (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Diefenbach</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maret</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Wdaqua-core0: A question answering component for the research community</article-title>
          .
          <source>In: ESWC, 7th Open Challenge on Question Answering over Linked Data (QALD-7)</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Unger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
          </string-name>
          , E.:
          <article-title>Evaluating question answering over linked data</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Pellissier</given-names>
            <surname>Tanon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Vrandecic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Schaffert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Steiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Pintscher</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>From Freebase to Wikidata: The great migration</article-title>
          .
          <source>In: Proc. of WWW</source>
          . pp.
          <volume>1419</volume>
          {
          <issue>1428</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Unger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forascu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.C.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cabrio</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Question answering over linked data (QALD-4)</article-title>
          .
          <source>In: Working Notes for CLEF 2014 Conference</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Unger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forascu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.C.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cabrio</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Answering over Linked Data (QALD-5)</article-title>
          .
          <source>In: Working Notes for CLEF 2015 Conference</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Unger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.C.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cabrio</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <article-title>Cimiano: 6th Open Challenge on Question Answering over Linked Data (QALD-6)</article-title>
          . In: The Semantic Web:
          <article-title>ESWC 2016 Challenges</article-title>
          .
          <article-title>(</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>