<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>QAKiS: an Open Domain QA System based on Relational Patterns</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elena Cabrio</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julien Cojan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessio Palmero Aprosio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bernardo Magnini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alberto Lavelli</string-name>
          <email>lavellig@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabien Gandon</string-name>
          <email>fabien.gandong@inria.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>FBK-Irst</institution>
          ,
          <addr-line>Via Sommarive 18, Povo-Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>INRIA</institution>
          ,
          <addr-line>2004 Route des Lucioles, Sophia Antipolis</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universita degli Studi di Milano</institution>
          ,
          <addr-line>Via Comelico 39/41, Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present QAKiS, a system for open domain Question Answering over linked data. It addresses the problem of question interpretation as a relation-based match, where fragments of the question are matched to binary relations of the triple store, using relational textual patterns automatically collected. For the demo, the relational patterns are automatically extracted from Wikipedia, while DBpedia is the RDF data set to be queried using a natural language interface.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        To enhance users interactions with the web of data, query interfaces providing
a exible mapping between natural language expressions, and concepts and
relations in structured knowledge bases are becoming particularly relevant. This
demonstration presents QAKiS (Question Answering wiKiframework-based
System), that allows end users to submit a query to an RDF triple store in English
and obtain the answer in the same language, hiding the complexity of the non
intuitive formal query languages involved in the resolution process. At the same
time, the expressiveness of these standards is exploited to scale to the huge
amounts of available semantic data. In its current implementation, QAKiS
addresses the task of QA over structured Knowledge Bases (KBs) (e.g.
DBpedia) where the relevant information is expressed also in unstructured form (e.g.
Wikipedia pages). Its major novelty is to implement a relation-based match for
question interpretation, to convert the user question into a query language (e.g.
SPARQL). Most of the current approaches (for an overview, see [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) base this
conversion on some form of exible matching between words of the question and
concepts and relations of a triple store, disregarding the relevant context around
a word, without which the match might be wrong. QAKiS tries instead rst
to establish a matching between fragments of the question and relational
textual patterns automatically collected from Wikipedia. The underlying intuition
is that a relation-based matching would provide more precision with respect to
matching on single tokens, as done by current QA systems.
      </p>
    </sec>
    <sec id="sec-2">
      <title>QAKiS system description</title>
      <p>QAKiS demo4 (Fig. 1) is based on Wikipedia for patterns extraction. DBpedia
is the RDF data set to be queried using a natural language interface.</p>
      <p>
        QAKiS makes use of relational patterns (automatically extracted from
Wikipedia and collected in the WikiFramework repository [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]), that capture di erent
ways to express a certain relation in a given language. For instance, the relation
crosses(Bridge,River) can be expressed in English, among the others, by the
following relational patterns: [Bridge crosses the River] and [Bridge spans over
the River]. Assuming that there is a high probability that the information in the
Infobox is also expressed in the same Wikipedia page, the WikiFramework
establishes a 4-step methodology to collect relational patterns in several languages
for the DBpedia ontology relations (similarly to [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]): i) a DBpedia relation is
mapped with all the Wikipedia pages in which such relation is reported in the
Infobox; ii) in such pages we collect all the sentences containing both the domain
and the range of the relation; iii) all sentences for a given relation are extracted
and the domain and range are replaced by the corresponding DBpedia ontology
classes; iv) the patterns for each relation are clustered according to the lemmas
between the domain and the range, and sorted according to their frequency.
QAKiS is composed of two main modules (Fig. 2): i) the query generator takes
the user question as input, generates the typed questions, and then generates the
SPARQL queries from the retrieved patterns; ii) the pattern matcher takes
as input a typed question, and retrieves the patterns (among those stored in the
pattern repository) matching it with the highest similarity.
4 Available at http://dbpedia.inria.fr/qakis/
The current version of QAKiS targets questions containing a NE related to the
answer through one property of the ontology, as Which river does the Brooklyn
Bridge cross?. Each question matches a single pattern (i.e. one relation).
Expected Answer Type (EAT) and NE identi cation. Before running
the pattern matcher component, we identify the target of the question with a
NER tool. We apply the Stanford Core NLP NE Recognizer together with a
set of strategies based on the comparison with the labels of the instances in the
DBpedia ontology. We plan to test the use of other NER tools in the future. At
the same time, simple heuristics are applied to infer the EAT from the question
keyword, e.g. if the question starts with \When", the EAT is [Date] or [Time],
with \Who", the EAT is [Person] or [Organisation] and so on.
Typed questions generation. We generate a typed question by replacing the
question keywords (e.g. who, where) and the NE by the types and supertypes.
Given the question \Who is the husband of Amanda Palmer?" 9 typed
questions are generated, since i) both [Person] or [Organisation] (subclasses of
[owl:Thing]) are considered as EAT, and ii) [MusicalArtist], [Artist] and
[owl:Thing] are the types of the NE Amanda Palmer.
      </p>
      <p>WikiFramework pattern matching. The typed questions are lemmatized,
tokenized, and stopwords are removed. A Word Overlap algorithm is then applied
to match such typed questions with the patterns for each relation. A similarity
score is provided for each match: the highest represents the most likely relation.
Query selector. A set of patterns (max. 5) is retrieved by the pattern matcher
component for each typed question, and sorted by decreasing matching score.
For each of them, one or two SPARQL queries are generated, either i) select ?s
wheref?s &lt;property&gt; &lt;NE&gt;g, ii) select ?s wheref&lt;NE&gt; &lt;property&gt; ?sg or
iii) both, according to the compatibility between their types and the property
domain and range. Such queries are then sent to the SPARQL endpoint for
answer retrieval. If the query produces no results, we try with the next pattern,
until a satisfactory query is found or no more patterns are retrieved.</p>
      <p>Experimental evaluation
train
test</p>
      <p>Precision Recall F-measure # answered # right answ. # partially right
0.476 0.479 0.477 40/100 17/40 4/40
0.39 0.37 0.38 35/100 11/35 4/35</p>
      <p>Most of QAKiS' mistakes concern wrong relation assignment (i.e. wrong
pattern matching). We plan to replace the Word Overlap algorithm with approaches
considering the syntactic structure of the question. Another issue concerns
questions ambiguity, i.e. the same surface forms can in fact refer to di erent relations
in the DBpedia ontology. We plan to cluster relations with several patterns in
common, to allow QAKiS to search among all the relations in the cluster.
The partially correct answers concern questions involving more than one relation:
the actual version of the algorithm detects indeed only one of them. We plan to
target questions as Give me all people that were born in Vienna and died in Berlin
in a short time, since the two relations are easily separable. On the contrary, we
need more complex strategies to answer questions with nested relations.</p>
    </sec>
    <sec id="sec-3">
      <title>Future perspectives</title>
      <p>
        We are currently considering to publish the WikiFramework relational patterns
as RDF triples, organized according to a newly de ned RDF vocabulary
(similarly to [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]). We are also planning improvements on: i) the WikiFramework
pattern extraction algorithm, following [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]; ii) the question-pattern matching
algorithm; iii) the system coverage, addressing boolean and n-relation questions.
We are also exploring QAKiS applicability in real application scenarios.
5 http://greententacle.techfak.uni-bielefeld.de/~cunger/qald/
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Gerber</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Ngonga</given-names>
            <surname>Ngomo</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.C.</surname>
          </string-name>
          (
          <year>2011</year>
          ),
          <article-title>Bootstrapping the Linked Data Web</article-title>
          ,
          <source>in 1st Workshop on Web Scale Knowledge Extraction @ ISWC</source>
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Lopez</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uren</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sabou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2011</year>
          ),
          <article-title>Is Question Answering t for the Semantic Web?: a Survey, in Semantic Web journal</article-title>
          , vol.
          <volume>2</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>125</fpage>
          -
          <lpage>155</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Mahendra</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wanzare</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernardi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magnini</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2011</year>
          ),
          <article-title>Acquiring Relational Patterns from Wikipedia: A Case Study</article-title>
          ,
          <source>in Proc. of LTC2011.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Wu</surname>
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weld</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          (
          <year>2010</year>
          ),
          <article-title>Open information extraction using Wikipedia</article-title>
          ,
          <source>in Proc. of ACL2010</source>
          , pp.
          <fpage>118</fpage>
          -
          <lpage>127</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>