<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ISOFT at QALD-5: Hybrid question answering system over linked data and text data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Seonyeong Park</string-name>
          <email>sypark322@postech.ac.kr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Soonchoul Kwon</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Byungsoo Kim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gary Geunbae Lee</string-name>
          <email>gblee@postech.ac.kr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Engineering, Pohang University of Science and Technology</institution>
          ,
          <addr-line>Pohang, Gyungbuk</addr-line>
          ,
          <country country="KR">South Korea</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We develop a question answering system over linked data and text data. We combine knowledgebase-based question answering (KBQA) approach and information retrieval based question answering (IRQA) approach to solve complex questions. To solve this kind of complex question using only knowledgebase and SPARQL query, we use various methods to translate natural language (NL) phrases in question to entities and properties in knowledgebase (KB). However, converting NL phrases to entities and properties in KB many times usually has low accuracy in most KBQA. To reduce the number of converting NL phrases to words in KB, we extract clues of answers using IRQA and generate one SPARQL based on the extracted clues and analyses of the question and the semantic answer type.</p>
      </abstract>
      <kwd-group>
        <kwd>Question answering</kwd>
        <kwd>Hybrid QA</kwd>
        <kwd>SPARQL</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Question answering (QA) systems extract short and preprocessed answers to natural
language questions. QA is the fundamental goal of information extraction. In the big
data era, this property of QA is increasingly gaining importance. Two popular types of
QAs are knowledgebase-based question answering (KBQA) and information
retrievalbased question answering (IRQA).</p>
      <p>
        Recently, structured and semantically-rich KBs have been released; examples
include Yago [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], DBpedia [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], and Freebase [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. As an increasing quantity of resource
description framework (RDF) data are published in linked form, finding intuitive ways
to access the data is becoming increasingly important. Several QAs use RDF data;
examples include Aqualog [4], template-based SPARQL learner (TBSL) [5] and
Parasempre [6]. Because it is structured in linked form, the semantic web inference is
possible [7], and KB is curated and verified by human, KBQA can provide more accurate
answers than IRQA. However KBQAs require that the user’s phrase be translated to
entities and properties that exist in a KB. Many previous works use PATTY1, pattern
based matching, and other methods, but it is still not sufficient to achieve high accuracy.
Especially in complex question such as ‘who is the architect of the tallest building in
Japan?’, system needs to map the predicate and generate queries in sequential way, and
errors in mappings result in propagating errors.
      </p>
      <p>To solve the problem, we combine IRQA approach and KBQA approach. We extract
clues from the question using sequential phrase queries which are generated
segmentation of the question using NLP tools such as chunking and dependency parsing. We
define the answer clue as extracted entities to find final answer. We use multi-source
tagged text database which is tagged with co-reference resolved and disambiguated
form using Stanford co-reference resolution tool2 and DBPedia Spotlight3 [8]. To
generate query, we don’t need to detect entity from every text in database in runtime
because the entities have been already detected in document processing time. We
concatenate next query and the answer of the first query and the next rightmost phrase. We
repeat this process to find the answer of the given question. If we failed to find
appropriate answer clue, we generate SPARQL query for one triple.</p>
      <p>We use semantic similarity based on explicit semantic analysis (ESA) [9] to map
predicates in the NL question to uniform resource identifiers (URIs) of properties in the
KB. ESA converts target strings to semantic vectors that can convey their explicit
meaning as weighted vectors of Wikipedia concepts, and calculating the similarity of
two vectors reveals the semantic relatedness of the two strings from which the vectors
were generated. We also increase the effectiveness of mapping the NL words to URIs
by concatenating additional information to the predicate. If the property is related to
arithmetic measurement (e.g. length or height), we map the predicate by pattern
matching and rules.</p>
      <p>After the final sequential phrase query is processed, we extract answer candidates
and select the answer by using answer types and rules in the case of questions which
requires to arithmetic comparison (e.g. ‘highest mountain’ and ‘tallest building’).
 Answer clue: entities to find final answer.
 Answer clue sentence: sentence which include answer clue.</p>
      <p> Predicate URI: property in DBpedia.</p>
    </sec>
    <sec id="sec-2">
      <title>2 http://nlp.stanford.edu/projects/coref.shtml 3 https://github.com/dbpedia-spotlight/dbpedia-spotlight</title>
      <p>2.1</p>
      <sec id="sec-2-1">
        <title>System Description</title>
        <sec id="sec-2-1-1">
          <title>Overall system architecture</title>
          <p>In our system (Fig 1), first we analyze question, extract the sequential phrase query
and classify the semantic answer type (SAT). For example, when the question is ‘who
is the architect of the tallest building in Japan?’, the question is divided into three
phrases; ‘is the architect’, ‘of the tallest building’, and ‘in Japan’. Because most given
hybrid questions in the QALD-5 are quite long to find answer at one time. We first
search the query containing ‘the tallest building’ and ‘in Japan’ from the
multi-information tagged text data. Then, we extract entities such as ‘Tokyo_Skytree’,
‘Tokyo_Tower’ and so on. If the query contains comparative form such as ‘deeper’ or
superlative form such as ‘tallest’, we map these indicator to properties in KB. To extract
the height of each entities, we generate SPARQL query such as SELECT DISTINCT ?y
WHERE { res:Tokyo_Skytree dbo:height ?y }. Then, we compare the heights among
the entities to get the tallest entity. If we failed to find answer, we generate SPARQL
query to get answer from the tallest entity and the property which is in the KB mapped
from the dependent of main verb or main verb itself. The SPARQL query is as follow,
SELECT DISTINCT ?y WHERE { res:Tokyo_Skytree dbo:architect ?y }.</p>
          <p>To find the correct answer, we should analyze the input of the system, the question,
carefully and thoroughly in various aspects (Fig 2). We use both statistical and
rulebased approaches to analyze the questions. These analyses include ordinary NLP
techniques such as tokenization and part of speech (PoS) tagging, and QA system oriented
techniques such as question to statement (Q2S) analysis, lexical answer type (LAT)
extraction, and SAT classification.</p>
          <p>The ordinary NLP techniques include tokenization, part of speech tagging,
dependency parsing, keyword extraction, term extraction, and named entity (NE) extraction.
This information is not only important features for SAT extraction and answer
selection, but also a basis for further question processing. We use ClearNLP4 for
tokenization, PoS tagging and dependency parsing. Keyword extraction is simply removing stop
words and a few functional words, such as ‘the’, ‘of’ and ‘please’, from the question.
Term extraction is finding nouns and verbs and their synonyms which exist in
WordNet5 dictionary. NE extraction uses Spotlight to map NEs in the question to entities in
DBPedia6. The keywords, terms, and NEs are further used for SAT classification and
query generation.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 https://github.com/clir/clearnlp 5 https://wordnet.princeton.edu/ 6 http://wiki.dbpedia.org/</title>
      <p>The QA system oriented techniques include Q2S analysis, LAT extraction and
phrase extraction. The Q2S analysis is a rule-based analysis that recovers the
corresponding declarative sentence from the interrogative or imperative sentence, the
question. This analysis uses the result of the previous analyses, LAT, modal verbs,
preceding preposition (e.g. “For whom does …”), usage of ‘be’ or ‘do’, and interrogative, to
match the rules built for this system. The missing information which is asked through
the question, which is called focus, is added to make a complete declarative sentence.
LAT is a part of the question that limits the type of the answer, and gives a strong hint
for SAT classification. LAT extraction is done along with Q2S analysis. Phrase
extraction is to extract predicate phrase and prepositional phrase for query generation.
2.3</p>
      <sec id="sec-3-1">
        <title>Query Generation</title>
        <p>We have to generate Apache Lucene7 queries to find sentences that contain the
answer of the question from multi-information tagged text database. Some questions in
QALD task cannot be solved with a single query. For example, a question ‘who is the
architect of the tallest building in Japan?’ should be queried twice; ‘Tokyo Skytree’
from ‘the tallest building in Japan’ and ‘Nikken Sekkei’ from ‘the architect of Tokyo
Skytree’ (Fig 3).</p>
        <p>We devise sequential phrase queries to answer such questions. A unit of queries is a
prepositional phrase or a predicate phrase. We generate the first query with
concatenating the two rightmost phrases and find the answer. The next query is concatenation of
the answer of the first query and the next rightmost phrase. We repeat this process to
find the answer of the given question. The result of query: “tallest building in Japan”
have many entities. If we failed to generate query in rightmost phrase, we used chunker
and generate sequential query including more than one named entity. We filter many
entities and select one answer cue. We can know whether the sentence including answer
clues is related to query or not. We measure cosine similarity, Jaccard similarity
between answer clue sentence and question statement. Also, we check whether named
entities in question are in the answer clue statements or not. If the question include
polarity such as “tallest”, we select the entities satisfy the condition. If we failed to find
answer clue, we generate SPARQL query to get answer candidates.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>7 https://lucene.apache.org/core/</title>
      <p>Who is the architect of the tallest building in Japan?
Is the architect
of the tallest building</p>
      <p>In Japan</p>
      <p>Tokyo Skytree</p>
      <p>Nikken Sekkei</p>
      <p>Semantic answer type (SAT) is a very important feature in reducing wrong answer
candidates. We can infer the SAT from the question before finding the answer
candidates. For example, the answer of ‘what did Bill Gates found?’ can’t be found before
searching the database, but the SAT, ‘ORGANIZATION’ can be inferred from the
question itself. Instead of other typesets built for SAT classification such as UIUC
typeset [10], we use open typeset from DBPedia because the DBPedia typeset covers most
entities properly and it can be extended as more entities are added to Wikipedia.</p>
      <p>To classify the SAT, we used features from previous analyses such as keywords and
LAT. The instances of the entities are not used because type rather than instance of
each entity is important. For example, SAT of ‘what is the capital of France?’ is more
similar than that of ‘what is the capital of Germany?’ than ‘who is the president of
France?’, even though the first and the third question shares the same NE, ‘France’.
Thus we replaced the NE instances to type of each NE for features. We used other
features such as interrogative wh-word, predicate and its arguments.</p>
      <p>We train the SAT classifier with libSVM8 and train data from four years of previous
QALD challenges. We achieve 71.42 % accuracy for 3-level type ontology and 84.62 %
for 2-level type ontology.
2.5</p>
      <sec id="sec-4-1">
        <title>Multi-information Tagged Text Database</title>
        <p>Even though most classical IR systems search the answer from plain text, we search
the answer from multi-informational tagged text database. Texts on web are ambiguous
and not structured, such as in ‘Obama is the president of America. He is born in
Hawaii.’, where ‘He’ should be co-reference resolved to ‘Obama’, and ‘Obama’ should
be disambiguated to ‘Barack Obama’. We use Stanford coreference tool for
co-reference resolution and DBPedia Spotlight for disambiguation.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>8 http://www.csie.ntu.edu.tw/~cjlin/libsvm/</title>
      <p>However, such processing is not a fast job at runtime: It takes more than three weeks
to process all texts in Wikipedia. That is why we store such information along plain
text in multi-information tagged text database. For each sentence in Wikipedia, we have
stored 1. plain text, 2. tagged text with co-reference resolution and disambiguation
information, 3. title from which Wikipedia page the sentence is, 4. PoS tagging,
dependency parsing, and SRL result. We used Apache Lucene to store and index these
attributes.</p>
      <p>When the system throws a SPARQL query, the Apache Lucene search engine finds
the related ‘tagged texts’. Our system selects all NEs from all ‘tagged texts’ as the
answer candidates. Of course most of them are irrelevant to the question, we use our SAT
and answer selection module to prune out such answer candidates.
2.6</p>
      <sec id="sec-5-1">
        <title>SPARQL query template generator</title>
        <p>We detect words from each question to extract the appropriate SPARQL template.
 Questions including arithmetic information
─ If the query contains a comparative word such as ‘deeper’ or a superlative word
such as ‘deepest’, we map these indicators to properties in KB based on mapping
rules such as PATTY. We generate SPARQL query to find the answer candidates
for further comparison. If the question contains comparative word ‘deeper’, we
select answer candidates those are ‘deeper’ than the criterion in the question. If
the question contains superlative word ‘deepest’, we sort the answer candidates
searched from the query and get the candidate which is the ‘deepest’. We used
polarity information of adjectives. If the adjectives “highest”, we sorted entities
in descent order, and the adjectives “lowest”, we sorted entities in ascending
order.
 Yes/No questions
─ Using lexical information, we detect whether the question is a ‘wh-question’ or
‘yes/no question’. If the question does not include ‘wh-keywords’ nor
‘list-question keywords’ (e.g., ‘Give’, ‘List’), we regard the question as a ‘yes/no question’.
 Simple question
─ If the question is not including arithmetic information nor yes/no question, we
generate SPARQL query for one triple. In this case, we map predicates to
properties in KB using lexical matching. If we fail to map using lexical, we try to use
semantic similarity same as our previous work [11]. The proposed system extract
important words which is related to predicate meaning. For example, just verb
such as “start” did not assign a high score to the desired predicate URI, but
concatenating additional information which make up for meaning of predicate works
well. The proposed system used ESA lib which converts strings to semantic
weighted vectors of Wikipedia concepts.</p>
        <sec id="sec-5-1-1">
          <title>Experiment</title>
          <p>We use the QALD-5 hybrid question test dataset for task 1 to evaluate the proposed
method and use multi-information tagged text database built from Wikipedia as the text
database and DBpedia 3.10 (DBPedia 2014) as the KB. We mainly compare two
approaches. One is without semantic answer type and the other uses sematic answer type
to filter answer candidates. We follow the QALD measurement. Correct states for how
many of questions were answered with an F-1 measure of 1. Partially correct specifies
how many of the questions were answered with an F-1 measure strictly between 0 and
1. Recall, Precision report the measures with respect to the number of processed
questions. F-1 Global reports the F-1 measure with respect to the total number of questions.
4</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>Results and Discussion</title>
          <p>To evaluate the QA system, we use global precision, recall and F-measure (Table 1).
With S_AT yield a higher F-measure than Without S_AT. For example, semantic
answer type detector extract “City” as answer type when the sentence “In which city
where Charlie Chaplin’s half brothers born?” is processed, so answer candidates such
as United Kingdom and England can be filtered.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Method</title>
      <p>b. Query generation: “tallest building in Japan”-&gt; result:
Tokyo_Skytree/ architect + Tokyo_Skytree
c. Success: The proposed system find tallest building in Japan is
“Tokyo_Skytree” and we mapped “architect” to “architect” which
is DBpedia predicate URI using only lexical information. Because
the lexicals in NL predicate and predicate URI are same.</p>
      <p>What is the name of the Viennese newspaper founded by the creator of the
croissant?
a. Semantic answer type: PERSON
b. Query generation: “creator of croissant/ Viennese newspaper
founded /
c. Fail: cannot map the founded to predicate URI in DBpedia.
5. In which city where Charilie Chaplin’s half brothers born?
a. Semantic answer type: City
b. Query generation: Charlie Chaplin half brother/born city
c. We find the half brother of Charlie Chaplin in multi-tagged text
database using IR approach, and using KB sparql such as
“select ?uri {Sydney_Chaplin birthplace?uri}, Finally we extract
answer such as “England” and “London”. Using Semantic answer
type, we can filter “England”.
6. Which German mathematicians were members of the von Braun rocket
group?
a.
b.</p>
    </sec>
    <sec id="sec-7">
      <title>Semantic answer type:Person</title>
      <p>Query generation: member von Braun rocket group/German
mathematician
c. Fail: cannot find relevant answer clue in multi-tagged text
database and KB.</p>
      <p>Which writers converted to Islam?
a. Semantic answer type: Person
b. Query generation: writer converted Islam
c. Fail: cannot find relevant answer clue in multi-tagged text
database and KB.</p>
      <p>Are there man-made lakes in Australia that are deeper than 100 meters?
a. Semantic answer type: Place
b. Query generation: man-made lake Australia, deeper 100
c. Success: extract answer clues of “man -made lake Australia”. The
proposed system compare the length of each named entities and
check the more than one river is deeper than 100 meters. The
proposed system find the river which is deeper than 100 meters.
9.</p>
      <p>Which movie by the Coen brothers stars John Turturro in the role of a New
York City playwright?
a. Semantic answer type: PLACE
b. Query generation: role of a New York City playwright/Coen
brother stars
c. Fail: cannot find relevant answer clue in multi-tagged text
database and KB.
10. Which of the volcanoes that erupted in 1550 is still active?
a. Semantic answer type: place
b. Query generation: volcanoes/erupted in 1550/active
c. Fail: cannot generate appropriate query to extract answer clue.
5</p>
      <sec id="sec-7-1">
        <title>Conclusion and Future work</title>
        <p>To process the complex question with reducing error in mapping NL-predicate to KB
property, we use both KBQA approach and IRQA approach. First, we search query
from multi-information tagged text data. If the results are not appropriate or related to
arithmetic, we generate SPARQL query and search KB. This combined approach works
reduce error from mapping predicate in NL to predicate URI. Furthermore, our
semantic answer type detector filter answer candidates.</p>
        <p>However, still many questions are answered wrong because of the failure of not
finding relevant answer clue, mapping NL predicate to predicate URI and query
segmentation. We will focus on extending query to find relevant answer clue well and mapping
NL-predicate to predicate URI to improve QA system in the future work. Also, we
developed semantic parsing and open information extraction for question processing
well.</p>
        <p>Acknowledgments. This work was supported by the ICT R&amp;D program of MSIP/IITP
[R0101-15-0176, Development of Core Technology for Human-like Self-taught
Learning based on a Symbolic Approach] and ATC(Advanced Technology Center) Program
- “Development of Conversational Q&amp;A Search Framework Based On Linked Data:
Project No. 10048448.
4. Lopez, V., Uren, V., Motta, E., &amp; Pasin, M. (2007). AquaLog: An ontology-driven question
answering system for organizational semantic intranets. Web Semantics: Science, Services
and Agents on the World Wide Web, 5(2), 72-105.
5. Unger, C., Bühmann, L., Lehmann, J., Ngonga Ngomo, A. C., Gerber, D., &amp; Cimiano, P.
(2012, April). Template-based question answering over RDF data. In Proceedings of the
21st international conference on World Wide Web (pp. 639-648). ACM.
6. Berant, Jonathan, and Percy Liang. "Semantic parsing via paraphrasing."Proceedings of</p>
        <p>ACL. Vol. 7. No. 1. 2014.
7. Sintek, M., &amp; Decker, S. (2002). TRIPLE – A query, inference, and transformation language
for the semantic web. In Semantic Web – ISWC 2002 (pp. 364-378). Springer Berlin
Heidelberg.
8. Mendes, Pablo N., et al. "DBpedia spotlight: shedding light on the web of documents."
Proceedings of the 7th International Conference on Semantic Systems. ACM, 2011.
9. Gabrilovich, E., &amp; Markovitch, S. (2007, January). Computing Semantic Relatedness Using</p>
        <p>Wikipedia-based Explicit Semantic Analysis. In IJCAI (Vol. 7, pp. 1606-1611).
10. Xin Li and Dan Roth. 2002. Learning Question Classifiers: The Role of Semantic
Information. Proceedings of the 19th international conference on Computational
linguistics-Volume 1. 1-7
11. Park, S., Shim, H., &amp; Lee, G. G. (2014). ISOFT at QALD-4: Semantic similarity-based
question answering system over linked data. In CLEF.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasneci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Weikum</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2007</year>
          , May).
          <article-title>Yago: a core of semantic knowledge</article-title>
          .
          <source>In Proceedings of the 16th international conference on World Wide Web</source>
          (pp.
          <fpage>697</fpage>
          -
          <lpage>706</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isele</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakob</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jentzsch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kontokostas</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendes</surname>
            ,
            <given-names>P. N.</given-names>
          </string-name>
          , ... &amp;
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>DBpedia-a large-scale, multilingual knowledge base extracted from wikipedia</article-title>
          .
          <source>Semantic Web Journal.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bollacker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paritosh</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sturge</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , J. (
          <year>2008</year>
          , June).
          <article-title>Freebase: a collaboratively created graph database for structuring human knowledge</article-title>
          .
          <source>In Proceedings of the 2008 ACM SIGMOD international conference on Management of data</source>
          (pp.
          <fpage>1247</fpage>
          -
          <lpage>1250</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>