<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Do it your own (DIY) Jeopardy Question Answering System</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andre Freitas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edward Curry</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Digital Enterprise Research Institute (DERI) National University of Ireland</institution>
          ,
          <addr-line>Galway</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The evolution and maturity of semantic technologies techniques and frameworks are bringing functionalities which were once considered academic or prototypical into real-life applications. Products such as IBM Watson [1] and Siri are examples of applications which are heavily leveraged on state-of-the-art semantic technologies. These systems provide a synthesis of the functionalities which are available for general applications today such as: natural language search and queries over large-scale data, semantic exibility and integration between structured and unstructured resources. The success of these projects in demonstrating the potential of existing technologies lies on the fact that they bring into a single system approaches from Natural Language Processing (NLP), Semantic Web (SW), Information Retrieval (IR) and Databases. This work demonstrates Treo, a framework which converges elements from NLP, IR, SW and Databases, to create a semantic search engine and question answering (QA) system for heterogeneous data. Jeopardy and Question Answering queries over open domain structured and unstructured data are used to demonstrate the approach. In this work, Treo is extended to cope with unstructured text in addition to structured data. The setup of the framework is done in 3 steps and can be adapted to other datasets in a simple DIY process.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Motivation</title>
      <p>and languages, not requiring the manual creation e ort of ontologies (Treo does
not rely on ontology-based reasoning for semantic approximation).</p>
      <p>In addition to queries over structured data, this work extends the query
mechanism for searching entities in unstructured text. Both structured and
unstructured data are linked in an entity-centric semantic index (Figure 1 (B)).
The elements of the query processing approach are depicted in Figure 1 (A).</p>
      <p>
        Two di erent query processing strategies are used:
- Query processing over structured data: In the query pre-processing phase,
the natural language query is analyzed by the Interpreter component, where a
set of query triple patterns and features are detected in the user query. The
second phase consists of the vocabulary independent query processing approach
which de nes a sequence of search and data transformation operations over the
structured data graph embedded in the Space [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], targeting the
maximization of the semantic matching with the query. The Query Planner generates
the sequence of semantic search, navigation and transformation operations over
the graph data, which de nes the query processing plan, based on a set of query
features which are determined in the pre-processing phase. The third phase
consists in the execution of the query processing plan operations over the Space
index.
- Query processing over structured &amp; unstructured data: In case the
query is not addressed by the available structured data, the query can be
processed against both structured data and unstructured text in the entity-centric
index. The query pre-processing approach for this query type consists on the
detection of the query focus by the application of POS Tag based rules and by the
detection and resolution of named entities in the query. The query plan consists
of the composition of keyword-search operations over the text segments
associated with entities, distributional search operations over structured data, and
keyword search over associated entities. A ranking function weights the results of
all operations, also taking into account the cardinality for each entity (number of
associated entities, facts and text segments). The initial top-20 entity results are
re-ranked based on the computation of the distributional semantic relatedness
scores between the query focus phrase and the associated entity types.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>DIY Setup Process</title>
      <p>
        The setup of the Treo platform for a new dataset consists in the creation of a
semantic index for both structured and unstructured data, which requires three
steps:
1. Construction of the distributional semantic model: Consists on the use of a
large-scale reference corpora to build the distributional semantic reference
model [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In this demonstration Wikipedia 2006 is used as the reference
corpus and Explicit Semantic Analysis (ESA) is the distributional semantic
model.
2. Semantic indexing of structured data: Consists in the indexing of structured
data using the distributional semantic reference model [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The framework
Dependency
      </p>
      <p>Parser
Entity Search
Querry
Processor</p>
      <p>Reference</p>
      <p>Corpora
Distributional
Compositional</p>
      <p>Index
(Ƭ-Space)
Entity+-Text</p>
      <p>Index
search operations
Distributional</p>
      <p>Search</p>
      <p>DBpedia
:company :Bad_Robot_Productions
:creator :J._J._Abrams
:format :Action_(fiction)
:location :Walt_Disney_Studios_(Burbank)
:location :Burbank,_California
:network :American_Broadcasting_Company
:numberOfEpisodes 105
:numberOfSeasons 5
:releaseDate 2001-09-30
:starring :Amy_Acker
:starring :Jennifer Garner
...</p>
      <p>Datasets :Alias(TV Series)</p>
      <p>YAGO
:type :2006AmericanTelevisionSeriesEndings
:type :2001AmericanTelevisionSeriesDebuts
:type :BadRobotProductions</p>
      <p>Wikipedia
:hasSentence :Jack Bristow (:Victor Garber) is</p>
      <p>Sydney's father and also works for
:SD-6 as a double agent for the :CIA.
:hasSentence It stars :Jennifer Garner as :Sydney</p>
      <p>Bristow, a CIA agent.
...
Pre-Processing</p>
      <p>Query</p>
      <p>Interpreter
feeudsbearck Disambiguation
pre-processed query
[[:Bil Clinton]] - daughter
married</p>
      <p>Query Processing</p>
      <p>Query</p>
      <p>Planner
user
feedback</p>
      <p>Disambiguation</p>
      <p>Operators
Answer:
Yes</p>
      <p>Triples:
Margareth Thatcher’s type is English Chemists</p>
      <p>Margareth Thatcher’s profession is chemist</p>
      <p>takes as input data any dataset following an Entity-Attribute-Value (EAV)
format. DBpedia 3.7 and YAGO are used as the demonstration datasets.
3. Unstructured data entity-centric indexing: This step takes as input a text
collection, recognizes the named entities based on the structured data
previously indexed, aligning it with the indexed structured data. The
demonstration uses Wikipedia 2013 as the test collection.</p>
      <p>The steps are executed by calling one script, which takes as input the three
types of resources (reference corpora, structured datasets and unstructured texts).
After the setup, natural language queries can be executed against the structured
and unstructured data indexes. Figure 1 shows the components of the Treo
architecture (A) and an example of the entity-centric linking between structured
and unstructured data (B).
4</p>
    </sec>
    <sec id="sec-3">
      <title>Demonstration</title>
      <p>The system is demonstrated over the open-domain DBpedia 3.7 /YAGO RDF
datasets and Wikipedia 2013 text data. The RDF datasets consist of 128,071,259
triples (17GB) loaded into the Treo index for structured data. A set of natural
language queries from the Jeopardy challenge1 and from the Question Answering
over Linked Data challenge2 are used to demonstrate the system. In the
demonstration, users input free natural language queries and the system returns two
1 http://j-archive.com/
2 QALD-1, http://www.sc.cit-ec.uni-bielefeld.de/qald-1, 2011
types of results: (i) a list of highly related triples or (ii) post-processed results,
depending on the query type.</p>
      <p>Figure 2 (2) shows the output of a query over the structured data index
for the query `Was Margaret Thatcher a chemist?'. In addition to the
postprocessed answer, which provides a direct (QA-style) answer for the query, the
mechanism shows the justi cation for the answer with the supporting triples.
Figure 2 (1) shows a query over structured data with a complex query plan
(`Which cities in New Jersey have more than 10000 inhabitants?' ). Figure 2 (3)
and (4) show examples of Jeopardy queries, which typically provide a natural
language description of a named entity or concept (for example: `Sydney's dad,
Jack, was a CIA double agent working against SD-6 on this Jennifer Garner
show' ). Further examples can be found online3.</p>
      <p>Acknowledgments. This work was funded by SFI Ireland (SFI/08/CE/I1380).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>D.</given-names>
            <surname>Ferrucci</surname>
          </string-name>
          et al.,
          <string-name>
            <surname>Building</surname>
            <given-names>Watson</given-names>
          </string-name>
          :
          <article-title>An Overview of the DeepQA Project</article-title>
          ,
          <source>AI Magazine</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>A.</given-names>
            <surname>Freitas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Curry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Oliveira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. O</given-names>
            <surname>'Riain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Distributional</given-names>
            <surname>Structured</surname>
          </string-name>
          <article-title>Semantic Space for Querying RDF Graph Data</article-title>
          .
          <source>International Journal of Semantic Computing (IJSC)</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>