<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Closed Domain Question Answering for Cultural Heritage</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bernardo Cuteri</string-name>
          <email>cuteri@mat.unical.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DEMACS, University of Calabria</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper I present my research goals and what I have obtained so far into my rst year of PhD. In particular this paper is about a novel architecture for closed domain question answering and a possible application in the cultural heritage context. Unlike open domain question answering, which makes intensive use of Information Retrieval (IR) techniques, closed domain question answering systems might be built on top of a formal model with the possibility to apply formal logics and reasoning. Natural language question answering poses some nontrivial problems to tackle. We investigate such problems and propose some solutions based on AI techniques, picking the Cultural Heritage domain as a target application.</p>
      </abstract>
      <kwd-group>
        <kwd>Closed domain question answering</kwd>
        <kwd>AI</kwd>
        <kwd>NLP</kwd>
        <kwd>ASP</kwd>
        <kwd>cultural heritage</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The information need of a user often resolves in a simple question where it would
be useful to have brief answers instead of whole documents to look into. IR
techniques have proven to be very successful at locating relevant documents to the
user query into large collections, but the e ort of looking for a speci c desired
information into such documents is then left to the user. Question answering
attempts to nd direct answers to user questions.</p>
      <p>As the intuition says, answering to any kind of question, with no linguistic and
no domain restriction is a very hard task. When no restriction is made on the
domain of the questions we are talking about open domain question answering.
Instead, when questions are bound to a speci c domain we are talking about
closed (or restricted) domain question answering (CDQA).</p>
      <p>
        In open domain QA, most systems are based on a combination of Information
Retrieval and NLP techniques[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Such techniques are applied to a large corpora
of documents: rst attempting to retrieve the best documents to look into for the
answer, then selecting the paragraphs which are more likely to bear the desired
answer and nally processing the extracted paragraphs by means of NLP. Such
approach is also behind many closed domain question answering systems, but in
this context we might bene t of existing structured knowledge.
      </p>
      <p>
        Some of the very early question answering systems were designed for closed
domains and they were essentially conceived as natural language interfaces to
databases [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ][
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The idea of studying and applying closed domain question answering for the
cultural heritage domain comes from the PIUCULTURA project, which is a project
of which my university is a research partner. This project aims at
implementing a mobile system for cultural heritage fruition. My university is in charge of
research and develop techniques for the implementation of a question
answering prototype for cultural heritage. For what concerns closed domains, cultural
heritage can bene t of structured data sources: in this context, information has
already started to be saved and shared with common standards. One of the most
successful standard is the CIDOC Conceptual Reference Model. The
CIDOCcrm provides a common semantic framework for the mapping of cultural heritage
information and can be adopted by museums, libraries and archives.
Our idea is to design and implement a system capable of interpreting natural
language questions regarding cultural heritage objects and facts, map the input
questions into formal queries compliant to the CIDOC-crm model and execute
such queries to retrieve the desired information.</p>
      <p>
        In closed domains, question structures are more predictable than in open domain
and we propose to design a sophisticated module of template matching based
on a declarative formalism (Answer Set Programming [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]) for question classi
cation and query extraction. A particular feature we want to introduce is the
possibility to have dialogues instead of only atomic questions. This might help
when the initial question is ambiguous or the system needs more clari cations
to provide an accurate answer. Also, the fact that the system is based on
formal queries rather than statistical methods might lead to a more robust answer
creation, with the possibility to obtain a step-by-step justi cation of the answer
and an easier validation. In the following sections we present an architectural
model of the system and provide some more details about the tasks involved in
the question answering process.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>System Architecture and Working Principles</title>
      <p>In the following subsections we are going to break through the process and
analyze it step by step.</p>
      <p>Fig. 1. Simpli ed architecture with single question interaction
2.1</p>
      <sec id="sec-2-1">
        <title>Question Processing</title>
        <p>This is the main NLP step. Fortunately, tokenization, POS tagging and natural
language parsing have decades of research behind and there are plenty of tools
around that are able to solve such problems e ciently. With respect to the
Cultural Heritage domain, something we can not overlook is the importance of
entity recognition as we might have proper nouns of artefacts or persons that
must not be mistakenly treated by NLP tools (e.g. splitting the title of a painting
into distinct grammatical parts).</p>
        <p>First the question is tokenized and tagged with part-of-speech (POS) tags. Then
a natural language parser is in charge of extracting grammatical relations (a.k.a.
typed dependencies) from the text (e.g. who is the subject of what verb, what
is the object and so on).
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Template Matching</title>
        <p>Questions are classi ed and transformed into formal queries by means of
template matching. In this context, templates represent the structure of typical
questions. If a certain template is matched we can infer something about the
question type. Every question template is accompanied with a formal query in
which some slots are empty and are lled with terms extracted from the question
that matches the template. For example, imagine we have the template for
questions of the type Who verb object : the question Who painted Guernica? matches
such pattern and a corresponding query can be created.Guernica and painted
might then be used as constants in the query, lling the empty slots mentioned
before.</p>
        <p>
          We want to investigate the possibility to implement template matching with
Answer Set Programming [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ](ASP). ASP evolved from deductive databases, logic
programming and nonmonotonic reasoning. It is a exible language for
knowledge representation and reasoning, and for declarative problem solving, and e
cient systems are available[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. ASP is thus a concrete tool for developing complex
applications by just specifying a set of logic rules of the form Head : Body,
where Body is a conjunction of possibly negated atoms, and Head is a
disjunction of atoms. The stable models, or answer sets, of an ASP program correspond
to the solutions of the modelled problem. The programmer does not need to
provide an algorithm for solving a problem with ASP; rather, she speci es the
properties of the desired solution for its computation by means of a collection
of logic rules called logic program. The stable models or answer sets of an ASP
program correspond to the solutions of the modelled problem. The logic rule is
an expression that looks like Head : Body, where Body is a logic conjunction
possibly involving negation, and Head is either an atomic formula or a logic
disjunction. The language of ASP, besides disjunction in rule heads and
nonmonotonic negation in rule bodies, features also special atoms for de ning aggregates,
strong constraints for selecting solutions, and weak constraints for solving
optimization problems. The implementation details go beyond the scope of this
paper, but we can say that ASP is a good candidate for a fast and declarative
implementation of template matching. This step requires a small preprocessing
step in which the input (i.e. words and associated grammatical relations and
parts-of-speech) is transformed into ASP facts. A simple example of a possible
template for matching questions of the type Who-verb-object is the following:
template(1, bt(W1,W2)):- textWord(1, who), gr(2,3,dobj),
textWord(2,W1), textWord(3,W2).
        </p>
        <p>Where the textWord predicate denotes the presence of a certain word in a certain
position and the gr predicate denotes a grammatical relation between two words.
gr(2,3,dobj) means that word in position 3 is the object of word in position 2.
This template is a bit simpli ed and does not take POS tags into account, but
gives an idea of how to implement a template in ASP.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Query Expansion and Contextualization</title>
        <p>The template matching result is a formal query. Sometimes, to be e ective,
the query has to be expanded with context information and/or word semantic
information. We can try to understand the importance of those two by providing
an example.</p>
        <p>
          Let's say that we asked When was the monalisa created?. An admissible following
question could be And who did it?. The pronoun it clearly stands for the painting,
but in order to understand it, the system has to get context information or at
least store question histories. Another problem that we can analyze with the
previous example is the following: let's say that our knowledge base contains the
information Leonardo da Vinci painted the monalisa. We know that, if someone
painted something we can also say that they did it. The question answering
system has to deal with such and other similar problems. A possible solution is
to expand the query by using synonyms, hyperonyms and other word semantic
relations. Fortunately, there are some available encyclopedic dictionaries that
are able to provide such relations. Among them there is BabelNet[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] which has
also the desirable property of being multilingual and this might help in case we
want to extend the work to di erent languages.
2.4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Query Execution and Answer Creation</title>
        <p>
          In our model, the query is executed against a structured knowledge base. Query
results (if any) can then be used to build a natural language answer with a
mechanism similar to template matching, but in the inverse direction. A possible
approach is that each question template is paired to an answer template. The
answer template may have empty slots for answer terms and it is used by the
answer creation module to build the NL answer once the query has been executed
successfully.
In this paper we presented an architecture and some implementation ideas for a
closed domain question answering system and discussed about the tasks involved
in the process. The work is currently under development, studies have been
conducted to investigate current research trends in question answering and available
solutions. At this moment we have developed a small QA prototype capable of
answering simple questions. It uses the Stanford parser[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] for tokenization,
POStagging and parsing, integrates Babelnet[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] for query expansion, and supports
some common question types. For example it is possible to ask who performed
a certain action on a certain object, or where a certain object is located.
Templates cover di erent ways to express the same question like where is Guernica
located? or in what museum is Guernica located?. We are investigating on how
to cope with question nuances, trying to design some more general templates for
questions that are not perfectly matched by a simpler template. This is where
ASP (with disjunction, weak constraints and aggregates) might play a crucial
role as opposed to less expressive languages like Datalog.
        </p>
        <p>We started to create a broad catalogue of possible questions in order to
create more complex templates and extend the system to adapt to more di cult
questions. We are now planning to extend this approach to its limits, trying to
manage a broad set of questions on cultural heritage. If it works ne we also plan
to add multilingual support checking if the template system is easy to extend
to di erent languages. We also want to investigate how to implement non-trivial
dialogues centered around questions instead of only single atomic questions.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Woods</surname>
            ,
            <given-names>W. A.</given-names>
          </string-name>
          (
          <year>1973</year>
          , June).
          <article-title>Progress in natural language understanding: an application to lunar geology</article-title>
          .
          <source>In Proceedings of the June 4-8</source>
          ,
          <year>1973</year>
          , national computer conference and exposition (pp.
          <fpage>441</fpage>
          -
          <lpage>450</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Green</given-names>
            <surname>Jr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. F.</given-names>
            ,
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            ,
            <surname>Chomsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Laughery</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          (
          <year>1961</year>
          , May).
          <article-title>Baseball: an automatic question-answerer</article-title>
          .
          <source>In Papers presented at the May 9-11</source>
          ,
          <year>1961</year>
          ,
          <article-title>western joint IRE-AIEE-ACM computer conference</article-title>
          (pp.
          <fpage>219</fpage>
          -
          <lpage>224</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hirschman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaizauskas</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Natural language question answering: the view from here</article-title>
          .
          <source>natural language engineering</source>
          ,
          <volume>7</volume>
          (
          <issue>04</issue>
          ),
          <fpage>275</fpage>
          -
          <lpage>300</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S. P.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network</article-title>
          .
          <source>Arti cial Intelligence</source>
          ,
          <volume>193</volume>
          ,
          <fpage>217</fpage>
          -
          <lpage>250</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Gelfond</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lifschitz</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          (
          <year>1991</year>
          ).
          <article-title>Classical negation in logic programs</article-title>
          and disjunctive databases.
          <source>New generation computing</source>
          ,
          <volume>9</volume>
          (
          <issue>3-4</issue>
          ),
          <fpage>365</fpage>
          -
          <lpage>385</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Calimeri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ianni</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ricca</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alviano</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bria</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Catalano</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , ...
          <string-name>
            <surname>Manna</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2011</year>
          , May).
          <article-title>The third answer set programming competition: Preliminary report of the system competition track</article-title>
          .
          <source>In International Conference on Logic Programming and Nonmonotonic Reasoning</source>
          (pp.
          <fpage>388</fpage>
          -
          <lpage>403</lpage>
          ). Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C. D.</given-names>
          </string-name>
          (
          <year>2003</year>
          ,
          <article-title>July)</article-title>
          .
          <article-title>Accurate unlexicalized parsing</article-title>
          .
          <source>In Proceedings of the 41st Annual Meeting on Association for Computational LinguisticsVolume</source>
          <volume>1</volume>
          (pp.
          <fpage>423</fpage>
          -
          <lpage>430</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>