<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>miraQA: Initial experiments in Question Answering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>C. de Pablo</string-name>
          <email>cdepablo@uc3m.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J.L. Martínez-Fernández</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>P. Martínez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. Villena</string-name>
          <email>jvillena@daedalus.es</email>
          <email>jvillena@it.uc3m.es</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. M. García-Serrano</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. M. Goñi</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>C. González</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Advanced Databases Group, Computer Science Department, Universidad Carlos III de Madrid</institution>
          ,
          <addr-line>Avda. Universidad 30, 28911 Leganés, Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Artificial Intelligence Department, Universidad Politécnica de Madrid.</institution>
          <addr-line>Campus de Montegancedo s/n, Boadilla del Monte 28660</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>DAEDALUS - Data, Decisiond and Language, S.A. Centro de Empresas “La Arboleda”</institution>
          ,
          <addr-line>Ctra. N-III km. 7,300 Madrid 28031</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Department of Mathematics Applied to Information Techmologies, E.T.S.I. Telecomunicación, Universidad Politécnica de Madrid</institution>
          ,
          <addr-line>Avda. Ciudad Universitaria s/n, 28040 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Department of Telematic Engineering, Universidad Carlos III de Madrid</institution>
          ,
          <addr-line>Avda. Universidad 30, 28911 Leganés, Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present the miraQA system that constitutes MIRACLE first experience in Question Answering for monolingual Spanish and has been developed for QA@CLEF 2004. The architecture of the system is described and details of our approach to Statistical Answer Extraction based on Hidden Markov Models are presented. One run that uses last year question set for training purposes has been submitted. The results are presented together with ideas for improvement.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>System Description</title>
      <p>miraQA, the system developed for QA@CLEF 2004 by the MIRACLE group represents our first attempt to face</p>
      <sec id="sec-2-1">
        <title>Question Answering . As Spanish is our mother tongue, we have developed the system for the monolingual</title>
      </sec>
      <sec id="sec-2-2">
        <title>Spanish task, where the group is familiar with available tools. Despite the system was developed for Spanish, we had in mind that it should be easily adapted to other target languages. For that reason, we have explored the potential of statistical models for Answer Extraction. Besides, most of the tools that we are using, like POS taggers or partial parsers, are available for almost every other european language.</title>
      </sec>
      <sec id="sec-2-3">
        <title>The general architecture of miraQA system for QA follows the classical structure in three modules and is presented in the following figure:</title>
        <sec id="sec-2-3-1">
          <title>Question Analysis</title>
        </sec>
        <sec id="sec-2-3-2">
          <title>Answer Extraction</title>
          <p>Answer
ranking
Answer
Recog.</p>
          <p>Anchor
searching
POS
+Parsing
Sen
tence</p>
        </sec>
        <sec id="sec-2-3-3">
          <title>Document retrieval</title>
          <p>IR
engine
Sentence
extractor
EFE94/95
Qclass
Answer
model</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>A fourth module for Answer Evaluation would be required to address this year novelty of providing a confidence measure for every answer. Although we appreciate the usefulness of that feature for the final user, we have not been able to include that module in our system due to time constraints.</title>
      </sec>
      <sec id="sec-2-5">
        <title>In addition to the QA system, we have also developed another system to train the used Hidden Markov Models</title>
        <p>in our answer extraction phase. The system uses questions and answers to build queries that are posed to</p>
      </sec>
      <sec id="sec-2-6">
        <title>Google. Snippets of the results are extracted and used to build a model for the co-ocurrence of question terms and answers. In order to build the models we have used CLEF 2003 evaluation question set.</title>
        <p>
          Question analysis
This module classifies the questions according to a manual taxonomy shown in Table 1 and composed of 17
classes. The taxonomy was decided considering mainly answer types. For some of them we decided to split or
clonflate the classes depending on the frequency of appearance question-answer in QA@CLEF 2003 evaluation
set. Questions are partially analysed using ms-tools [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. We used MACO tagger and TACAT parser (slightly
modified to avoid attacchment of PP chunks). Once the questions are partially parsed, a set of simple rules is
applied to classify the questions determining its type, the type of the answer that is expected and assign a set of
semantic tags to some of the chunks according to the relations they have with the answer. A simple example for
the question "¿Cual es la capital de Croacia?" is shown in Figure 2:
Fia
pt
vsip
da
        </p>
        <p>nc
¿</p>
        <p>Cuál
es
la
capital
sps
de</p>
        <p>Fit
?</p>
      </sec>
      <sec id="sec-2-7">
        <title>Name</title>
      </sec>
      <sec id="sec-2-8">
        <title>Person</title>
      </sec>
      <sec id="sec-2-9">
        <title>Group</title>
      </sec>
      <sec id="sec-2-10">
        <title>Count</title>
        <p>espec grup-nom
prep
grup-nom
Document retrieval</p>
      </sec>
      <sec id="sec-2-11">
        <title>The IR module retrieves the top most relevant documents for a query and extracts the sentences that contain any</title>
        <p>
          of the words expressed in a query. After the question is analyzed, words that have a semantic tag assigned, are
used in the query. For robustness purposes, the semantic tags are scanned again to remove stopwords and a
query with all the terms is built and given to the IR engine. Our system uses Xapian [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] probabilistic engine to
search for the most relevant documents. The last step consists on tokenizing the document using Daedalus
        </p>
      </sec>
      <sec id="sec-2-12">
        <title>Tokenizer [4] to extract those sentences that contain any of the words or stems that appeared in the query. The system assigns two scores to every sentence, the relevance measure provided by Xapian to the document and another figure proportional to the number of terms that were found in the sentence.</title>
        <p>Answer extraction
The answer extraction module uses a statistical approach to answer pinpointing that is based on a
syntacticsemantic context model of the answer built for any of the question-answer types. The following operations are
performed:
1. Parsing and Anchor Searching. The sentences provided by the IR modules containing terms
from the questions are parsed in a similar way as questions and training sentences using the
ms-tools. Once parsed, the chunks containing the question terms are substituted by their
semantic tags and constitute what we have called anchor terms. Finally, sentences are chunked
in pieces that form a window of words around anchor terms and passed to the next module.</p>
      </sec>
      <sec id="sec-2-13">
        <title>2. Answer Recognition. Pieces built in this way are passed to the answer extraction module that</title>
        <p>uses the HMM model. A variant of the N-best recognition strategy is used to identify the most
probable sequence of states that originated the POS sequence and identifies an answer as the
sequence of words that has been generated from the answer state. The recognition algorithm is
guided by the semantic information in order to find a path that passes through the answer state.</p>
      </sec>
      <sec id="sec-2-14">
        <title>Besides, the algorithm provides a score for every path computed as the sum of the log of the</title>
        <p>probabilities for that path and sequence.</p>
      </sec>
      <sec id="sec-2-15">
        <title>3. Ranking. Candidate answers are conflated when they present small differences in the outer form due to stopwords, for example. Finally, the candidate answers are ranked attending to a weighted score that takes into account the score of the document, the sentence, the path followed in recognition and their lengths.</title>
        <p>Training of Answer Context</p>
      </sec>
      <sec id="sec-2-16">
        <title>Models that are used in the answer extraction phase are previously trained from examples. For the training of the</title>
        <p>models we used the question-answer set provided for QA at CLEF 2003. Questions are analyzed in the same
way that they are in the main QA system. Question terms and the answers strings are combined in queries that
we send to Google using the Google API. Snippets for the top 100 results are retrieved and stored to build the
model. They are splitted in sentences, analyzed and chunks having questions and answer terms are retagged.</p>
      </sec>
      <sec id="sec-2-17">
        <title>The tag is either the semantic class assigned to that term in the question or the answer tag (##ANSWER##).</title>
      </sec>
      <sec id="sec-2-18">
        <title>Only sentences containing the answer and at least one of the other semantic tags are selected to train the model.</title>
        <sec id="sec-2-18-1">
          <title>Question Analysis</title>
        </sec>
      </sec>
      <sec id="sec-2-19">
        <title>The machine that we built to extract answers is a Hidden Markov Model in wich the states are the syntactic</title>
        <p>semantic tags assigned to the chunks while the emitted symbols are the POS tags assigned to the classes. To
estimate the transition and emission probabilities we have counted the frequencies of the bigrams for POS-POS
and POS-CHUNKS. In order to account for states or symbols that were not seen in the Google sentences we
have used the simple add-one smoothing technique. We build a model like this one for every question-answer
type that uses a closed set of one to three semantic tags.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results analysis</title>
      <sec id="sec-3-1">
        <title>We submitted one run for the monolingual Spanish task (mira041eses) that provides one exact answer to every</title>
        <p>question. Our system is unable to compute the confidence measure and we limited us to assign the default value
of 0. There are two main kinds of questions, factoid and definition and we have tried the same approach for both
of them. Besides, the question set contains some questions whose answer could not be found in the document
corpus and the valid answer in that case is the NIL string.</p>
      </sec>
      <sec id="sec-3-2">
        <title>The results obtained for our run mira041eses are outlined in Table 2</title>
        <p>Question Type
Right</p>
        <p>Wrong</p>
        <p>IneXact Unsupported</p>
      </sec>
      <sec id="sec-3-3">
        <title>Factoid</title>
      </sec>
      <sec id="sec-3-4">
        <title>Definition</title>
        <p>TOTAL
18</p>
        <p>0
18
157
17
174
4
3
7
Results are fairly low if we compare them with other systems. We attribute these bad results to the fact that the
system is in a very early stage of development and tuning. We have obtained several conclusions from the
analysis of correct and wrong answers that will guide our future work. The extraction algorithm is working
better for factoid questions than definitional. Obviusly, among factoid questions results are also better for certain
question-answer classes (DATE,NAME...) which are found often in the training set of questions. This is
remarkable as the algorithm extract answers of the proper type even if they are incorrect. We were aware that
such effect could appear as the amount of questions in each of the question-answers classes were unevenly
distributed in QA@CLEF 2003 question set. There were specially few questions that we could classify as
definitions wich diminish the amount of training data available and therefore the accuracy of the probabilities.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Another fact noteworthy in our HMM algorithm is that is somewhat greedy when trying to identify answer and in that case shows some preference for words appearing near anchor terms . Finally, the algorithm is actually doing two jobs at a time as it identifies answers and, in some way, analyses entities according to patterns that were present in training answers of the same kind.</title>
      </sec>
      <sec id="sec-3-6">
        <title>Another important source of errors in our system is induced by the document retrieval process and the way we</title>
        <p>posed questions and score documents. In our system all the terms that have been assigned a semantic tag will be
used in queries and as anchors. Some terms are not very discriminating, specially if they are considered against
proper names, and therefore lot of noisy documents are retrieved. As well, the simple scoring schema that we
used for sentences (one token-one point) contributes to mask some of the useful fragments.</p>
      </sec>
      <sec id="sec-3-7">
        <title>Errors are also generated by the question classification step as it is unable to handle some of the new surface forms introduced in this year question set. For that reason a catch all classification was also defined and used as a ragbag, but results were not expected to be good for that class.</title>
      </sec>
      <sec id="sec-3-8">
        <title>The evaluation also provides results for the percentage of NIL answers that we have returned. In our case we returned 74 NIL answers and only 11 of them were correct (14.86%). NIL values were returned when the process did not provided any answer and their high value is due to the chaining of the other problems mentioned above.</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Future work</title>
      <sec id="sec-4-1">
        <title>Several lines for further research are open along with the deficiencies that we have detected in our system. We</title>
        <p>are also intending to extend the same approach to other languages both in the source and the target language.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Some attempts to address different language for the question have already been done by translating questions, but the low quality of the translations would have obligued us to extend the set of question patterns or to develop correction mechanisms. These problems as well as the errors caused by new questions not addressed in our schema are claiming for a more robust approach to question classification and analysis.</title>
      </sec>
      <sec id="sec-4-3">
        <title>One of the most straightful improvements we should introcuce in miraQA is a module for specific answer type</title>
        <p>recognition that address Named Entity Recognition but also other common answer types as dates, time, amounts,
etc. With such extension the answer extraction task would be reduce to identify proper units based on context.</p>
      </sec>
      <sec id="sec-4-4">
        <title>We believe that with this improvement the method would be able to reduced the inexact ratio and address short definitional questions.</title>
      </sec>
      <sec id="sec-4-5">
        <title>Results show that for the answer extraction mechanism to work properly, a thorough training is needed. We are</title>
        <p>already carrying out experiments to determine the amount of training data that would be needed in order to
improve recognition results. We would likely need to acquire or generate larger question-answer corpus. Several
improvements in the learning and recognition machine would be definetily beneficial and therefore several
extensions to Hidden Markov Model and other statistical finite state approaches are under study, as well as more
effective methods for learning the structure and parameters of these machines.</p>
      </sec>
      <sec id="sec-4-6">
        <title>Besides the previous improvements a more careful look at the interfaces and dependencies between the different</title>
        <p>subsystems is also needed. In that sense, the main work involves developing better strategies to query de
document database and retrieve the most meaningful passages. We also need to estimate more precisely the
contribution of any of the modules and elaborate a method to combine this information in a succesful measure of
answer confidence as this would greatly increase the acceptance of QA systems by the final user.</p>
      </sec>
      <sec id="sec-4-7">
        <title>Finally, as we have stated from the beginning, we believe that an statistical approach could be practical from a software engineering approach and would allow the rapid development of baseline QA systems for different languages and domains.</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <sec id="sec-5-1">
        <title>This work has been partially supported by the projects OmniPaper (European Union, 5th Framework</title>
      </sec>
      <sec id="sec-5-2">
        <title>Programme for Research and Technological Development, IST-2001-32174) and MIRACLE (Regional</title>
      </sec>
      <sec id="sec-5-3">
        <title>Government of Madrid, Regional Plan for Research, 07T/0055/2003).</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Abney</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Collins</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singhal</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <source>Answer Extraction Applied Natural Language Processing (ANLP): Proceedings of the Conference</source>
          .
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Atserias</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carmona</surname>
          </string-name>
          , I. Castellón,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cervell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Civit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Màrquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.A.</given-names>
            <surname>Martí</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Padró</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Placer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Rodríguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Taulé</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Turmo</surname>
          </string-name>
          <article-title>Morphosyntactic Analysis and Parsing of Unrestricted Spanish Text</article-title>
          .
          <source>Proceedings of the 1st International Conference on Language Resources and Evaluation (LREC'98)</source>
          . Granada, Spain,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Baeza-Yates R. Ribeiro-Neto</surname>
            <given-names>B.</given-names>
          </string-name>
          (
          <year>1999</year>
          )
          <article-title>Modern Information Retrieval</article-title>
          . Addison Wesley.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Brill E. Lin J. Banco</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2001</year>
          )
          <article-title>Data-Intensive Question Answering</article-title>
          .
          <source>Proceedings of TREC 2001</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>5. Daedalus Website: http;//www.daedalus.es</mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Jurafsky D. Martin</surname>
            <given-names>J.H.</given-names>
          </string-name>
          (
          <year>2000</year>
          )
          <article-title>Speech</article-title>
          and
          <string-name>
            <given-names>Language</given-names>
            <surname>Processing</surname>
          </string-name>
          . Prentice Hall.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Manning</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schütze</surname>
            <given-names>H.</given-names>
          </string-name>
          (
          <year>1999</year>
          )
          <article-title>Foundations of Statistical Natural Language Processing</article-title>
          .. MIT Press
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Magnini</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romagnoli</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vallin</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herrera</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peñas</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peinado</surname>
            <given-names>V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verdejo F and de Rijke M. The Multiple Language Question Answering Track at CLEF 2003</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Vicedo</surname>
            <given-names>J.L.</given-names>
          </string-name>
          (
          <year>2003</year>
          )
          <article-title>Recuperando información de alta precisión</article-title>
          . Los sistemas de Búsqueda de Respuestas.
          <source>Phd Thesis</source>
          . Universidad de Alicante.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>10. Xapian Website: http://www.xapian.org/</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>