<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <pub-date>
        <year>2009</year>
      </pub-date>
      <fpage>275</fpage>
      <lpage>281</lpage>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>arg max P (A | X ) · P (W | A).</p>
      <p>A
| an{swzer } | an{swzer }
retrieval filter</p>
      <p>model model
|CE|
P (W | A) = X P (W | ceW ) · P (ceA | A),
e=1
|CE| lAe
P (W | A) = X P (W | ce) · Y P (aje | aj).</p>
      <p>e=1 j=1
Two main types of questions were considered: Factoid questions and Denition questions. The
NIL. Details are given in Table 2.</p>
      <p>Factoid questions were further divided into the following types: Person, Organisation, Location,
For QAST 2009, two sets of questions were given: 100 written questions and manual transcriptions
European Parliament Plenary sessions in English (TC-STAR05 EPPS English corpus), which
shown in in Table 1.</p>
      <p>Time and Measure. The Denition questions were of the following types: Person, Organisation
We cleaned the data by automatically removing llers and pauses, and performed simple text
and Other. Questions where an answer cannot be found in the corpus, were to be answered by
were available: manual transcriptions and 3 ASR transcriptions. There were one run for each of
processing of abbreviations and numerical expressions to ensure consistency between the dieren t
consists of 6 spoken documents, transcribed from 3 hours of recordings. 4 versions of the corpus
the possible combinations of question sets and transcriptions, thus there were 2 4 = 8 runs, as ×
of the corresponding spoken questions. The answers were to be extracted from transcriptions of
(except in the pre-processing stage), we restrict ourselves to further analyzing run a_m, in which
results by answer type for this run.
zero. Furthermore, since the system is a factoid QA system, we could in advance predict a low
corpus, thus we chose never to return a NIL response. Therefore the score for NIL questions is
written questions and manual transcriptions are used. Figure 2 shows the break-down of the
The results of each team’s best submission for each run are plotted in Figure 1. These results
as automatic transcriptions of speeches or questions are used instead of manual transcriptions.</p>
      <p>Since our system does not treat automatic transcriptions dieren tly from manual transcriptions
Thus the only run in which we are not placed last, is the most dicult task: b_c.</p>
      <p>Our system is not able to identify whether the answer to a question can be found in the
show that we end up last of the four teams in all but one run. Our performance does not regress
5 answer candidates, as ranked by the AE module, were submitted for evaluation. Although each
tool. A set of stop words was also used. The top 100 sentences and their contexts (the
immediquestion sets and transcriptions. The ASR transcriptions lacked sentence boundaries, unlike the
scriptions by automatically aligning the text with the manual transcriptions using the GNU sdiff
team was allowed to submit two answer sets for each run, we decided to submit only one per run.
manual transcriptions, where punctuation was provided. We sentence segmented the ASR
tranately preceding and succeeding sentence), were passed to the answer extraction module. The top
of non-factoid questions, which our system is not built to answer, in addition to our inability
small size of the corpus, which means our system cannot take advantage of redundant answer
Answering on Speech Transcriptions evaluation. The results show that, using our QA system, we
are not able to achieve good performance on this task. Obvious explanations are the presence
In this paper we have given an overview of our methods and results for the CLEF 2009 Question
to identify questions which have no answer in the given corpus. Another possible reason is the
information.
such redundancy, which might be an explanatory factor for our low performance.
answer since there is little restrictions on the format of organisation names. The zero score for
answers and the eorts we made on date normalization. Organisation questions are dicult to
questions, which has an MRR which is more than double that of the MRR for the Person type,
the second best question type. This might be explained by the fairly restricted format of Time
score for Denition questions. For Factoid questions we achieve the highest performance on Time
correct answer. However, in this task the corpus was small, thus we are not able to benet from
questions.</p>
      <p>Normally our QA system utilizes a large corpus, such as the Web, and the more often an
answer candidate occurs in the context of query terms, the more likely it is to be considered a
Location score is more disappointing, since those answers mostly consist of a single geographical
term. The score for Measure questions yields little information, since there were only 2 such</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>