<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jordi Turmo</string-name>
          <email>turmo@lsi.upc.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pere Comas</string-name>
          <email>pcomas@lsi.upc.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sophie Rosset</string-name>
          <email>rosset@limsi.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lori Lamel</string-name>
          <email>lamel@limsi.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicolas Moreau</string-name>
          <email>moreau@elda.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Djamel Mostefa</string-name>
          <email>mostefa@elda.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Question Answering, Spontaneous Speech Transcripts</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ELDA/ELRA.</institution>
          <addr-line>Paris.</addr-line>
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Experimentation</institution>
          ,
          <addr-line>Performance, Measurement</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LIMSI.</institution>
          <addr-line>Paris.</addr-line>
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>TALP Research Centre (UPC). Barcelona.</institution>
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2008</year>
      </pub-date>
      <abstract>
        <p>This paper describes the experience of QAST 2008, the second time a pilot track of CLEF has been held aiming to evaluate the task of Question Answering in Speech Transcripts. Five sites submitted results for at least one of the five scenarios (lectures in English, meetings in English, broadcast news in French and European Parliament debates in English and Spanish). In order to assess the impact of potential errors of automatic speech recognition, for each task contrastive conditions are with manual and automatically produced transcripts. The QAST 2008 evaluation framework is described, along with descriptions of the five scenarios and their associated data, the system submissions for this pilot track and the official evaluation results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Question Answering (QA) technology aims at providing relevant answers to natural language
questions. Most Question Answering research has focused on mining document collections
containing written texts to answer written questions [
        <xref ref-type="bibr" rid="ref3 ref6">3, 6</xref>
        ]. Documents can be either open domain
(newspapers, newswire, Wikipedia...) or restricted domain (biomedical papers...) but share, in
general, a decent writing quality, at least grammar-wise. In addition to written sources, a lot
(and growing amount) of potentially interesting information appears in spoken documents, such
as broadcast news, speeches, seminars, meetings or telephone conversations. The QAST track
aims at investigating the problem of question answering in such audio documents.
Current text-based QA systems tend to use technologies that require texts to have been written
in accordance with standard norms for written grammar. The syntax of speech is quite different
than that of written language, with more local but less constrained relations between phrases,
and punctuation, which gives boundary cues in written language, is typically absent. Speech also
contains disfluencies, repetitions, restarts and corrections. Moreover, any practical application
of search in speech requires the transcriptions to be produced automatically, and the Automatic
Speech Recognizers (ASR) introduce a number of errors. Therefore current techniques for
textbased QA need substantial adaptation in order to access the information contained in audio
documents. Preliminary research on QA in speech transcriptions was addressed in QAST 2007, a
pilot evaluation track at CLEF 2007 in which systems attempted to provide answers to written
factual questions by mining speech transcripts of seminars and meetings [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>This paper provides an overview of the second QAST pilot evaluation. Section 2 describes the
principles of this evaluation track. Sections 3 present the evaluation framework and section 4 the
systems that participated. Section 5 reports and discusses the achieved results, followed by some
conclusions in Section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The QAST 2008 task</title>
      <p>The objective of this pilot track is to develop a framework in which QA systems can be evaluated
when the answers have to be found in speech transcripts, these transcripts being either produced
manually or automatically. There are five main objectives to this evaluation:
• Motivating and driving the design of novel and robust QA architectures for speech
transcripts;
• Measuring the loss due to the inaccuracies in state-of-the-art ASR technology;
• Measuring this loss at different ASR performance levels given by the ASR word error rate;
• Comparing the performance of QA systems on different kinds of speech data (prepared
speech such as broadcast news (BN) or parliamentary hearings vs. spontaneous in meeting
for instance);
• Motivating the development of monolingual QA systems for languages other than English.
In the 2008 evaluation, as in the 2007 pilot evaluation, an answer is structured as a simple [answer
string, document id] pair where the answer string contains nothing more than the full and exact
answer, and the document id is the unique identifier of the document supporting the answer. In
2008, for the tasks on automatic speech transcripts, the answer string consisted of the
&lt;starttime&gt; and the &lt;end-time&gt; giving the position of the answer in the signal. Figure 1 illustrates
this point comparing the expected answer to the question What is the Vlaams Blok? in a manual
transcript (the text criminal organisation) and in an automatic transcript (the time segment
1019.228 1019.858). A system can provide up to 5 ranked answers per question.
Question: What is the Vlaams Blok?
Manual transcript: the Belgian Supreme Court has upheld a previous ruling that declares
the Vlaams Blok a criminal organization and effectively bans it .</p>
      <p>Answer: criminal organisation</p>
      <p>A total of ten tasks were defined for this second edition of QAST covering five main task scenarios
and three languages: lectures in English about speech and language processing (T1), meetings in
English about design of television remote controls (T2), French broadcast news (T3) and European
Parliament debates in English (T4) and Spanish (T5). The complete set of tasks are:
• T1a: QA in manual transcriptions of lectures in English.
• T1b: QA in automatic transcriptions of lectures in English.
• T2a: QA in manual transcriptions of meetings in English.
• T2b: QA in automatic transcriptions of meetings in English.
• T3a: QA in manual transcriptions of broadcast news for French.
• T3b: QA in automatic transcriptions of broadcast news for French.
• T4a: QA in manual transcriptions of European Parliament Plenary sessions in English.
• T4b: QA in automatic transcriptions of European Parliament Plenary sessions in English.
• T5a: QA in manual transcriptions of European Parliament Plenary sessions in Spanish.
• T5b: QA in automatic transcriptions of European Parliament Plenary sessions in Spanish.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Evaluation protocol</title>
      <p>3.1</p>
      <sec id="sec-3-1">
        <title>Data collections</title>
        <p>
          The data for this second edition of QAST is derived from five different resources, covering
spontaneous speech, semi-spontaneous speech and prepared speech: The first two are the same as were
used in QAST 2007 [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
        <p>
          • The CHIL corpus1 (as used for QAST 2007): The corpus contains about 25 hours of
speech, mostly spoken by non native speakers of English, with an estimated ASR Word
Error Rate (WER) of 20%.
• The AMI corpus2 (as used for QAST 2007): This corpus contains about 100 hours of
speech, with an ASR WER of about 38%.
• French broadcast news: The test portion of the ESTER corpus [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] contains 10 hours
of broadcast news in French, recorded from different sources (France Inter, Radio France
International, Radio Classique, France Culture, Radio Television du Maroc). There are 3
different automatic speech recognition outputs with different error rates (WER = 11.0%,
23.9% and 35.4%). The manual transcriptions were produced by ELDA.
• Spanish parliament: The TC-STAR05 EPPS Spanish corpus [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] is comprised of three
hours of recordings from the European Parliament in Spanish. The data was used to evaluate
recognition systems developed in the TC-STAR project. There are 3 different automatic
speech recognition outputs with different word error rates (11.5%, 12.7% and 13.7%). The
manual transcriptions were done by ELDA.
• English parliament: The TC-STAR05 EPPS English corpus [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] contains 3 hours of
recordings from the European Parliament in English. The data was used to evaluated speech
recognizers in the TC-STAR project. There are 3 different automatic speech recognition
outputs with different word error rates (10.6%, 14% and 24.1%) . The manual transcriptions
were done by ELDA.
        </p>
        <p>
          The spoken data cover a broader range of types, both in terms of content and in speaking style.
The Broadcast News and European Parliament date are less spontaneous than the lecture and
meeting speech as they are typically prepared in advance and are closer in structure to written
texts. While meetings and lectures are representative of spontaneous speech, Broadcast News and
European Parliament sessions are usually referred to as prepared speech. Although they typically
have few interruptions and turn-taking problems when compared to meeting data, many of the
characteristics of spoken language are still present (hesitations, breath noises, speech errors, false
starts, mispronunciations and corrections). One of the reasons for including the additional types
of data was to be closer to the textual data used to assess written QA, and to benefit from the
availability of multiple speech recognizers that have been developed for these languages and tasks
in the context of European or national projects [
          <xref ref-type="bibr" rid="ref1 ref2 ref4">2, 1, 4</xref>
          ].
3.1.1
        </p>
        <p>Questions and answer types
For each of the five scenarios, two sets of questions have been provided to the participants, the
first for development purposes and the second for the evaluation.
– Lectures: 10 seminars and 50 questions.
– Meetings: 50 meetings and 50 questions.
– French broadcast news: 6 shows and 50 questions.
– English EPPS: 2 sessions and 50 questions.</p>
        <p>– Spanish EPPS: 2 sessions and 50 questions.
1http://chil.server.de
2http://www.amiproject.org
– Lectures: 15 seminars and 100 questions.
– Meetings: 120 meetings and 100 questions.
– French broadcast news: 12 shows and 100 questions.
– English EPPS: 4 sessions and 100 questions.</p>
        <p>– Spanish EPPS: 4 sessions and 100 questions.</p>
        <p>Two types of questions were considered this year: factual questions and definitional ones. For each
corpus (CHIL, AMI, ESTER, EPPS EN, EPPS ES) roughly 70% of the questions are factual, 20%
are definitional, and 10% are NIL (i.e., questions having no answer in the document collection).
The question sets are formatted as plain text files, with one question per line (see the QAST 2008
Guidelines3). The factual questions similar to those used in the 2007 evaluation. The expected
answer to these questions is a Named Entity (person, location, organization, language, system,
method, measure, time, color, shape and material). The definition questions are questions such
as What is the Vlaams Blok? and the answer can be anything. In this example, the answer would
be a criminal organization. The definition questions are subdivided into the following types:
• Person: question about someone</p>
        <p>Q: Who is George Bush?</p>
        <p>R: The President of the United States of America.
• Organisation: question about an organisation</p>
        <p>Q: What is Cortes?</p>
        <p>R: Parliament of Spain.
• Object: question about any kind of objects</p>
        <p>Q: What is F-15?</p>
        <p>R: combat aircraft.
• Other: questions about technology, natural phenomena, etc.</p>
        <p>Q: What is the name of the system created by AT&amp;T?</p>
        <p>R: The How can I help you system.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Human judgment</title>
        <p>As in QAST 2007, the answer files submitted by participants have been manually judged by native
speaking assessors, who considered the correctness and exactness of the returned answers. They
also checked that the document labeled with the returned docid supports the given answer. One
assessor evaluated the results, and another assessor manually checked each judgment of the first
one. Any doubts about an answer was solved through various discussions. The assessors used
the QASTLE4 evaluation tool developed in Perl (at ELDA) to evaluate the responses. A simple
window-based interface permits easy, simultaneous access to the question, the answer and the
document associated with the answer.</p>
        <p>For T1b, T2b, T3b, T4b and T5b (QA on automatic transcripts) the manual transcriptions were
aligned to the automatic ASR outputs to find associate times with the answers in the automatic
3http://www.lsi.upc.edu/˜qast: News
4http://www.elda.org/qastle/
transcripts. The alignments between the automatic and the manual transcription were done using
time information. Unfortunately, for some documents time information were not available and
only word alignments were used.</p>
        <p>
          After each judgment the submission files were modified, adding a new element in the first column:
the answer’s evaluation (or judgment). The four possible judgments (also used at TREC[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ])
correspond to a number ranging between 0 and 3:
• 0 correct: the answer-string consists of the relevant information (exact answer), and the
answer is supported by the returned document.
• 1 incorrect: the answer-string does not contain a correct answer.
• 2 inexact: the answer-string contains a correct answer and the docid supports it, but the
string has bits of the answer missing or contains additional texts (longer than it should be).
• 3 unsupported: the answer-string contains a correct answer, but is not supported by the
docid.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Measures</title>
        <p>The two following metrics (also used in CLEF) were used in the QAST evaluation:
1. Mean Reciprocal Rank (MRR): This measures how well the right answer is ranked in the
list of 5 possible answers..
2. Accuracy: The fraction of correct answers ranked in the first position in the list of 5 possible
answers.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Submitted runs</title>
      <p>A total of five groups from four different countries submitted results for one or more of the proposed
QAST 2008 tasks. Due to various reasons (technical, financial, etc.), three other groups registered
but were not be able to submit any results.</p>
      <p>The five participating groups were:
• CUT, Chemnitz University of Technology, Germany;
• INAOE, Instituto Nacional de Astrofica, Optica y Electrica, Mexico;
• LIMSI, Laboratoire d’Informatique et de M´ecanique des Sciences de l’Ing´enieur, France;
• UA, Universidad de Alicante, Spain;
• UPC, Universitat Polit`ecnica de Catalunya, Spain.</p>
      <p>All groups participated to task T4 (English EPPS). Only LIMSI participated to task T3 (French
broadcast news). Table 1 shows the number of submitted runs per participant and task. Each
participant could submit up to 32 submissions (2 runs per task and transcription). The number
of submissions ranged from 2 to 20. The characteristics of the systems used in the submissions
are summarized in Table 2. A total of 49 submissions were evaluated with the distribution across
tasks shown in the bottom row of Table 2.</p>
      <p>The results for the ten QAST 2008 tasks are presented in Tables 3 to 12, according to factual
questions, definitional questions, and all questions.</p>
      <p>For manual transcriptions, the accuracy ranges from 45% (LIMSI1 on task T3a) down to 7%
(UPC1 on task T5a). For automatic transcriptions, the accuracy goes from 41% (LIMSI1 on task
T3b and ASR a) to 2% (UPC1 on task T5b and ASR c). Generally speaking, a loss in accuracy
is observed when dealing with automatic transcriptions. Comparing the best accuracy results on
manual transcription and automatic transcriptions, the loss of accuracy goes from 15% for task
T2 to 4% for tasks T3 and T4 tasks. This difference is larger for tasks where the ASR word error
rate is higher.
cut1
cut2
limsi1
upc1</p>
      <p>Acc</p>
      <p>Acc
0.16
0.20
0.45
0.38
16.0
17.0
41.0
34.0</p>
      <p>All
MRR
cut1
cut2
inaoe1
limsi1
ua1
upc1
12
12
41
44
32
38
0.21
0.22
0.38
0.42
0.27
0.37
21.0
21.0
33.0
33.0
20.0
34.0</p>
      <p>Another observation concerns the loss of accuracy when dealing with different word error rates.
Generally speaking higher WER results in lower accuracy (e.g. from 30% for T4b A to 20% for
T4b B). Strangely enough this is not completely true for the T5b task where results for ASR C
(13.7% WER) are 4% higher than for ASR B (12.7% WER). The WER being rather close, it is
probable that ASR C errors had a smaller impact on the named entities present in the questions.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In this paper, the QAST 2008 evaluation has been described. Five groups participated in this track
with a total of 49 submitted runs, across ten tasks that included dealing with different types of
speech (spontaneous or prepared), different languages (English, Spanish and French) and different
word error rates for automatic transcriptions (from 10.5% to 35.4%). For the tasks where the word
error rate was low enough (around 10%) the loss in accuracy compared to manual transcriptions
was under 5%, suggesting that QA in such documents is potentially feasible. However, even
where ASR performance is reasonably good, there remain outstanding challenges in dealing with
spoken language and the earlier mentioned differences from written language. The results from
the QAST evaluation indicate that if a QA system which performs well on manual transcriptions
it also performs reasonably well on high quality automatic transcriptions. The performance on
spoken language have not yet reached the level of those in the main QA track.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work has been jointly funded by the Spanish Ministry of Science (TEXTMESS project) and
OSEO under the Quaero program.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Galliano</surname>
          </string-name>
          , E. Geoffrois, G. Gravier,
          <string-name>
            <given-names>J.F.</given-names>
            <surname>Bonastre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mostefa</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Choukri</surname>
          </string-name>
          .
          <article-title>Corpus description of the ESTER Evaluation Campaign for the Rich Transcription of French Broadcast News</article-title>
          .
          <source>In Proceedings of LREC'06</source>
          ,
          <string-name>
            <surname>Genoa</surname>
          </string-name>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G.</given-names>
            <surname>Gravier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.F.</given-names>
            <surname>Bonastre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Galliano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Geoffrois</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>McTait</surname>
          </string-name>
          , , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Choukri</surname>
          </string-name>
          .
          <article-title>The ESTER evaluation campaign of Rich Transcription of French Broadcast News</article-title>
          .
          <source>In Proceedings of LREC'04</source>
          , Lisbon,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.C.</given-names>
            <surname>Gey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlgren</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          <article-title>M;agnini</article-title>
          ,
          <string-name>
            <given-names>D.W.</given-names>
            <surname>Oard</surname>
          </string-name>
          , M. de Rijke, and M. Stempfhuber, editors.
          <source>Evaluation of Multilingual and Multi-modal Information Retrieval</source>
          . Springer-Verlag.,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>TC-Star</surname>
          </string-name>
          . http://www.tc-star.org, 2004-
          <fpage>2008</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Turmo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.R.</given-names>
            <surname>Comas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ayache</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mostefa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosset</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Lamel</surname>
          </string-name>
          .
          <article-title>Overview of qast 2007</article-title>
          . In C. Peters,
          <string-name>
            <given-names>V.</given-names>
            <surname>Jijkoun</surname>
          </string-name>
          , Th. Mandl, H. Mu¨ller,
          <string-name>
            <given-names>D.W.</given-names>
            <surname>Oard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Peas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Petras</surname>
          </string-name>
          , and D. Santos, editors,
          <source>8th workshop of the Cross Language Evaluation Forum (CLEF</source>
          <year>2007</year>
          ).
          <source>Revised Selected Papers. LNCS</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.M.</given-names>
            <surname>Voorhees</surname>
          </string-name>
          and L.L. Buckland, editors.
          <source>The Fifteenth Text Retrieval Conference Proceedings (TREC</source>
          <year>2006</year>
          ),
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>