<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cross-Language French-English Question Answering using the DLT System at CLEF 2004</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Question Type</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>353 61 202734 Fax</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Documents and Linguistic Technology Group Department of Computer Science and Information Systems University of Limerick Limerick</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article outlines the participation of the Documents and Linguistic Technology (DLT) Group in the Cross Language French-English Question Answering Task of the Cross Language Evaluation Forum (CLEF). Following our experiences last year (Sutcliffe, Gabbay and O'Gorman, 2003), our aim was to improve the system particularly in the early stages of processing, and to make further refinements to other components.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>what_capital</p>
      <p>company
what_country
mountain</p>
      <p>where
how_did_die
who
when
unknown</p>
      <p>Example Question
130 Quelle est la capitale du Vénézuela?
149 Qui fabrique Invirase?
37 Dans quel pays européen est située la
ville de Galway?
162 Quelle est la plus haute montagne du
monde?
166 Où se trouve Halifax?
47 Comment est mort River Phoenix?
40 Qui a réalisé "Braveheart"?
30 Quand est-ce que le prince Charles et</p>
      <sec id="sec-1-1">
        <title>Diana se sont mariés? 36 Citez une unité de radioactivité</title>
        <sec id="sec-1-1-1">
          <title>Translation</title>
        </sec>
      </sec>
      <sec id="sec-1-2">
        <title>What is the capital of Venezuela?</title>
      </sec>
      <sec id="sec-1-3">
        <title>Who manufactures Invirase?</title>
      </sec>
      <sec id="sec-1-4">
        <title>In which European country is the town of</title>
      </sec>
      <sec id="sec-1-5">
        <title>Galway located?</title>
      </sec>
      <sec id="sec-1-6">
        <title>What is the highest mountain of the world?</title>
      </sec>
      <sec id="sec-1-7">
        <title>Where is Halifax?</title>
      </sec>
      <sec id="sec-1-8">
        <title>How did River Phoenix die?</title>
      </sec>
      <sec id="sec-1-9">
        <title>Who directed “Braveheart”?</title>
      </sec>
      <sec id="sec-1-10">
        <title>When did prince Charles and Diana get married?</title>
      </sec>
      <sec id="sec-1-11">
        <title>Name a unit of radioactivity</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Architecture of the CLEF 2004 DLT System</title>
      <sec id="sec-2-1">
        <title>2.1 Outline</title>
        <p>The basic architecture of our system is standard in nature and comprises query type identification, query analysis
and translation, retrieval query formulation, document retrieval, text file parsing, named entity recognition and
answer entity selection.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Query Type Identification</title>
        <p>As last year, simple keyword combinations and patterns were used to classify the query. This was accomplished
by using the CLEF 03 queries and translated TREC queries (RALI, 2004) as a model.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3 Query Analysis and Translation</title>
        <p>This stage differed greatly from last year. We started off by tagging the query for part-of-speech using XeLDA
(2004). We then carried out shallow parsing looking for various types of phrase. Each phrase was then translated
using three different methods. Two translation engines and one dictionary were used. The engines were Reverso
(2004) and WorldLingo (2004) which were chosen because we had found them to give the best overall
performance in various experiments. The dictionary used was the Grand Dictionnaire Terminologique (GDT,
2004) which is a very comprehensive terminological database for Canadian French with detailed data for a large
number of different domains. The three candidate translations were then combined – if a GDT translation was
found then the Reverso and WorldLingo translations were ignored. The reason for this is that if a phrase is in
GDT the translation for it is nearly always correct. It is an excellent resource. For example 'equipe de football'
becomes 'football team' and not 'team of football', 'salle d'opera' (Canadian dialect for 'opera') becomes 'opera
house' not 'room of opera' and so on. In the case where words or phrases are not in GDT, then the Reverso and</p>
        <sec id="sec-2-3-1">
          <title>WorldLingo translations were simply combined.</title>
          <p>The types of phrase recognised were determined after a study of the constructions used in French queries together
with their English counterparts. The aim was to group words together into sufficiently large sequences to be
independently meaningful but to avoid the problems of structural translation, split particles etc which tend to
occur in the syntax of a question, and which the engines tend to analyse incorrectly.</p>
          <p>The structures used were number, quote, cap_nou_prep_det_seq, all_cap_wd, cap_adj_cap_nou,
cap_adj_low_nou, cap_nou_cap_adj, cap_nou_low_adj, low_nou_low_adj, low_nou_prep_low_nou,
low_adj_low_nou, nou_seq and wd. These were based on our observations that (1) Proper names usually only
start with a capital letter with subsequent words uncapitalised, unlike English; (2) Adjective-Noun combinations
either capitalised or not can have the status of compounds in French and hence need special treatment; (3) Certain
noun-preposition-noun phrases are also of significance.</p>
          <p>As part of the translation and analysis process, weights were assigned to each phrase in an attempt to establish
which parts were more important in the event of query simplification being necessary.</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>2.4 Retrieval Query Formulation</title>
        <p>The starting point for this stage was a set of possible translations for each of the phrases recognised above. For
each phrase, a boolean query was created comprising the various alternatives as disjunctions. In addition,
alternation was added at this stage to take account of morphological inflections (e.g., 'go'&lt;-&gt;'went',
'company'&lt;&gt;'companies' etc) and European English vs. American English spelling ('neighbour'&lt;-&gt;'neighbor',
'labelled'&lt;&gt;'labeled' etc). The reason for this last step was the addition for this year of the Glasgow Herald collection to the
existing LA Times. The list of the above components was then ordered by the weight assigned during the
previous stage and the ordered components were then connected with AND operators to make the complete
boolean query.</p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5 Document Retrieval</title>
        <p>During document retrieval, the boolean query was submitted to the DTSearch search engine (DTSearch, 2000)
which had previously been indexed on the LA Times and Glasgow Herald collections, with each sentence in the
collection being considered as a separate document for indexing purposes. This followed our observation that in
most cases the search keywords and the correct answer appear in the same sentence.</p>
        <p>In the event that no documents were found, the conjunction in the query (corresponding to one phrase recognised
in the query) with the lowest weight was eliminated and the search was repeated. Some attempts were made this
year to avoid the situation in which the query is inadvertently simplified to something insufficiently selective and
highly frequent in the corpus (e.g. United States).</p>
      </sec>
      <sec id="sec-2-6">
        <title>2.6 Text File Parsing</title>
        <p>This stage is straightforward and simply involves retrieving the matching 'documents' (i.e. sentences) from the
corpus and extracting the text from the markup.</p>
      </sec>
      <sec id="sec-2-7">
        <title>2.7 Named Entity Recognition</title>
        <p>Named Entity recognition was carried out in the standard way using a mixture of grammars and lists. The number
of types was increased to 75 by studying previous CLEF and TREC question sets and these were incorporated
into the query categoriser also.</p>
      </sec>
      <sec id="sec-2-8">
        <title>2.8 Answer Entity Selection</title>
        <p>We used the highest-scoring method of answer selection. In this, the named-entity instance is selected which
occurs in the vicinity of the maximum number of keywords taken from the translated query, across all document
passages. We also experimented with Google re-ordering using a Magnini-type method (Magnini, Negri, Prevete
and Tanev, 2002).
animal
colour
company
def_org
def_person
distance
how_did_die
how_many3
how_old
name_part
nationality
pol_party
population
team
what_capital
what_country
what_mountain
what_river
when
when_wk_day
when_year
where
who
unknown
Totals</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3.2 Results</title>
    </sec>
    <sec id="sec-4">
      <title>3. Runs and Results</title>
      <sec id="sec-4-1">
        <title>3.1 Two Experiments</title>
        <p>Query Type</p>
        <sec id="sec-4-1-1">
          <title>We submitted two runs which differed slightly in their term translation strategy.</title>
          <p>
            Results are summarised by query type in Table 2. Concerning query classification it shows for each query type
the number of queries assigned to that type which were correctly categorised along with the number incorrectly
categorised. The overall rate of categorisation success was 85% which is identical to the one achieved in TREC
            <xref ref-type="bibr" rid="ref2">(Sutcliffe, Gabbay, Mulcahy and White, 2004)</xref>
            . The number of queries classified as unknown was 70.
The performance of question answering in Run 1 can be summarised as follows. Out of the 170 queries classified
correctly, 35 were answered correctly. Out of the remaining 30 queries classified incorrectly a further three were
answered correctly. Overall performance was thus 38 / 200 i.e. 19%. Results for Run 2 were as follows. 28 of the
170 queries were answered correctly along with two of the 30 queries giving a total of 30 / 200 i.e. 15%. In both
runs 63 questions were answered NIL.
          </p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3.3 Platform</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusions</title>
      <p>We used a Dell PC running Windows 2000 and having 256 Mb RAM. The whole system was ported this year to</p>
      <sec id="sec-5-1">
        <title>SICStus Prolog 3.11.1 (SICStus, 2004) which is much faster than Quintus.</title>
        <p>The overall performance this year was 19% compared to the 11.5% we achieved last year. We can attribute this
improvement mostly to a superior translation strategy although much further work on this is required.
If we exclude query types which were represented by just one example, the best performance (100%) was on
how_did_die queries. However, there were only three queries of this type. On the more common types the best
performance was achieved when answering ‘when’ and ‘where’ queries (47% and 41%, respectively). The
performance on the relatively common types how_many and when_year was poor (9%, 10%, respectively). A
more sophisticated answer selection strategy may improve performance on these. It is hard to assess performance
on many of the other query types due to the small number of each.</p>
        <p>Query categorisation for this year stood at 85% compared to 79.5% last year. The reduction in the number of
(correctly classified) unknown queries from 58 last year to 47 in the current run may reflect an improvement in
the coverage of categorisation. However, among the unknown queries several categories emerged which should
be added in future systems. These include queries which are similar in nature to list questions but ask for a single
hyponym (e.g., Query 17 ‘name a cetacean’), queries about materials (e.g. Query 70 ‘What are fibre-optic cables
made of?’), queries of the type ‘What does company X sell/produce?’ and queries about diseases (treatment, way
of transmission), newspaper names, wars, and musical bands. This year the test set included ‘how’ and ‘why’
queries (e.g., Query 8 ‘How is the pope?’, Query 87 ‘Tell me a reason for teenage suicide’, Query 107 ‘How does
acupuncture work?’) which our system cannot answer
A significant proportion of the queries are likely to remain unknown in increasingly difficult evaluations. The
poor performance on unknown queries highlights the need to develop an alternative to our current strategy of
answering such queries (i.e., finding a sequence of capitalised words). Less than a quarter of the correctly
classified unknown queries this year can be answered by the current strategy.</p>
        <p>Simple heuristics may improve performance on existing categories. For example, year numbers could often be
eliminated from answers to how_many queries.</p>
        <p>Our boolean search query formulation strategy was a big improvement on last year but was not without its
problems. In particular the combination for each phrase of translation alternatives, inflection alternatives and
spelling alternatives could result on occasion in highly complex queries which were a problem for our relatively
lightly engineered search engine.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. References</title>
      <sec id="sec-6-1">
        <title>DTSearch (2000). www.dtsearch.com</title>
      </sec>
      <sec id="sec-6-2">
        <title>GDT (2004) http://w3.granddictionnaire.com/btml/fra/r_motclef/index1024_1.asp Magnini, B., Negri, M., Prevete, R., &amp; Tanev H. (2002). Is it the Right Answer? Exploiting Web Redundancy for</title>
        <p>Answer Validation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics,
July 6-12, 2002, Philadelphia, 425-432.</p>
      </sec>
      <sec id="sec-6-3">
        <title>RALI (2004) http://www-rali.iro.umontreal.ca/LUB/qabilingue.en.html</title>
      </sec>
      <sec id="sec-6-4">
        <title>Reverso (2004) http://grammaire.reverso.net/textonly/default.asp</title>
      </sec>
      <sec id="sec-6-5">
        <title>SICStus (2004) http://www.sics.se/isl/sicstuswww/site/index.html</title>
      </sec>
      <sec id="sec-6-6">
        <title>XeLDA (2004) http://www.temis-group.com/temis/XeLDA.htm</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Sutcliffe</surname>
            ,
            <given-names>R. F. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gabbay</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>O'Gorman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Cross-Language French-English Question Answering using the DLT System at CLEF 2003</article-title>
          .
          <article-title>Proceedings of the Cross Language Evaluation Forum</article-title>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2003</year>
          , ,
          <year>August</year>
          21-
          <issue>22</issue>
          ,
          <year>2003</year>
          , Trondheim, Norway,
          <fpage>373</fpage>
          -
          <lpage>378</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Sutcliffe</surname>
            ,
            <given-names>R. F. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gabbay</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mulcahy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>White</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Question Answering using the DLT system at TREC 2003</article-title>
          . In E. M. Voorhees and
          <string-name>
            <surname>L. P.</surname>
          </string-name>
          Buckland (Eds)
          <article-title>Proceedings of the Twelfth Text Retrieval Conference</article-title>
          ,
          <source>TREC 2003, November 18-21</source>
          ,
          <year>2003</year>
          , Gaithersburg, Maryland. NIST Special Publication 500-
          <fpage>255</fpage>
          . Gaithersburg, MD: Department of Commerce,
          <source>National Institute of Standards and Technology</source>
          ,
          <volume>686</volume>
          -
          <fpage>692</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>