<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cross Lingual Question Answering using CINDI_QA for QA@CLEF 2007</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Chedid Haddad, Bipin C. Desai Department of Computer Science and Software Engineering Concordia University 1455 De Maisonneuve Blvd. W. Montreal</institution>
          ,
          <addr-line>Quebec H3G 1M8</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article presents the first participation of the CINDI group in the Multiple Language Question Answering Cross Language Evaluation Forum (QA@CLEF). We participated in a track using French as the source language and English as the target language. CINDI_QA first uses an online translation tool to convert the French input question into an English sentence. Second, a Natural Language Parser extracts keywords such as verbs, nouns, adjectives and capitalized entites from the query. Third, synonyms of those</p>
      </abstract>
      <kwd-group>
        <kwd>Question answering</kwd>
        <kwd>Questions beyond factoids</kwd>
        <kwd>Bilingual</kwd>
        <kwd>French</kwd>
        <kwd>English</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>System Overview</title>
      <p>CINDI_QA is made up of one central unit called the Processor, a couple of peripheral components, four integrated
tools and a template module. These are illustrated in the figure below and elaborated on in the next section.</p>
      <p>The system’s “brain” is located in the CINDI_QA Processor where all the logical inferences are made. Since
CINDI_QA communicates with several tools, the information it retrieves from those modules - which is
meaningless on its own - is analyzed in the Processor in a structured way that helps build the eventual answer.
This process is done in a pre-defined order, with the Processor getting data from one module, sorting it out then
using it to probe the next module.</p>
      <p>One of the peripheral components is the PHP Script that acts as an interface between the Online Translator and
the Processor. Its purpose is to send the French question to the Online Translator and bring its English equivalent
back. The other peripheral unit is the Graphical User Interface (GUI) which is the façade of our system for the
user; it will prompt for a French question and return the English answer.</p>
      <p>In a typical scenario, the question inputted in French by the user is translated to English then parsed to extract
the keywords. Afterward, the synonyms of the keywords are obtained to form an internal query sent to the Search
engine that already has the CLEF data indexed. The answer candidates are localized at which point they are
matched against a pre-existing set of templates that enables the selection of the best answer. That answer, which is
in English, is sent back to the user.</p>
      <p>In order to improve performance, CINDI_QA allows user interaction to direct its flow of operations. Actually,
CINDI_QA can be run in two modes. In the automatic mode, it acts as a complete black box by only taking a
question and returning an answer. In the feedback mode, it has the possibility of disambiguating the query by
prompting the user after each stage. For example, since certain synonyms of a word often don’t suit a particular
context, the user can choose to eliminate them and only retain relevant ones.</p>
      <p>The process flow of CINDI_QA in the feedback mode is shown in the following diagram.</p>
    </sec>
    <sec id="sec-3">
      <title>Tools Integration</title>
      <sec id="sec-3-1">
        <title>3.1 Online Translation</title>
        <p>Since we are working in a bilingual environment, the system is queried in a language different from the data
collection it is using as reference. A translation tool is needed for this reason. After researching the available tools,
we noticed that the Google [2], Babel Fish [3] and Systran [4] translators appear to be powered by the same engine
because they offer the same result when asked to translate from French to English.</p>
        <p>We chose to use Google Translate in our system due to its better interface and speed of processing. A PHP
script is responsible of delivering the French question typed in by the user to the Google Translate webpage and
bring back the translated English equivalent to the CINDI_QA Processor.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Natural Language Parsing</title>
        <p>We need a way to understand the question asked by the user in order to single out the keywords. This was
achieved thanks to the Link Grammar Parser [5], a syntactic parser of English based on link grammar, an original
theory of English syntax. Given a sentence, the Link Parser assigns to it a syntactic structure which consists of a
set of labeled links connecting pairs of words. Each structure is called a linkage and several linkages are generated
for the same input. This tool is written in generic C code, but since CINDI_QA has been programmed in Java, we
used the Java Native Code Link Grammar Interface [6].</p>
        <p>The Link Parser is plugged into our system to generate linkages of the English question. Using one linkage, we
are able to determine which words form nouns, verbs, adjectives and capitalized entities. If those keywords appear
wrong or incomplete, we go on to the next linkage. Remember that the user has the option of choosing the most
appropriate linkage.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3 Lexical Reference</title>
        <p>To increase the chances of finding candidate answers among the data collection, we include synonyms of the
keywords in addition to the keywords themselves. WordNet [7] is a lexical database for the English language
developed at Princeton University and has been used in other CINDI related projects [10] so its selection was
pretty obvious.</p>
        <p>WordNet was used in concordance with Lucene. Lucene enables us to create an index composed strictly of
synonyms defined by WordNet that can be queried like a regular index so we can actually get a list of synonyms
of a specific word. This strategy is highlighted in pages 292-296 of Lucene in Action [8].</p>
        <p>After defining the keywords using the Link Parser, we queried the WordNet index to obtain the synonyms of
each keyword, except for capitalized entities. Since some of those synonyms are irrelevant or out of context, the
user has the choice to discard them and select only the appropriate ones.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4 Document Indexing and Searching</title>
        <p>
          We needed a tool that could not only index all the documents that make up QA@CLEF’s data collection but also
search the created index with some level of intelligence such as ranking results and highlighting query terms. A
perfect match for this requirement is the Apache Software Foundation’s Lucene [
          <xref ref-type="bibr" rid="ref2">9</xref>
          ], a high-performance,
fullfeatured text search engine library written entirely in Java.
        </p>
        <p>CINDI_QA makes extensive use of Lucene. As mentioned before, it is used in concordance with WordNet to
get synonyms of keywords. Lucene also creates the CLEF index and ranks the results found. The following
features are of great importance to our system.</p>
        <sec id="sec-3-4-1">
          <title>3.4.1 Lucene Query Building using Proximity Search</title>
          <p>Once we have identified the keywords and their synonyms, the building of the query takes place. The query is
constructed by putting together each keyword or its synonym with the other keywords or their synonyms. The
crucial point here is to add the proximity search flag to the built query so that Lucene will not look for a sentence
that has our keywords adjacent to each other, but rather one where the keywords are close to each other but spread
out in a paragraph. This is done by adding the tilde character ‘~’ and a number to the end of the query.</p>
        </sec>
        <sec id="sec-3-4-2">
          <title>3.4.2 Highlighting query terms in Answer Candidates</title>
          <p>Once the Lucene query is built, it is searched against the index of the document collection. Lucene then returns a
list of filenames ranked according to the frequency of occurrence of the words in the query. At this point, we take
advantage of the Lucene Highlighter, a wonderful tool that actually displays snippets of text with the query terms
highlighted. This allows us not only to know which document has the answer, but also to obtain a sentence in that
document that displays the actual answer. This mechanism is mentioned on page 300 of Lucene in Action [8].
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Template Matching</title>
      <p>CINDI_QA’s template module comes into play in the final stages, between the identification of the answer
candidates and the return of an answer to the user. Indeed, after the Lucene component’s job is done, we are left
with a set of sentences one of which holds the actual answer. To determine which one to choose, we match them
against our set of pre-defined templates. We chose to use templates because a previous project [10] done by a
member of our group was pretty successful at parsing English questions using mainly templates.</p>
      <p>
        According to the CLEF Working Notes [
        <xref ref-type="bibr" rid="ref3">11</xref>
        ], there are three different types of questions: Factoid, Definition
and List. Factoid and Definition kinds are further broken down into sub-categories: Person, Time, Location,
Organization, Measure, Count, Object and Other for the former and Person, Organization, Object and Other for
the latter. Our templates take advantage of this approach to differentiate among the questions that are inputted in
our system, and also of the keywords we identified in the Link Parser module. The capitalized entity, made up of a
word whose first letter is uppercase, identifies proper nouns. It is an important keyword because of its high
frequency of occurrence; this is due to the fact that most questions at QA@CLEF ask about famous people,
locations or events.
      </p>
      <sec id="sec-4-1">
        <title>4.1 The Person Definition Template: Who is X?</title>
        <p>This is without a doubt the easiest template to consider. To discover that a question is of that type, it must start
with the word “who”, have the verb “is” or its derivative “was” and must be followed by a capitalized entity X. X
could be one word or a composition of words; as long as adjacent words are capitalized, CINDI_QA will define
them as belonging to one and the same capitalized entity. The following illustrates the query given to CINDI_QA,
the candidate returned by the Highlighter of the Lucene Module and the answer returned after template matching.</p>
        <p>Example: - query: Who is Robert Altman?
- candidate: Robert Bernard Altman (February 20, 1925 – November 20, 2006) was an American
film director known for making films that are highly naturalistic, but with a stylized
perspective.</p>
        <p>- answer: American film director.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2 The Count Factoid Template: How many N V?</title>
        <p>This template is selected if the question starts with the words “how many” and has at least one verb V and one
noun N, capitalized entities being optional. In the following example, had the expression “Solar System” not been
capitalized, the Link Parser module would have identified “solar” as an adjective and “system” as a noun, yet the
template would still be selected because “planets” is a noun.</p>
        <p>Example: - query: How many planets does the Solar System have?
- candidate: The Solar System or solar system consists of the Sun and the other celestial objects
gravitationally bound to it: the eight planets, their 165 known moons, three currently
identified dwarf planets (Ceres, Eris, and Pluto) and their four known moons, and
billions of small bodies.</p>
        <p>- answer: Eight.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3 The Time Factoid Template: What year V1 X V2? / When V1 X V2?</title>
        <p>This template is selected if the following occurs: the question starts with the words “what year” or “when”, then
has two verbs, V1 and V2 as well as one capitalized entity X. The Link Parser module will identify two separate
verbs even if they belong to the participate form of the same verb, i.e. it will flag “was” as V1 and “murdered” as
V2, even though technically they both define the verb “murder”. In the following example, notice how the word
“assassinate”, synonym of the inputted verb “murder”, is used to return the correct answer.
Example: - query: What year was Martin Luther King murdered?
- candidate: On April 4, 1968, King was assassinated in Memphis, Tennessee.</p>
        <p>- answer: 1968.</p>
        <p>In the feedback mode, the user will be prompted one last time to approve the template chosen by CINDI_QA
before being shown an answer. In the black-box mode, the system chooses the least costly linkage from the Link
Parser, retains only the first two synonyms of a keyword, uses the best ranked document from Lucene with the
most highlighted sentence and constructs the answer based on the first template matched.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>The CINDI group participated in the FR to EN track of the QA@CLEF edition of 2007. We produced two runs
that were sent: cind071fren and cind072fren. The two runs ended up with the same overall accuracy value of 13%.
The following figure details the different assessments both runs obtained.</p>
      <p>cind071fren
cind072fren</p>
      <sec id="sec-5-1">
        <title>Right</title>
        <p>26
26</p>
      </sec>
      <sec id="sec-5-2">
        <title>Wrong</title>
      </sec>
      <sec id="sec-5-3">
        <title>Inexact</title>
      </sec>
      <sec id="sec-5-4">
        <title>Unsupported</title>
      </sec>
      <sec id="sec-5-5">
        <title>Unassessed Accuracy</title>
        <p>171
170
1
2
2
2
0
0
13%
13%</p>
        <p>In addition to the usual news collections, articles from Wikipedia were also considered as answer source for the
first time this year. But we became aware of that very late in our answering process and were only able to use a
small part of the corpora available. This is the main reason for our system’s low performance.</p>
        <p>Since our two runs are equivalent, the following figure shows the accuracy by question type only of
cind072fren. CINDI_QA’s score on definition questions is much higher than any other type of question. This
result is attributed to the robustness of template 4.1.</p>
      </sec>
      <sec id="sec-5-6">
        <title>Factoids</title>
      </sec>
      <sec id="sec-5-7">
        <title>Lists</title>
      </sec>
      <sec id="sec-5-8">
        <title>Definitions</title>
      </sec>
      <sec id="sec-5-9">
        <title>Total</title>
      </sec>
      <sec id="sec-5-10">
        <title>Right</title>
      </sec>
      <sec id="sec-5-11">
        <title>Wrong</title>
      </sec>
      <sec id="sec-5-12">
        <title>Unsupported</title>
      </sec>
      <sec id="sec-5-13">
        <title>Inexact Accuracy</title>
        <p>161
18
140
2
1
9
1
8
0
0
30
7
22
0
1
11.18%
11.11%
23.33%</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Works</title>
      <p>Although the approach we used to tackle bilingual question-answering at CLEF looks good in theory, the results
obtained are a bit disappointing, as is shown by our 13% success rate. Because this is our first participation in
QA@CLEF and the CINDI group only invested 1 man/year on the project, we have high hope for the future,
especially since we missed a large part of the source corpora.</p>
      <p>We learned that because CINDI_QA relies on so many external tools, it is only as strong as its weakest link.
For instance, if from the start, the translation of the question isn’t a successful one, there is nothing the system can
do after that stage to come up with a correct answer.</p>
      <p>Future works include the addition of new templates to handle the multitude of question sub-categories as well
as a mechanism to identify questions whose answer isn’t located in the CLEF data collection and return NIL for
those questions. We hope we can have these enhancements ready for the next edition of QA@CLEF.</p>
      <p>Babel Fish Translation, http://babelfish.altavista.com/</p>
      <p>WordNet, a lexical database for the English language, http://wordnet.princeton.edu/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] The CINDI System</article-title>
          , http://cindi.encs.concordia.ca/about_cindi.html Google Translate, http://translate.google.com/translate_t [4]
          <string-name>
            <given-names>Systran</given-names>
            <surname>Box</surname>
          </string-name>
          , http://www.systransoft.com/ [5]
          <string-name>
            <given-names>The</given-names>
            <surname>Link Grammar Parser</surname>
          </string-name>
          , http://www.link.cs.cmu.edu/link [6]
          <string-name>
            <given-names>Java</given-names>
            <surname>Native Code Link Grammar Interface</surname>
          </string-name>
          , http://chrisjordan.ca/projects
          <string-name>
            <given-names>O.</given-names>
            <surname>Gospodnetic</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Hatcher</surname>
          </string-name>
          , Lucene in Action, Manning,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Lucene</surname>
            , http://lucene.apache.org/ [10]
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Stratica</surname>
            , NLPQC:
            <given-names>A Natural</given-names>
          </string-name>
          <string-name>
            <surname>Language Processor for Querying</surname>
            <given-names>CINDI</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Master</surname>
            <given-names>Thesis</given-names>
          </string-name>
          , Concordia University,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B.</given-names>
            <surname>Magnini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Giampiccolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Former</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ayache</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Jijkoun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Osenova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Peñas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rocha</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Sacaleanu</surname>
            and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Sutcliffe</surname>
          </string-name>
          ,
          <article-title>Overview of the CLEF 2006 Multilingual Question Answering Track</article-title>
          ,
          <source>in: Cross Language Evaluation Forum: Working Notes for the CLEF 2006 Workshop (CLEF</source>
          <year>2006</year>
          ), Alicante, Spain,
          <fpage>20</fpage>
          -
          <issue>22</issue>
          <year>September 2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>