<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Hybrid Question Answering System based on Information Retrieval and Answer Validation</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Partha Pakray</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The article presents the experiments carried out as part of the participation in the main task of QA4MRE@CLEF 2011. We have submitted total five unique runs in the main task: two runs from systems based on Answer Validation (AV) machine reading techniques, one run from systems based on Question Answering (QA) techniques while the last two runs are hybrid systems where the decision is taken based on the outputs from the AV and QA based systems. In the AV system, we first combine the question and each answer option to form the Hypothesis (H). Stop words are removed from each H and query words are identified to retrieve the most relevant sentences from the associated document using Lucene. Relevant sentences are retrieved from the associated document based on the TF-IDF of the matching query words along with n-gram overlap of the sentence with the H. Each retrieved sentence defines the Text T. Each T-H pair is assigned a ranking score in the AV system that works on textual entailment principle. The answer option for which the TH pair gets the maximum score is selected as the possible answer. The two unique runs differ in the way in which the relevant sentences are retrieved from the associated document. The second system is based on Question Answering (QA) technique. Each question along with each answer option generates the possible answer patterns. Each sentence in the associated document is assigned an inference score with respect to each answer pattern. The sentence that receives the highest inference score corresponding to the answer patterns is identified as the relevant sentence in the document and the corresponding answer option is selected as the answer to the given question.</p>
      </abstract>
      <kwd-group>
        <kwd>QA4MRE Data Sets</kwd>
        <kwd>Named Entity</kwd>
        <kwd>Textual Entailment</kwd>
        <kwd>Question Answering technique</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        After the success of ResPubliQA 2009 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and ResPubliQA 2010 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], the third
evaluation campaign on Question Answering system is Question Answering for
Machine Reading Evaluation (QA4MRE)1 at CLEF 2011.
      </p>
      <p>
        The main objective of QA4MRE [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is to develop a methodology for evaluating
Machine Reading systems through Question Answering and Reading Comprehension
Tests. Machine Reading task obtains an in-depth understanding of just one or a small
number of texts. The task focuses on the reading of single documents and
identification of the correct answer to a question from a set of possible answer
options. The identification of the correct answer requires various kinds of inference
and the consideration of previously acquired background knowledge. Ad-hoc
collections of background knowledge have been provided for each of the topics in all
the languages involved in the exercise so that all participating systems work on the
same background knowledge. Texts have been included from a diverse range of
sources, e.g. newspapers, newswire, web, blogs, Wikipedia entries.
      </p>
      <p>We have submitted total five unique runs in the main task: two runs from systems
based on Answer Validation (AV) techniques, another one run from systems based on
Question Answering (QA) techniques while the last two runs are hybrid system where
the decision is taken based on the outputs from the AV and QA based systems.</p>
      <p>
        Answer Validation Exercise (AVE) is a task introduced in the QA@CLEF
competition. AVE task is aimed at developing systems that decide whether the answer
of a Question Answering system is correct or not. There were three AVE
competitions AVE 2006 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], AVE 2007 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and AVE 2008 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. AVE systems receive
a set of triplets (Question, Answer and Supporting Text) and return a judgment of
“SELECTED”, “VALIDATED” or “REJECTED” for each triplet. We have
participated in the Paragraph Selection (PS) Task and Answer Selection (AS) Task [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
in QA@CLEF 2010 – ResPubliQA.
      </p>
      <p>Section 2 describes the corpus statistics. Section 3 describes the Answer Validation
based Machine Reading System Architecture. Section 4 details the Question
Answering based Machine Reading System Architecture. Section 5 details the Hybrid
Machine Reading System Architecture. The experiments carried out on test data sets
are discussed in Section 6 along with the results. The conclusions are drawn in
Section 7.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Corpus Statistics</title>
      <p>The data set is made up of a series of tests. Each test consists of one single document
(Test Document) with several questions and a set of choices per question. So, the task
is a Reading Comprehension test of the given document. Each participating system is
provided with:
- 6 Test Documents (2 documents for each of the three topics)
- 10 questions per document with 5 choices for each question.</p>
      <sec id="sec-2-1">
        <title>1 http://celct.fbk.eu/ResPubliQA/</title>
        <p>2 http://members.unine.ch/jacques.savoy/clef/</p>
        <p>Topics, documents and questions were made available in English, German, Italian,
Romanian, and Spanish. We worked only with English language data. The
Background Collections (one for each topic) are comparable (but not identical)
topicrelated collections created in all the different languages.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Answer Validation based Machine Reading System Architecture</title>
      <p>
        The architecture of the Answer Validation (AV) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] based machine reading system is
described in Figure 1. The various components of the AV system are Pattern
Generation Module, Hypothesis Generation Module, Document Parsing Module,
Question Type Analysis Module, Named Entity Recognition (NER) Module, Textual
Entailment (TE) Module, Chunk Boundary, Syntactic Similarity Module, Answer
Scoring Module and Answer Ranking Module. Each of these modules is now being
described in subsequent subsections.
      </p>
      <sec id="sec-3-1">
        <title>3.1 Pattern Generation Module</title>
        <p>At first we convert each question into an affirmative sentence that denotes the answer
pattern and place the &lt;/answer&gt; template in place of the appropriate answer. The
pattern generation module is rule based.</p>
        <p>For example, let us consider the question id 7 in doc id 2 of the QA4MRE train set,
Question: Where is the U.S. nuclear waste repository located?
The generated pattern is The U.S. nuclear waste repository is located &lt;/answer&gt;.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Hypothesis Generation Module</title>
        <p>After Pattern generation the &lt;/answer&gt; template is replaced by each answer option
string forming the generated Hypothesis. The generated hypothesis is termed as the
query. For example, for question id 7 (QA4MRE Train set), the following hypotheses
(or queries) are generated for each of the answer options:
H_1: The U.S. nuclear waste repository is located at Oklo.</p>
        <p>H_2: The U.S. nuclear waste repository is located in Morsleben.</p>
        <p>H_3: The U.S. nuclear waste repository is located in New Mexico.</p>
        <p>H_4: The U.S. nuclear waste repository is located in a suitable geological formation.
H_5: The U.S. nuclear waste repository is located in the U.S. State of Nevada.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3 Document Processing and Indexing</title>
        <p>The web documents are full of noises mixed with the original content. It is very
difficult to identify and separate the noises from the actual content. First of all the
documents had to be preprocessed. The document structure is checked and
reformatted according to the system requirements.</p>
        <p>The corpus is in XML format. All the XML test data has been parsed before
indexing using our XML Parser. The XML Parser extracts the sentences from the
document. After parsing the documents, they are indexed using Lucene, an open
source full text search tool.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4 Query Word Identification</title>
        <p>After indexing has been done, the queries have to be processed to retrieve relevant
sentences from the associated documents. Each answer pattern or query is processed
to identify the query words for submission to Lucene.</p>
        <p>Certain key characters in the query cause implicit query handling during searching.
For example, the dot character between two query words denotes AND of the two
query words. Such key characters are thus removed from the question before
submission to Lucene. For example, http://wt.jrc.it/ and doug@nutch.org are
rephrased as “http wt jrc it” and “doug nutch org” respectively.</p>
        <p>The Stop words (using the stop word list2) and question words (what, when,
where, which etc.) are removed from each question. The remaining words in the
question are identified as the query words. Query words may appear in inflected
forms in the question. For English, standard Porter Stemming algorithm3 has been
used to stem the query words.</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5 Sentence Retrieval</title>
        <p>After searching each query into the Lucene index, a set of sentences in ranked order
for each query is retrieved.</p>
        <p>First of all, all query words are fired with AND operator. If at least one sentence is
retrieved using the query with AND operator then the query is removed from the
query list and need not be searched again. The rest of the queries are fired again with
OR operator. OR searching retrieves at least one sentence for each query. Now, the
top ranked relevant ten sentences for each query are considered for further processing
In case of AND search only the top ranked sentence is considered. Sentence retrieval
is the most crucial part of this system. We take only the top ranked relevant sentences
assuming that these are the most relevant sentences in the associated document for the
question from which the query has been generated.</p>
        <p>Each retrieved sentence is considered as the Text (T) and is paired with each
generated hypothesis (H). Each T-H pair identified for each answer option
corresponding to a question is now assigned a score based on the NER module,
Textual Entailment module, Chunking module, Syntactic Similarity module and
Question Type module.</p>
      </sec>
      <sec id="sec-3-6">
        <title>3.6 NER Module</title>
        <p>
          It is based on the detection and matching of Named Entities (NEs) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] in the Retrieved
Sentence (T) - generated Hypothesis (H) pair. Once the NEs of the hypothesis and the
text have been detected, the next step is to determine the number of NEs in the
hypothesis that match in the corresponding retrieved sentence. The measure
NE_Match is defined as NE_Match = number of common NEs between T and
H/Number of NEs in Hypothesis.
        </p>
        <p>If the value of NE_Match is 1, i.e., 100% of the NEs in the hypothesis match in the
text, then the T-H pair is considered as an entailment. The T-H pair is assigned the
value “1”, otherwise, the pair is assigned the value “0”.</p>
      </sec>
      <sec id="sec-3-7">
        <title>3.7 Textual Entailment Module (TE)</title>
        <p>
          This TE module [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] is based on three types of matching, i.e., WordNet based Unigram
Match and Bigram Match and Skip-bigram Match.
2 http://members.unine.ch/jacques.savoy/clef/
3 http://tartarus.org/~martin/PorterStemmer/java.txt
a. WordNet based Unigram Match. In this method, the various unigrams in the
hypothesis for each Retrieved Sentence (T) - generated Hypothesis (H) pair are
checked for their presence in the retrieved text. WordNet synsets are identified for
each of the unmatched unigrams in the hypothesis. If any synset for the H unigram
match with any synset of a word in the T then the hypothesis unigram is considered as
a successful WordNet based unigram match. If the value of
Wordnet_Unigram_Match is 0.75 or more, i.e., 75% or more unigrams in the H match
either directly or through WordNet synonyms, then the T-H pair is considered as an
entailment. The T-H pair is then assigned the value “1”, otherwise, the pair is
assigned the value “0”.
b. Bigram Match. Each bigram in the hypothesis is searched for a match in the
corresponding text part. The measure Bigram_Match is calculated as the fraction of
the hypothesis bigrams that match in the corresponding text, i.e.,
Bigram_Match=(Total number of matched bigrams in a T-H pair /Number of
hypothesis bigrams). If the value of Bigram_Match is 0.5 or more, i.e., 50% or more
bigrams in the H match in the corresponding T, then the T-H pair is considered as an
entailment. The T-H pair is then assigned the value “1”, otherwise, the pair is
assigned the value “0”.
c. Skip-grams. A skip-gram is any combination of n words in the order as they
appear in a sentence, allowing arbitrary gaps. In the present work, only
1-skipbigrams are considered where 1-skip-bigrams are bigrams with one word gap between
two words in a sentence. The measure 1-skip_bigram_Match is defined as
1_skip_bigram_Match = skip_gram(T,H) / n,
where skip_gram(T,H) refers to the number of common 1-skip-bigrams (pair of words
in order with one word gap) found in T and H and n is the number of 1-skip-bigrams
in the hypothesis H. If the value of 1_skip_bigram_Match is 0.5 or more, then the
TH pair is considered as an entailment. The text-hypothesis pair is then assigned the
value “1”, otherwise, the pair is assigned the value “0”.
        </p>
      </sec>
      <sec id="sec-3-8">
        <title>3.8 Question-Answer Type Analysis Module</title>
        <p>
          The original questions are pre-processed using Stanford Dependency parser [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The
question type and its expected answer type are generally identified by looking at the
question keyword. Table 1 lists the questions and the expected answer types. For
example, if the question type is “When”, the expected answer type is a
“DATE/TIME”. The answer string “&lt;a_str&gt;” is parsed by the RASP Parser [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. If the
RASP parser generates the tag “&lt;timex type=date&gt;” then the answer string is “1”,
otherwise it is “0”. For “What” type questions we look for the keyword (e.g.,
Company) that is related to “What” through a dependency relation. If the keyword is
“Company” the expected answer type is “Organization”. If the corresponding answer
string is tagged by the RASP parser as “Organization”, the answer string is marked as
“1”, otherwise it is “0”. If the question type is “How” and the answer string is tagged
as “CD” by the RASP parser, the answer string is marked as “1”, otherwise it is “0”.
        </p>
      </sec>
      <sec id="sec-3-9">
        <title>3.9 Chunk Module</title>
        <p>
          The question sentences are pre-processed using Stanford dependency parser. The
words along with their part of speech (POS) information are passed through a
Conditional Random Field (CRF) based chunker [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] to extract phrase level chunks
of the questions. A rule-based module is developed to identify the chunk boundaries.
The question-retrieved text pairs that achieve the maximum weight are identified and
the corresponding answers are tagged as “1”. The question-retrieved text pair that
receives a zero weight is tagged as “0”.
        </p>
      </sec>
      <sec id="sec-3-10">
        <title>3.10 Syntactic Similarity Module</title>
        <p>
          This module is based on the Stanford dependency parser [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], which normalizes data
from the corpus of text and hypothesis pairs, accomplishes the dependency analysis
and creates appropriate structures.
        </p>
      </sec>
      <sec id="sec-3-11">
        <title>3.10.1 Matching Module</title>
        <p>After dependency relations are identified for both the retrieved sentence and the
hypothesis in each pair, the hypothesis relations are compared with the retrieved text
relations. The different features that are compared are noted below. In all the
comparisons, a matching score of 1 is considered when the complete dependency
relations along with all of its arguments match in both the retrieved sentence and the
hypothesis. In case of a partial match for a dependency relation, a matching score of
0.5 is assumed.
a. Subject-Verb Comparison. The system compares hypothesis subject and verb
with retrieved sentence subject and verb that are identified through the nsubj and
nsubjpass dependency relations. A matching score of 1 is assigned in case of a
complete match. Otherwise, the system considers the following matching process.
b. WordNet Based Subject-Verb Comparison. If the corresponding hypothesis and
sentence subjects do match in the subject-verb comparison, but the verbs do not
match, then the WordNet distance between the hypothesis and the sentence is
compared. If the value of the WordNet distance is less than 0.5, indicating a closeness
of the corresponding verbs, then a match is considered and a matching score of 0.5 is
assigned. Otherwise, the subject-subject comparison process is applied.
c. Subject-Subject Comparison. The system compares hypothesis subject with
sentence subject. If a match is found, a score of 0.5 is assigned to the match.
d. Object-Verb Comparison. The system compares hypothesis object and verb with
retrieved sentence object and verb that are identified through dobj dependency
relation. In case of a match, a matching score of 0.5 is assigned.
e. WordNet Based Object-Verb Comparison. The system compares hypothesis
object with text object. If a match is found then the verb corresponding to the
hypothesis object with retrieved sentence object's verb is compared. If the two verbs
do not match then the WordNet distance between the two verbs is calculated. If the
value of WordNet distance is below 0.5 then a matching score of 0.5 is assigned.
f. Cross Subject-Object Comparison. The system compares hypothesis subject and
verb with retrieved sentence object and verb or hypothesis object and verb with
retrieved sentence subject and verb. In case of a match, a matching sc ore of 0.5 is
assigned.
g. Number Comparison. The system compares numbers along with units in the
hypothesis with similar numbers along with units in the retrieved sentence. Units are
first compared and if they match then the corresponding numbers are compared. In
case of a match, a matching score of 1 is assigned.
h. Noun Comparison. The system compares hypothesis noun words with retrieved
sentence noun words that are identified through nn dependency relation. In case of a
match, a matching score of 1 is assigned.
i. Prepositional Phrase Comparison. The system compares the prepositional
dependency relations in the hypothesis with the corresponding relations in the
retrieved sentence and then checks for the noun words that are arguments of the
relation. In case of a match, a matching score of 1 is assigned.
j. Determiner Comparison. The system compares the determiner in the hypothesis
and in the retrieved sentence that are identified through det relation. In case of a
match, a matching score of 1 is assigned.
k. Other relation Comparison. Besides the above relations that are compared, all
other remaining relations are compared verbatim in the hypothesis and in the retrieved
sentence. In case of a match, a matching score of 1 is assigned.</p>
        <p>API for WordNet Searching RiWordnet4 provides Java applications with the ability
to retrieve data from the WordNet database.</p>
      </sec>
      <sec id="sec-3-12">
        <title>3.11 Answer Scoring Module</title>
        <p>In this module, we have got the weight from Named Entity Recognition (NER)
Module (Section 3.6), Textual Entailment (TE) Module (Section 3.7), Question Type
Analysis Module (Section 3.8), Chunk Boundary (Section 3.9) and Syntactic
Similarity Module (Section 3.10).
4 http://www.rednoise.org/rita/wordnet/documentation/index.htm</p>
      </sec>
      <sec id="sec-3-13">
        <title>3.12 Answer Ranking Module</title>
        <p>For each question has five hypothesis. Hypothesis are ranked by using NER Module,
Textual Entailment Module, Chunking Module, Syntactic Similarity Module,
Question type analysis Module. The highest weight hypothesis is the final answer.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 Question</title>
    </sec>
    <sec id="sec-5">
      <title>Architecture</title>
      <sec id="sec-5-1">
        <title>4.1 Document Processing Module</title>
      </sec>
      <sec id="sec-5-2">
        <title>4.1.1 XML parser</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Answering based</title>
    </sec>
    <sec id="sec-7">
      <title>Machine</title>
    </sec>
    <sec id="sec-8">
      <title>Reading</title>
    </sec>
    <sec id="sec-9">
      <title>System</title>
      <p>The given XML corpus has been parsed using XML parser. The XML parser extracts
the document and associated questions. After parsing, the documents and the
associated questions are extracted from the given XML documents and stored in the
system.</p>
      <sec id="sec-9-1">
        <title>4.1.2 Answer Pattern Generation</title>
        <p>Each question has a number of answer options and the task is to identify the best
answer to the question given an associated document. Each question in the system is
identified as the (question, document) pair represented as {qi, d_id} where i=1…10.
There are 10 questions corresponding to each document. The “WH” word in the
question is substituted by the given answer option to generate the answer pattern. The
set of WH words include WHP ={‘Who’, ‘What’, ‘Where’, ‘Name’, ‘Which’,
‘Whom’, ‘Why’}. Each answer pattern is represented in the system as {d_id, q_idi,
a_idj}, where, d_id=document id, q_idi= i th query, where i=1…10, a_idj= j th answer
option, where j=1…5.</p>
        <sec id="sec-9-1-1">
          <title>Let us consider an example.</title>
          <p>Question: Who is the founder of the SING campaign?
Answer Option: Nelson Mandela
WHP: who
Generated Answer Pattern: Nelson Mandela is the founder of the SING campaign
Each answer pattern is stored in the system as the pair (PAT, KL) where,
PAT= Probable Answer Text, which is the generated answer pattern and
KL= Keyword List, is a list of words after removing the stop words.
For example, the above generated answer pattern is stored as
PAT= “Nelson Mandela is the founder of the SING campaign”
KL= “Nelson”, “Mandela”, “founder”, “SIGN”, “campaign”.</p>
        </sec>
      </sec>
      <sec id="sec-9-2">
        <title>4.1.3 Named Entity (NE) Identification</title>
        <p>For each question, the system must identify the correct answer among the proposed
alternative answer options. Each generated answer pattern corresponding to a question
is compared with each sentence in the document to assign an inference score. The
score assignment module requires that the named entities in each sentence and in each
answer pattern are identified. The CRF-based Stanford Named Entity Tagger5 (NE
Tagger) has been used to identify and mark the named entities in the documents and
queries. The tagged documents and queries are passed to the lexical inference
submodule.</p>
      </sec>
      <sec id="sec-9-3">
        <title>4.1.4 Anaphora Resolution</title>
        <p>It has been observed that resolving the anaphors in the sentences in the documents
improves the inference score of the sentence with respect to each associated answer
pattern. The following basic anaphora resolution techniques have been applied in the
present task.
(a) Each first person personal pronoun in the set PNI = {‘I’, ‘me’, ‘my’, ‘myself’}
generally refers the author of the document as describer. For example, the anaphors in
the following sentence can be resolved in the following steps:
I am going to share with you the story as to how I have become an HIV/AIDS
campaigner.</p>
        <sec id="sec-9-3-1">
          <title>5 http://nlp.stanford.edu/ner/index.shtml</title>
          <p>Step1: &lt; PNI &gt; am going to share with you the story as to how &lt; PNI &gt; have become
an HIV/AIDS campaigner.</p>
          <p>Step2: &lt; PNI= “author” value= “Annie Lennox”&gt; am going to share with you the
story as to how &lt; PNI= “author” value= “Annie Lennox” &gt; have become an
HIV/AIDS campaigner.</p>
          <p>Step3: &lt;NE=Person value= Annie Lennox&gt; am going to share with you the story as
to how &lt;NE=Person value= Annie Lennox&gt; have become an HIV/AIDS
campaigner.</p>
          <p>In direct speech sentences PNI refers to the first named entity (speaker) of that
sentence. For example, the anaphors in the following sentence can be resolved in the
following steps:
Frankie said, "I am number 22 in line, and I can see the needle coming down towards
me, and there is blood all over the place.</p>
          <p>Step 1: &lt;NE=Person value=Frankie&gt; said, " &lt; PNI &gt; am number 22 in line, and &lt;
PNI &gt; can see the needle coming down towards &lt; PNI &gt;, and there is blood all over
the place.</p>
          <p>Step 2: &lt;NE=Person value=Frankie&gt; said, "&lt; PNI= “NE” value= “Frankie” &gt;” am
number 22 in line, and &lt; PNI= “NE” value= “Frankie” &gt; can see the needle coming
down towards &lt; PNI= “NE” value= “Frankie” &gt;, and there is blood all over the
place.</p>
          <p>Step 3: &lt;NE= “Person” value= “Frankie”&gt; said, " &lt;NE= “Person” value=
“Frankie”&gt; am number 22 in line, and &lt;NE=Person value=Frankie&gt; can see
the needle coming down towards &lt;NE=Person value=Frankie&gt;, and there is
blood all over the place.
(b) Each second person personal pronoun in the set PNHe/She = {‘he’, ‘his’, ‘him’, ‘her’,
‘she’} generally refers the last NE of the previous sentence. For example, the
anaphors in the following sentence can be resolved in the following steps:
I was invited to take part in the launch of Nelson Mandela's 46664 Foundation. That
is his HIV/AIDS foundation.</p>
          <p>Step 1: &lt; PNI &gt; was invited to take part in the launch of &lt;NE= “Person” value=
“Nelson Mandela”&gt;'s 46664 Foundation. That is &lt;PNHe/She &gt; HIV/AIDS foundation.
Step 2: &lt; PNI= “author” value= “Annie Lennox”&gt;was invited to take part in the
launch of &lt;NE=Person value= “Nelson Mandela”&gt;'s 46664 Foundation. That is
&lt;PNHe/She = “PREV_NE” value= “Nelson Mandela”&gt; HIV/AIDS foundation.
Step 3: &lt; PNI= “author” value= “Annie Lennox”&gt;was invited to take part in the
launch of &lt;NE= “Person” value= “Nelson Mandela”&gt;'s 46664 Foundation. That is
&lt;NE= “Person” value= “Nelson Mandela”&gt; HIV/AIDS foundation.
But, in indirect speech sentences, PNHe/She refers to the first named entity (speaker) of
that sentence. For example, the anaphors in the following sentence can be resolved in
the following steps:
Alexander Graham Bell famously said that on his first successful telephone Call.
Step 1: &lt;NE= “Person” value= “Alexander Graham Bell”&gt; famously said that on
&lt; PNHe/She &gt; first successful telephone Call.</p>
          <p>Step 2: &lt;NE= “Person” value= “Alexander Graham Bell”&gt; famously said that on
&lt; PNHe/She = “SEN_NE” value= “Alexander Graham Bell”&gt; first successful
telephone Call.</p>
          <p>Step 3: &lt;NE= “Person” value= “Alexander Graham Bell”&gt; famously said that on
&lt;NE=“Person” value=“Alexander Graham Bell”&gt;first successful telephone Call.</p>
        </sec>
      </sec>
      <sec id="sec-9-4">
        <title>4.2 Answering Module</title>
      </sec>
      <sec id="sec-9-5">
        <title>4.2.1 Scoring Assignment</title>
        <p>Each sentence in the associated document is assigned an inference score with respect
to each generated answer pattern.</p>
        <p>This module takes query frame as input and returns score as output. The algorithm
AnswerScore describes the scoring procedure.
keywordmatched = 0 // count no of matched keyword
Algorithm AnswerScore (sentence, PAT, KL)
Step 1: [Initialization]</p>
        <p>score = 0
Step 2: [Check whether PAT matches in a sentence]</p>
        <p>If PAT matches in a sentence then
Score = 1
goto step 5
Step 3: [Check each keyword in KL]</p>
        <p>For each keyword in KL</p>
        <p>If keyword matches in a sentence then</p>
        <p>Score = score + 1 / (number of keywords -1)</p>
        <p>Keywordmatched = keywordmatched + 1
Step 4: [Check whether all the keywords have matched]</p>
        <p>If (keywordmatched = = total keywords – 1) then</p>
        <p>Score =1
Step 5: Return score</p>
        <p>End
Now, for each given answer option a score is calculated and the answer option with
highest score is taken as correct answer for the given query. The algorithm
SelectAnswerOption describes the option selection procedure.
correct_option= index of maximum AQ={ AQ1, AQ2, AQ3, AQ4, AQ5 }</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>5 Hybrid Machine Reading System Architecture</title>
      <p>We have got the ranking score for each answer option from the AV System and also
from the QA system. Now we have used voting technique to select the final answer
using the selected answer from our IR (Nutch) system. The Nutch system also
provides a ranking score for each answer option. We have passed all the three answers
from the different systems, i.e., the highest ranked answer option, to the voting
module, which selects the final answer option based on the highest vote of the
answers. The system architecture is shown in figure 3. When any two systems identify
the same answer then this answer gets automatically selected as the final answer as it
has at least two votes. But when three different answers have been selected from the
three systems for a question, then this voting module has to give priority to a system.
We have submitted two runs from our two different hybrid system based on this
voting module. In the hybrid system of run id 4 we have given priority to the QA
system and in the hybrid system of run id 5 we have given priority to the AV system.</p>
    </sec>
    <sec id="sec-11">
      <title>6 Evaluation</title>
      <p>In QA4MRE track, we have submitted total seven runs. But run 1 and run 2 are
identical because same run has been submitted twice. Run 4 and run 5 also are
identical because the same run has been submitted twice. So we have five unique run
from three different systems. Evaluation results are shown in table 4.</p>
      <p>For System 1: Run 1, Run 2 (Based on Answer Validation System)
For System 2: Run 3 (Based on Question Answering System)</p>
      <p>For System 3: Run 4, Run 5 (Hybrid system)
The main measure used in this evaluation campaign is c@1, which is defined in
equation 1.
(1)
Where, nR: the number of correctly answered questions, nU: number of unanswered
questions and n: the total number of questions</p>
      <sec id="sec-11-1">
        <title>Evaluation at question-answering level:</title>
        <sec id="sec-11-1-1">
          <title>C1: Number of questions ANSWERED</title>
          <p>C2: Number of questions UNANSWERED
C3: Number of questions ANSWERED with RIGHT candidate answer
C4: Number of questions ANSWERED with WRONG candidate answer
C5: Number of questions UNANSWERED with RIGHT candidate answer
C6: Number of questions UNANSWERED with WRONG candidate answer
C7: Number of questions UNANSWERED with EMPTY candidate</p>
        </sec>
      </sec>
      <sec id="sec-11-2">
        <title>Evaluation at reading-test level:</title>
      </sec>
    </sec>
    <sec id="sec-12">
      <title>7 Conclusion</title>
      <p>The question answering system has been developed as part of the participation in the
QA4MRE track as part of the CLEF 2011 evaluation campaign. The overall system
has been evaluated using the evaluation metrics provided as part of the QA4MRE
2011 track. The evaluation results are satisfactory considering that this is the second
participation in the track. Future works will be motivated towards improving the
performance of the system.</p>
      <p>Acknowledgements. We acknowledge the support of the IFCPAR funded
IndoFrench project “An Advanced Platform for Question Answering Systems” and the
DIT, Government of India funded project “Development of Cross Lingual
Information Access (CLIA) System Phase II”.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Anselmo</given-names>
            <surname>Peñas</surname>
          </string-name>
          , Pamela Forner, Richard Sutcliffe, Álvaro Rodrigo, Corina Forăscu, Iñaki Alegria, Danilo Giampiccolo, Nicolas Moreau, Petya Osenova.: Overview of ResPubliQA 2009:
          <article-title>Question Answering Evaluation over European Legislation</article-title>
          .
          <source>In Working Notes for the CLEF 2009 Workshop</source>
          , 30 September-2
          <string-name>
            <surname>October</surname>
          </string-name>
          ,
          <year>2009</year>
          , Corfu, Greece.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Anselmo</given-names>
            <surname>Peñas</surname>
          </string-name>
          , Pamela Forner, Álvaro Rodrigo, Richard Sutcliffe, Corina Forăscu and
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Mota</surname>
          </string-name>
          .: Overview of ResPubliQA 2010:
          <article-title>Question Answering Evaluation over European Legislation</article-title>
          .
          <source>In Working Notes for the CLEF 2010 Workshop</source>
          , Padua, Italy,
          <fpage>20</fpage>
          -
          <issue>23</issue>
          <year>September 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Anselmo</given-names>
            <surname>Peñas</surname>
          </string-name>
          , Eduard Hovy, Pamela Forner, Álvaro Rodrigo, Richard Sutcliffe, Corina Forascu,
          <string-name>
            <given-names>Caroline</given-names>
            <surname>Sporleder</surname>
          </string-name>
          .
          <source>Overview of QA4MRE at CLEF</source>
          <year>2011</year>
          :
          <article-title>Question Answering for Machine Reading Evaluation</article-title>
          , Working Notes of CLEF
          <year>2011</year>
          .
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Peñas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodrigo</surname>
          </string-name>
          , Á. ,
          <string-name>
            <surname>Sama</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verdejo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Overview of the answer validation exercise 2006</article-title>
          .
          <source>Working Notes of CLEF</source>
          <year>2006</year>
          .
          <article-title>(</article-title>
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Peñas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodrigo</surname>
            ,
            <given-names>Á</given-names>
          </string-name>
          , Verdejo,
          <string-name>
            <surname>F.</surname>
          </string-name>
          :
          <source>Overview of the Answer Validation Exercise 2007. Working Notes of CLEF</source>
          <year>2007</year>
          .
          <article-title>(</article-title>
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Rodrigo</surname>
          </string-name>
          , Á.,
          <string-name>
            <surname>Peñas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verdejo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Overview of the answer validation exercise 2008</article-title>
          .
          <source>Working Notes of CLEF</source>
          <year>2008</year>
          .
          <article-title>(</article-title>
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Partha</given-names>
            <surname>Pakray</surname>
          </string-name>
          , Pinaki Bhaskar, Santanu Pal,
          <string-name>
            <surname>Dipankar Das</surname>
          </string-name>
          ,
          <article-title>Sivaji Bandyopadhyay and Alexander Gelbukh: JU_CSE_TE: System Description QA@CLEF 2010 - ResPubliQA</article-title>
          .
          <source>CLEF 2010 Workshop on Multiple Language Question Answering (MLQA</source>
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Pakray</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelbukh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <source>Answer Validation using Textual Entailment. 12th CICLing, Lecture Notes in Computer Science</source>
          ,
          <year>2011</year>
          , Volume
          <volume>6609</volume>
          /
          <year>2011</year>
          ,
          <fpage>353</fpage>
          -
          <lpage>364</lpage>
          , DOI: 10.1007/978-3-
          <fpage>642</fpage>
          -19437-5_
          <fpage>29</fpage>
          . (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>E.</given-names>
            <surname>Briscoe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carroll</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Watson</surname>
          </string-name>
          .:
          <article-title>The Second Release of the RASP System</article-title>
          .
          <source>In Proceedings of the COLING/ACL 2006 Interactive Presentation Sessions.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Marie-Catherine de Marneffe</surname>
          </string-name>
          , Bill MacCartney, and
          <string-name>
            <surname>Christopher</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          .:
          <article-title>Generating Typed Dependency Parses from Phrase Structure Parses</article-title>
          .
          <source>In 5th International Conference on Language Resources and Evaluation (LREC)</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Xuan-Hieu Phan</surname>
          </string-name>
          .:
          <string-name>
            <surname>CRFChunker: CRF English Phrase</surname>
            <given-names>Chunker. PACLIC</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>(</article-title>
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>