<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Hybrid Tweet Contextualization System using IR and Summarization</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pinaki Bhaskar</string-name>
          <email>pinaki.bhaskar@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Somnath Banerjee</string-name>
          <email>s.banerjee1980@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sivaji Bandyopadhyay</string-name>
          <email>sivaji_cse_ju@yahoo.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Engineering, Jadavpur University</institution>
          ,
          <addr-line>Kolkata - 700032</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The article presents the experiments carried out as part of the participation in the Tweet Contextualization (TC) track of INEX 2012. We have submitted three runs. The INEX TC task has two main sub tasks, Focused IR and Automatic Summarization. In the Focused IR system, we first preprocess the Wikipedia documents and then index them using Nutch with NE field. Stop words are removed and all NEs are tagged from each query tweet and all the remaining tweet words are stemmed using Porter stemmer. The stemmed tweet words form the query for retrieving the most relevant document using the index. The automatic summarization system takes as input the query tweet along with the title from the most relevant text document. Most relevant sentences are retrieved from the associated document based on the TF-IDF of the matching query tweet, NEs text and title words. Each retrieved sentence is assigned a ranking score in the Automatic Summarization system. The answer passage includes the top ranked retrieved sentences with a limit of 500 words. The three unique runs differ in the way in which the relevant sentences are retrieved from the associated document.</p>
      </abstract>
      <kwd-group>
        <kwd>Information Retrieval</kwd>
        <kwd>Automatic Summarization</kwd>
        <kwd>Question Answering</kwd>
        <kwd>Information Extraction</kwd>
        <kwd>INEX 2012</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>With the explosion of information in Internet, Natural language Question Answering
(QA) is recognized as a capability with great potential. Traditionally, QA has
attracted many AI researchers, but most QA systems developed are toy systems or
games confined to laboratories and to a very restricted domain. Several recent
conferences and workshops have focused on aspects of the QA research. Starting in
1999, the Text Retrieval Conference (TREC)1 has sponsored a question-answering
track, which evaluates systems that answer factual questions by consulting the
documents of the TREC corpus. A number of systems in this evaluation have
successfully combined information retrieval and natural language processing</p>
      <sec id="sec-1-1">
        <title>1 http://trec.nist.gov/</title>
        <p>
          techniques. More recently, Conference and Labs of Evaluation Forums (CLEF)2 are
organizing QA lab from 2010. INEX3 has also started Question Answering track. Last
year, INEX 2011 designed a QA track [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] to stimulate the research for real world
application. The Question Answering (QA) task performed by the participating
groups of INEX 2011 is contextualizing tweets, i.e., answering questions of the form
"what is this tweet about?" using a recent cleaned dump of the Wikipedia (April
2011). This year they renamed this task as Tweet Contextualization.
        </p>
        <p>Current INEX 2012 Tweet Contextualization (TC) track gives QA research a new
direction by fusing IR and summarization with QA. The TC track of INEX 2012 had
two major sub tasks. The first task is to identify the most relevant document from the
Wikipedia dump, for this we need a focused IR system. And the second task is to
extract most relevant passages from the most relevant retrieved document. So we need
an automatic summarization system. The general purpose of the task involves tweet
analysis, passage and/or XML elements retrieval and construction of the answer, more
specifically, the summarization of the tweet topic.</p>
        <p>
          Automatic text summarization [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] has become an important and timely tool for
assisting and interpreting text information in today’s fast-growing information age.
Text Summarization methods can be classified into abstractive and extractive
summarization. An Abstractive Summarization ([
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] and [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]) attempts to develop an
understanding of the main concepts in a document and then expresses those concepts
in clear natural language. Extractive Summaries [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] are formulated by extracting key
text segments (sentences or passages) from the text, based on statistical analysis of
individual or mixed surface level features such as word/phrase frequency, location or
cue words to locate the sentences to be extracted. Our approach is based on Extractive
Summarization.
        </p>
        <p>
          In this paper, we describe a hybrid Tweet Contextualization system of focused IR
and automatic summarization for TC track of INEX 2012. The focused IR system is
based on Nutch architecture and the automatic summarization system is based on
TFIDF based sentence ranking and sentence extraction techniques. The same sentence
scoring and ranking approach of [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] has been followed. We have submitted
three runs in the QA track (177, 191 and 192).
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2 Related Works</title>
      <p>
        Recent trend shows hybrid approach of tweet contextualization using Information
Retrieval (IR) can improve the performance of the TC system. Reference [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] removed
incorrect answers of QA system using an IR engine. Reference [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] successfully used
methods of IR into QA system. Reference [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] used the IR system into QA and [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
proposed an efficient hybrid QA system using IR in QA.
      </p>
      <p>
        Reference [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] presents an investigation into the utility of document
summarization in the context of IR, more specifically in the application of so-called
query-biased summaries: summaries customized to reflect the information need
      </p>
      <sec id="sec-2-1">
        <title>2 http://www.clef-initiative.eu//</title>
        <p>3 https://inex.mmci.uni-saarland.de/
expressed in a query. Employed in the retrieved document list displayed after retrieval
took place, the summaries’ utility was evaluated in a task-based environment by
measuring users’ speed and accuracy in identifying relevant documents. This was
compared to the performance achieved when users were presented with the more
typical output of an IR system: a static predefined summary composed of the title and
first few sentences of retrieved documents. The results from the evaluation indicate
that the use of query-biased summaries significantly improves both the accuracy and
speed of user relevance judgments.</p>
        <p>
          A lot of research work has been done in the domain of both query dependent and
independent summarization. MEAD [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] is a centroid based multi document
summarizer, which generates summaries using cluster centroids produced by topic
detection and tracking system. NeATS [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] selects important content using sentence
position, term frequency, topic signature and term clustering. XDoX [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] identifies
the most salient themes within the document set by passage clustering and then
composes an extraction summary, which reflects these main themes. Graph based
methods have been also proposed for generating summaries. A document graph based
query focused multi-document summarization system has been described by [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]
and [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>
          In the present work, we have used the IR system as described in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]
and the automatic summarization system as discussed in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. In the later
part of this paper, section 3 describes the corpus statistics and section 4 shows the
system architecture of combined TC system of focused IR and automatic
summarization for INEX 2012. Section 5 details the Focused Information Retrieval
system architecture. Section 6 details the Automatic Summarization system
architecture. The evaluations carried out on submitted runs are discussed in Section 7
along with the evaluation results. The conclusions are drawn in Section 8.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 Corpus statistics</title>
      <p>The training data is the collection of documents that has been rebuilt based on recent
English Wikipedia dump (November 2011). All notes and bibliographic references
have been removed from Wikipedia pages to prepare plain xml corpus for an easy
extraction of plain text answers. Each training document is made of a title, an abstract
and sections. Each section has a sub-title. Abstract and sections are made of
paragraphs and each paragraph can have entities that refer to Wikipedia pages.
Therefore, the resulting corpus has this simple DTD as shown in table 1.</p>
      <p>Test data is made up of 1142 tweets from Twitter. There are two different formats
of tweets, one is the full JSON format with all tweet metadata as shown in the table 2
and another is the two-column text format with only tweet id and tweet text as shown
in the table 3.</p>
    </sec>
    <sec id="sec-4">
      <title>4 System Architecture</title>
      <p>In this section the overview of the system framework of the current INEX system has
been shown. The current INEX system has two major sub-systems; one is the Focused
IR system and the other one is the Automatic Summarization system. The Focused IR
system has been developed on the basic architecture of Nutch4, which use the
architecture of Lucene5. Nutch is an open source search engine, which supports only
the monolingual Information Retrieval in English, etc. The Higher-level system
architecture of the combined Tweet Contextualization system of Focused IR and
Automatic Summarization is shown in the Figure 1.</p>
    </sec>
    <sec id="sec-5">
      <title>5 Focused Information Retrieval (IR)</title>
      <p>The web documents are full of noises mixed with the original content. In that case it is
very difficult to identify and separate the noises from the actual content. INEX 2012
corpus, i.e., Wikipedia dump, had some noise in the documents and the documents are</p>
      <sec id="sec-5-1">
        <title>4 http://nutch.apache.org/</title>
        <p>5 http://lucene.apache.org/
in XML tagged format. So, first of all, the documents had to be preprocessed. The
document structure is checked and reformatted according to the system requirements.
XML Parser. The corpus was in XML format. All the XML test data has been parsed
before indexing using our XML Parser. The XML Parser extracts the Title of the
document along with the paragraphs.</p>
        <p>Noise Removal. The corpus has some noise as well as some special symbols that are
not necessary for our system. The list of noise symbols and the special symbols is
initially developed manually by looking at a number of documents and then the list is
used to automatically remove such symbols from the documents. Some examples are
“&amp;quot;”, “&amp;amp;”, “'''”, multiple spaces etc.</p>
        <p>Named Entity Recognizer (NER). After cleaning the corpus, the named entity
recognizer identifies all the named entities (NE) in the documents and tags them
according to their types, which are indexed during the document indexing.
Document Indexing. After parsing the Wikipedia documents, they are indexed using
Lucene, an open source indexer.</p>
        <sec id="sec-5-1-1">
          <title>5.2 Tweets Parsing</title>
          <p>After indexing has been done, the tweets had to be processed to retrieve relevant
documents. Each tweet / topic was processed to identify the query words for
submission to Lucene. The tweets processing steps are described below:
Stop Word Removal. In this step the tweet words are identified from the tweets. The
stop words and question words (what, when, where, which etc.) are removed from
each tweet and the words remaining in the tweets after the removal of such words are
identified as the query tokens. The stop word list used in the present work can be
found at http://members.unine.ch/jacques.savoy/clef/.</p>
          <p>Named Entity Recognizer (NER). After removing the stop words, the named entity
recognizer identifies all the named entities (NE) in the tweet and tags them according
to their types, which are used during the scoring of the sentences of the retrieved
document.</p>
          <p>Stemming. Query tokens may appear in inflected forms in the tweets. For English,
standard Porter Stemming algorithm6 has been used to stem the query tokens. After
stemming all the query tokens, queries are formed with the stemmed query tokens.</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>6 http://tartarus.org/~martin/PorterStemmer/java.txt</title>
        <sec id="sec-5-2-1">
          <title>5.3 Document Retrieval</title>
          <p>After searching each query into the Lucene index, a set of retrieved documents in
ranked order for each query is received.</p>
          <p>First of all, all queries were fired with AND operator. If at least one document is
retrieved using the query with AND operator then the query is removed from the
query list and need not be searched again. The rest of the queries are fired again with
OR operator. OR searching retrieves at least one document for each query. Now, the
top ranked relevant document for each query is considered for Passage selection.
Document retrieval is the most crucial part of this system. We take only the top
ranked relevant document assuming that it is the most relevant document for the
query or the tweet from which the query had been generated.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6 Automatic Summarization</title>
      <sec id="sec-6-1">
        <title>6.1 Sentence Extraction</title>
        <p>The document text is parsed and the parsed text is used to generate the summary. This
module will take the parsed text of the documents as input, filter the input parsed text
and extract all the sentences from the parsed text. So this module has two sub
modules, Text Filterization and Sentence Extraction.</p>
        <p>Text Filterization. The parsed text may content some junk or unrecognized character
or symbol. First, these characters or symbols are identified and removed. The text in
the query language are identified and extracted from the document using the Unicode
character list, which has been collected from Wikipedia7. The symbols like dot (.),
coma (,), single quote (‘), double quote (“), ‘!’, ‘?’ etc. are common for all languages,
so these are also listed as symbols.</p>
        <p>Sentence Extraction. In Sentence Extraction module, filtered parsed text has been
parsed to identify and extract all sentences in the documents. Sentence identification
and extraction is not an easy task for English document. As the sentence marker ‘.’
(dot) is not only used as a sentence marker, it has other uses also like decimal point
and in abbreviations like Mr., Prof., U.S.A. etc. So it creates lot of ambiguity. A
possible list of abbreviation had to created to minimize the ambiguity. Most of the
times the end quotation (”) is placed wrongly at the end of the sentence like .”. These
kinds of ambiguities are identified and removed to extract all the sentences from the
document.</p>
        <sec id="sec-6-1-1">
          <title>7 http://en.wikipedia.org/wiki/List_of_Unicode_characters</title>
        </sec>
      </sec>
      <sec id="sec-6-2">
        <title>6.2 Key Term Extraction</title>
        <p>Key Term Extraction module has three sub modules like Query Term, i.e., tweet term
extraction, tweet text extraction and Title words extraction. All these three sub
modules have been described in the following sections.</p>
        <p>Query/Tweet Term Extraction. First the query generated from the tweet, is parsed
using the Query Parsing module. In this Query Parsing module, the Named Entities
(NE) are identified and tagged in the given query using the Stanford NER8 engine.
Title Word Extraction. The title of the retrieved document is extracted and
forwarded as input given to the Title Word Extraction module. After removing all the
stop words from the title, the remaining tile words are extracted and used as the
keywords in this system.</p>
      </sec>
      <sec id="sec-6-3">
        <title>6.3 Top Sentence Identification</title>
        <p>All the extracted sentences are now searched for the keywords, i.e., query terms,
tweet’s text keywords and title words. Extracted sentences are given some weight
according to search and ranked on the basis of the calculated weight. For this task this
module has two sub modules: Weight Assigning and Sentence Ranking, which are
described below.</p>
        <p>Weight Assigning. This sub module calculates the weights of each sentence in the
document. There are three basic components in the sentence weight like query term
dependent score, tweet’s text keyword dependent score and title word dependent
score. These three components are calculated and added to get the final weight of a
sentence.</p>
        <p>Query/Tweet Term dependent score: Query/Tweet term dependent score is the most
important and relevant score for summary. Priority of this query/tweet dependent
score is maximum. The query dependent scores are calculated using equation 1.</p>
        <p>nq " " "
Q = ( Fq $ 20 + (nq ! q + 1)$ ( $1 !
s
q=1 # # p #
fpq ! 1% % %</p>
        <p>' ' ) p'
Ns &amp; &amp; &amp;
(1)
where, QS is the query/tweet term dependent score of the sentence s, q is the no. of the
query/tweet term, nq is the total no. of query terms, fpq is the possession of the word
which was matched with the query term q in the sentence s, Ns is the total no. of
words in sentence s,</p>
        <sec id="sec-6-3-1">
          <title>8 http://www-nlp.stanford.edu/ner/ A Hybrid Tweet Contextualization System using IR and Summarization and</title>
          <p>Fq =
p =
0; if query term q is not found .
1; if query term q is found
5; if query term is NE
3; if query term is not NE
(2)
(3)
(4)
(5)
(6)</p>
          <p>At the end of the equation 1, the calculated query term dependent score is
multiplied by p to give the priority among all the scores. If the query term is NE and
contained in a sentence then the weight of the matched sentence are multiplied by 5 as
the value of p is 5, to give the highest priority, other wise it has been multiplied by 3
(as p=3 for non NE query terms).</p>
          <p>Title Word dependent score: Title words are extracted from the title field of the top
ranked retrieved document. A title word dependent score is also calculated for each
sentence. Generally title words are also the much relevant words of the document. So
the sentence containing any title words can be a relevant sentence of the main topic of
the document. Title word dependent scores are calculated using equation 4.</p>
          <p>nt " "
Ts = ( Ft (nt ! t + 1)$ ( $1 !
t=0 # p #
fpt ! 1% %</p>
          <p>Ns '&amp; '&amp;
where, TS is the title word dependent score of the sentence s, t is the no. of the title
word, nt is the total number of title words, fpt is the position of the word which
matched with the title word t in the sentence s, Ns is the total number of words in
sentence s and</p>
          <p>Ft =
0; if title word t is not found .</p>
          <p>1; if title word t is found</p>
          <p>After calculating all the above three scores the final weight of each sentence is
calculated by simply adding all the two scores as mentioned in the equation 6.</p>
          <p>Ws = Qs + Ts
where, WS is the final weight of the sentence s.</p>
          <p>Sentence Ranking. After calculating weights of all the sentences in the document,
sentences are sorted in descending order of their weight. In this process if any two or
more than two sentences get equal weight, then they are sorted in the ascending order
of their positional value, i.e., the sentence number in the document. So, this Sentence
Ranking module provides the ranked sentences.</p>
        </sec>
      </sec>
      <sec id="sec-6-4">
        <title>6.4 Summary Generation</title>
        <p>
          This is the final and most critical module of this system. This module generates the
Summary from the ranked sentences. As in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] using equation 9, the module selects
the ranked sentences subject to maximum length of the summary.
        </p>
        <p>∑ li Si &lt; L (9)</p>
        <p>i
where li is the length (in no. of words) of sentence i, Si is a binary variable representing
the selection of sentence i for the summary and L (=500 words) is the maximum length
of the summary.</p>
        <p>Now, the selected sentences along with their weight are presented as the INEX
output format.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7 Evaluation</title>
      <sec id="sec-7-1">
        <title>7.1 Informative Content Evaluation</title>
        <p>
          The organizers did the Informative Content evaluation [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] by selecting relevant
passages. 50 topics were evaluated which was the pool of 14 654 sentences, 471 344
tokens, vocabulary of 59 020 words. Among them, 2801 sentences, 103889 tokens,
vocabulary of 19037 words, are relevant. There are 8 topics with less than 500
relevant tokens. The evaluation measures of Information content divergences over
{1,2,3,4gap}-grams (FRESA package) because it was too sensitive to smoothing on
the qa-rels. So simple log difference of equation 10 was used:
        </p>
        <p>! max (P (t / reference), P (t / summary))$
' log #
" min (P (t / reference), P (t / summary)) %&amp;
(10)</p>
        <p>We have submitted three runs (177, 191 and 192). The evaluation scores with the
baseline system scores of informativeness by organizers of all topics are shown in the
table 4.</p>
      </sec>
      <sec id="sec-7-2">
        <title>7.2 Readability Evaluation</title>
        <p>
          For Readability evaluation [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] all passages in a summary have been evaluated
according to Syntax (S), Anaphora (A), Redundancy (R) and Trash (T). If a passage
contains a syntactic problem (bad segmentation for example) then it has been marked
as Syntax (S) error. If a passage contains an unsolved anaphora then it has been
marked as Anaphora (A) error. If a passage contains any redundant information, i.e.,
an information that have already been given in a previous passage then it has been
marked as Redundancy (R) error. If a passage does not make any sense in its context
(i.e., after reading the previous passages) then these passages must be considered as
trashed, and readability of following passages must be assessed as if these passages
were not present, so they were marked as Trash (T). The readability evaluation scores
are shown in the table 5.
        </p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>8 Conclusion and Future Works</title>
      <p>The tweet contextualization system has been developed as part of the participation in
the Tweet Contextualization track of the INEX 2012 evaluation campaign. The
overall system has been evaluated using the evaluation metrics provided as part of this
track of INEX 2012. Considering that this is the second participation in the track, the
evaluation results are satisfactory, which will really encourage us to continue work on
it and participate in this track in future.</p>
      <p>Future works will be motivated towards improving the performance of the system
by concentrating on co-reference and anaphora resolution, multi-word identification,
para phrasing, feature selection etc. In future, we will also try to use semantic
similarity, which will increase our relevance score.</p>
      <p>Acknowledgements. We acknowledge the support of the IFCPAR funded
IndoFrench project “An Advanced Platform for Question Answering Systems” and the
DIT, Government of India funded project “Development of Cross Lingual
Information Access (CLIA) System Phase II”.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>SanJuan</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Moriceau</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tannier</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bellot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mothe</surname>
          </string-name>
          , J.:
          <article-title>Overview of the INEX 2011 Question Answering Track (QA@INEX)</article-title>
          .
          <article-title>In: Focused Retrieval of Content and Structure, 10th International Workshop of the Initiative for the Evaluation of XML Retrieval (INEX), Geva</article-title>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Kamps</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Schenkel</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (Eds.). Lecture Notes in Computer Sc., Springer (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Jezek</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steinberger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Automatic Text summarization</article-title>
          . In: Snasel,
          <string-name>
            <surname>V</surname>
          </string-name>
          . (ed.)
          <article-title>Znalosti 2008</article-title>
          .
          <source>ISBN 978-80-227-2827-0</source>
          , pp.
          <fpage>1</fpage>
          --
          <lpage>12</lpage>
          . FIIT STU Brarislava,
          <article-title>Ustav Informatiky a softveroveho inzinierstva (</article-title>
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Erkan</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.:</given-names>
          </string-name>
          <article-title>LexRank: Graph-based Centrality as Salience in Text Summarization</article-title>
          .
          <source>In: Journal of Artificial Intelligence Research</source>
          , vol.
          <volume>22</volume>
          , pp.
          <fpage>457</fpage>
          --
          <lpage>479</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hahn</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romacker</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The SYNDIKATE text Knowledge base generator</article-title>
          .
          <source>In: the first International conference on Human language technology research, Association for Computational Linguistics</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , Morristown, NJ, USA (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Kyoomarsi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khosravi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eslami</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dehkordy</surname>
            ,
            <given-names>P.K.</given-names>
          </string-name>
          :
          <article-title>Optimizing Text Summarization Based on Fuzzy Logic</article-title>
          . In: Seventh IEEE/ACIS International Conference on Computer and Information Science, pp.
          <fpage>347</fpage>
          --
          <lpage>352</lpage>
          . IEEE, University of Shahid Bahonar Kerman, UK (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bhaskar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A Query Focused Multi Document Automatic Summarization</article-title>
          .
          <source>In: the 24th Pacific Asia Conference on Language, Information and Computation (PACLIC 24)</source>
          , Tohoku University, Sendai, Japan (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Bhaskar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <string-name>
            <given-names>A Query</given-names>
            <surname>Focused</surname>
          </string-name>
          <article-title>Automatic Multi Document Summarizer</article-title>
          .
          <source>In: the International Conference on Natural Language Processing (ICON)</source>
          , pp.
          <fpage>241</fpage>
          --
          <lpage>250</lpage>
          . IIT, Kharagpur, India (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Rodrigo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iglesias</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          , Pe˜nas,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Garrido</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Araujo</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>A Question Answering System based on Information Retrieval and Validation</article-title>
          ,
          <source>ResPubliQA</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Schiffman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McKeown</surname>
            ,
            <given-names>K.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grishman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allan</surname>
          </string-name>
          , J.:
          <article-title>Question Answering using Integrated Information Retrieval and Information Extraction</article-title>
          .
          <source>In: NAACL HLT</source>
          , pp.
          <fpage>532</fpage>
          --
          <lpage>539</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Pakray</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhaskar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelbukh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>JU_CSE_TE: System Description QA@CLEF 2010 - ResPubliQA</article-title>
          . In:
          <article-title>Multiple Language Question Answering (MLQA</article-title>
          <year>2010</year>
          ), CLEF-2010, Padua, Italy (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pakray</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhaskar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pal</surname>
            ,
            <given-names>B.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelbukh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A Hybrid Question Answering System based on Information Retrieval and Answer Validation</article-title>
          .
          <source>In: Question Answering for Machine Reading Evaluation (QA4MRE)</source>
          ,
          <source>CLEF-2011</source>
          , Amsterdam (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Tombros</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanderson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Advantages of Query Biased Summaries in Information Retrieval</article-title>
          . In: SIGIR (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jing</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Styś</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tam</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Centroid- based summarization of multiple documents</article-title>
          .
          <source>J. Information Processing and Management</source>
          .
          <volume>40</volume>
          ,
          <fpage>919</fpage>
          -
          <lpage>938</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
            ,
            <given-names>E.H.</given-names>
          </string-name>
          :
          <article-title>From Single to Multidocument Summarization: A Prototype System and its Evaluation</article-title>
          . In: ACL, pp.
          <fpage>457</fpage>
          --
          <lpage>464</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Hardy</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shimizu</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strzalkowski</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ting</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wise</surname>
            ,
            <given-names>G. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          . X.:
          <article-title>Cross-document summarization by concept classification</article-title>
          .
          <source>In: SIGIR</source>
          , pp.
          <fpage>65</fpage>
          --
          <lpage>69</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Paladhi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A Document Graph Based Query Focused MultiDocument Summarizer</article-title>
          .
          <source>In: the 2nd International Workshop on Cross Lingual Information Access (CLIA)</source>
          , pp.
          <fpage>55</fpage>
          -
          <lpage>62</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Bhaskar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neogi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A Hybrid QA System with Focused IR and Automatic Summarization for INEX 2011</article-title>
          . In: Geva,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Kamps</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Schenkel</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          .(eds.):
          <article-title>Focused Retrieval of Content and Structure: 10th International Workshop of the Initiative for the Evaluation of XML Retrieval</article-title>
          ,
          <source>INEX 2011. Lecture Notes in Computer Science</source>
          , vol.
          <volume>7424</volume>
          . Springer Verlag, Berlin, Heidelberg (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>