<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bengali and Hindi to English Cross-language Text Retrieval under Limited Resources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Debasis Mandal</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandipan Dandapat</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mayank Gupta</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pratyush Banerjee</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sudeshna Sarkar IIT Kharagpur</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>India</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>This paper describes our experiment on two cross-lingual and one monolingual English text retrievals at CLEF1 in the ad-hoc track. The cross-language task includes the retrieval of English documents in response to queries in two most widely spoken Indian languages, Hindi and Bengali. For our experiment, we had access to a HindiEnglish bilingual lexicon, 'Shabdanjali', consisting of approx. 26K Hindi words. But neither we had any effective Bengali-English bilingual lexicon nor any parallel corpora to build the statistical lexicon. Under this limited resources, we mostly depended on our phoneme-based transliterations to generate equivalent English query from Hindi and Bengali topics. We adopted Automatic Query Generation and Machine Translation approach for our experiment. Other language-specific resources included a Bengali morphological analyzer, a Hindi stemmer and a set of 200 Hindi and 273 Bengali stopwords. Lucene framework was used for stemming, indexing, retrieval and scoring of the corpus documents. The CLEF results suggested the need for a rich bilingual lexicon for CLIR involving Indian languages. The best MAP values for Bengali, Hindi and English queries for our experiment were 7.26, 4.77 and 36.49 respectively.</p>
      </abstract>
      <kwd-group>
        <kwd>Bengali</kwd>
        <kwd>Hindi</kwd>
        <kwd>Transliteration</kwd>
        <kwd>Cross-language Text Retrieval</kwd>
        <kwd>CLEF Evaluation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Cross-language (or cross-lingual) Information Retrieval (CLIR) involves the study of retrieving
the documents in a language other than the query language. Since the language of query and
documents to be retrieved are different, either the documents or queries need to be translated
in CLIR. But this translation step tends to cause a reduction in the retrieval performance of</p>
    </sec>
    <sec id="sec-2">
      <title>1Cross Language Evaluation Forum. http://clef-campaign.org</title>
      <p>
        CLIR as compared to monolingual information retrieval. A study in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] showed that missing
specialized vocabulary, missing general terms, wrong translation due to ambiguity and correct
identical translation are the four most important factors for the difference in performance for
over 70% queries between monolingual and cross-lingual retrievals. This puts the importance on
effective translation in CLIR research. Again, the document translation requires a lot of memory
and processing capacity than its counterpart and therefore the query translation is more popular
in the IR research community involving multiple languages[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Oard [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] presents an overview of the Controlled Vocabulary and Free Text retrieval approaches
followed in CLIR research within the query translation framework. But the present research in
CLIR are mainly concentrated around three approaches: Dictionary based Machine Translation
(MT), Parallel Corpora based statistical lexicon and Ontology-based methods. The basic idea in
Machine Translation is to replace each term in the query with an appropriate term or a set of
terms from the lexicon. In current MT systems the quality of translations is very low and the high
quality is achieved only when the application is domain-specific [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The Parallel Corpora-based
method utilizes the broad repository of multi-lingual corpora to build the statistical lexicon from
the simliar training data as of the target collection. Knowledge-based approaches use ontology or
thesauri to replace the source language word by all of its target language equivalents. Some of the
CLIR models built on these approaches or on their hybrids can be found in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ][
        <xref ref-type="bibr" rid="ref6">6</xref>
        ][
        <xref ref-type="bibr" rid="ref8">8</xref>
        ][
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>This paper presents two cross-lingual and one English monolingual text retrieval. The
crosslanguage task includes English document retrieval in response to queries in two Indian languages:
Hindi and Bengali. Although Hindi is mostly spoken in north India and Bengali in the Eastern
India and Bangladesh only, the former is the fifth most widely spoken language in the world and
Bengali the seventh. This requires attention on CLIR involving these languages. In this paper,
we restrict ourselves to Cross-Language text retrieval applying Machine Translation approach.</p>
      <p>The rest of the paper is structured as follows. Section 2 briefly presents some of the works
on CLIR involving Indian languages. The next section provides the language specific and open
source resources used for our experiment. Section 4 builds our CLIR model on the resources and
explains our approach. CLEF evaluations of our results and their discussions are presented in the
subsequent section. We conclude this paper with a set of inferences and scope of future works.
2</p>
      <sec id="sec-2-1">
        <title>Related Work</title>
        <p>
          Cross-language retrieval is a budding field in India and the works are still in its primitive state.
The first major work involving Hindi occurred during TIDES Surprise Language exercise in a
one month period. The objective of the exercise was to retrieve Hindi documents, provided by
LDC (Linguistic Data Consortium), in response to English queries. The participants used parallel
corpora based approach to build the statistical lexicon [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ][
          <xref ref-type="bibr" rid="ref4">4</xref>
          ][
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] assigned statistical weightage
on query and expansion terms using the training corpora and this improved their cross-lingual
results over monolingual runs. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ][
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] indicated some of the language-specific obstacles for Indian
languages, viz., propritary encodings of much of the web text, lack of availability of parallel
corpora, variability in Unicode encoding etc. But all of these works were the reverse of our
problem statement for CLEF. The related work of Hindi-English retrieval can be found in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
3
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Resources used</title>
        <p>
          We used various language specific resources and open source tools for our Cross Language
Infomation Retrieval (CLIR) experiments. For the processing of English query and corpus, we used
the stop-word list (33 words) and porter stemmer of Lucene framework. For Bengali query, a
Bengali-English transliteration (ITRANS) tool2 [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], a set of Bengali stop-words3 (273 words),
2ITRANS is an encoding standared specifically for Indian languages. It converts the Indian language letters into
Roman (English) mostly using its phoneme structure.
        </p>
        <p>3The list was provided by Jadavpur University, Kolkata.
an open source Bengali-English bio-chemical lexicon ( 9k Bengali words) and a Bengali
morphological analyzer of moderate performance were used. Hindi language specific resources included
a Hindi-English Transliteration tool (wx and ITRANS), a Hindi stop-word list of 200 words, a
Hindi-English bilingual lexicon ’Shabdanjali’ containing approximately 26K Hindi words and a
Hindi Stemmer4. We also manually built a named entity list of 1510 entries mainly drawn from
the names of countries and cities, abbreviations, companies, medical terms, rivers, seven wonders,
global awards, tourist spots, diseases, events of 2002 from wiki etc. Finally, the open source Lucene
framework was used for indexing and retrieval of the documents with their corresponding scores.
4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Experimental Model</title>
        <p>The objective of Ad-Hoc Bilingual (X2EN) and Monolingual English tasks was to retrieve the
relevant documents from English target collection and submit the results in ranked order. The
topic sets for these two tasks consist of 50 topics and the participant is asked to retrieve at least
1000 documents from the corpus per query for each of the source languages. Each topic consists
of three fields: a brief ’title’, almost equivalent to a query provided by the end-user to a search
engine; a one-sentence ’description’, specifying more accurately what kind of documents the user is
looking for from the search and a ’narrative’ for relevance judgements, describing what is relevant
to the the topic and what is not. Our approach to the problem can be broken into 3 phases:
corpus processing, query generation and document retrieval.
4.1</p>
        <sec id="sec-2-3-1">
          <title>Corpus Processing</title>
          <p>The English news corpus of LA Times 2002, provided by CLEF, contained 1,35,153 documents of
433.5 MB size. After removing stop words and stemming the documents, they were indexed using
the Lucene indexer to obtain the index terms corresponding to the documents.
4.2</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>Query Generation</title>
          <p>
            We adopted Automatic Query Generation method to immitate the possible application of CLIR
on the web. The language-specific stop-words were first removed from the topics. To remove the
most frequent suffixes, we used a morphological analyzer for Bengali, a stemmer for Hindi and
the Lucene stemmer for English topics. We considered all possible stems for a single term as no
training data was available to pick the most relevant stem. This constitutes the final query for
English monolingual run. For Indian languages, the stemmed terms were then looked up in the
bilingual lexicon for their translations into English. All the translations for the term were used for
the query generation (Structured Query Translation), if the term was found in the lexicon. But
many terms did not occur in the lexicon due to its limitation in size or the improper stemming
or as the term is a named entity [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]. Those terms were first transliterated into ITRANS and
then matched against the named entity list with the help of an approximate string matching
algorithm, edit-distance algorithm. The algorithm returns the best match of the term for the
pentagram statistics. This produces the final query terms for cross-lingual runs. The queries were
constructed from the topics consisting of one or more of the topic fields.
          </p>
          <p>
            Note that we did not expand the query using the Pseudo Relevance Feedback (PRF). This
is due to the fact that it does not improve the retrieval significantly for CLIR, rather hurts by
increasing noise [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ], or increases queries in which no relevant documents are returned [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ].
4.3
          </p>
        </sec>
        <sec id="sec-2-3-3">
          <title>Document Retrieval</title>
          <p>The query generated in the above phase is fed into Lucene search engine and the documents were
retrieved along with their normalized scores. Lucene scorer follows the Vector Space Model (VSM)
of Information Retrieval.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4’Shabdanjali’ and the Hindi stemmer were built by IIIT, Hyderabad.</title>
      <sec id="sec-3-1">
        <title>CLEF Evaluation and Discussions</title>
        <p>
          The evaluation document set for the ad-hoc bilingual and monolingual tracks consists of 1,35,153
documents from LA Times 2002. For the 50 topics originally provided by CLEF, there were
manually selected 2247 relevant documents which were matched against the retrieved documents
of the participants. We provided the set of 50 queries to the system for each run of our experiments.
Six official runs were submitted for the Indian langauges to English bilingual retrieval, three for
Hindi queries and three for Bengali queries. Three monolingual English runs were also submitted
to compare the results between bilingual and monolingul retrievals. The runs were performed
using only &lt;title&gt; field, &lt;title+desc&gt; fields and &lt;title+desc+narr&gt; fields per topic for each of
these languages. The performance metrics for the nine runs of our experiments are presented in
the following tables.
&lt;title&gt;
&lt;title+desc&gt;
Table 1 presents four basic primary metrics for CLIR, viz., MAP (Mean Average Precision),
GMAP (Geometric Mean Average Precision), B-Preference and Precision at 10 retrieved
documents (P@10) for all of our official runs. The lower values of the GMAP corresponding to MAP
clearly specifies the poor performance of our retrievals in the lower end of the average precision
scale. Also, lower values of the MAP for Hindi than the work of [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] clearly suggests the need
for query expansion at the source language end. It is evident from the monolingual English and
bilingual Bengali runs that adding extra information to query through &lt;title+desc&gt; increases the
performance of the system. But adding the &lt;narr&gt; field has not improved the result significantly.
This is probably due to the fact that this field was meant for the relevance judgement in the
retrieval and we have not made any effort in preventing the retrieval of irrelevant documents in
our IR model. This, in turn, has also affected the MAP value for all the runs. However, the
improvement in the result for &lt;title+desc&gt; run over &lt;title&gt; run is not significant for Hindi. This is
probably due to the fact that using Structured Query Translation (SQT) increased too much noise
in the query to compensate the effect of a better lexicon. Also, we used morphological analyzer for
bengali rather than stemmer (for hindi) which was suggested by [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and this may have contributed
to the better result for Bengali.
        </p>
        <p>Table 2 shows the results of the topicwise score breakup for the relevant 2247 documents. As
seen from the table, number of failed topics (with no relevant retrieval) and topics with MAP
≤ 10% gradually decreased with more fields from the topic, thus establishing the fact again
mentioned in the earlier paragraph. Also, the better result for Hindi than Bengali is clearly
attributed to its better lexicon. But when it comes to the number of topics with MAP ≥ 50%,
Bengali clearly outperforms Hindi due to the n´oise factor´, mentioned in the previous paragraph.
A careful analysis of the queries revealed that the queries with named entities provided better
results for all the runs, whereas the queries without named entities performed very poor due to
poor bilingual lexicons and thus brininging down the overall performance metrics. This clearly
implies the importance of a very good bilingual lexicon and transliteration tool in the CLIR for
Indian languages. Recall is a very important performance metric for CLIR specifically for the
case when the number of relevant documents is significantly low compared to the target collection
(in this case, it is 1.66% only). It is noteworthy that the recall value has improved even for the
&lt;title+desc+narr&gt; field compared to other runs and Bengali has again outperformed Hindi due
to n´oise factor´.</p>
        <p>The recall vs average precision graphs in Figure 1 and retrieved documents vs precision graphs
in Figure 2 suggest the need for refinement of important query terms (e.g. named entity) and
weigh them more than the translated terms. Also, we used Structured Query Translation and
assigned uniform weight on all of them. But this has affected the precision values for some queries
even when the recall is significantly high. Again, we used all possible stems for a term when
multiple stems are possible and this has also added to the lower precision values for some queries.
A proper named entity recognizer is also important to prune out the named entities from other
non-lexical terms. All of these will decrease the noise in the final query and thereby help the
lucene ranking algorithm to push the relevant documents to the top.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Conclusions and Future Works</title>
        <p>
          This was our first participation in CLEF and we performed our experiment under a limited resource
scenario with a very basic Machine Translation approach. But the experiment pointed out the
necessity of good language-specific resources, specifically a rich bilingual lexicon. A close analysis
between cross-lingual and monolingual retrievals clarly pointed out the importance of four factors
in CLIR, mentiond earlier [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Apart from the above language-specific requirements, a number of
good computational approaches like query expansion by Pseudo Relevance Feedback (PRF) both
at the source and target language ends, query refinement by assigning various weightage to query
terms, proper stemming, irrelevance judgement to improve ranking, parallel corpus to build the
statistical lexicon, named entity recognizer to prune them out of non-lexical terms, Multi-word
Expression (MWE) detection, Word Sense disambiguation to avoid multiple translations (SQT)
will also increase the performance of the system. Also, pruning out the irrelevant documents from
retrieveing will increase the precision of the results. We will make attempt to experiment and
verify their effects in CLIR involving Indian languages in our future works.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Anna</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Diekema</surname>
          </string-name>
          .
          <article-title>Translation Events in Cross-Language Information Retrieval</article-title>
          .
          <source>In ACM SIGIR Forum</source>
          , Vol.
          <volume>38</volume>
          , No. 1,
          <string-name>
            <surname>June</surname>
          </string-name>
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Prasad</given-names>
            <surname>Pingali</surname>
          </string-name>
          and
          <string-name>
            <given-names>Vasudeva</given-names>
            <surname>Varma</surname>
          </string-name>
          .
          <article-title>Hindi and Telegu to English Cross Language Information Retrieval at CLEF 2006</article-title>
          .
          <article-title>In Cross Language Evaluation Forum (CLEF</article-title>
          ),
          <year>2006</year>
          , Spain.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Leah</surname>
            <given-names>S Larkey</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Margaret E</given-names>
            <surname>Connell</surname>
          </string-name>
          and
          <string-name>
            <given-names>Nasreen</given-names>
            <surname>Abduljaleel</surname>
          </string-name>
          .
          <article-title>Hindi CLIR in Thirty Days</article-title>
          .
          <source>In ACM Transactions on Asian Language Information Processing</source>
          ,
          <year>2003</year>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>130</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>jinxi</surname>
            <given-names>J</given-names>
          </string-name>
          and
          <string-name>
            <surname>Ralph Weischedel</surname>
          </string-name>
          .
          <article-title>Cross-Lingual Retrieval for Hindi</article-title>
          .
          <source>In ACM Transactions on Asian Language Information Processing</source>
          , Vol
          <volume>2</volume>
          , No.1,
          <year>March 2003</year>
          , pp.
          <fpage>164</fpage>
          -
          <lpage>168</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D</given-names>
            <surname>Hull</surname>
          </string-name>
          and
          <string-name>
            <given-names>G</given-names>
            <surname>Grefenstette</surname>
          </string-name>
          .
          <article-title>Querying across languages: A dictionary-based approach to multilingual informaion retrieval</article-title>
          .
          <source>In Proceedings of the 19th Annual International ACM SIGIR Conference on Research and Development in Information retrieval, Zurich</source>
          , Switzerland, pp.
          <fpage>49</fpage>
          -
          <lpage>57</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Victor</given-names>
            <surname>Lavrenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Choquette</surname>
          </string-name>
          and
          <string-name>
            <given-names>W. Bruce</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>Cross-Lingual Relevance Models</article-title>
          .
          <source>In SIGIR'02, August 11-15</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Douglas</surname>
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Oard</surname>
          </string-name>
          .
          <article-title>Alternative Approaches for Cross-Language Text Retrieval</article-title>
          .
          <source>In AAAI Symposium on Cross-Language Text and Speech Retrieval</source>
          , American Association for Artificial Intelligence,
          <year>March 1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Jinxi</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ralph</given-names>
            <surname>Weischedel</surname>
          </string-name>
          and
          <string-name>
            <given-names>Chanh</given-names>
            <surname>Nguyen</surname>
          </string-name>
          .
          <article-title>Evaluating a Probablistic Model for Crosslingual Information Retrieval</article-title>
          .
          <source>In SIGIR'01, September</source>
          <volume>9</volume>
          -
          <issue>12</issue>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Prasad</given-names>
            <surname>Pingali</surname>
          </string-name>
          , Jagadeesh Jagarlamudi and
          <string-name>
            <given-names>Vasudeva</given-names>
            <surname>Varma</surname>
          </string-name>
          . Webkhoj:
          <article-title>Indian language IR from Multiple Character Encodings</article-title>
          .
          <source>In Internatioanl World Wide Web Conference, May 23-26</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Lisa</given-names>
            <surname>Ballesteros</surname>
          </string-name>
          and
          <string-name>
            <given-names>W. Bruce</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>Resolving Ambiguity for Cross-language Retrieval</article-title>
          .
          <source>In SIGIR'98</source>
          ,
          <string-name>
            <surname>Melbourne</surname>
          </string-name>
          , Australia.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Avinash</given-names>
            <surname>Chopde</surname>
          </string-name>
          .
          <source>ITRANS version 5</source>
          .30. http://www.aczone.com/itrans, July,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>James</surname>
            <given-names>Allan</given-names>
          </string-name>
          , Victor Lavrenko and
          <string-name>
            <given-names>Margaret E.</given-names>
            <surname>Connell</surname>
          </string-name>
          .
          <article-title>A Month to Topic Detection and Tracking in Hindi</article-title>
          .
          <source>In ACM Transactions on Asian Language Processing</source>
          , Vol.
          <volume>2</volume>
          , No.2, June 2003, pp.
          <fpage>85</fpage>
          -
          <lpage>100</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Paul</given-names>
            <surname>Clough</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Sanderson</surname>
          </string-name>
          .
          <source>SIGIR'04, July 25-29</source>
          ,
          <year>2004</year>
          , UK.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>