<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of FIRE-2015 Shared Task on Mixed Script Information Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Royal Sequiera</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Monojit Choudhury Microsoft Research Lab India</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>a-rosequ</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>monojitc}@microsoft.com</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shubham Kumar IIT</string-name>
          <email>amitava.das@iiits.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patna shubh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>@gmail.com</string-name>
          <email>gokulchittaranjan@gmail.com</email>
          <email>kchax4377@gmail.com</email>
          <email>sb.cse.ju@gmail.com</email>
          <email>sudip.naskar@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Amitava Das IIIT</institution>
          ,
          <addr-line>Sri City</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Gokul Chittaranjan QuaintScience</institution>
          ,
          <addr-line>Bangalore</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Kunal Chakma NIT</institution>
          ,
          <addr-line>Agartala</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Parth Gupta, Paolo Rosso Technical University of Valencia</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Somnath Banerjee, Sudip Kumar Naskar, Sivaji Bandyopadhyay Jadavpur University</institution>
          ,
          <addr-line>Kolkata</addr-line>
        </aff>
      </contrib-group>
      <fpage>19</fpage>
      <lpage>25</lpage>
      <abstract>
        <p>The Transliterated Search track has been organized for the third year in FIRE-2015. The track had three subtasks. Subtask I was on language labeling of words in code-mixed text fragments; it was conducted for 8 Indian languages: Bangla, Gujarati, Hindi, Kannada, Malayalam, Marathi, Tamil, Telugu, mixed with English. Subtask II was on ad-hoc retrieval of Hindi film lyrics, movie reviews and astrology documents, where both the queries and documents were either in Hindi written in Devanagari or in Roman transliterated form. Subtask III was on transliterated question answering where the documents as well as questions were in Bangla script or Roman transliterated Bangla. A total of 24 runs were submitted by 10 teams, of which 14 runs were for subtask I and 10 runs for subtask II. There were no participation for Subtask III. The overview presents a comprehensive report of the subtasks, datasets, runs submitted and performances.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        A large number of languages, including Arabic, Russian, and
most of the South and South East Asian languages, are written
using indigenous scripts. However, often websites and user generated
content (such as tweets and blogs) in these languages are written
using Roman script due to various socio-cultural and technological
reasons [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This process of phonetically representing the words of
a language in a non-native script is called transliteration [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. A
lack of standard keyboards, a large number of scripts, as well as
familiarity with English and QWERTY keyboards has given rise to
a number of transliteration schemes for generating Indian language
text in Roman transliteration. Some of these are an attempt to
standardise the mapping between the Indian language script and the
Roman alphabet, e.g., ITRANS1 but mostly the users define their
own mappings that the readers can understand given their
knowledge of the language. Transliteration, especially into Roman script,
is used abundantly on the Web not only for documents, but also for
user queries that intend to search for these documents.
      </p>
      <p>A challenge that search engines face while processing
transliterated queries and documents is that of extensive spelling
varia1http://www.aczoom.com/itrans/
tion. For instance, the word dhanyavad (“thank you" in Hindi and
many other Indian languages) can be written in Roman script as
dhanyavaad, dhanyvad, danyavad, danyavaad, dhanyavada, dhanyabad,
and so on. The aim of this shared task is to systematically
formalize several research problems that one must solve to tackle this
unique situation prevalent in Web search for users of many
languages around the world, develop related data sets, test benches
and most importantly, build a research community around this
important problem that has received very little attention till date.</p>
      <p>In this shared task track, we have hosted three subtasks. Subtask
1 was on language labeling of short text fragments; this is one of
the first steps before one can tackle the general problem of mixed
script information retrieval. Subtask 2 was on ad-hoc retrieval of
Hindi film lyrics, movie reviews and astrology documents, which
are some of the most searched items in India, and thus, are
perfect and practical examples of transliterated search. We introduced
a third subtask this year on mixed script question answering. In
the first subtask, participants had to classify words in a query as
English or a transliterated form of an Indian language word.
Unlike last year though, we did not ask for the back-transliteration
of the Indic words in the native scripts. This decision was made
due to the observation that the most successful runs from previous
years had used off-the-shelf transliteration APIs (e.g. Google Indic
input tool) which beats the purpose of a research shared task. In
the second subtask, participants had to retrieve the top few
documents from a multi-script corpus for queries in Devanagari or
Roman transliterated Hindi. Last two years, this task was run only
for Hindi film lyrics. This year, movie reviews and astrology
documents, both transliterated Hindi and in Devanagari, were also added
to the document collection. The queries apart from being in
different scripts were also in mixed languages (e.g. dil chahta hai lyrics).
In the third subtask, the participants were required to provide
answers to a set of factoid questions written in Romanized Bangla.</p>
      <p>This paper provides the overview and datasets of the Mixed Script
Information Retrieval track at the seventh Forum for Information
Retrieval Conference 2015 (FIRE ’15). The track was coordinated
jointly by Microsoft Research India, Technical University of
Valencia, and Jadavpur University and was supported by BMS
College of Engineering, Bangalore. The track on mixed script IR
consists of three subtasks. Therefore, the task descriptions, results,
and analyses are divided into three parts in the rest of the paper.
Details of these tasks can also be found on the website http:
//bit.ly/1G8bTvR. We have received participation from 10
teams. A total of 24 runs were submitted in total for subtask 1 and
subtask 2.</p>
      <p>The next three sections describe the three subtasks including the
definition, creation of datasets, description of the submitted
systems, evaluation of the runs and other interesting observations. We
conclude with a summary in Sec. 5.</p>
    </sec>
    <sec id="sec-2">
      <title>SUBTASK 1: LANGUAGE LABELING</title>
      <p>
        Suppose that q :&lt; w1w2w3 : : : wn &gt;, is a query is written in
Roman script. The words, w1, w2, w3, : : :, wn, The words, w1
w2 etc., could be standard English(en) words or transliterated from
another language L={Bengali(bn), Gujarati(gu), Hindi(hi),
Kannada(kn), Malayalam(ml), Marathi(mr), Tamil(ta), Telugu(te)}.
The task is to label the words as en or L or Named Entity
depending on whether it is an English word, or a transliterated
Llanguage word [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], or a named-entity. Named Entities(NE) could
be sub-categorized as person(NE_P), location (NE_L),
organization(NE_O), abbreviation(NE_PA,NE_LA,NE_OA), inflected named
entities and other. For instance, the word USA is tagged as NE_LA
as the name entity is both a location and an abbreviation.
Sometimes, the mixing of languages can occur at the word level. In
other words, when two languages are mixed at word level, the root
of the word in one language, say Lr, is inflected with a suffix that
belongs to another language, say Ls. Such words should be tagged
as MIX. A further granular annotation of the mixed tags can be
done by identifying the languages Lr and Ls and thereby tagging
the word as M IX_Lr Ls.
      </p>
      <p>The subtask differs greatly from the previous years’ language
labeling task. While the previous years’ subtask required one to
identify the language at the word level of a text fragment given the
two languages contained in the text (in other words, the language
pair was known a priori). This year all the text fragments
containing monolingual or code-switched (multilingual) data were mixed
in the same file. Hence, an input text could belong to any of the 9
languages or a combination of any two out of the 9. This clearly
is a more challenging task from last years’, but also is more
appropriate because in real world, a search engine would not know the
languages contained in a document to begin with.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Datasets</title>
      <p>This subsection describes the datasets that have been used for the
subtasks this year.</p>
      <p>
        Table 1 shows the language-wise distribution of training and test
data for subtask 1. The training data set was composed of 2908
utterances and 51,513 tokens. The number of utterances, tokens
for each language pair in the training set is given in the . The
data for various languages of subtask 1 was collected from various
sources. In addition to the previous years’ training data, newly
annotated data from [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ] was used for Hindi-English language pair.
Similarly, for Bangla-English language pairs, data was collected by
combining previous year’s data with data from [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The data
collection and annotation project for six language pairs viz.
GujaratiEnglish, Kannada-English, Telugu-English, Marathi-English,
TamilEnglish and Malayalam-English was conducted at BMS College of
Engineering supported by a research grant by Microsoft Research
Lab India. This year, we introduced two new language pairs viz.
Marathi-English and Telugu-English, the data for which was
obtained from the aforementioned project. Marathi-English data was
contributed by ISI, Kolkata and Telugu-English data was obtained
      </p>
      <p>Utterances Tokens</p>
      <p>
        L-tags
Bangla
Gujarati
Hindi
Kannada
Malayalam
Marathi
Tamil
Telugu
Bangla
Gujarati
Hindi
Kannada
Malayalam
Marathi
Tamil
Telugu
from MSR India [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
      </p>
      <p>The labeled data from all language pairs was collated into a
single file to form a training data set. The training set was provided
in two files viz. input.txt and annotation.txt. The input.txt file
consisted of only the utterances where tokens are white space separated
and each utterance was assigned a unique id. The annotation.txt
file consisted of the annotations or labels of the tokens, exactly in
the same order as the corresponding input utterance, which can be
identified using the utterance id. Both the input and annotation files
are XML formatted.</p>
      <p>We used the unlabeled data set from the previous years’ shared
task in addition to the data that was procured from the Internet.
Similar to the training set, the test set contained utterances
belonging to different language pairs. The test set contained 792
utterances with 12,000 tokens. The number of utterances, tokens for
each language pair in the training set is given in the table 1. Unlike
the training set, only the input.txt was provided to the participants
and the participants were asked submit annotation.txt file which
was used for evaluation purposes.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Submissions</title>
      <p>A total of 10 teams made 24 submissions for subtask 1 and
subtask 2. Subtask 1 received 14 submissions from 9 unique teams.
Majority of the teams in subtask 1 made single run submissions.
Three teams viz. WISC, Watchdogs and JU_NLP submitted
multiple runs. A total of 8 runs were submitted for subtask-2 by 4
teams. Since subtask 3 is a newer subtask, it did not witness any
participation.</p>
      <p>All the submissions made by the teams for subtask 1 have used
supervised machine learning techniques with character n-grams and
character features to identify the language of the tokens.
However, WISC and ISMD teams have not used any character features
to train the classifier. TeamZine has used word normalization has
one of the features, Watchdogs converted the words into vectors
using Word2Vec techniques, clustering the vectors using k-means
algorithm and then using cluster IDs as the features. Three teams,
Watchdogs, JU and JU_NLP have gone beyond using token and
character level features, by using contextual information or a
seAmritaCEN (Amrita Vishwa Vidyapeetham)
Hrothgar (PESIT)
IDRBTIR (IDRBT)
ISMD (Indian School of Mines)
JU (Jadavpur University)
JU_NLP (Jadavpur University)
TeamZine (MNIT)
Watchdogs (DAIICT)
WISC (BITS,Pilani)
Team
quence tagger. A brief summary of all the systems is given in table
2.
2.3</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>In this section, we define the metric used to evaluate the runs
submitted to the subtasks. For subtask 1, we used the standard
precision, recall and f-measure values for evaluation. In addition, we
also used the average f-measure and weighted f-measure metrics to
compare the performance of teams. As there were some
discrepancy in the training data with respect to the X tag, two separate
versions of the aforementioned metrics were released: one
considering the X tags liberally and the other version where X tags were
considered strictly.</p>
      <p>We used the following metrics for evaluating Subtask 1. For each
category, we compute the precision, recall and F-score as shown
below.</p>
      <p>Precision (P) =
#(Correct category)
#(Generated category)</p>
      <p>Rrecall (R) =
#(Correct category)
#(Reference category)
(1)
(2)
SVM
Naive Bayes
SVM + Logistic
Regression
MaxEnt
CRF
CRF
Linear SVM +
Logistic Regression +
Random Forest
CRF
Linear Regression +
Naive Bayes
+Logistic Regression
0.683
0.692
0.680
0.615
0.538
0.610
0.423
0.618
0.576
0.623
0.525
0.387
0.387
0.387
0.538
0.807
0.899
0.371
0.278
for each language and corresponding number of tokens provided in
the training file. It can be seen that the mean score decreases as the
number of the tokens available for that language decreases.
However, the score for te has been considerably low in spite of having
a large number of tokens in the training file. This discrepancy can
be attributed to the fact that the te data provided in the training file
was not naturally generated. As some of the teams have used
additional data sets, which might have affected their performance, such
a correlation does not exist between the Max Score and number of
scores.</p>
      <p>We also infer that the most confusing language pairs from the
confusion matrices of the individual submissions. For a language
pair L1-L2, we calculate the number of times L1 is confused with
L2 and also the number of times L2 is confused with L1. We
average both the counts over an average of all the submissions. Table
6 illustrates the results obtained. It was found that the gu-hi is
the most confusing language pair. We also observe that apart from
gu-hi and ta-ml language pairs all the other Indian languages are
mostly confused with en. This may not be surprising, given the
presence of large amount of en tokens in the training set.</p>
    </sec>
    <sec id="sec-6">
      <title>SUBTASK 2: MIXED-SCRIPT AD HOC</title>
    </sec>
    <sec id="sec-7">
      <title>RETRIEVAL FOR HINDI</title>
      <p>
        This subtask uses the terminology and concepts defined in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In
this subtask, the goal was to retrieve mixed-script documents from
a corpus for a given mixed-script query. This year, the documents
and queries were written in Hindi language but using either
Roman or Devanagari script. Given a query in Roman or Devanagari
script, the system has to retrieve the top-k documents from a
corpus that contains mixed script (Roman and Devanagari). The input
is a query written in Roman (transliterated) or Devanagari script.
The output is a ranked list of ten (k = 10 here) documents both
in Devanagari and Roman scripts, retrieved from a corpus. This
year there were three different genres or documents: i) Hindi songs
lyrics, ii) movie reviews, and iii) astrology. Total 25 queries were
used to prepare the test collection for various information needs.
Queries related to lyrics documents were expressing the need to
retrieve relevant song lyric while queries related to movie reviews
and astrology were informational in nature.
3.1
      </p>
    </sec>
    <sec id="sec-8">
      <title>Datasets</title>
      <p>We first released a development (tuning) data for the IR
system – 15 queries, associated relevance judgments (qrels) and the
corpus. The queries were related to three genres: i) Hindi songs
lyrics, ii) movie reviews, and iii) astrology. The corpus consisted
of 63; 334 documents in Roman (ITRANS or plain format),
Devanagari and mixed scripts. The test set consisted of 25 queries
in either Roman or Devanagari script. On an average, there were
47:52 qrels per query with average relevant documents per query to
be 5:00 and cross-script2 relevant documents to be 3:04. The mean
query length was 4:04 words. The song lyrics documents were
created by crawling several popular domains like dhingana,
musicmaza and hindilyrix. The movie reviews data was crawled from
http://www.jagran.com/ while astrology data was crawled
from http://astrology.raftaar.in/.
3.2</p>
    </sec>
    <sec id="sec-9">
      <title>Submissions</title>
      <p>Total 5 teams submitted 12 runs. Most of the submitted runs
handled the mixed-script aspect using some type of transliteration
approach and then different matching techniques were used to
retrieve documents.</p>
      <p>BIT-M system consisted of two modules, the transliteration
module, and the searching module. The transliteration module was
trained using transliteration pairs data provided. The module was a
statistical model which used the relative frequency of letter group
mappings. The search module used the transliteration module to
treat everything in devanagari script. LCS based similarity was
used to resolve erroneous and ambiguous transliterations.
Documents were preprocessed and indexed on the content bigrams. Parts
of queries were expanded centred around high idf terms.</p>
      <p>Watchdogs used Google transliterator to transliterate every
Roman script word in the documents and queries to Devanagari word.
They submitted 4 runs with these settings: 1. Indexed the
individual words using simple analyser in lucene and then fired the query,
2. Indexed using word level 2 to 6 grams and then fired a query, 3.
Removed all the vowel signs and spaces from the documents and
queries and indexed the character level 2-6 grams of the documents,
and 4. Removed the spaces and replaced vowel signs with actual
characters in the documents and queries and indexed the character
level 2-6 grams of the documents.</p>
      <p>ISMD also submitted four runs. First two runs were using simple
indexing, with and without query expansion. Third and fourth runs
were using block indexing, with and without query expansion.</p>
      <p>The other teams did not share their approaches.
3.3</p>
    </sec>
    <sec id="sec-10">
      <title>Results</title>
      <p>The test collection for Subtask 2 contained 25 queries in Roman
and Devanagari script. The queries were of different difficulty
levels: world level joining and splitting, ambiguous short queries,
different script queries and inclusion of different language keywords.
We received total 12 runs from 5 different teams and the
performance evaluation of the runs is presented in Table 3.1. We also
present performance of systems in cross-script setting: where query
and relevant documents are strictly in different scripts. Cross-script
results are reported in Table 3.1.
4.</p>
    </sec>
    <sec id="sec-11">
      <title>SUBTASK 3: MIXED-SCRIPT QUESTION</title>
    </sec>
    <sec id="sec-12">
      <title>ANSWERING</title>
      <p>Nowadays, among many other things in social networks, people
share their travel experiences gathered during their visits to popular
2Those documents which contain duplicate content in both the
scripts are ignored.</p>
      <sec id="sec-12-1">
        <title>Team</title>
        <p>tourist spots. Often social media users seek suggestions as
guidance from their social networks before traveling, such as mode of
communication, travel fair, places to visit, accommodation, foods,
etc. Similarly, sports events are among the mostly discussed topics
in social media. People post live updates on scores, results and
fixtures of ongoing sports events such as Football leagues (e.g.
Champions League, Indian Super League, English Premier League etc.),
Cricket Series (e.g. ODI, T20, Test), Olympic games, Tennis
tournaments (e.g. Wimbledon, US Open, etc.), etc. Though question
answering (QA) is a well addressed research problem and several
QA systems are available with reasonable accuracy, there has been
hardly any research on QA on social media text mainly due to the
challenges social media content presents to NLP research. Here we
introduce a subtask on mixed-script QA considering the
information need of the bilingual speakers particularly on the tourist and
sports domains.
4.1</p>
      </sec>
    </sec>
    <sec id="sec-13">
      <title>Task</title>
      <p>Let, Q = fq1; q2; : : : ; qng, be a set of factoid questions
associated with a document corpus C in domain D and topic T ,
written in Romanized Bengali. The document corpus C consists of a
set of Romanized Bengali social media messages which could be
code-mixed with English (i.e., it also contains English words and
phrases). The task is to build a QA system which can output the
exact answer, along with the message/posts identification number
(msg_ID) and message segment (S_ans) that contains the exact
answer. An example is given in Table 9. This task deals with factoid
questions only. For this subtask,
domain D = {Sports, Tourism}</p>
      <sec id="sec-13-1">
        <title>Domain</title>
        <p>msg ID</p>
        <sec id="sec-13-1-1">
          <title>Message/Post QID</title>
        </sec>
        <sec id="sec-13-1-2">
          <title>Question</title>
        </sec>
        <sec id="sec-13-1-3">
          <title>Exact Answer S_Ans M_ans</title>
        </sec>
      </sec>
      <sec id="sec-13-2">
        <title>Tourism</title>
        <p>T818
Howrah station theke roj 3 train diyeche
ekhon Digha jabar jonno...just chill chill!!
TQ8008
Howrah station theke Digha jabar
koiti train ache?
3
Howrah station theke roj 3 train diyeche
ekhon Digha jabar jonno
T818</p>
        <p>Being the most likely potential source of code-mixed cross-script
data, we procured all the data from social media, e.g., Facebook,
Twitter, blogs, forums, etc. Initially, we released a small dataset
which was made available to all the registered participants after
signing the agreement analogous to subtask-1. The code-mixed
messages/posts related to ten popular tourist spots in India were
selected for tourism domain. For sports domain, posts/messages
related to recently held ten exciting cricket matches were selected.
We split the corpus in two sets namely, development and test sets.
The distribution of public posts/messages and questions in the
corpus for the two different domains, namely Sports and Tourism, are</p>
      </sec>
    </sec>
    <sec id="sec-14">
      <title>Submission Overview</title>
      <p>A total of 11 teams registered for subtask-3. However, no runs
were submitted by the registered participants. In this scenario, we
produced a baseline system which is presented in the following
section.
4.4</p>
    </sec>
    <sec id="sec-15">
      <title>Baseline</title>
      <p>
        A baseline system has been developed to confront the challenges
of this task. At first, a language identifier [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] was applied to
identify the languages in the dataset, i.e., Bengali and English. Then
interrogatives were identified using an interrogative list. Word level
translation was applied to English words using an in-house
English to Bengali dictionary which is prepared as part of the ELIMT
project. The detected Bengali words are transliterated using a
phrasebased statistical machine transliteration system. Then a named
entity (NE) system [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is applied only to the transliterated Bengali
words to identify the NEs. The aforementioned steps are applied
both to the messages and questions. Then separate procedures are
applied for code-mixed messages and questions. For each
question, heuristically based on the interrogative an expected NE of
answer type is assigned and a query is formed after removing the stop
words. The highest ranked message is selected by measuring the
semantic similarity to the query words. The exact answer is
extracted from the highest ranked messages using the suggested NE
type of the answer in query formulation step. In case of missing
the expected NE type, the system chooses the NOA option. This
baseline only can output exact answer with message identification
number. i.e., partial supported answer. Baseline results are reported
in Table 11.
      </p>
      <sec id="sec-15-1">
        <title>Train</title>
      </sec>
      <sec id="sec-15-2">
        <title>Test</title>
      </sec>
      <sec id="sec-15-3">
        <title>Domain</title>
        <sec id="sec-15-3-1">
          <title>Sports Tourism Sports Tourism</title>
          <p>N
Acc
0.4894
0.5148
0.4457
0.5588
ASP
0.3670
0.3861
0.3343
0.4199</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-16">
      <title>Evaluation Matrices</title>
      <p>In this task, an answer is basically structured as [Answer String
(AS), Message Segment (M S), Message ID (M Id)] triplet,
where– AS is the one of the exact answers (EA) and must be an NE
in this case,</p>
      <p>– M S is the supported text segment for the extracted answer,
and</p>
      <p>– M Id is the unique identifier of the message that justifies the
answer.</p>
      <p>While answering the questions, one has to consider the
following:</p>
      <p>i) The QA system has the provision of not answering, i.e., no
answer option (NAO).</p>
      <p>ii) The answer is the exact answer to the question.
iii) The exact answer must be a Named Entity.</p>
      <p>iv) The system has to return a single exact answer. In case there
exists more than one correct answer to a question, the system needs
to provide only one of the correct answers.</p>
      <p>While evaluating, the primary focus remains on
“responsiveness” and “usefulness” of each answer. Each answer is manually
judged by native speaking assessors. Each [AS, M S, M Id] triplet
is assessed in a five-valued scale (Table 12) and marked with
exactly one of the following judgments:</p>
      <p>Incorrect: The AS does not contain EA (i.e., not responsive)
Unsupported: The AS contains correct EA, but MS and
MIid do not support the EA (i.e., missing usefulness)
Partial-supported: The AS contains the correct EA with
correct Mid, but MS does not support EA
Correct: The AS provides the correct EA with correctly
supporting MS and MIid (i.e., “responsive” as well as “useful”).
Inexact: The supporting MS and MIid are correct, but the
AS is wrong.</p>
      <sec id="sec-16-1">
        <title>Judgment</title>
        <p>Incorrect (W)
Inexact(I)
Unsupported (U)
Partial-supported (P)
Correct (C)</p>
        <p>AS
X
X
X
X
X</p>
        <p>MS
X
X
X
X
X</p>
        <p>MId</p>
        <p>X
X
X
X
X</p>
        <p>In order to maintain consistency with previous QA shared tasks,
we have chosen accuracy and c@1 as evaluation metrics. MRR has
not been considered since a QA system is supposed to return an
exact answer, i.e. not list. Just as in the past ResPubliQA3
campaigns, systems are allowed to have the option of withholding the
answer to a question because they are not sufficiently confident that
it is correct (i.e., NAO). As per ResPubliQA, inclusion of NAO
improves the system performance by reducing the number of incorrect
answers.</p>
        <p>Now, C@1 = N1 (Nr + Nu: NNur )
Accuracy = Nr</p>
        <p>N
C@1 = Accuracy; if Nu = 0
Where, Nr = number of right answers.</p>
        <p>Nu = number of unanswered questions
N = total questions</p>
        <p>Correct, Partially-supported and Unsupported answers provide
the exact answers only.</p>
        <p>Therefore, Nr = (#C + #U + #P )</p>
        <p>Considering the importance of supporting segment, we introduce
a new metric “Answer-Support performance” (ASP) which
measures the answer correctness.</p>
        <p>ASP = N1 (c 1:0 + p 0:75 + i 0:25)
where, c, p and i denote total number of correct, partially-supported
and inexact answers.
3http://nlp.uned.es/clef-qa/repository/resPubliQA.php</p>
      </sec>
    </sec>
    <sec id="sec-17">
      <title>Discussion</title>
      <p>In spite of a significant number of registrations in subtask-3, no
run was received. Personal communication with registered
participants revealed that the time provided for this subtask was not
sufficient to develop the required mixed-script QA system. Next year
we could simplify the task load by asking participants to solve
various subtasks of the said QA system, such as question
classification, question focus identification, etc. Introducing more
IndianEnglish language pairs could encourage this subtask across other
Indian languages speakers.</p>
    </sec>
    <sec id="sec-18">
      <title>SUMMARY</title>
      <p>In this overview, we elaborated on the various subtasks of the
Mixed Script Information Retrieval track at the seventh Forum for
Information Retrieval Conference (FIRE’15). The overview is
divided into three major parts one for each subtask, where the dataset,
evaluations metric and results are discussed in detail.</p>
      <p>There were a total of 14 submissions from a total of 9 teams.
In subtask 1, we noted that gu-hi was the most confused language
pair. It was also found that the performance of the system for a
category is positively correlated to the number of tokens for that
category. Subtask 2 received 12 submissions from 5 teams. The
subtask consisted of three different genres viz. Hindi songs lyrics,
movie reviews and astrology documents. A third subtask on mixed
script question answering was introduced this year. However, there
were no participants for subtask 3.</p>
    </sec>
    <sec id="sec-19">
      <title>Acknowledgments</title>
      <p>We would like to thank Prof Shambhavi Pradeep, BMS College of
Engineering for data creation for subtask 1 in 6 Indian languages
and Dnyaneshwar Patil, ISI Kolkata for contributing Marathi data.
We are also grateful to Shruti Rijhwani and Kalika Bali, Microsoft
Research Lab India, for helping with reviewing the working notes
of this shared task.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Ahmed</surname>
            ,
            <given-names>U.Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bali</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choudhury</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.V.</surname>
          </string-name>
          :
          <article-title>Challenges in designing input method editors for indian languages: The role of word-origin and context</article-title>
          .
          <source>Advances in Text Input Methods (WTIM</source>
          <year>2011</year>
          )
          <article-title>(</article-title>
          <year>2011</year>
          )
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Knight</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Graehl</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Machine transliteration</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>24</volume>
          (
          <issue>4</issue>
          ) (
          <year>1998</year>
          )
          <fpage>599</fpage>
          -
          <lpage>612</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Antony</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soman</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Machine transliteration for indian languages: A literature survey</article-title>
          .
          <source>International Journal of Scientific &amp; Engineering Research</source>
          , IJSER
          <volume>2</volume>
          (
          <year>2011</year>
          )
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>King</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abney</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Labeling the languages of words in mixed-language documents using weakly supervised methods</article-title>
          .
          <source>In: Proceedings of NAACL-HLT</source>
          . (
          <year>2013</year>
          )
          <fpage>1110</fpage>
          -
          <lpage>1119</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Vyas</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gella</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bali</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choudhury</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Pos tagging of english-hindi code-mixed social media content</article-title>
          . In: EMNLP'
          <fpage>14</fpage>
          . (
          <year>2014</year>
          )
          <fpage>974</fpage>
          -
          <lpage>979</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Barman</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wagner</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Foster</surname>
          </string-name>
          , J.:
          <article-title>Code mixing: A challenge for language identification in the language of social media</article-title>
          . (
          <year>2014</year>
          )
          <fpage>13</fpage>
          -
          <lpage>23</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Sowmya</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choudhury</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bali</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dasgupta</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Resource creation for training and testing of transliteration systems for indian languages</article-title>
          .
          <source>In: LREC</source>
          . (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bali</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Banchs</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choudhury</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Query expansion for mixed-script information retrieval</article-title>
          .
          <source>In: The 37th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , SIGIR '14,
          <string-name>
            <surname>Gold</surname>
            <given-names>Coast</given-names>
          </string-name>
          ,
          <string-name>
            <surname>QLD</surname>
          </string-name>
          ,
          <source>Australia - July 06 - 11</source>
          ,
          <year>2014</year>
          . (
          <year>2014</year>
          )
          <fpage>677</fpage>
          -
          <lpage>686</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kuila</surname>
          </string-name>
          , Roy,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.N.P.</given-names>
            ,
            <surname>Bandyopadhyay</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.:</surname>
          </string-name>
          <article-title>A hybrid approach for transliterated word-level language identification: Crf with post-processing heuristics</article-title>
          .
          <source>In: FIRE</source>
          , ACM Digital Publishing (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naskar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Bengali named entity recognition using margin infused relaxed algorithm</article-title>
          .
          <source>In: TSD</source>
          , Springer International Publishing (
          <year>2014</year>
          )
          <fpage>125</fpage>
          -
          <lpage>132</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>