<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the Mixed Script Information Retrieval (MSIR) at FIRE-2016</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Somnath Banerjee</string-name>
          <email>sb.cse.ju@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kunal Chakma</string-name>
          <email>kchax4377@gmail.com</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sudip Kumar Naskar</string-name>
          <email>sudip.naskar@cse.jdvu.ac.in</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amitava Das</string-name>
          <email>amitava.das@iiits.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Rosso</string-name>
          <email>prosso@dsic.upv.es</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sivaji Bandyopadhyay</string-name>
          <email>sbandyopadhyay@cse.jdvu.ac.in</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Monojit Choudhury</string-name>
          <email>monojitc@microsoft.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IIIT Sri City</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Jadavpur University</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Microsoft Research india</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>NIT Agartala</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Technical University of Valencia</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The shared task on Mixed Script Information Retrieval (MSIR) was organized for the fourth year in FIRE-2016. The track had two subtasks. Subtask-1 was on question classi cation where questions were in code mixed Bengali-English and Bengali was written in transliterated Roman script. Subtask-2 was on ad-hoc retrieval of Hindi lm song lyrics, movie reviews and astrology documents, where both the queries and documents were in Hindi either written in Devanagari script or in Roman transliterated form. A total of 33 runs were submitted by 9 participating teams, of which 20 runs were for subtask-1 by 7 teams and 13 runs for subtask-2 by 7 teams. The overview presents a comprehensive report of the subtasks, datasets and performances of the submitted runs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        A large number of languages, including Arabic, Russian,
and most of the South and South East Asian languages like
Bengali, Hindi etc., have their own indigenous scripts.
However, the websites and the user generated content (such as
tweets and blogs) in these languages are written using
Roman script due to various socio-cultural and technological
reasons[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This process of phonetically representing the
words of a language in a non-native script is called
transliteration. English being the most popular language of the
web, transliteration, especially into the Roman script, is
used abundantly on the Web not only for documents, but
also for user queries that intend to search for these
documents. This situation, where both documents and queries
can be in more than one scripts, and the user expectation
could be to retrieve documents across scripts is referred to
as Mixed Script Information Retrieval.
      </p>
      <p>The MSIR shared task was introduced in 2013 as
\Transliterated Search" at FIRE-2013 [13]. Two pilot subtasks on
transliterated search were introduced as a part of the
FIRE2013 shared task on MSIR. Subtask-1 was on language
identi cation of the query words and subsequent back
transliteration of the Indian language words. The subtask was
conducted for three Indian languages - Hindi, Bengali and
Gujarati. Subtask-2 was on ad hoc retrieval of Bollywood
song lyrics - one of the most common forms of transliterated
search that commercial search engines have to tackle. Five
teams participated in the shared task.</p>
      <p>
        In FIRE-2014, the scope of subtask-1 was extended to
cover three more South Indian languages - Tamil, Kannada
and Malayalam. In subtask-2, (a) queries in Devanagari
script, and (b) more natural queries with splitting and
joining of words, were introduced. More than 15 teams
participated in the 2 subtasks [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Last year (FIRE-2015), the shared task was renamed from
\Transliterated Search" to \Mixed Script Information
Retrieval (MSIR)" for aligning it to the framework proposed
by [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In FIRE-2015, three subtasks were conducted [16].
Subtask-1 was extended further by including more Indic
languages, and transliterated text from all the languages were
mixed. Subtask-2 was on searching movie dialogues and
reviews along with song lyrics. Mixed script question
answering (MSQA) was introduced as subtask-3. A total of
10 teams made 24 submissions for subtask-1 and subtask-2.
In spite of a signi cant number of registrations, no run was
received in subtask-3.
      </p>
      <p>
        This year, we hosted two subtasks in the MSIR shared
task. Subtask-1 was on classifying code-mixed cross-script
question; this task was the continuation of last year's
subtask3. Here Bengali words were written in Roman
transliterated Bengali. Here Bengali words were written in Roman
transliterated Bengali. The subtask-2 was on information
retrieval of Hindi-English code-mixed tweets. The objective
of subtask-2 was to retrieve the top k tweets from a corpus
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] for a given query consisting of Hind-English terms where
the Hindi terms are written in Roman transliterated form.
      </p>
      <p>This paper provides the overview of the MSIR track in
the Eighth Forum for Information Retrieval Conference 2016
(FIRE-2016). The track was coordinated jointly by
Microsoft Research India, Jadavpur University, Technical
University of Valencia, IIIT Sriharikota and NIT Agartala.
Details of these tasks can also be found on the website https:
//msir2016.github.io/.</p>
      <p>The rest of the paper is organized as follows. Section 2
and 3 describe the datasets, present and analyze the run
submissions for the Subtask-1 and Subtask-2 respectively.
We conclude with a summary in Section 4.</p>
    </sec>
    <sec id="sec-2">
      <title>2. SUBTASK-1: CODE-MIXED CROSS</title>
    </sec>
    <sec id="sec-3">
      <title>SCRIPT QUESTION ANSWERING</title>
      <p>
        Being a classic application of natural language
processing, question answering (QA) has practical applications in
various domains such as education, health care, personal
assistance, etc. QA is a retrieval task which is more
challenging than the task of common search engines because the
purpose of QA is to nd accurate and concise answer to
a question rather than just retrieving relevant documents
containing the answer [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Recently, the code-mixed
crossscript QA research problem was formally introduced in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
The rst step of understanding a question is to perform
question analysis. Question classi cation is an important task in
question analysis which detects the answer type of the
question. Question classi cation helps not only lter out a wide
range of candidate answers but also determine answer
selection strategies [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Furthermore, it has been observed that
the performance of question classi cation has signi cant
inuence on the overall performance of a QA system.
      </p>
      <p>Let, Q = fq1; q2; : : : ; qng be a set of factoid questions
associated with domain D. Each question q : hw1w2w3 : : : wpi,
is a set of words where p denotes the total number of words
in a question. The words, w1; w2; w3; : : : ; wp, could be
English words or transliterated from Bengali in the code mixed
scenario. Let C = fc1; c2; : : : ; cmg be the set of question
classes. Here n and m refer to the total number of questions
and question classes respectively.</p>
      <p>The objective of this subtask is to classify each given
question qi 2 Q into one of the prede ned coarse-grained classes
cj 2 C. For example, the question \last volvo bus kokhon
chare?" (English gloss: \When does the last volvo bus
depart?") should be classi ed to the class `TEMPORAL'.
2.1</p>
    </sec>
    <sec id="sec-4">
      <title>Datasets</title>
      <p>
        We prepared the datasets for subtask-1 from the dataset
described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] which is the only dataset available for
codemixed cross-script question answering research. The dataset
described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] contains questions, messages and answers
from the sports and tourism domains in code-mixed
crossscript English{Bengali. The dataset contains a total of 20
documents from two domains, namely sports and tourism.
There are 10 documents in the sports domain which consist
of 116 informal posts and 192 questions, while the 10
documents in the tourism domain consist of 183 informal posts
and 314 questions. We initially provided 330 labeled factoid
questions as the development set to the participants after
accepting the data usage agreement. The testset contains 180
unlabeled factoid questions. Table 1 and Table 2 provide
statistics of the dataset. Question class speci c distribution
of the datasets is given in Figure 1.
      </p>
      <p>A total of 15 research teams registered for subtask-1.
However, only 7 teams submitted runs and a total of 20 runs
were received. All the teams submitted 3 runs except
AMRITA CEN who submitted 2 runs.</p>
      <p>
        AMRITA CEN [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] team submitted 2 runs. They used
bag-of-words (BoW) model for the Run-1. The Run-2 was
based on Recurrent Neural Network (RNN). The initial
embedding vector was given to RNN and the output of RNN
was fed to logistic regression for training. Overall, the BoW
model outperformed the RNN model by almost 7% ons
F1measure.
      </p>
      <p>
        AMRITA-CEN-NLP [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] team submitted 3 runs. They
approached the problem using Vector Space Model (VSM).
Weighted term based on the context was applied to overcome
the shortcomings of VSM. The proposed approach achieved
upto 80% accuracy in terms of F1-measure.
      </p>
      <p>ANUJ [15] also submitted 3 runs. The author used term
frequency ^aAS inverse document frequency (TF-IDF)
vector as feature. A number of machine learning algorithms,
namely Support Vector Machines (SVM), Logistic
Regression (LR), Random Forest (RF) and Gradient Boosting were
applied using Grid Search to come up with the best
parameters and model. The RF model performed the best among
the 3 runs.</p>
      <p>
        BITS PILANI [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] submitted 3 runs. Instead of
applying the classi ers on the code-mixed cross-script data, they
convert the data into English. The translation was
performed using Google translation API 1. Then they applied
three machine learning classi ers for each run, namely
Gaussian Nave Bayes, LR and RF Classi er. However, Gaussian
Naive Bayes classi er outperformed the other two classi ers.
      </p>
      <p>
        IINTU [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] was the best performing team. The team
submitted 3 runs which were based on machine learning
approaches. They trained three separate classi ers namely RF,
One-vs-Rest and k-NN, followed by building an ensemble
classi er using these 3 classi ers for the classi cation task.
The ensemble classi er took the output label by each of the
1https://translate.google.com/
individual classi ers and selected the majority label as
output. In case of tie any one label was chosen at random as
output.
      </p>
      <p>
        NLP-NITMZ [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] submitted 3 runs of which 2 runs were
rule based - a rst set of direct rules were applied for the
Run-1 while a second set of dependent rules were used for the
Run-3. A total of 39 rules were identi ed for the rule based
runs. Nave Bayes classi er was used in Run-2 whereas
Nave Bayes updateable classi er was used in Run-3.
      </p>
      <p>IIT(ISM)D used three di erent machine learning based
classi cation models - Sequential Minimal Optimization, Nave
Bayes Multimodel and Decision Tree FT to annotate the
question text. This team submitted the runs after the
deadline.
2.3</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>
        In this section, we de ne the evaluation metrics used to
evaluate the runs submitted to the subtask-1. Typically, the
performance of a question classi er is measured by
calculating the accuracy of that classi er on a particular test set
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. We also used this metric for evaluating the code-mixed
cross-script question classi cation performance.
      </p>
      <p>accuracy =
number of correctly classi ed samples</p>
      <p>total number of testset samples</p>
      <p>In addition, we also computed the standard precision,
recall and F1-measure to evaluate the class speci c
performances of the participating systems. The precision, recall
and F1-measure of a classi er on a particular class c are
de ned as follows:
number of samples correctly classi ed as c</p>
      <p>number of samples classi ed as c
precision(P ) =
recall(R) =
number of samples correctly classi ed as c</p>
      <p>total number of samples in class c
F 1
measure =
2:P:R
P + R</p>
      <p>Table 3 presents the performance of the submitted runs in
terms of accuracy. Class speci c performances are reported
in Table 4. A baseline system was also developed for the
sake of comparison using the BoW which obtained 79.444%
accuracy. It can be observed from Table 3 that the highest
accuracy (83.333%) was achieved by the IINTU team. The
classi cation performance on the temporal (TEMP) class
was very high for almost all the teams. However, Table 4 and
Figure 2 suggest that the miscellaneous (MISC) question
class was very di cult to identify. Most of the teams could
not identify the MISC class. The reason could be very low
presence(2%) of MISC class in the training dataset.</p>
    </sec>
    <sec id="sec-6">
      <title>3. SUBTASK-2:INFORMATION RETRIEV</title>
    </sec>
    <sec id="sec-7">
      <title>AL ON CODE-MIXED HINDI-ENGLISH</title>
    </sec>
    <sec id="sec-8">
      <title>TWEETS</title>
      <p>
        This subtask is based on the concepts discussed in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
In this subtask, the objective was to retrieve Code-Mixed
Hindi-English tweets from a corpus for code-mixed queries.
The Hindi components in both the tweets and the queries
are written in Roman transliterated form. This subtask
did not consider cases where both Roman and Devanagari
scripts are present. Therefore, the documents in this case are
tweets consisting of code-mixed Hindi-English texts where
the Hindi terms are in Roman transliterated form. Given a
query consisting of Hindi and English terms written in
Roman script, the system has to retrieve the top-k documents
(i.e., tweets) from a corpus that contains Code-Mixed
HindiEnglish tweets. The expected output is a ranked list of the
top twenty (k=20 here) tweets retrieved from the given
corpus.
3.1
      </p>
    </sec>
    <sec id="sec-9">
      <title>Datasets</title>
      <p>Initially we released 6,133 code-mixed Hindi-English tweets
with 23 queries as the training dataset. Later we released
a document collection containing 2,796 code-mixed tweets
along with with 12 code-mixed queries as the testset. Query
terms are mostly named entities with Roman transliterated
Hindi words. The average length of the queries in the
training set is 3.43 words and in the testset it is 3.25 words. The
tweets in the training set cover 10 topics whereas the testset
cover 3 topics.
3.2</p>
    </sec>
    <sec id="sec-10">
      <title>Submissions</title>
      <p>
        This year total 7 teams have submitted 13 runs. The
submitted runs for the retrieval task of Code-Mixed tweets
mostly adopted preprocessing of the data and then applying
di erent techniques for retrieving the desired tweets. Team
Amrita CEN [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] removed some Hindi/English stop words to
declutter useless words. After that, they have tokenized all
the tweets. The cosine distance was used to score the
relevance of tweets to the query. After that, the top 20 tweets
based on the scores were retrieved. Team CEN@Amrita[14]
used a Vector Space Model based approach. Team UB [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
has adopted three di erent techniques for the retrieval task.
First, they have used Named Entity boosts where the
purpose was to boost the documents based on their NE matches
from the query, i.e., the query was parsed to extract NEs and
each document (tweet) that matched the given NE was
provided a small numeric boost. At the second level of
boosting, phrase matching was done , i.e., documents that more
closely matched the input query phrase were ranked higher
than those that did not. The UB team used Synonym
Expansion and Narrative based weighting as the second and third
techniques. Team NITA NITMZ performed stop words
removal followed by query segmentation and nally merging.
Team IIT(ISM) D considered every tweet as a document
and indexed using uniword indexing on Terrier
implementation. Query terms were expanded using soundex coding
scheme. Terms with identical soundex code were selected
as candidate query and included in nal queries to retrieve
the relevent tweets (documents). Further, they have used
three di erent retrieval models BM25, DFR and TF-IDF to
measure the similarity. However, this team submitted the
runs after the deadline.
3.3
      </p>
    </sec>
    <sec id="sec-11">
      <title>Results</title>
      <p>The retrieval task requires that the retrieved documents
at higher ranks be more important than the retrieved
documents at lower ranks for a given query and we want our
measures to account for that. Therefore, set based
evaluation metrics such as Precision, Recall and F-measure are
not suitable for this task. Therefore, we used Mean Average
Precision (MAP) as the performance evaluation metric for
subtask-2. MAP is also referred to as \average precision at
seen relevant documents". The idea is that, rst, average
precision is computed for each query and subsequently the
average precisions are averaged over the queries. MAP is
represented as</p>
      <p>M AP =</p>
      <p>N
j=1 Qj i=1</p>
      <p>N Qj
1 X 1 X P (doci)
where Qj refers to the number of relevant documents for
query j, N indicates the number of queries and P (doci)
represents precision at the ith relevant document.</p>
      <p>The evaluation results of the submitted runs are reported
in Table 5. The highest MAP (0.0377) was achieved by team
Amrita@CEN which is still very low. The signi cantly low
MAP values in Table 5 suggest that the task of retrieving
Code-Mixed tweets against query terms comprising
codemixed Hindi and English words is a di cult task and the
techniques proposed by the teams do not produce
satisfactory results. Therefore, the problem of retrieving relevant
code-mixed tweets requires better techniques and
methodologies to be developed for improving system performance.</p>
    </sec>
    <sec id="sec-12">
      <title>4. SUMMARY</title>
      <p>In this overview, we elaborated on the two subtasks of
the MSIR-2016 at FIRE-2016. The overview is divided into
two major parts one for each subtask, where the dataset,
evaluations metric and results are discussed in detail. A
total of 33 runs were submitted from 9 unique teams.</p>
      <p>In subtask-1, 20 runs were received from 7 teams. The
best performing team achieved 83.333% accuracy. The
average question classi cation performance obtained in terms
of accuracy was 78.19% which was quite satisfactory
considering this new research problem. The subtask-1 deals
with code-mixed Bengali-English language. In the coming
years, we would like to include more Indian languages. The
participation was encouraging and we plan to continue the
subtask-1 in subsequent FIRE conferences.</p>
      <p>Subtask-2 received a total of 13 run submissions from 7
teams out of which one team submitted after the deadline.
The best MAP value achieved was 0.0377 which is
considerably low. From the results of the run submissions it can
be inferred that information retrieval of code-mixed
informal micro blog texts such as tweets is a very challenging
task. Therefore, the stated problem opens and calls for
new avenues of research for developing better techniques and
methodologies.</p>
    </sec>
    <sec id="sec-13">
      <title>ACKNOWLEDGMENTS</title>
      <p>Somnath Banerjee, Sudip Kumar Naskar and Sivaji
Bandyopadhyay acknowledge the support of the Ministry of
Electronics and Information Technology (MeitY), Government
of India, through the project \CLIA System Phase II".</p>
      <p>The work of Paolo Rosso has been partially funded by
SomEMBED MINECO TIN2015-71147-C2-1-P research project
and by the Generalitat Valenciana under the grant
ALMAMATER (PrometeoII/2014/030).</p>
      <p>We would also like to thank everybody who helped spread
awareness about this track in their respective capacities and
the entire FIRE team for giving us the opportunity and
platform for conducting this new track smoothly.
6.
Classi cation. In Working notes of FIRE 2016
Forum for Information Retrieval Evaluation, Kolkata,
India, December 7-10, 2016, CEUR Workshop</p>
      <p>Proceedings. CEUR-WS.org, 2016.
[13] R. S. Roy, M. Choudhury, P. Majumder, and</p>
      <p>K. Agarwal. Overview and datasets of FIRE 2013
track on transliterated search. In Fifth Forum for</p>
      <p>Information Retrieval Evaluation, 2013.
[14] S. Singh and Anand Kumar, M and Soman, KP.</p>
      <p>CEN@Amrita: Information Retrieval on CodeMixed
Hindi-English Tweets Using Vector Space Models. In
Working notes of FIRE 2016 - Forum for Information
Retrieval Evaluation, Kolkata, India, December 7-10,
2016, CEUR Workshop Proceedings, December 2016.
[15] A. Saini. Code Mixed Cross Script Question</p>
      <p>Classi cation. In Working notes of FIRE 2016
Forum for Information Retrieval Evaluation, Kolkata,
India, December 7-10, 2016, CEUR Workshop</p>
      <p>Proceedings. CEUR-WS.org, 2016.
[16] R. Sequiera, M. Choudhury, P. Gupta, P. Rosso,
S. Kumar, S. Banerjee, S. K. Naskar,
S. Bandyopadhyay, G. Chittaranjan, A. Das, and
K. Chakma. Overview of FIRE-2015 Shared Task on
Mixed Script Information Retrieval.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>U. Z.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Choudhury</surname>
          </string-name>
          , and
          <string-name>
            <surname>S. VB.</surname>
          </string-name>
          <article-title>Challenges in designing input method editors for indian languages: The role of word-origin and context</article-title>
          .
          <source>Advances in Text Input Methods (WTIM</source>
          <year>2011</year>
          ), pages
          <fpage>1</fpage>
          <issue>{9</issue>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M. Anand</given-names>
            <surname>Kumar</surname>
          </string-name>
          and
          <string-name>
            <given-names>K. P.</given-names>
            <surname>Soman.</surname>
          </string-name>
          Amrita-CEN@
          <article-title>MSIR-FIRE2016: Code-Mixed Question Classi cation using BoWs and RNN Embeddings</article-title>
          .
          <source>In Working notes of FIRE 2016 - Forum for Information Retrieval Evaluation</source>
          , Kolkata, India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          ,
          <string-name>
            <given-names>CEUR</given-names>
            <surname>Workshop</surname>
          </string-name>
          <article-title>Proceedings</article-title>
          . CEUR-WS.org,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Naskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Bandyopadhyay. The First</surname>
          </string-name>
          Cross-Script
          <string-name>
            <surname>Code-Mixed Question</surname>
          </string-name>
          Answering Corpus.
          <source>Proceedings of the workshop on Modeling, Learning and Mining for Cross/Multilinguality (MultiLingMine</source>
          <year>2016</year>
          ), co-located
          <source>with The 38th European Conference on Information Retrieval (ECIR)</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H. B.</given-names>
            <surname>Barathi Ganesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Anand</given-names>
            <surname>Kumar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K. P.</given-names>
            <surname>Soman</surname>
          </string-name>
          .
          <article-title>Distributional Semantic Representation for Text Classi cation and Information Retrieval</article-title>
          .
          <source>In Working notes of FIRE 2016 - Forum for Information Retrieval Evaluation</source>
          , Kolkata, India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          ,
          <string-name>
            <given-names>CEUR</given-names>
            <surname>Workshop</surname>
          </string-name>
          <article-title>Proceedings</article-title>
          . CEUR-WS.org,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Bhargava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhatia</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sharmai</surname>
          </string-name>
          .
          <article-title>Modeling Classi er for Code Mixed Cross Script Questions</article-title>
          .
          <source>In Working notes of FIRE 2016 - Forum for Information Retrieval Evaluation</source>
          , Kolkata, India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          ,
          <string-name>
            <given-names>CEUR</given-names>
            <surname>Workshop</surname>
          </string-name>
          <article-title>Proceedings</article-title>
          . CEUR-WS.org,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bhattacharjee</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          .
          <article-title>Ensemble Classi er based approach for Code-Mixed Cross-Script Question Classi cation</article-title>
          .
          <source>In Working notes of FIRE 2016 - Forum for Information Retrieval Evaluation</source>
          , Kolkata, India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          ,
          <string-name>
            <given-names>CEUR</given-names>
            <surname>Workshop</surname>
          </string-name>
          <article-title>Proceedings</article-title>
          . CEUR-WS.org,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Chakma</surname>
          </string-name>
          and
          <string-name>
            <surname>A. Das.</surname>
          </string-name>
          <article-title>CMIR: A Corpus for Evaluation of Code Mixed Information Retrieval of Hindi-English Tweets</article-title>
          .
          <source>In In the 17th International Conference on Intelligent Text Processing and Computational Linguistics (CICLING)</source>
          ,
          <year>April 2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Choudhury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. C.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Das</surname>
          </string-name>
          .
          <source>Overview of FIRE 2014 Track on Transliterated Search</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Banchs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Choudhury</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <article-title>Query expansion for mixed-script information retrieval</article-title>
          .
          <source>In Proceedings of the 37th international ACM SIGIR conference on Research &amp; development in information retrieval</source>
          , pages
          <volume>677</volume>
          {
          <fpage>686</fpage>
          . ACM,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Roth</surname>
          </string-name>
          .
          <article-title>Learning question classi ers</article-title>
          .
          <source>In Proceedings of the 19th international conference on Computational linguistics-Volume</source>
          <volume>1</volume>
          , pages
          <fpage>1</fpage>
          <lpage>{</lpage>
          7. Association for Computational Linguistics,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Londhe</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Srihari. Exploiting Named Entity Mentions Towards Code Mixed</surname>
          </string-name>
          <string-name>
            <surname>IR</surname>
          </string-name>
          :
          <article-title>Working Notes for the UB system submission for MSIR@FIRE'16.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Majumder</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Pakray.</surname>
          </string-name>
          NLP-NITMZ @
          <article-title>MSIR 2016 System for Code-Mixed Cross-Script Question</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>