<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NLP-NITMZ @ MSIR 2016 System for Code-Mixed Cross-Script Question Classification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Goutam Majumder</string-name>
          <email>goutam.nita@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Partha Pakray</string-name>
          <email>parthapakray@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Institute of Technology Mizoram, Deptt. of Computer Science &amp; Engg.</institution>
          ,
          <addr-line>Mizoram</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes our approach on Code{Mixed Cross{ Script Question Classi cation task, which is a subtask 1 of MSIR 2016. MSIR is a Mixed Script Information Retrieval event in conjunction with FIRE 2016, which is the 8th meeting of Forum for Information Retrieval Evaluation. For this task, our team NLP{NITMZ submitted three system runs such as: i) using a direct feature set; ii) using direct and dependent feature set and iii) using Naive Bayes classi er. The rst system is our baseline system, which is based direct feature sets and we used a group of keywords to generate this direct feature set. To identify question classes our baseline system falls in ambiguity (means one question is tagged with multiple classes). To deal with this ambiguity, we developed another set of feature and we consider this feature set as dependent feature set, because keywords from this set is worked with direct feature set. The highest accuracy of our system is 78.88% using method{2 and we submitted as run{3. Our other two runs have same accuracy as 74.44%.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Question Answering (QA) concerned with the building
system, which can answer the questions automatically posed
by human. The QA is a common discipline within the elds
of Information Retrieval (IR) and Natural Language
Processing (NLP). It is a computer program, querying a
structured or unstructured database of knowledge or information
and constructs its answer [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        The current research of QA deals with a wide range of
question type such as fact-based, hypothetical, semantically
constrained, and cross-lingual questions. The need and
importance of a QA system was rst introduced in 1999, by
the rst QA task in TREC 8 (Text REtrieval Conference).
It was revealed the need of a sophisticated search engines,
which is able to retrieve the speci c piece of information
that could be considered as the best possible answer to a
user question. In present days such QA system works as a
backbone for successful of any E-Commerce business. In this
type of systems, many frequently asked question (FAQ) les
are generated based on most frequently asked use questions
and types of those questions [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>Being a classic application of NLP, QA has practical
applications in various domains such as education, health care,
MNY</p>
    </sec>
    <sec id="sec-2">
      <title>Airport theke Howrah Station volvo bus</title>
      <p>fare koto?
Volvo bus howrah station jete koto time nei?
Airport theke howrah station distance koto?
Airport theke textit kothai jabar bus nei?
Prepaid taxi counter naam ki?
Murshidabad kon nodir tire obosthito?
Hazarduari te koto dorja ache?
Ke Hazarduari toiri kore?</p>
      <p>
        Early morning journey hole kon service valo?
personal assistance, etc. QA is a retrieval task, which is more
challenging than the task of common search engine because
the purpose of QA is to nd out accurate and concise answer
to a question rather than just retrieving relevant documents
containing the answer [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        In this paper, the participation of subtask 1 is reported,
which is a code-mixed cross-script Question Classi cation
task [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] of MSIR 2016 (Shared Task on Mixed Script
Information Retrieval) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The rst step of understanding
a question is to perform question analysis (QA). Question
classi cation is an important task of QA, which detects the
answer type of the question. Question classi cation not only
helps to lter out a wide range of candidate answers but also
determines answer selection strategies [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>In this subtask 1, given two sets as Q = fq1; q2; :::; qng and
C = fc1; c2; :::; cng, where Q be a set of factoid questions
written in Romanized Bengali along with English and C be
the set of question classes. The task is to classify each given
questions into one of the prede ned coarse-grained classes.
Total of 9 question classes are given as classi cation task
and example of each question class with speci c tag is listed
in Table 1.</p>
      <p>Rest of the paper is organized as follows: in Section 2
we have discussed the three methods in detail. Performance
of the three systems is analysed in Section 3 and we also
compared with other submitted systems. Finally, conclusion
of the report is drawn in Section 4.
2.</p>
      <sec id="sec-2-1">
        <title>THE PROPOSED METHOD</title>
        <p>Three methods are developed for MSIR16 subtask1 to
identify the question classes. Two systems are based on
feature sets and identi cation stages for questions are
dependent on these features. Two feature sets are identi ed rst
using the training dataset and we consider one set as direct
and other set as dependent. For the third system we
combined these two sets and machine learning features are build
using Naive Bayes classi er. Details of the three methods
are discussed next.
2.1</p>
      </sec>
      <sec id="sec-2-2">
        <title>Method–1 (using direct feature set)</title>
        <p>1. MNY: To identify the 'MNY' class 10 features/
keywords from the training dataset is identi ed and
question contains these keywords are tagged as 'MNY' class.
In Table 2, we have listed out all of these features with
questions. With these 10 features we also identi ed
another keyword koto, to tag questions as 'MNY' class.
But it was analysed that, if we consider koto as direct
feature for 'MNY' tag, then questions of other classes
are also tagged as 'MNY' class. Such as 'Shankarpur
Digha theke koto dure?', which is a 'DIST' class type
question.</p>
        <p>Like koto, taka feature is also unable to tag some 'MNY'
questions. So we identi ed two other keywords as
dependent feature set, which is discussed in Section 2.2.
2. DIST: Seven keywords are identi ed as direct feature
set to tag the 'DIST' question class. We also
consider same set of features for second method. All the
identi ed keywords having the meaning as distance in
Bengali as well as in English language. For this class
no keyword is found for dependent set and in Table 3,
we have listed out all the features with questions.
3. TEMP: Question contains any temporal unit such as
somoi, time, month, year etc. are tagged as a TEMP
class. For the temporal question class eight keywords
are identi ed and all of these keywords are considered
as direct feature set and no dependent features are
con4. LOC: To tag the location class only one direct feature
as kothai is identi ed and this keyword is also used
in second method. Examples for this class are given
below:
kothai ras mela hoi ?
train r jonno kothai advice nite hobe ?
5. ORG: For organization class four direct features are
identi ed and these features with questions are listed
in Table 4. Among these four features, the ki
feature has ambiguity with other question classes such as
'OBJ' and 'PER'. Examples of questions with multiple
tags using ki feature is listed below:</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>OBJ{Ekhon ki museum hoye geche ?</title>
    </sec>
    <sec id="sec-4">
      <title>PER{Rabindranath er babar naam ki ?</title>
      <p>The kon feature of 'ORG' class also having ambiguity
issues with other classes such as 'TEMP' and 'OBJ'.
Questions with kon feature of other classes are listed
below:</p>
      <p>TEMP{Kon month e vasanta utsob hoi
shantiniketan e ?</p>
    </sec>
    <sec id="sec-5">
      <title>OBJ{Kon mountain er upor Ooty ache ?</title>
      <p>Issues related to ambiguity are addressed in Section
2.2 with the help of dependent feature set.
6. NUM: Two direct features such as koiti and koto are
identi ed to tag the questions as 'NUM' class. But koto
keyword having ambiguity and questions of 'MNY'
classes are tagged as 'NUM' . So a dependent feature
set is identi ed, which merge with koto feature and this
issues are discussed in Section 2.2. So in this method,
to identi ed questions as 'NUM' we only consider the
keyword koiti as feature.</p>
    </sec>
    <sec id="sec-6">
      <title>Bishnupur e koiti gate ache ?</title>
      <p>Leie koiti wicket niyeche ?
7. PER: To identify the 'PER' class ve direct features
are identi ed. Among these three features such as ke,
kake, and ki are worked with dependent feature sets,
which is discussed in Section 2.2 and two other
features such as kar and kader consider for direct feature
set. These direct features with example are listed
bellow:
Kar wall e sri krishna er life dekte paoya jabe ?</p>
    </sec>
    <sec id="sec-7">
      <title>Jagannath temple e kader dekha paben ? 8. OBJ: Two direct rules are found, but these rules are not able to identify the questions of 'OBJ' class. These rules are as follows:</title>
    </sec>
    <sec id="sec-8">
      <title>Ekhon ki museum hoye geche ?</title>
    </sec>
    <sec id="sec-9">
      <title>Hazarduari er opposit e kon masjid ache ?</title>
      <p>From these questions it is clearly understood that, the
'OBJ' class is in ambiguity with 'ORG' class. So these
two features are used with other dependent features for
question classi cation.
9. MISC: If no rules are satis ed then, questions are
classi ed as 'MISC' class.
2.2</p>
      <sec id="sec-9-1">
        <title>Method–2 (using direct and dependent feature set)</title>
        <p>This dependent feature set is identi ed to improve the
efciency of the rst method. The name of the second set is
given as dependent, because some features of the direct sets
are not able to identify the questions and those features work
correctly when the dependent feature set is also available in
the questions.</p>
        <p>1. MNY: To identify the MNY class two dependent
features such as charge and koto along with ten direct
features are identi ed. If any questions contains any
of these two features, it also look for the direct feature
such as taka else it will not consider the 'MNY' class
for this question and examples are listed below:
koto{Digha te Veg meal koto taka ?
charge{Semi-o cial guide koto taka charge nei ?
2. DIST: Same as method{1 using all set of direct
features only.
3. TEMP: All direct features are used to tagged the
'TEMP' class.
4. LOC: No dependent features, same set of direct
features of method{1 is used.
5. ORG: To handle the ambiguity issues with 'ORG'
class questions three sets of dependent features are
identi ed, which improves the system accuracy. In
the rst feature set, ambiguity with 'OBJ' class is
addressed. By identifying a term such as museum,
mondir, mosque which can qualify a question as 'OBJ'
class. Examples of these features are as follows:
museum{ekhon ki museum hoye geche ?
lake{murshidabad e ki lake ache ?
All such questions are not identi ed as 'ORG' class,
instead of these questions are forwarded to the other
feature set for prediction. In the second set, features are
identi ed to handle the issues related to 'PER' class
and the features are as follows:
(*eche){ke jiteche, ke hereche, ke hoyeche
team{kon team Ashes hereche ?
In questions, if tokens ends with the format eche and
questions also have 'ORG' features then those
questions are classi ed as 'PER' class. The third feature
are identi ed not to deal with ambiguity, these set is
worked with kon keyword of direct feature set used in
method{1. In this set, we explicitly identi ed those
words, which means an organization such as shop,
hotel, city, town etc. and examples are listed below:
shope{kon shop e tea kena jete pare ?
town{rat 9 PM kon town ghumiye pore ?
6. NUM: The koto keyword of direct feature set for 'NUM'
class, is also in ambiguity with other classes such as
'DIST', 'TEMP', and 'MNY'. So the direct features
are not used here to tag the questions of NUM class,
but are used to check rules, present in the questions or
not, if yes then questions will not tag or else is tagged
with 'NUM' class. Example of each ambiguity is listed
in Table 5.
7. PER: We have used two dependent features to predict
the questions of 'PER' class. In this method,
dependent features are also worked with direct features for
question prediction. Examples of questions of 'PER'
class using direct and dependent features is listed
below:</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Chilka lake jaber tour conduct ke kore ?</title>
    </sec>
    <sec id="sec-11">
      <title>Woakes kake run out koreche ?</title>
      <p>Bangladesh r leading T20 wicket-taker r naam
ki ?
8. OBJ: A dependent feature set is identi ed, which
contains all such words those quali ed as a object name to
handle the ambiguity issues of 'OBJ' class with 'ORG'
class. These object names are combined with two
direct features such as ki and kon. Example of such
ambiguity issues are listed below:
ekhon ki museum hoye geche ?
bengal r sobcheye boro mosjid ki ?</p>
      <p>Nawab Wasef Ali Mirza r residence ki chilo ?
From these dependent features, it was clear that, the
direct features look for a token in the questions those
have an entity of object type and these entities are
represented as bold face in the examples. So for kon
direct feature, same set of dependent features are used
to identify the 'OBJ' class. Examples are as follows:
Berhampore-Lalgola Road e kon mosjid ache ?</p>
    </sec>
    <sec id="sec-12">
      <title>Murshidabad kon nodir tire obosthito ?</title>
    </sec>
    <sec id="sec-13">
      <title>9. MISC same as method{1.</title>
      <p>In this method, Naive Bayes classi er is used to train the
model. For training, a feature matrix with probable class
tags is input to the Bayes classi er. For each question in
training set, one feature is considered and the last column
of the feature matrix represents the question classes. This
feature matrix is generated using the sets of direct and
dependent features used in Method{1 and 2.</p>
      <sec id="sec-13-1">
        <title>EXPERIMENT RESULTS</title>
      </sec>
      <sec id="sec-13-2">
        <title>Data and Resources</title>
        <p>
          Two datasets as training and testing data set are released
for this task [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. It was allowed that, participants can use
any number of resource for this task. Each entry in dataset
has the following format as q no q string q class and is
referred as question number, code{mixed cross{script question
string and the class of the question respectively. Training
data set contains a total no. of 330 questions and is tagged
among 9 question classes and details of the training data set
with question classes is given in Table 6.
3.2
        </p>
      </sec>
      <sec id="sec-13-3">
        <title>Results</title>
        <p>For this task, NLP{NITMZ team submitted 3 system runs.
Among these runs, two of them are feature/ keyword based
and the third run is based on machine learning features.
For the rst run, di erent set of rules are identi ed for each
question classes.
3.2.1</p>
        <p>Run–1</p>
        <p>The rst run, is conducted using method{1, which is used
all the direct rules. In this run, questions are identi ed
and tagged with the class based on these direct rules. We
achieved a success rate of 74.44% using Method{1 and the
accuracy of identi cation of question classes are listed in
Table 7 using the performance parameters.
3.2.2</p>
        <p>Run–2</p>
        <p>Naive Bayes classi er is used for this run. After being
trained the model using training dataset, model is tested
with the test data and class levels are predicted as classi er
output. For this run it has the accuracy of 74.44%, which
is same as run{1. In Table 3.2.2, the precision, recall and
F{1 score for each class labels is listed.
3.2.3</p>
        <p>Run–3</p>
        <p>For this run the direct and dependent feature set is used
together to address the ambiguity issues among the question
classes. 78.89% is the success rate achieved in this run using
Method{2 and the accuracy is listed in Table 3.2.3.
3.3</p>
      </sec>
      <sec id="sec-13-4">
        <title>Comparative Analysis</title>
        <p>In this subtask-1, total of 20 system runs is submitted
by 7 teams and as an average 140 questions are successfully
tagged by these teams with an average of 40 unsuccessful
tags. For this task, IINTU team achieved 83.33333% as
highest accuracy and our team NLP{NITMZ got the
highest accuracy as 78.88889%. Among these 9 question classes,
'DIST' class has the highest precession value as 0.9903 and
'NUM' class got the highest recall value as 0.9961 and
temporal class achieved 0.9612 as highest F{1 score.
4.</p>
      </sec>
      <sec id="sec-13-5">
        <title>CONCLUSIONS</title>
        <p>We submitted 3 system runs and accuracy of our systems
are 74.44%, 78.89%, and 74.44% respectively using three
methods. For this subtask{1 of MSIR16, our system has
given the best performance using Method{2 and we
submitted the output of this method as run{3. For this run, two
types of features are working together and we have given
direct and dependent as the name of two sets. Between these
two sets, dependent features are mainly worked for tagging
the questions without ambiguity and we got the 7th and 9th
rank with the system run 3 and 2.
5.</p>
      </sec>
      <sec id="sec-13-6">
        <title>ACKNOWLEDGMENTS</title>
        <p>This work presented here under the research project Grant
No. YSS/2015/000988 and supported by the Department of
Science &amp; Technology (DST) and Science and Engineering
Research Board (SERB), Govt. of India. Authors have also
acknowledged the Department of Computer Science &amp;
Engineering of National Institute of Technology Mizoram, India
for proving infrastructural facilities.
6.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chakma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Naskar</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Bandyopadhyay</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Choudhury</surname>
          </string-name>
          .
          <article-title>Overview of the Mixed Script Information Retrieval (MSIR) at FIRE</article-title>
          .
          <source>In Working notes of FIRE 2016 - Forum for Information Retrieval Evaluation</source>
          , Kolkata, India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          ,
          <string-name>
            <given-names>CEUR</given-names>
            <surname>Workshop</surname>
          </string-name>
          <article-title>Proceedings</article-title>
          . CEUR-WS.org,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Naskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Bandyopadhyay. The First</surname>
          </string-name>
          Cross-Script
          <string-name>
            <surname>Code-Mixed Question</surname>
          </string-name>
          Answering Corpus.
          <source>Proceedings of the workshop on Modeling, Learning and Mining for Cross/Multilinguality (MultiLingMine</source>
          <year>2016</year>
          ), co-located
          <source>with The 38th European Conference on Information Retrieval (ECIR)</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B.</given-names>
            <surname>Gamba</surname>
          </string-name>
          <article-title>ck and A. Das. Comparing the level of code-switching in corpora</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ), May
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Banchs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Choudhury</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <article-title>Query expansion for mixed-script information retrieval</article-title>
          .
          <source>In The 37th Annual ACM SIGIR Conference, SIGIR-2014</source>
          , pages
          <fpage>677</fpage>
          {
          <fpage>686</fpage>
          ,
          <string-name>
            <surname>Gold</surname>
            <given-names>Coast</given-names>
          </string-name>
          , Australia,
          <year>June 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Roth</surname>
          </string-name>
          .
          <article-title>Learning question classi ers</article-title>
          .
          <source>In Proceedings of the 19th international conference on Computational linguistics-Volume</source>
          <volume>1</volume>
          , pages
          <fpage>1</fpage>
          <lpage>{</lpage>
          7. Association for Computational Linguistics,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Pakray</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Majumder.</surname>
          </string-name>
          <article-title>Nlp{nitmz:part-of-speech tagging on italian social media text using hidden markov model. techreport, (Acceptd) In the SHARED TASK ON PoSTWITA { POS tagging for Italian Social Media Texts</article-title>
          , EVALITA -
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>