<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>BMSCE_ISE@INLI-FIRE-2017: A simple n-gram based approach for Native Language Identification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sowmya Lakshmi B S.</string-name>
          <email>sowmyalakshmibs.ise@bmsce.ac.in1</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dr. Shambhavi B R.</string-name>
          <email>shambhavibr.ise@bmsce.ac.in2</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of ISE, BMS College of Engineering</institution>
          ,
          <addr-line>Bangalore</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Native Language Identification (NLI) aims to identify native language L1 of an author by analysing the text written by him/her in other language L2. NLI is often implemented as a supervised classification problem. In this paper, we report a NLI system implemented using character tri-grams, word uni-grams and bigrams methods using linear classifier, Support Vector Machines (SVM). The work demonstrated is a participant of Indian Native Language Identification@FIRE 2017, achieving 0.27 overall accuracy for the corpus with 6 native languages. Furthermore, with subsequent evaluations, the best accuracy score obtained was 0.73 with 10 fold cross-validation on training data. We were able to achive above accuracy by incorporating uni-grams and bigrams of words.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 INTRODUCTION</title>
      <p>Classification; Feature
Recently, author profiling is gaining more importance to improve
performance of certain applications like forensics, security and
marketing. Author profiling aims to detect author’s details like
age, educational level and native language. Native Language
Identification (NLI) is a sub-class of author profiling where,
native language L1 of a writer is automatically detected by
analysing the text written in the second language L2. NLI is often
implemented as a multiclass supervised classification task.</p>
      <p>The applications of NLI are categorised into two categories:
security related applications and Second Language Acquisition
(SLA)- related applications. Security related applications are
identifying phishing sites or spam e-mails that usually consist of
strange sentences that might be written by non-native persons.
SLA applications are to analyse the effect of L1 on later learned
languages.</p>
      <p>As proved by preceding work in this area there exist quite a
few linguistic hints that helps in predicting native language. With
the impact of their native language, authors tend to make common
mistakes in spelling, punctuation and grammar while using other
languages.</p>
      <p>In this work, we examine the possibility of building native
language classifiers by ignoring grammatical errors and semantic
analysis of the text written in L2. A naive set of features using
ngrams of words and characters are explored to develop NLI
system.</p>
    </sec>
    <sec id="sec-2">
      <title>2 PREVIOUS WORK</title>
      <p>The work presented in this study was a participant of Indian
Native Language Identification@FIRE 2017 shared task. Several
researchers have investigated NLI and similar problems. An
overview of few common methods used for NLI prior to this
shared task is provided.</p>
      <p>Most of the researchers have featured NLI as a supervised
classification task, where classifiers were trained on data from
different L1. Most commonly included features for NLI are
character n-grams, POS n-grams, content words, function words
and spelling mistakes. An SVM model [1-3] was trained on these
features and obtained an accuracy of 60%-80%.</p>
      <p>In the recent past, word embedding and document embedding
has gained much attention along with other features. Continuous
Bag of Words (CBOW) and Skip Grams were used to obtain
vectors of word embedding. Vector representations for documents
were generated with distributed bag-of-words architectures using
Doc2Vec tool. In [4], authors developed a native language
classifier using document and word embedding with an accuracy
of 82% for essays and 42% on speech data.</p>
      <p>LIBSVM2, variant of SVM was verified to be efficient for text
classification. In [5], authors developed a NLI algorithm for
Arabic language with LIBSVM2. They combined production
rules, function words and POS bi-grams to perform machine
learning process and obtained an accuracy of 45%.</p>
      <p>First NLI shared task was organized with BEA workshop in
2013. System participated in closed training task was presented in
[6]. The model was trained on 11 L1 languages of TOEFL11
corpus and cross-validation testing was performed for unseen
essays resulted in accuracy of about 84.55%. Authors adopted
features like n-grams of words, characters and POS and spelling
errors with TF-IDF weighing to train SVM model.</p>
      <p>In [7], author reported the work participated in essay track of
the Second NLI Shared Task 2017 held at BEA-12 workshop. A
novel 2-stacked sentence-document architecture was introduced
by considering lexical and grammatical features of text. A stack of
two SVM classifiers were used, where first and second classifier
were sentence and document classifiers respectively. First
classifier aimed at predicting the native language of each sentence
of a document whereas, these predictions were adopted as features
by document classifier. Finally, system was used to predict native
language of unseen documents which resulted F1-score of 0.88.</p>
    </sec>
    <sec id="sec-3">
      <title>3 TASK DESCRIPTION AND DATA</title>
      <p>NLI has drawn the attention of many researchers in recent years.
With the influx of new researchers, the most substantive study in
this field has led to INLI@FIRE 2017 shared task [8]. Task
focuses on identifying native language of a writer based on his
writing in other language. In this case, the second language was
English. The task was, native language prediction of a writer from
the given Text/XML file which contains Facebook comments in
English language. Six Indian languages were proposed to consider
for this task. They were Tamil, Hindi, Kannada, Malayalam,
Bengali and Telugu.</p>
    </sec>
    <sec id="sec-4">
      <title>Dataset</title>
      <p>The training dataset for the task was xml files, which contains a
set of Facebook comments in English by different native language
speakers. Xml files were annotated as BE, HI, KA, MA, TE, and
TA for Bengali, Hindi, Kannada, Malayalam, Telugu and Tamil
language respectively. Table 1 shows the training data statistics
that was used for the task.</p>
    </sec>
    <sec id="sec-5">
      <title>4 FEATURES</title>
      <p>Files
202
211
203
207
200
210
NLI has been formulated as a multiclass classification task. We
used language-independent features such as character tri-grams
and word n-grams for NLI as described in [1]. From the previous
works we observed that character tri-grams were useful for NLI,
and they suggested that this might be due to the impact of author’s
native language. To reflect this, we calculate character n-grams
and word n-grams as features. For characters, we consider
trigrams. The features are generated over the entire training data,
i.e., every tri-gram in the training dataset is used as a feature.
Similarly, uni-grams and bi-grams of words were used as separate
features.</p>
    </sec>
    <sec id="sec-6">
      <title>5 APPROACH</title>
      <p>Training dataset provided were xml files which contained
Facebook comments in English written by different native
language speakers and files were annotated w.r.t native language
of the speaker. As a part of preprocessing, these xml files were
scraped to extract Facebook comments and comments related to
similar native language were saved in a text file. We extracted</p>
      <sec id="sec-6-1">
        <title>Sowmya Lakshmi et al. features from the text files generated and developed two methods for NLI using python as explained below.</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Character tri-grams method</title>
      <p>The tri-gram model reads text files and extracts all tri-grams
(sequence of three bytes) and their corresponding counts from the
text. Frequencies of tri-grams are pursued for every training
language separately. For every language, frequencies are
relativized by dividing individual tri-gram counts through the
number of all tri-grams in the training corpus and are sorted based
on the relative frequency (the probability of the tri-gram in the
given corpus of a language) to create language model of that
language. A language model for each language in the corpus
provided was created.</p>
      <p>Relative frequencies of the tri-grams for test dataset is
calculated and compared with the tri-grams in language models.
Intuitively, we would say that the tri-gram frequencies of
trigrams extracted from two different texts of the same native
language speaker should be very similar. The absolute difference
was calculated by subtracting the relative frequency of individual
tri-gram in the test dataset from the relative frequency of
corresponding tri-gram in each language model. The absolute
differences were summed up. For instance, if we compare test
data with 5 language models, we would have 5 different values for
the sum of absolute differences. The minimum value represents
the best match for test data. Algorithm 1 describes the algorithm
for character tri-gram approach.</p>
      <sec id="sec-7-1">
        <title>Algorithm 1: Character tri-grams method</title>
        <p>Input: Train Dataset for each language, Test Dataset
Output: Native Language Identification of Test Dataset
begin
for each language in language set
for each document in Train Dataset of language</p>
        <p>Derive all possible tri-grams
end for
Language model&lt;- Frequency of each tri-gram in Train
Dataset of language
end for
for every document in Test Dataset</p>
        <p>Derive all possible tri-grams</p>
        <p>Calculate relative frequency of each tri-gram in the document
end for
for each language in language model
for every tri-gram in the Test Dataset document</p>
        <p>Calculate absolute difference
Absolute difference &lt;- (Relative frequency of tri-gram
in Test Data) – (Relative frequency of corresponding
trigram in language model)
end for
Sum up the calculated absolute differences of each language
model
end for
Best match &lt;- Among all the computed absolute differences select
the one with minimum value.
end
A simple n-gram based approach for Native Language Identification</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Word n-grams method</title>
      <p>Frequencies of word uni-grams and bi-grams were collected for
each language irrespective of their meanings and order of words
in the document. We instantiated countvectorizer module in
python to achieve word uni-grams and bi-grams. A Document
Term Matrix X [i, j] was formed, where i is the document id, j
represents dictionary index of each word and Wij is the frequency
of occurrence of each word w in document i. Each uni-gram and
bi-gram of test data was compared with their frequencies of
occurrences in the documents of all languages.</p>
      <p>In this experiment, we used Document Term Matrix with
ngrams and applied linear SVM from scikit-learn as a classification
algorithm for NLI.
6</p>
    </sec>
    <sec id="sec-9">
      <title>RESULT ANALYSIS</title>
      <p>We submitted the output of the system for test data provided to
INLI@FIRE 2017 shared task workshop. A single run of each
method for six different languages was submitted and the results
of native language classification for all the languages are
recapitulated in Table 2 and Table 3. Character tri-gram model
achieved 22% accuracy and word n-grams model achieved an
overall accuracy of 27%.</p>
      <p>The combined features of uni-grams and bi-grams on the
training data was used to perform 10 fold cross-validation. With
these features an improved accuracy of 73% was achieved.
0.456
0.127
0.266
0.261
0.237
0.137
0.27
7</p>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSION</title>
      <p>In this paper, a supervised system for Indian Native Language
Identification has been presented. We describe character
trigrams, word uni-grams and bi-grams features, which are the
subset of frequently used features for NLI task. Results of the
supervised classification using these features on a test data set
consisting of 6 languages were reported as part of INLI@FIRE</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Nicolai</surname>
            , Garrett, Md Asadul Islam, and
            <given-names>Russ</given-names>
          </string-name>
          <string-name>
            <surname>Greiner</surname>
          </string-name>
          .
          <article-title>"Native Language Identification using probabilistic graphical models</article-title>
          .
          <source>" Electrical Information and Communication Technology (EICT)</source>
          ,
          <source>2013 International Conference on. IEEE</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Abu-Jbara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jha</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morley</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          (
          <year>2013</year>
          , June).
          <article-title>Experimental Results on the Native Language Identification Shared Task</article-title>
          .
          <source>In BEA@ NAACLHLT</source>
          (pp.
          <fpage>82</fpage>
          -
          <lpage>88</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Mizumoto</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayashibe</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sakaguchi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Komachi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Matsumoto</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          (
          <year>2013</year>
          , June).
          <article-title>NAIST at the NLI 2013 Shared Task</article-title>
          .
          <source>In BEA@ NAACL-HLT</source>
          (pp.
          <fpage>134</fpage>
          -
          <lpage>139</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Vajjala</surname>
            , Sowmya, and
            <given-names>Sagnik</given-names>
          </string-name>
          <string-name>
            <surname>Banerjee</surname>
          </string-name>
          .
          <article-title>"A study of N-gram and Embedding Representations for Native Language Identification."</article-title>
          <source>In Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications</source>
          , pp.
          <fpage>240</fpage>
          -
          <lpage>248</lpage>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Mechti</surname>
          </string-name>
          , Seifeddine, Lamia Hadrich Belguith, Ayoub Abbassi, Rim Faiz, and
          <string-name>
            <surname>Carthage IHEC</surname>
          </string-name>
          .
          <article-title>"An empirical method using features combination for Arabic native language identification."</article-title>
          <string-name>
            <surname>Gebre</surname>
            ,
            <given-names>B. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zampieri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wittenburg</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Heskes</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>” Improving native language identification with tf-idf weighting”</article-title>
          .
          <source>In the 8th NAACL Workshop on Innovative Use of NLP for Building Educational Applications (BEA8)</source>
          (pp.
          <fpage>216</fpage>
          -
          <lpage>223</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Cimino</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Dell'Orletta</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>“Stacked Sentence-Document Classifier Approach for Improving Native Language Identification”</article-title>
          .
          <source>In Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications</source>
          (pp.
          <fpage>430</fpage>
          -
          <lpage>437</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Anand Kumar</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barathi Ganesh</surname>
            <given-names>HB</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shivkaran</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soman</surname>
            <given-names>K P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paolo Rosso</surname>
          </string-name>
          .
          <article-title>Overview of the INLI PAN at FIRE-2017 Track on Indian Native Language Identification</article-title>
          .
          <source>In: Notebook Papers of FIRE</source>
          <year>2017</year>
          , FIRE-2017, Bangalore, India, December 8-10, CEUR Workshop Proceedings.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>