<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ASE@DPIL-FIRE2016: Hindi Paraphrase Detection using Natural Language Processing Techniques &amp; Semantic Similarity Computations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vani K</string-name>
          <email>k_vani@blr.amrita.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Deepa Gupta</string-name>
          <email>g_deepa@blr.amrita.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Paraphrase Detection; Semantic Concepts; POS Tags; Weka</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science &amp;, Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Amrita University</institution>
          ,
          <addr-line>Bangalore</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Mathematics, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Amrita University</institution>
          ,
          <addr-line>Bangalore</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The paper reports the approaches utilized and results achieved for our system in the shared task (in FIRE-2016) for paraphrase identification in Indian languages (DPIL). Since Indian languages have a complex inherent nature, paraphrase identification in these languages becomes a challenging task. In the DPIL task, the challenge is to detect and identify whether a given sentence pairs paraphrased or not. In the proposed work, natural language processing with semantic concept extractions is explored for paraphrase detection in Hindi. Stop word removal, stemming and part of speech tagging are employed. Further similarity computations between the sentence pairs are done by extracting semantic concepts using WordNet lexical database. Initially, the proposed approach is evaluated over the given training sets using different machine learning classifiers. Then testing phase is used to predict the classes using the proposed features. The results are found to be promising, which shows the potency of natural language processing techniques and semantic concept extractions in detecting paraphrases.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Computing methodologies-Natural language processing</title>
    </sec>
    <sec id="sec-2">
      <title>Information systems -Document analysis and feature selection; Near-duplicate and paraphrase detection</title>
      <sec id="sec-2-1">
        <title>1. INTRODUCTION</title>
        <p>
          Paraphrasing is the process of restating the meaning of a text
using other words or adopting the idea and completely rewriting
the text information. Paraphrase detection is widely explored in
English language. The Microsoft Research Paraphrase corpus
(MSRP) is most commonly used benchmark database in English
paraphrase detections [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].Vector space models (VSM), Latent
Semantic Analysis (LSA), graph structures and semantic
similarity based paraphrase detections are explored in English
language [
          <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5 ref6 ref7">2-7</xref>
          ]. Even in English language, detection of
paraphrasing becomes more complex when the idea is adopted
and rewritten. Effective techniques incorporating syntax-semantic
techniques, deeper NLP techniques and soft computing
approaches may be required to tackle such scenarios [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. But
when it comes to Indian languages, the task becomes more
intricate. A paraphrase detection approach using deep learning for
Tamil language was proposed in [9]. Paraphrase detection in
twitter data and for SMS messages were explored in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]
respectively. A method for paraphrasing Hindi sentences by
synonym and antonym replacements and substitutions was
proposed in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].In FIRE 2016;a shared task for Detecting
Paraphrases in Indian Languages (DPIL) [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] is organized. The
tasks are defined for 4 Indian languages: Tamil, Malayalam, Hindi
and Punjabi. For each language two subtasks are defined as
follows:


        </p>
        <p>Task -1: Given a pair of sentences from news paper domain
in the specific language, the task is to classify them as
paraphrases (P) or not paraphrases (NP).</p>
        <p>Task-2: Given two sentences from news paper domainin the
specific language, the task is to identify whether they are
completely equivalent (E) or roughly equivalent (RE) or not
equivalent (NE). It is defined with three classes, viz.,
paraphrases (P), semi-paraphrase (SP) or not paraphrases
(NP) respectively.</p>
        <p>Task-1 is a binary classification problem, while Task-2 is a
multiclass problem. The proposed work is carried out for identification
of paraphrases in Hindi language. An approach that utilizes
natural language processing (NLP) techniques with semantic
similarity computations is adopted. The main focus is given to
Task-1 and the same model is applied for evaluating Task-2.
The paper is organized as follows. Section 2 describes the
proposed approach in detail. In Section 3, data statistics and
evaluation measures are discussed. Section 4 discuss and analyze
the results obtained. Section 4 concludes the paper with some
insights to future work.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2. PROPOSED APPROACH</title>
        <p>Fig.1 depicts the general work-flow of proposed approach. The
three main modules and the sub modules are described in the
following subsections.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2.1. Feature Extraction</title>
      <p>Initially feature extraction is applied for extracting the traits from
given sentence pairs. These features are given as the input to the
classifier. For extracting the feature, sentences are processed using
various pre-processing procedures with the incorporation of NLP
techniques.</p>
    </sec>
    <sec id="sec-4">
      <title>2.1.1. Pre-processing</title>
      <p>Initially the input sentence pairs are tokenized. Then part of
speech tagging is carried out.</p>
      <p>
        POS Tagging &amp;POS based Pruning: The word tokens are
tagged with their respective classes using NLTK1 POS tagger
[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The word classes include noun, verb, adjective, adverb,
preposition, conjunction etc. This is followed by POS based
pruning. In this pruning process, the tags that can possibly convey
some meaning or semantics are only retained while others are
pruned out. The retained tags include Noun, Verb, Adjective and
Adverb. The tag for cardinality which includes numbers and
indicates years, cost etc. are also retained. The remaining tags are
pruned out and not considered in further proceedings. This is
followed by stop word removal and stemming.
      </p>
      <p>Stop Word Removal: Stop words are the frequent and irrelevant
words appearing within the document. The Hindi stop word list
used in reported work is given in Table 1.Prior to stemming,
punctuation removal is done. As punctuations play an important
role in structural composition of documents, their removal can
alter the results of NLP applications. The scenario becomes more
affected with NLP techniques that operate at document level such
as parsing, chunking, semantic role labeling (SRL) etc.
Considering the dependence of NLP techniques on these
structures of a document, punctuation removal is applied after
POS tagging in our approach.</p>
      <p>Stemming: Stemming is the process of removal of affixes from
the given word. Stemming of the words is done using the suffix
list given in Table 2. An example illustration for all these
processing’s is done using a sample sentence S.</p>
      <p>S:</p>
    </sec>
    <sec id="sec-5">
      <title>After</title>
    </sec>
    <sec id="sec-6">
      <title>Tokenization:</title>
    </sec>
    <sec id="sec-7">
      <title>After</title>
    </sec>
    <sec id="sec-8">
      <title>Tagging: POS</title>
    </sec>
    <sec id="sec-9">
      <title>After POS based</title>
    </sec>
    <sec id="sec-10">
      <title>Pruning:</title>
    </sec>
    <sec id="sec-11">
      <title>After Stop</title>
    </sec>
    <sec id="sec-12">
      <title>Word Removal: After Punctuation</title>
    </sec>
    <sec id="sec-13">
      <title>2.1.3. Semantic Similarity Computations</title>
      <p>
        Once the basic pre-processing and NLP techniques based
processing is done, the semantic similarity between the processed
sentences pairs are computed. The metric used extracts the
semantic concepts in the form of synonyms of given word. Instead
of considering just surface-level word matching, synonym–level
matching is also done. This facilitates paraphrase detection, since
in many cases paraphrasing is done by replacing the words with
their synonyms. The synonyms are extracted using Word Net2
lexical database [
        <xref ref-type="bibr" rid="ref15 ref16 ref17 ref18">15-18</xref>
        ].The steps for computing the semantic
similarity is explained in following steps.
1. For all processed sentence pair,(S1, S2) Repeat steps 2 to 8.
2. Initialize Count =0.
3.
4.
5.
6.
7.
8.
      </p>
      <sec id="sec-13-1">
        <title>For each word win S1, do steps 4 to 7.</title>
        <p>If w is in S2, then Count = Count +1, else go to step 5.
Extract synonyms of the word from WordNet.</p>
        <p>For each synonym syn for word w, do step 7.</p>
        <p>If syn is in S2, then Count = Count +1 and goto step 6, else
go to step 3.</p>
        <p>Compute similarity, sim using Equation (1).
sim </p>
        <p>Count
max S1, S 2 </p>
        <p>(1)
Equation (1) computes similarity between the processed sentences
(S1, S2) as the ratio of Count value, to the maximum among the
lengths of given sentence pair. For illustration consider two
sentences S1 and S2.</p>
        <p>S1:43 कह ुएसचिनतंद िुकरज्मददनम
ुबारकहो ,द िजएबधाई|
S2:किकटकभगवानसचिनकोज्मददवसम
ुबारकहो , द िजएबधाई|
The sentences after doing tokenization, POS tagging, pruning,
stop word removal and stemming are given below. The procedure
is same as explained in subsection 2.1.1.</p>
      </sec>
    </sec>
    <sec id="sec-14">
      <title>Processed S1:</title>
    </sec>
    <sec id="sec-15">
      <title>Processed S2:</title>
      <p>[43(CD),सचिन (NNP),तंदिुकर (NNP),ज्मददन (
NNP),मुबारक (NNP),द</p>
      <p>ज(VP), बधाई(NN)]
[किकट (NN),भगवान(NNP),सचिन (NNP),ज्मदद व
स(NNP),मुबारक (NNP), द (जVP), बधाई(NN)]
In these sentences, each word in S1 in checked for its presence in
S2. If word is not present, then synonym checking is done. In the
given example, 4 exact matches are found, viz.,
सचिन ,मुबारक ,द</p>
      <p>जand बधाई. One word is identified as synonym;
viz. ज्मददवस is a synonym of ज्मददन . Thus the count value
1http://www.nltk.org/
2http://wordnet.princeton.edu/
will be, Count =5. The similarity is computed using Equation (1)
which will be:=5/(max(7,7)=0.7142.</p>
      <p>The similarity output obtained is considered as the feature input
from a sentence pair. This is the input to machine learning
classifier.</p>
    </sec>
    <sec id="sec-16">
      <title>2.2. Machine Learning Classifiers</title>
      <p>Machine learning (ML) based classifiers are used for the
paraphrase identification task. The similarity score which is the
feature input is fed to the classifier and classification task is done.
In the proposed work, different classifiers are tested and the best
among them is selected based on accuracy.</p>
    </sec>
    <sec id="sec-17">
      <title>2.3. Decision making</title>
      <p>Using the training data, initially training phase is implemented.
This is followed by testing, where decision making is done.
Decision is made on whether a given sentence pair is paraphrased
or not in Task-1. In Task-2,multi-class classification is done to
decide whether the sentence pair is paraphrased, semi-paraphrased
or non-paraphrased.</p>
      <p>Section 3 describes the data statistics used in evaluation (training
and test data) and the evaluation measures.</p>
      <sec id="sec-17-1">
        <title>3. DATA STATISICS &amp; EVALUATION</title>
      </sec>
      <sec id="sec-17-2">
        <title>MEASURES</title>
        <p>In DPIL,Task-1 provides 2500 sentence pairs for training. The
sentences are labeled as either paraphrased (P) or
NonParaphrased (NP). The set include 1000 instances for P class and
1500 instances for NP class.Task-2 provides 3500 sentence pairs
out of which 1000 are paraphrased (P), 1000 semi paraphrased
(SP) and 1500 non-paraphrased (NP). Our main focus was Task-1
while we implemented the same model for Task-2 as well. The
feature input is the semantic similarity computed, i.e., sim , using
Equation (1). Result evaluation is carried out using the
classification measures, viz., recall, precision, F-measure and %
accuracy.</p>
        <p>P
P</p>
        <p>Confusion matrix is mainly used to evaluate classification
problems. The true positives (TP), false negatives (FN), true
negatives (TN) and false positives (FP) are obtained from this
matrix. General confusion matrix for binary class problem is
shown in Equation (2). In the proposed work, TP indicates the
number of paraphrased documents correctly classified as
paraphrased. FN indicates the number of paraphrased documents
misclassified as non-paraphrased. TN is the number of
nonparaphrased documents correctly classified as non-paraphrased
and FP indicates the number of non-paraphrased documents
misclassified as paraphrased. The total population is computed
using Equation (3). Accuracy is measured using Equation (4)
which is the fraction of number of correctly classified instances to
the total number of instances in the population. Precision, Recall,
and F-measure are computed using Equation (5), (6) and (7)
respectively. Recall is defined as the number of correctly
classified documents to the actual number of correct documents to
be identified with respect to a particular class. Precision is defined
as the number of correctly classified documents to the total
number of documents identified as belonging to that class by the
system. F-measure defines the harmonic mean of precision and
recall.</p>
        <p>Receiver Operating Characteristic Curve (ROC) is also plotted for
better understanding. ROC curve plots sensitivity Vs 1-specificity.
Sensitivity is same as recall or true positive rate (TPR) while
specificity is the true negative rate (TNR) which is defined by the
fraction of documents correctly rejected to the total number of
documents to be rejected. 1-specificity is termed fall-out, which is
the false positive rate (FPR) defined as the fraction of documents
misclassified or incorrectly rejected to the total number of
documents to be rejected.ROC curves help us to understand the
discriminative power of the classifier. Using these measures, the
performance of proposed approach is evaluated over Task-1 and
Task-2.</p>
      </sec>
      <sec id="sec-17-3">
        <title>4. EXPERIMENTAL RESULTS &amp;</title>
      </sec>
      <sec id="sec-17-4">
        <title>ANALYSIS</title>
        <p>Initially the proposed approach is evaluated using different
classifiers in Weka. Weka3 is an open source machine learning
suite. The accuracy obtained using 10 fold cross-validations over
Task-1 ad Task-2 by the tested classifiers is reported in Table 3.It
is observed that decision tree exhibits the maximum accuracy in
both tasks. Thus for further evaluations decision tree is
considered. The Weka implementation of C4.5 decision tree, viz.,
J48 is used in proposed work.</p>
        <p>For better understanding, the ROC curves obtained using J48
onTask-1 and Task-2 is also plotted.Figure.2 and 3 plots the ROC
curve for class P in Task-1 and Task-2 respectively. From Figure
2 and 3, it is observed that area under ROC curve (AUC) is 0.9
and 0.799 respectively for Task-1 and 2. The values show that the
J48 classifier is able to discriminate the classes significantly in
Task-1 and it is not so low in Task-2.
3http://www.cs.waikato.ac.nz/ml/weka/</p>
        <p>NP</p>
      </sec>
    </sec>
    <sec id="sec-18">
      <title>Weighted</title>
    </sec>
    <sec id="sec-19">
      <title>Average</title>
    </sec>
    <sec id="sec-20">
      <title>Recall</title>
      <p>Compared to the training results, during the testing phase, our
results exhibited significant variation. Figure 4 plots the test
results obtained using the proposed approach. Test results
presented a considerable drop. In contrast to the 90.52% accuracy
(Task-1) on training set, test set presented only 35.88% accuracy.
Similar drop is noted in Task-2 also (Run-1 in Figure 4). On
rechecking the submission, we found that the results were
submitted wrongly. In Run-1 submission, the first sentence pair
was not written to the final output file and hence making the
second pair as first, third as second etc. and thus completely
altering our results.</p>
      <p>On request to DPIL, our results were reevaluated. The results of
Run 2 are the final results of proposed approach. It is observed
from Figure 4, that an accuracy of 89% in Task-1 and 66.6% in
Task-2 is obtained on test sets for Run-2. This matches the
training results also.</p>
      <p>The proposed approach was originally developed for plagiarism
identification and classification in English language. The results
obtained in Task-1 reflect the potency of our model to be
extended to other languages also. Task-2 can be further improved
by extracting significant features and using advanced NLP
techniques.</p>
      <sec id="sec-20-1">
        <title>5. CONCLUSIONS &amp; FUTURE WORK</title>
        <p>In the proposed approach natural language processing techniques
and semantic similarity computations are used to classify a Hindi
sentence pair as paraphrased or not. Part of speech tagging is used
for comparing only relevant tags within each sentence pair. A
semantic similarity metric is employed which extracts the word
synonyms from WordNet to check whether the compared words
are synonyms or not. This facilitates in detailed analysis and
comparisons and helps in unmasking paraphrasing imposed by
synonym replacements. The metric as a whole helps in detecting
and classifying paraphrased and non-paraphrased sentence pairs
effectively. In future, deeper natural language processing
techniques and intelligent computing techniques can be explored.
These advanced techniques are very less explored in Indian
language paraphrase detections.</p>
        <sec id="sec-20-1-1">
          <title>Natural [15] Miller, G.A.1995. WordNet: A lexical database for English,</title>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Dolan</surname>
            ,
            <given-names>W.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quirk</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Brockett</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>Unsupervised construction of large paraphrase corpora: Exploiting massively parallel news sources</article-title>
          .
          <source>In Proceedings of the 20th International Conference on Computational Linguistics</source>
          , Geneva, Switzerland.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Mihalcea</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corley</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Strapparava</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>Corpusbased and knowledge-based measures of text semantic similarity</article-title>
          ,
          <source>Proceedings of the National Conference on Artificial Intelligence (AAAI</source>
          <year>2006</year>
          ), Boston, Massachusetts,
          <fpage>775</fpage>
          -
          <lpage>780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Rus</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>McCarthy</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lintean</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>McNamara</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Graesser</surname>
            ,
            <given-names>A.C.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Paraphrase identification with lexico-syntactic graph subsumption</article-title>
          ,
          <source>FLAIRS</source>
          <year>2008</year>
          ,
          <volume>201</volume>
          -
          <fpage>206</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Fernando</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Stevenson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>A semantic similarity approach to paraphrase detection, Computational Linguistics UK (</article-title>
          <year>CLUK 2008</year>
          )
          <article-title>11th Annual Research Colloquium</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Blacoe</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lapata</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>A comparison of vectorbased representations for semantic composition</article-title>
          ,
          <source>Proceedings of EMNLP</source>
          ,
          <string-name>
            <surname>Jeju</surname>
            <given-names>Island</given-names>
          </string-name>
          , Korea,
          <fpage>546</fpage>
          -
          <lpage>556</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Islam</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Inkpen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Semantic similarity of short texts</article-title>
          ,
          <source>Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP</source>
          <year>2007</year>
          ), Borovets, Bulgaria,
          <fpage>291</fpage>
          -
          <lpage>297</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Ul-Qayyum</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Altaf</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Paraphrase Identification using Semantic Heuristic Features</article-title>
          .
          <source>Res. J. of Appld. Sci</source>
          .,
          <string-name>
            <surname>Engg</surname>
          </string-name>
          . and
          <string-name>
            <surname>Tech</surname>
          </string-name>
          .,
          <volume>4</volume>
          (
          <issue>22</issue>
          ),
          <fpage>4894</fpage>
          -
          <lpage>4904</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Vani</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Study on Extrinsic Text Plagiarism Detection Techniques and Tools</article-title>
          ,
          <source>J. of Engg. Sci. and Tech. Review.</source>
          ,
          <volume>9</volume>
          (
          <issue>4</issue>
          ),
          <fpage>150</fpage>
          -
          <lpage>164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Mahalakshmi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anand</surname>
            <given-names>Kumar</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            , and
            <surname>Soman</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.P</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Paraphrase detection for Tamil language using deep learning algorithm</article-title>
          .
          <source>Int. J. of Appld. Engg. Res.</source>
          ,
          <volume>10</volume>
          (
          <issue>17</issue>
          ),
          <fpage>13929</fpage>
          -
          <lpage>13934</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Mahalakshmi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anand</surname>
            <given-names>Kumar</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            , and
            <surname>Soman</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.P.</surname>
          </string-name>
          <year>2015</year>
          .AMRITA CEN@ SemEval-2015:
          <article-title>Paraphrase Detection for Twitter using Unsupervised Feature Learning with Recursive Autoencoders</article-title>
          , SemEval-2015,
          <fpage>45</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Wei</surname>
            <given-names>Wu.</given-names>
          </string-name>
          , Yun-Cheng Ju.,
          <string-name>
            <given-names>Xiao</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <surname>Ye-Yi Wang</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Paraphrase detection on SMS messages in automobiles</article-title>
          .
          <source>2010. In Acoustics Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2010</year>
          ,
          <fpage>5326</fpage>
          -
          <lpage>5329</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Sethi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madaan</surname>
            ,Vishu and
            <given-names>Kumar Singh S.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>A Novel Approach to Paraphrase Hindi Sentences using Natural Language Processing</article-title>
          .
          <source>Ind. J. of Sci. and Tech.</source>
          ,
          <volume>9</volume>
          (
          <issue>28</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Anand</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Kavirajan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Soman</surname>
          </string-name>
          ,KP.
          <year>2016</year>
          .
          <article-title>DPIL@FIRE2016: Overview of shared task on DetectingParaphrases in Indian Languages</article-title>
          .
          <source>In Working notes of FIRE 2016-Forum for Information Retrieval Evaluation</source>
          , Kolkata, India.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Charniak</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <year>1997</year>
          .
          <article-title>Statistical Techniques for Language Parsing</article-title>
          ,
          <source>AI</source>
          Magazine
          <volume>18</volume>
          (
          <issue>4</issue>
          ),
          <fpage>33</fpage>
          -
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          <year>1995</year>
          .
          <article-title>WordNet: A lexical database for English, Commun</article-title>
          . of the ACM,
          <volume>38</volume>
          (
          <issue>11</issue>
          ),
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Bhingardive</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shukla</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saraswati</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kashyap</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Bhattacharyya</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Synset Ranking of Hindi WordNet</article-title>
          .
          <source>In Proceedings of theLanguage Resources and Evaluation Conference</source>
          , Portorož, Slovenia.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vani</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>C.K.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Using Natural Language Processing techniques and fuzzy-semantic similarity for automatic external plagiarism detection</article-title>
          .
          <source>In Proceedings of theInternational Conference on Advances in Computing, Communication and Informatics</source>
          , Noida,
          <fpage>2694</fpage>
          -
          <lpage>2699</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Vani</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,andGupta,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Investigating the Impact of Combined Similarity Metrics and POS tagging in Extrinsic Text Plagiarism Detection System</article-title>
          .
          <source>In Proceedings of the International Conference on Advances in Computing, Communication and Informatics</source>
          , Kochi, India,
          <fpage>1578</fpage>
          -
          <lpage>1584</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>