<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NLP-NITMZ@DPIL-FIRE2016: Language Independent Paraphrases Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sandip Sarkar</string-name>
          <email>sandipsarkar.ju@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saurav Saha</string-name>
          <email>me@sauravsaha.in</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jereemi Bentham</string-name>
          <email>jereemibentham@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Partha Pakray</string-name>
          <email>parthapakray@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dipankar Das</string-name>
          <email>dipankar.dipnil2005@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Gelbukh</string-name>
          <email>gelbukh@gelbukh.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CIC, Instituto Politécnico Nacional</institution>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Computer Science and Engineering, Jadavpur University</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Computer Science and Engineering</institution>
          ,
          <addr-line>NIT Mizoram</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we describe the detailed information of NLP-NITMZ system on the participation of DPIL1 shared task at Forum for Information Retrieval Evaluation (FIRE 2016). The main aim of DPIL shared task is to detect paraphrases in Indian Languages. Paraphrase detection is an important part in the field of Information Retrieval, Document Summarization, Question Answering, Plagiarism Detection etc. In our approach, we used language independent feature-set to detect paraphrases in Indian languages. Features are mainly based on lexical based similarity. Our system's three features are: Jaccard Similarity, length normalized Edit Distance and Cosine Similarity. Finally, these feature-set are trained using Probabilistic Neural Network (PNN) to detect the paraphrases. With our feature-set, we achieved 88.13% average accuracy in Sub-Task 1 and 71.98% average accuracy in Sub-Task 2.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Ambiguity is one of major difficulties in Natural Language
Processing (NLP). In ambiguity, one text can be represented using
many forms like lexical and semantic. This is known as
paraphrasing. Here we consider only lexical level similarity for
paraphrase detection. Paraphrase detection is a very important and
challenging task in Information Retrieval, Question Answering,
Text Simplification, Plagiarism Detection, Text summarization
and even paraphrase detection on SMS [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In Information
Retrieval, relevant documents are retrieved using paraphrase
detection. Similarly, in Question Answering System, the best
answer is identified using paraphrase detection. Paraphrase
detection is also used in plagiarism detection to detect the
sentences which are paraphrases of each other.
Researcher used different type of approaches [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] like
Lexical Similarity, Syntactic Similarity [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and other approaches
to detect paraphrases. Research problem based on paraphrasing
1 http://nlp.amrita.edu/dpil_cen/
can be divided into three categories: Paraphrase generation,
Paraphrase extraction and Paraphrase recognition.
      </p>
      <p>
        This paper describes the NLP-NITMZ system which participated
in DPIL shared task [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. DPIL (Detecting Paraphrases in Indian
Languages) task is focused on sentence level paraphrase
identification for Indian languages (Tamil, Malayalam, Hindi and
Punjabi). DPIL shared task is divided into two sub-tasks.
In Sub-Task 1, the participants have to classify sentences into two
categories viz. Paraphrase (P) and Non-Paraphrase (NP).
      </p>
      <p>Pair of Sentences
പിുഞ്ചകുഞ്ുളങ്െ
ളകൊന് യുവതി</p>
      <p>വിഷം ളകൊ ടുുത്
ആഹത്മതയ ളെ യ ുത.
രുണ്ട മക്ളെ
ശ്നേഷം
യുവതി
വിഷം ളകൊടുുത്
ആത്മഹതയ</p>
      <p>ളകൊ
ളെ യ ുത.
மும்பை குண்டுவெடிப்பு வழக்கில்
மேலும் ஒருவர் கைது.
பிரசெல்ஸ் குண்டுவெடிப்பு முக்கிய கு
ற்றவாளி நஜீம் லாஷ்ராவி ஐ.எஸ் அ
மை ப்பில் ஜெயிலராக இருந்தார்.
ਹੁਣ ਵਿਭਾਗ ਨੂੰ ਬਣਦਾ ਕਿਰਾਇਆ ਅਦਾ ਕਰਨ ਲਈ ਕੇਸ
ਬਣਾ ਕੇ ਭੇਜ ਦਿੱਤਾ ਹੈ ਤੇ ਜਲਦ ਹੀ ਕਿਰਾਇਆ ਅਦਾ
ਕਰਦਿੱਤਾ ਜਾਵੇਗਾ।
ਹੁਣ ਵਿਭਾਗ ਨੂੰ ਬਣਦਾ ਕਿਰਾਇਆ ਅਦਾ ਕਰਨ ਲਈ ਕੇਸ
ਬਣਾ ਕੇ ਭੇਜ ਦਿੱਤਾ ਹੈ|
क्रिकेट के भगवान सचिन को जन्मदिन मुबारक हो,
दीजिए बधाई|
बधाई|
के हुए सचिन तेंदुलकर जन्मदिन मुबारक हो, दीजिए
Tag</p>
      <p>P
NP
SP
P
Similarly in Sub-Task 2, the participants have to classify
sentences into a three point scale i.e., three categories: Completely
Equivalent (E), Roughly Equivalent (RE) and Not Equivalent
(NE) i.e. (Paraphrase, Non-paraphrase, and Semi-paraphrase).
Table 1 describes the examples of DPIL training dataset.
probability of two sentences to be paraphrases is high when the
edit distance of those two sentences is small.</p>
      <p>( ,  ) =  −
In Section 2 we provide the detailed architecture of our system
like feature-set and machine-learning technique. Section 3
describes the detailed statistics of test and training data which are
used by our system. The result on test data is described in
Section 4. Section 5 describes the conclusion and future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. SYSTEM ARCHITECTURE</title>
      <p>In this section, we elaborate our proposed architecture. As shown
in Figure 1, our system NLP-NITMZ is based on three
languageindependent features: Jaccard Similarity, Levenshtein Ratio and
Cosine Similarity. To find the Jaccard Similarity, first we
calculate the number of similar unigram between two texts. After
that, similarity score is obtained by dividing the count by the total
unigram of those two sentences. Next one is Levenshtein Ratio
which calculates total number of operations required to
change one string to another form. Final feature is Cosine
Similarity where each word of sentences is represented using
Vector Space model.</p>
      <p>For machine learning portion we have used Probabilistic Neural
Network to predict the class. Probabilistic Neural Network (PNN)
is derived from Bayesian network. PNN is normally used in
classification problem and it has 4 layers. Those layers are namely
Input layer, Pattern layer, Summation layer and Output layer. The
advantage of PNN is that, that are much faster than feed forward
Neural Network.</p>
    </sec>
    <sec id="sec-3">
      <title>2.1 Features</title>
      <p>
        Our system NLP-NITMZ used three types of features which are
Language Independent. We used lexical based features which are
mainly used to find the similarity between sentences for all
Languages [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <sec id="sec-3-1">
        <title>2.1.1 Jaccard Similarity</title>
        <p>The similarity and difference of two sets is calculated using
Jaccard Similarity coefficient. For our task, Jaccard similarity
coefficient between two sentences is the ratio between the
numbers of unigram match to the total number of unique words in
those two sentences. If S1 and S2 are two sets, then the Jaccard
similarity is defined using following equation.</p>
        <p>(  ,   ) =
  ∩ 
  ∪</p>
      </sec>
      <sec id="sec-3-2">
        <title>2.1.2 Levenshtein Ratio</title>
        <p>
          The most common feature to compare two strings is the
Levenshtein Distance which is obtained by minimum number
of operations required (i.e. replacements, insertions, and
deletions) to convert one string to another [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. In our task we
assign same weight, e.g. 1 to all operations. Here we consider
character level distance between words of sentences. The
Example of Levenshtein Ratio is given in Table 3.
        </p>
        <p>Sentence 1
Sentence 2</p>
      </sec>
      <sec id="sec-3-3">
        <title>2.1.3 Cosine Similarity</title>
        <p>Cosine similarity is another widely used feature to measure the
similarity between two sentences. In this feature, each sentence is
represented using word vectors. Here word vectors are mainly the
frequency of words in the sentences. After that cosine similarity is
calculated using the dot product of those two word vectors divided
by the product of their lengths.</p>
        <p>
          ( ,  ) =
 . 
| || |
probabilistic neural network is illustrated in Figure 2. The PNN
[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] has four layers: the Input layer, Pattern layer, Summarization
layer and Output Layer. PNN have many advantages like it is
much faster than well-known back propagation algorithm and has
simple structure, PNN networks generate accurate predicted target
probability scores, PNN approach Bayes optimal classification
[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. In the same time, it is robust to noise examples.
        </p>
        <p>A simple probabilistic density function (pdf) for class k is as
follows where X = unknown (input), Xk = “Kth” sample, σ =
smoothing parameter and p = length of vectors
  ( ) =


) .  

−|| −  ||
  
The accuracy of PNN classification depends mainly on probability
density function. The probability density function for single
population is described using the following equation where n = no
of samples in the population.</p>
        <p>( ) =
(
) .      =

 
∑ 
−|| −  ||
  
If there are two classes i, j then classification criteria is decided
using the following comparison:
gi (X) &gt; gj(X) for all j ≠ i
(


The advantage of PNN networks is that the training process is
easy and quick. They can be used in real time. For our experiment
we used existing MATLAB toolkit to classify test data2.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Dataset</title>
      <p>DPIL shared task includes sentence pairs of four languages:
Tamil, Malayalam, Hindi, and Punjabi. This shared task is divided
into two sub-tasks. In Sub-Task 1, the main aim was to classify
those four sentences as paraphrases (P) or not paraphrases (NP).
Similarly Sub-Task 2 is to assign those sentences into three
categories completely equivalent (E) or roughly equivalent (RE)
or not equivalent (NE). Table 5 describes the details statistics of
training and test dataset.


2 http://in.mathworks.com/help/nnet/ref/newpnn.html</p>
      <sec id="sec-4-1">
        <title>LANGUAGE</title>
        <p>Hindi</p>
        <p>Hindi
Malayalam
Malayalam
Punjabi
Punjabi
Tamil
Tamil</p>
      </sec>
      <sec id="sec-4-2">
        <title>TASK</title>
        <p>Task 1
Task 2
Task 1
Task 2
Task 1
Task 2
Task 1
Task 2</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. CONCLUSION AND FUTURE WORK</title>
      <p>In this paper, we presented our NLP-NITMZ system used for
DPIL shared task. Overall, our approach looks promising, but
needs some improvement. There are some disadvantages of PNN
like: require large memory, slow execution. In future we want to
overcome those problems using better machine learning approach
and also want to implement semantic features for all languages to
increase performance. We can also identify stop words of all four
languages so that we can omit them from the corpus. Since our
approach is based on language independent feature set so our
methodology can be extended to various languages.</p>
    </sec>
    <sec id="sec-6">
      <title>6. ACKNOWLEDGMENTS</title>
      <p>This work presented here is under the Research Project Grant No.
YSS/2015/000988 and supported is by the Department of Science
&amp; Technology (DST) and Science and Engineering Research
Board (SERB), Govt. of India. Authors are also acknowledges the
Department of Computer Science &amp; Engineering of National
Institute of Technology Mizoram, India for providing
infrastructural facilities and support.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Wu</surname>
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ju</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            <given-names>X.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Wang</surname>
            <given-names>Y.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>Paraphrase detection on SMS messages in automobiles</article-title>
          .
          <source>In Acoustics Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2010</year>
          IEEE International Conference on (pp.
          <fpage>5326</fpage>
          -
          <lpage>5329</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Dolan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Brockett</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2005</year>
          .
          <article-title>Automatically Constructing a Corpus of Sentential Paraphrases</article-title>
          . In Third International Workshop on Paraphrasing (
          <year>IWP2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Sundaram</surname>
            ,
            <given-names>M. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madasamy</surname>
            ,
            <given-names>A. K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Padannayil</surname>
            ,
            <given-names>S. K.</given-names>
          </string-name>
          <year>2005</year>
          . AMRITA_CEN@SemEval-2015:
          <article-title>Paraphrase Detection for Twitter using Unsupervised Feature Learning with Recursive Autoencoders</article-title>
          .
          <source>In Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval</source>
          <year>2015</year>
          ),
          <fpage>45</fpage>
          -
          <lpage>50</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Mahalakshmi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Anand</given-names>
            <surname>Kumar</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Soman</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.P.</surname>
          </string-name>
          <article-title>Paraphrase detection for Tamil language using deep learning algorithm</article-title>
          ,
          <source>In (2015) International Journal of Applied Engineering Research</source>
          ,
          <volume>10</volume>
          (
          <issue>17</issue>
          ), pp.
          <fpage>13929</fpage>
          -
          <lpage>13934</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Socher</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennin</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning C. D.</surname>
            and
            <given-names>And Ng A.Y.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Dynamic pooling and unfolding recursive autoencoders for paraphrase detection</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          (pp.
          <fpage>801</fpage>
          -
          <lpage>809</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Anand Kumar</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kavirajan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Soman</surname>
            ,
            <given-names>K P</given-names>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>DPIL@FIRE2016: Overview of shared task on Detecting Paraphrases in Indian Languages</article-title>
          .
          <source>In Working notes of FIRE 2016 - Forum for Information Retrieval Evaluation</source>
          , Kolkata, India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Pakray</surname>
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Sojka</surname>
            <given-names>P.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>An Architecture for Scientific Document Retrieval Using Textual and Math Entailment Modules</article-title>
          .
          <source>In RASLAN 2014: Recent Advances in Slavonic Natural Language Processing</source>
          , Karlova Studánka,
          <source>Czech Republic, December 5-7</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Lynum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pakray</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gamback</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Jimenez</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>NTNU: Measuring semantic similarity with sublexical feature representations and soft cardinality</article-title>
          .
          <source>In Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval</source>
          <year>2014</year>
          ), pages
          <fpage>448</fpage>
          -
          <lpage>453</lpage>
          , Dublin, Ireland,
          <source>August 23- 24</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Sarkar</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Das</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pakray</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelbukh</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <year>2016</year>
          . JUNITMZ at SemEval
          <article-title>-2016 Task 1: Identifying Semantic Similarity Using Levenshtein Ratio</article-title>
          .
          <source>In Proceedings of SemEval-2016</source>
          , pages
          <fpage>702</fpage>
          -
          <lpage>705</lpage>
          , San Diego, California.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Specht</surname>
            ,
            <given-names>D.F.</given-names>
          </string-name>
          ,
          <year>1990</year>
          .
          <article-title>Probabilistic neural networks</article-title>
          .
          <source>Neural Networks</source>
          <volume>3</volume>
          ,
          <fpage>109</fpage>
          -
          <lpage>118</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Donald</surname>
            <given-names>F. S.</given-names>
          </string-name>
          <year>1990</year>
          .
          <article-title>Probabilistic Neural Networks and the Polynomial Adaline as Complementary Techniques for Classification</article-title>
          .
          <source>In IEEE Transactions on Neural Networks</source>
          ,
          <string-name>
            <surname>vol. I. No. I,</surname>
          </string-name>
          <article-title>march 1990</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Hajmeer</surname>
            ,
            <given-names>M</given-names>
          </string-name>
          and Basheer,
          <string-name>
            <surname>I.</surname>
          </string-name>
          <year>2002</year>
          .
          <article-title>A probabilistic neural network approach for modeling and classification of bacterial growth/no-growth data</article-title>
          .
          <source>In Journal of microbiological methods</source>
          , vol.
          <volume>51</volume>
          , No.
          <volume>2</volume>
          ,
          <fpage>217</fpage>
          -
          <lpage>226</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>