<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Text Alignment Corpus for Persian Plagiarism Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fatemeh Mashhadirajab</string-name>
          <email>f.mashhadirajab@mail.sbu.ac.ir</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mehrnoush Shamsfard</string-name>
          <email>m-shams@sbu.ac.ir</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Razieh Adelkhah</string-name>
          <email>r.adelkhah@yahoo.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fatemeh Shafiee</string-name>
          <email>f.shafiee@hotmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chakaveh Saedi</string-name>
          <email>Ch_saedi@sbu.ac.ir</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>NLP Research Lab, Faculty of Computer Science and Engineering, Shahid Beheshti University</institution>
          ,
          <country country="IR">Iran</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>NLP Research Lab, Faculty of Computer Science and, Engineering, Shahid Beheshti University</institution>
          ,
          <country country="IR">Iran</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>NLX Lab of university of Lisbon, Department of Informatics</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2009</year>
      </pub-date>
      <abstract>
        <p>This paper describes how a Persian text alignment corpus was constructed to evaluate plagiarism detection systems. This corpus is in PAN format and contains 11,089 documents and more than 11,603 plagiarism cases. Efforts were made to simulate various types of plagiarism manually, semi-automatically, or automatically in this large-scale corpus.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Plagiarism detection</kwd>
        <kwd>Text alignment corpus</kwd>
        <kwd>Types of plagiarism</kwd>
        <kwd>Corpus construction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Near-duplicate and plagiarism
corpus [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] includes 1000 documents in English and 250
plagiarism cases. In this corpus, in order to obfuscate texts, a
number of students of different academic courses were asked to
select and rewrite a number of texts related to their fields and put
them inside documents with the same subject such as Wikipedia
documents. Also A bilingual English-Urdu corpus that includes
1000 documents and 270 plagiarism cases sent to the PAN 2015
competitions by Hanif [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. In this corpus he used machine
translation with and without manual correction of results, with the
use of random-obfuscation strategy in some translation results to
obfuscate the text. Khoshnavataher [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] has presented a corpus in
Persian that includes 2111 documents and 823 plagiarism cases.
In order to obfuscate, he used Random obfuscation technique and
no-obfuscation technique where a piece of the source document is
added to suspicious document without any change. Kong [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] also
took part in the PAN 2015competition with 160 documents in
Chinese and 152 cases of plagiarism. In order to obfuscate text,
Kong asked a number of volunteers to write a paper for topics that
have been identified. Mohtaj’s corpus [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] also was submitted to
PAN 2015 with 4261 documents in English and 2781 plagiarism
cases. In this corpus, techniques of no-obfuscation,
randomobfuscation and simulated-obfuscation is used to obfuscate text.
Palkovskii [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] also makes use of PAN 2013-2014 corpus to
prepare a corpus that included 5057 documents in English and
4185 plagiarism cases. Obfuscation was made based on
techniques of random-obfuscation, no-obfuscation,
cyclictranslation-obfuscation and summary-obfuscation. In the rest of
this paper we will describe the construction method we employed
to develop a text alignment corpus to evaluate Persian plagiarism
detection systems.
      </p>
    </sec>
    <sec id="sec-2">
      <title>3. TEXT ALIGNMENT CORPUS</title>
    </sec>
    <sec id="sec-3">
      <title>CONSTRUCTION</title>
      <p>
        The goal in text alignment is to identify plagiarized segments for
each given source and suspicious document pairs [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        In this study, a text alignment corpus is created to evaluate
plagiarism detection systems on Persian scientific documents. The
conducted procedure to build this corpus is described herein.
a. Data Source Preparation
We use some documents of source documents collection in
Mahtab plagiarism detection system [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] to construct our text
alignment corpus. Mahtab plagiarism detector is developed at the
Shahid Beheshti University NLP Lab. The goal of Mahtab is
detecting plagiarized articles in the fields of computer science and
engineering. Our text alignment corpus in this study contains
11,089 documents. They are all articles or theses in the fields of
computer science and engineering and also electrical engineering
with the following distribution:
4,500 documents from Wikipedia articles;
      </p>
      <sec id="sec-3-1">
        <title>1,500 documents from CSICC1 articles (2004-2015);</title>
        <p>
          1,500 documents from articles and theses available from
online stores;
3,589 documents from free Persian resources including
magiran2, iran-doc3, SID4, prozhe5, and MatlabSite6.
b. Documents Clustering
Since all documents in the corpus are in the field of computer
science, there is a general similarity among them. The method
proposed for document clustering is to estimate cluster features
first, and then perform clustering based on the introduced features.
Finally, an optimization process improves the results. To extract
features, all words included in a document are extracted and
stemmed using STeP-1 [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Each word is then labeled based on
Table 1 which is introduced by Makrehchi [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. For each
document, an n-bit histogram vector is produced named V( , ,
…, ) where n is the number of features. If existed in a
document, = 1; otherwise, = 0. Afterwards, these vectors are
classified based on the K-means algorithm and Cosine similarity.
To optimize the extracted features in a cluster, the sum of all
vectors of a cluster is found and used to produce H ( , , …,
)), where h1 indicates the number of documents containing the
first feature. H is produced for all clusters; Equation (1) can be
used to calculate the weight of each feature in the corresponding
cluster.
fc indicates the number of clusters containing this feature. The
features are sorted in a descending order based on their weights.
Afterwards, the first 100 words of each cluster are considered as
the features for that cluster. To improve clustering accuracy, the
1 Computer Society of Iran Computer Conference, http://csi.org.ir
2 http://mag-iran.com
3 http://www.irandoc.ac.ir
4 http://sid.ir
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>5 http://www.prozhe.com</title>
      </sec>
      <sec id="sec-3-3">
        <title>6 www.MatlabSite.com</title>
        <p>membership degree to each cluster must be calculated, and
documents must be placed in the most similar cluster. The
membership degree for each document is calculated as follows:
√
In which is the number of all seen cluster features (the first
100 words of each cluster based on their weights are considered as
cluster features) in the corresponding document, is the
number of cluster features occurring in the document, and is
the document length.</p>
        <p>
          n
i
m
r
c. Suspicious Documents Selection
Some documents are randomly selected from each cluster as
suspicious documents. Almost half of the documents are
employed as source documents and the other half as suspicious
documents. Half of the suspicious documents are considered as
no-plagiarism documents, and the other half of the documents are
used to produce plagiarized documents.
d. Source Set for a plagiarism Document
For each plagiarism document in a cluster, a set of source
documents named Dsrc is selected in which there is no repeated
document or very similar document to suspicious document, a
source document can be used in many suspicious documents so
every time a suspicious document can select each source
document from the corresponding cluster therefore the selected
documents may be selected by this suspicious document before.
Moreover if the similarity between source document and
suspicious document is more than 50% before adding plagiarism
passages to suspicious document, then it is not a good selection
because even if a hard strategy is used to obfuscate, plagiarism
passages may be discovered by simple similarity detection
algorithms. To create Dsrc for each plagiarism document, a
document from the corresponding cluster is selected randomly; if
the similarity based on the SimHash method [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] between the
selected document and each document in Dsrc is more than 50%,
the document is considered repeated; otherwise, it is included in
Dsrc. This step is continued until there are at least 3 documents in
Dsrc. A Dsrc contains a suspicious document and at least 2 source
documents. The reason for employing the SimHash method is the
noticeable results achieved in [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. In this phase, the source and
suspicious document pairs are specified. In this way, 3,867 paired
documents (source- suspicious) are produced to be included in the
corpus.
e. Source Set for a No-plagiarism Document
For each no-plagiarism document, a source set is selected as
described in step d. However, in this step a similarity detection at
the sentence level for each randomly selected source and
suspicious document is considered based on the Jaccard similarity
measure and a threshold of 0.9; if there are no same sentences
between both mentioned files, the source document is added to
Dsrc. Using this method, 2,630 pairs of documents are produced in
this phase.
f. Source Documents Segmentation
In this step, first a document is divided into its paragraphs. Each
subsequence of paragraphs that contain at least 300 words is
considered a segment. If a paragraph contain less than 300 words
it is combined with the next paragraph. Ultimately, all segments
contain at least 300 words.
g. Determine the Length of Plagiarized
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Segments in each Suspicious Document</title>
      <p>The number of plagiarized segments which are employed in a
suspicious document depends on the source document length and
the length of any plagiarized segments. To decide the number of
segments to use from a source document in its paired suspicious
one, all paired documents are first labeled. Each randomly
selected pair of documents is labeled as “entirely,” “much,”
“medium,” or “hardly” as described below.
 Entirely: The length of the source document is more than 80%
of the length of the suspicious document.
 Much: The length of the source document is more than
50%80% of the length of the suspicious document.
 Medium: The length of the source document is about
20%50% of the length of the suspicious document.
 Hardly: The length of the source document is less than 20% of
the length of the suspicious document.</p>
      <p>
        If the number of paired documents with the same label is more
than one-fourth of the number of paired documents with a label of
smaller length that do not have enough paired documents, the
label with the lower length is assigned; thus, a uniform
distribution is obtained.
h. Segment Extraction
From each source document, some segments are randomly
selected. The number of selected segments is based on the
classification defrofrepin step g.
i. Segment Obfuscation
This study offers a strategy to manually, semi-automatically, or
automatically produce each type of plagiarism mentioned in
Alzahrani’s taxonomy of plagiarism. In this step, each segment is
obfuscated based on one strategy and add to one suspicious
document. It is noteworthy that all obfuscated segments included
in a document must be obfuscated using the same strategy because
according to PAN corpus format, there is no overlap between
suspicious documents in different strategies[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and only one type
of plagiarism should be employed in each suspicious document.
j. Obfuscated Segment Insertion
In this step each obfuscated segment is inserted into a suspicious
document in a randomly chosen space.
      </p>
    </sec>
    <sec id="sec-5">
      <title>4. STRATEGIES FOR PLAGIARISMS</title>
    </sec>
    <sec id="sec-6">
      <title>TYPES</title>
      <p>
         Exact Copy
In this strategy, the segments produced in step h were inserted into
a suspicious document with no obfuscation. Using this strategy,
324 paired documents were produced.
 Near Copy
According to Fig .1 a type of plagiarism is Near Copy [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] that
consists insertion, deletion, substitution and sentence split or join
methods. To create this kind of plagiarism, the segments produced
in step h are obfuscated through deletion, insertion, sentence
replacement, and sentence division. With this method, some
randomly selected sentences are deleted from the segment and
replaced with randomly selected sentences from the suspicious
document. Then, some randomly selected sentences are swapped.
Finally, complex sentences are identified and broken into main
simple sentences. To do this, the complex sentence identifier
developed at the Shahid Beheshti University NLP Lab is
employed. Each complex sentence in this segment is replaced
with its main and subordinate clauses, and 457 paired documents
are produced based on this strategy.
 Modified Copy
In the taxonomy of plagiarism [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] there is a type of plagiarism
called Modified Copy that to obfuscate a text using this strategy,
the Persian sentence understanding and generation system
introduced by Adelkhah et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] is employed. This system
performs a bidirectional conversion between Persian sentences
and their semantic representation. It changes each sentence to its
semantic representation and then generates the Persian sentence
using semantic representation. To clarify, this system is composed
of 2 sub-systems: 1) semantic representation production for
sentences (sentence understanding), and 2) sentence production
based on semantic representation (sentence generation). It is
noteworthy that in the sentence production phase, in addition to
structural changes, there might be samples of chunk relocations in
a sentence or samples of word relocations in a chunk. The aim of
this system is to produce sentences with the same meaning (deep
structure) but different surface structures and words. Using this
strategy, 465 paired documents are created.
 Text Manipulation (Paraphrasing)
Text Manipulation was performed as described earlier in Modified
Copy. The difference here is the word replacement in the sentence
generation phase. Each word is replaced with a synonym retrieved
from FarsNet (Persian WordNet) [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] or FavaNet (WordNet of
Computer domain) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Hence, in addition to structure
modification, different words are included in the sentence
compared to the main sentence, although the concept remains the
same. Chunks may be moved inside a sentence; however, there is
no movement for words in a chunk. Using this method, 604 paired
documents are produced.
 Text Manipulation (Summarizing)
The goal in this step is to obfuscate a text document using
summarization methods. To create such queries, the automatic
Persian summarizer introduced by Shafiee et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is used, and
506 paired documents are produced.
 Automatic Translation
According to types of plagiarism in Fig .1 translation is a type of
plagiarism that is divided into automatic and manual translation.
Hanif et al. [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] use the automatic translation strategy to
obfuscate documents in their corpus. Moreover in the PAN
20132014 corpus [
        <xref ref-type="bibr" rid="ref18 ref9">9, 18</xref>
        ] use cyclic translation strategy.
      </p>
      <p>We use all of above three strategies in our corpus (described in
Automatic Translation, Manual Translation and Cyclic
Translation stages). For the automatic translation strategy, the
selected sections are translated from Persian to English by Google
translate and the results are checked by Hunspell. Then they are
added to the English suspicious documents. 306 paired documents
are produced using this strategy.
 Manual Translation
The suspicious documents in this step are English articles in the
field of computer engineering, and the source documents are
Persian articles in the same field. The English articles are
clustered as described in step b, and an equivalent English cluster
is produced for each Persian one. Then, for each suspicious
document, a source document from its equivalent Persian cluster
is randomly selected. Based on what was described in steps f, g
and h, some sections of the source document are selected.
Afterwards, these sections are translated by experts in the fields of
computer engineering and are added to the suspicious documents
as described in step j. Seven hundred paired documents are
produced using this strategy which can be employed to evaluate
cross-language similarity detection systems (Persian-English).
 Cyclic Translation
With the cyclic translation strategy the selected sections are
translated from English to Persian using Google translate, and the
results are checked by Negar, a Persian spell checker developed at
the NLP Lab of Shahid Beheshti University. The selected sections
are then translated again from Persian to English. Finally, the
results are checked by Hunspell and add to the English suspicious
documents. Using this method, 388 paired documents are created.
 Idea Adoption (semantic-based meaning)
The goal in this step is to represent the main idea of a source
document using new words/wording. Since most source
documents are computer related theses and articles, automatic
idea extraction would be a complex task here for which no high
accurate system is yet available. Hence, the researchers asked
computer experts to rewrite the idea of each document in their
own words. To simplify the task, only important sections of
source documents, such as the abstract, were considered. Source
documents were distributed among three computer PhD
candidates and 30 computer MS students, and 109 paired
documents were produced.</p>
    </sec>
    <sec id="sec-7">
      <title>5. DATASET STATISTICS</title>
      <p>Overall, employing all the mentioned strategies, 11,603
plagiarism cases and 6,497 paired documents are produced, from
which 2,650 are no-plagiarism, 780 are no obfuscation, and 3,067
are obfuscated ones. The dataset statistics are shown in Table 2.</p>
    </sec>
    <sec id="sec-8">
      <title>6. CONCLUSION</title>
      <p>This article describes a methodology for building a Persian corpus
for evaluating plagiarism detection systems. This large-scale
corpus is in PAN format. To produce this corpus, the focus is on
the simulation of different types of plagiarism. Different strategies
are employed to create obfuscation in each plagiarism category;
hence, a variety of plagiarism types in large volume are created.</p>
    </sec>
    <sec id="sec-9">
      <title>7. REFERENCES</title>
      <sec id="sec-9-1">
        <title>7 A page is measured as 1500 chars.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Asghari</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohtaj</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fatemi</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faili</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <year>2016</year>
          .
          <article-title>Algorithms and Corpora for Persian Plagiarism Detection: Overview of PAN at FIRE 2016</article-title>
          . In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Alzahrani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salim</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Abraham</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Understanding plagiarism linguistic patterns, Textual features, and detection Methods</article-title>
          .
          <source>IEEE Trans. SYSTEMS</source>
          ,
          <string-name>
            <given-names>MAN</given-names>
            , AND
            <surname>CYBERNETICS-PART</surname>
          </string-name>
          <string-name>
            <surname>C</surname>
          </string-name>
          :
          <article-title>APPLICATIONS</article-title>
          AND REVIEWS, vol.
          <volume>42</volume>
          , no. 2.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and et.al.
          <year>2010</year>
          .
          <article-title>An Evaluation Framework for Plagiarism Detection</article-title>
          .
          <source>Proceedings of the 23rd International Conference on Computational Linguistics</source>
          , COLING 2010 Beijing,_c ACL.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Shamsfard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kiani</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Shahedi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <article-title>STeP-1: Standard Text Preparation for Persian Language</article-title>
          . CAASL3 Third Workshop on Computational Approaches to Arabic Script- Languages.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Makrehchi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kamel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>A fuzzy set approach to extracting keywords from abstracts</article-title>
          .
          <source>North American Fuzzy Information Processing Society- NAFIPS</source>
          <year>2003</year>
          , Banf, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Shafiee</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Shamsfard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>The automatic Persian summarizer</article-title>
          .
          <source>The 20st Computer Society of Iran Computer Conference.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Adelkhah</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sadeghi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Shamsfard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Persian sentence understanding and generation: a mutual conversion</article-title>
          .
          <source>The 21st Computer Society of Iran Computer Conference.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Göring</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>and et</article-title>
          .al.
          <year>2015</year>
          .
          <article-title>Towards Data Submissions for Shared Tasks: First Experiences for the Task of Text Alignment</article-title>
          .
          <source>Working Notes Papers of the CLEF 2015 Evaluation Labs, CEUR Workshop Proceedings</source>
          , (
          <year>September 2015</year>
          ), ISSN 1613-0073.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <article-title>and et</article-title>
          .al.
          <year>2013</year>
          .
          <article-title>Overview of the 5th International Competition on Plagiarism Detection”</article-title>
          ,
          <source>Working Notes Papers of the CLEF 2013Evaluation Labs and Workshop</source>
          , (
          <year>September 2013</year>
          ),
          <source>ISBN 978-88-904810-3-1.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Manku</surname>
            ,
            <given-names>G. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Sarma</surname>
            ,
            <given-names>A. D.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Detecting NearDuplicates for Web Crawling</article-title>
          .
          <source>Data mining.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Kamran</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ahmadi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kazemivanhari</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Plagiarism detection in Persian text using Fingerprint algorithms</article-title>
          .
          <source>The 21st Iranian Conference on Electrical Engineering.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Davarpanah</surname>
            ,
            <given-names>M. R.</given-names>
          </string-name>
          , sanji, M. and
          <string-name>
            <surname>Aramideh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>Farsi Lexical Analysis</article-title>
          and
          <string-name>
            <given-names>StopWord</given-names>
            <surname>List</surname>
          </string-name>
          .
          <source>Library Hi Tech</source>
          , vol.
          <volume>27</volume>
          , pp
          <fpage>435</fpage>
          -
          <lpage>449</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>[13] Iran Telecommunication Research Center (ITRC)</source>
          ,
          <year>2013</year>
          . Buali Sina University. http://217.218.62.234:
          <fpage>8080</fpage>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Shamsfard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hesabi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fadaei H</surname>
          </string-name>
          .
          <article-title>and et</article-title>
          .al
          <year>2010</year>
          .
          <article-title>Semi Automatic Development of FarsNet; The Persian WordNet</article-title>
          .
          <source>Proceedings of 5th Global WordNet Conference.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Mashhadirajab</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Shamsfard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Plagiarism Detection in Persian documents</article-title>
          .
          <source>Master's thesis</source>
          . Shahid Beheshti University.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eiselt</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          and et.al.
          <year>2011</year>
          .
          <article-title>Overview of the 3rd International Competition on Plagiarism Detection</article-title>
          .
          <source>Notebook Papers of CLEF 2011 Labs and Workshops</source>
          , (
          <year>September 2011</year>
          ),
          <source>ISBN 978-88-904810-1-7.</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <article-title>and et</article-title>
          .al.
          <year>2012</year>
          .
          <article-title>Overview of the 4th International Competition on Plagiarism Detection</article-title>
          .
          <article-title>CLEF 2012 Evaluation Labs</article-title>
          and Workshop - Working Notes Papers, (
          <year>September 2012</year>
          ),
          <source>ISBN 978-88-904810-3-1.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>and et</article-title>
          .al.
          <year>2014</year>
          .
          <article-title>Overview of the 6th International Competition on Plagiarism Detection</article-title>
          .
          <article-title>CLEF 2014 Evaluation Labs</article-title>
          and Workshop - Working Notes Papers, (
          <year>September 2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Joy</surname>
            ,
            <given-names>M. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sinclair</surname>
            ,
            <given-names>J. E.</given-names>
          </string-name>
          <article-title>and et</article-title>
          .al.
          <year>2013</year>
          .
          <article-title>Student perspectives on source-code plagiarism</article-title>
          .
          <source>International Journal for Educational Integrity</source>
          , Vol.
          <volume>9</volume>
          , No.
          <issue>1</issue>
          , pp.
          <fpage>3</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Joy</surname>
            ,
            <given-names>M. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cosma</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <article-title>and et</article-title>
          .al.
          <year>2009</year>
          .
          <article-title>A TAXONOMY OF PLAGIARISM IN COMPUTER SCIENCE</article-title>
          .
          <source>Proceedings of EDULEARN09 Conference</source>
          , (
          <year>July 2009</year>
          ), ISBN:
          <fpage>978</fpage>
          -
          <lpage>84</lpage>
          -612- 9802-0.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Naik</surname>
            ,
            <given-names>R. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landge</surname>
            ,
            <given-names>M. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mahender</surname>
            ,
            <given-names>C. N.</given-names>
          </string-name>
          <article-title>and et</article-title>
          .al
          <year>2015</year>
          .
          <article-title>A Review on Plagiarism Detection Tools</article-title>
          .
          <source>International Journal of Computer Applications</source>
          , vol.
          <volume>125</volume>
          - No.
          <year>11</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Alvi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stevenson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          and et.al
          <year>2015</year>
          .
          <article-title>The short stories corpus. Notebook for PAN at CLEF</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Cheema</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Najib</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ahmed</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>and et</article-title>
          .al
          <year>2015</year>
          .
          <article-title>A corpus for analyzing text reuse by people of different groups. Notebook for PAN at CLEF</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Hanif</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nawab</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arbab</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>and et</article-title>
          .al
          <year>2015</year>
          .
          <article-title>Crosslanguage urdu-english (clue) text alignment corpus. Notebook for PAN at CLEF</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Kong</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , Han,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <article-title>and et</article-title>
          .al
          <year>2015</year>
          .
          <article-title>Source retrieval and text alignment corpus construction for plagiarism detection. Notebook for PAN at CLEF</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Khoshnavataher</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zarrabi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohtaj</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Asghari</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Developing monolingual Persian corpus for extrinsic plagiarism detection using artificial obfuscation. Notebook for PAN at CLEF</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Asghari</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khoshnavataher</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fatemi</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Faili</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Developing bilingual plagiarism detection corpus using sentence aligned parallel corpus. Notebook for PAN at CLEF</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Mohtaj</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Asghari</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <article-title>and</article-title>
          <string-name>
            <surname>Zarrabi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Developing monolingual english corpus for plagiarism detection using human annotated paraphrase corpus. Notebook for PAN at CLEF</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Palkovskii</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Belov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Submission to the 7th international competition on plagiarism detection</article-title>
          . http://www.uni-weimar.de/medien/webis/events/pan-15, http://www.clef-initiative.eu/publication/working-notes, From the Zhytomyr State University and SkyLine LLC.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>