<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Developing Monolingual Persian Corpus for Extrinsic Plagiarism Detection Using Artificial Obfuscation</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>ICT Research Institute, Academic Center for Education</institution>
          ,
          <addr-line>Culture and Reseach (ACECR)</addr-line>
          ,
          <country country="IR">Iran</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Khadijeh Khoshnavataher</institution>
          ,
          <addr-line>Vahid Zarrabi, Salar Mohtaj, Habibollah Asghari</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>The task of text alignment corpus construction at PAN 2015 competition consists of preparing a plagiarism corpus so that it can provide various obfuscation types and versatile obfuscation degrees. Meanwhile, its format and metadata structure should follow previous PAN plagiarism corpora. In this paper, we describe our approach for construction of a monolingual Persian plagiarism corpus that can be used to evaluate the performance of Persian plagiarism detection systems.</p>
      </abstract>
      <kwd-group>
        <kwd>Plagiarism Detection</kwd>
        <kwd>Corpus Construction</kwd>
        <kwd>Text Alignment Corpus</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Plagiarism is the re-use of another person’s ideas, processes, results, or words without
explicitly acknowledging the original source [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Plagiarism detection is the
algorithms for retrieval and extraction of text reuse within a suspicious document and
corresponding source documents. The suspicious and source documents can be
written either in the same language named as monolingual (MLPD) or in different
languages named as cross lingual (CLPD) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        In the process of developing a system for plagiarism detection in natural language
texts, the system should be trained and tested on a text corpus containing known
plagiarized passage. The PAN is a major international competition for the task of
plagiarism detection, and provides corpora for the evaluation of plagiarism detection
systems. The corpus used in the first PAN competition known as PAN-PC-09 which,
contain only artificial plagiarism cases [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In later revisions of the corpus, a variety
of obfuscation strategies have been applied in corpora, including artificial
obfuscation, simulated (or manual) obfuscation, translation obfuscation and summary
obfuscation [
        <xref ref-type="bibr" rid="ref4 ref5 ref6 ref7 ref8">4, 5, 6, 7, 8</xref>
        ].
      </p>
      <p>In this paper, in order to construct a monolingual plagiarism detection corpus for
Persian language, we have deployed our approach based on PAN corpora strategies.
We have used an artificial obfuscation strategy to create plagiarism cases.</p>
      <p>The paper is organized as follow: In section 2, we describe our approach for corpus
construction. Then in section 3 we will discuss the results of the corpus. Finally, we
will conclude and discuss about some future works in section 4.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Our Approach</title>
      <p>In this section the overall procedure of our approach for building monolingual Persian
corpus is described. It is organized in five steps: preprocessing, documents clustering,
fragment extraction, fragment obfuscation and inserting plagiarism cases in
suspicious documents. In the following subsections, we describe the process of each step.
2.1</p>
      <sec id="sec-2-1">
        <title>Preprocessing</title>
        <p>
          Persian language belongs to the category of Arabic-Scripted based languages. There
are some problems dealing with preprocessing in this language [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. There are some
efforts to develop Persian preprocessing algorithms [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ]. In this paper, in the
preprocessing stage of the system, we have applied some algorithms such as
normalization, tokenization, stemming and part of speech (POS) tagging.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Documents Clustering</title>
        <p>Establishing topically similarity between suspicious passage and its corresponding
source documents is an important issue for corpus construction. By inserting
plagiarized passages into topically related surrounding text, the corpus may become more
realistic. Therefore, in this step, collection of Wikipedia documents clustered into
different topically related groups. A bipartite graph of documents-categories was
created to cluster the documents. In the next step, the info- map community detection
algorithm was applied to the graph and all communities were detected. Finally,
Documents within a community are considered as one cluster. Each suspicious document
and its corresponding source documents are selected from one cluster.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Fragments Extraction</title>
        <p>The documents used in the corpus are divided into two categories: 50% of the
documents selected as source documents and 50% are designated as suspicious documents.
Note that only 50% of suspicious documents contain plagiarism cases.</p>
        <p>The task of the fragment extraction is to extract fragments from source documents.
The length of fragments is evenly distributed between 30 and 500 words. The length
of fragments is shown in Table 1.
We have used an artificial obfuscation strategy for the plagiarism corpus. To create
artificial plagiarism, we have used five obfuscation strategies as follows:
 None (No Obfuscation)
 Random Change of Order
Source fragment without any change consider as obfuscated fragment. In other words,
obfuscation fragment is an exact copy of source fragment.</p>
        <p>Given source fragment, obfuscation fragment is created by shuffling words at random.
Since the tokenization of words in Persian is a challenging issue, so this task should
be done under supervision.
 POS-preserving Change of Order
In order to accomplish this obfuscation, the sequence of parts of speech (POS) in
source fragment is determined. Then, words are shuffling randomly, while retaining
the original POS sequence.
 Synonym Substitution
The plagiarized fragment is created in such a way to replace some words by one of
their synonyms.
 Addition / Deletion
Obfuscation fragment is created by inserting or removing words at random.</p>
        <p>The number of operations made on source fragment specifies the degree of
obfuscation. Different degrees of obfuscation are “None”, “Low”, “Medium”, and “High”
obfuscation.
2.5</p>
      </sec>
      <sec id="sec-2-4">
        <title>Insert Plagiarism Cases in Suspicious Documents</title>
        <p>In this step, according to suspicious document’s length, one or more plagiarism cases
which are in the same cluster of suspicious documents are selected. Then, each of
them inserted at random positions in suspicious document.</p>
        <p>The fraction of plagiarism in each document is not fixed. The percentage of
plagiarism in each suspicious document is distributed between 5% and 100% of its length.
The ratio of plagiarism per suspicious documents is shown in Table 2.</p>
        <p>Finally, for each pair of source and suspicious documents, an XML file is created
which contains Meta information about the plagiarism cases. These include:
─ this_length: Length of plagiarism case in suspicious document.
─ this_offset: Start offset of the plagiarism case in the suspicious document.
─ source_reference: Name of source file.
─ source_length: Length of Source fragment in source document.
─ source_offset: Start offset of source fragment in the source document.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>In this section, we have presented the result and statistics of our corpus. Because of
the lack of Persian plagiarism detection corpora, which contain manual cases of
plagiarism, we cannot compare plagiarism cases in our corpus to actual cases. An
overview of important corpus statistics is shown in Table 3. The corpus is based on 2114
documents from Wikipedia articles.</p>
      <p>For developing this corpus and other corpora in our laboratory, we have developed
a web based application that can process the input documents and construct various
plagiarism corpora based on corpus builder settings. The reason for implementing the
corpus builder in a web application is for crowd sourcing the simulated plagiarism
cases and inserting them in resulted corpus.</p>
      <p>The established Persian mono-lingual plagiarism detection corpus is available at
the website1 of “Research Institute for Information and Communication Technology”
for research purposes.
1 http://www.ictrc.ir/plaglab/corpora/Monolingual_Persian_Corpus(khoshnava15).zip</p>
      <p>Plagiarism per Document
The number of Little plagiarized documents:
The number of Medium plagiarized documents:
The number of Much plagiarized documents:
The number of Very much plagiarized documents:
1057
529
528
259
564
301
80
96
52</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Work</title>
      <p>We have discussed our approach to the task of text alignment in the context of PAN
2015 competition. We describe a system that generates a monolingual Persian
plagiarism detection corpus. This corpus is the first plagiarism detection corpus in Persian
language and is intended to be used for testing the performance of plagiarism
detection systems.</p>
      <p>In our future work, we plan to improve our corpus by implementing obfuscation
techniques such that simulated obfuscation and other obfuscation strategies using
plagiarism corpus builder.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgement</title>
      <p>This work has been accomplished in ICT research Institute, ACECR, under the
support of Vice Presidency for Science and Technology of Iran - grant No. 1164331. The
authors gratefully acknowledge the support of aforementioned organizations. Special
thanks go to the members of ITBM research group for their valuable collaboration.
The authors also express their gratitude to M.R. Ghahari and Samira Rezaei.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Barrón-Cedeño</surname>
            , Alberto, Marta Vila,
            <given-names>M. Antònia</given-names>
          </string-name>
          <string-name>
            <surname>Martí</surname>
            , and
            <given-names>Paolo</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
          </string-name>
          .
          <article-title>"Plagiarism meets paraphrasing: Insights for the next generation in automatic plagiarism detection</article-title>
          .
          <source>" Computational Linguistics</source>
          <volume>39</volume>
          , no.
          <issue>4</issue>
          (
          <year>2013</year>
          ):
          <fpage>917</fpage>
          -
          <lpage>947</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Barrón-Cedeno</surname>
            , Alberto, and
            <given-names>Paolo</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
          </string-name>
          .
          <article-title>"Monolingual and Crosslingual Plagiarism Detection</article-title>
          .
          <source>Towards the Competition@ SEPLN09."</source>
          (
          <year>2009</year>
          ):
          <fpage>29</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Eiselt</surname>
          </string-name>
          , Andreas, Martin Potthast, Benno Stein, and
          <article-title>Alberto Barrón-Cedeno Paolo Rosso. "Overview of the 1st international competition on plagiarism detection."</article-title>
          <source>In 3rd PAN Workshop</source>
          . Uncovering Plagiarism,
          <source>Authorship and Social Software Misuse</source>
          , p.
          <fpage>1</fpage>
          .
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Potthast</surname>
            , Martin,
            <given-names>Alberto</given-names>
          </string-name>
          <string-name>
            <surname>Barrón-Cedeño</surname>
            , Andreas Eiselt, Benno Stein, and
            <given-names>Paolo</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
          </string-name>
          .
          <article-title>"Overview of the 2nd International Competition on Plagiarism Detection." In CLEF (Notebook Papers</article-title>
          /LABs/Workshops).
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Potthast</surname>
            , Martin,
            <given-names>Alberto</given-names>
          </string-name>
          <string-name>
            <surname>Barrón-Cedeño</surname>
            , Andreas Eiselt, Benno Stein, and
            <given-names>Paolo</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
          </string-name>
          .
          <article-title>"Overview of the 3rd International Competition on Plagiarism Detection." In CLEF (Notebook Papers</article-title>
          /LABs/Workshops).
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Potthast</surname>
            , Martin,
            <given-names>Tim Gollub</given-names>
          </string-name>
          , Matthias Hagen, Jan Graßegger, Johannes Kiesel, Maximilian Michel, Arnd Oberländer, Martin Tippmann, Alberto Barrón-Cedeño, Parth Gupta, Paolo Rosso, and
          <string-name>
            <given-names>Benno</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <article-title>Overview of the 4th International Competition on Plagiarism Detection</article-title>
          .
          <source>In Working Notes Papers of the CLEF 2012 Evaluation Labs</source>
          ,
          <year>September 2012</year>
          .
          <source>ISBN 978-88-904810-3-1. ISSN</source>
          <year>2038</year>
          -
          <volume>4963</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Potthast</surname>
            , Martin,
            <given-names>Matthias Hagen</given-names>
          </string-name>
          , Tim Gollub, Martin Tippmann, Johannes Kiesel, Paolo Rosso, Efstathios Stamatatos, and
          <string-name>
            <given-names>Benno</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <article-title>"Overview of the 5th international competition on plagiarism detection."</article-title>
          <source>In CLEF Conference on Multilingual and Multimodal Information Access Evaluation</source>
          , pp.
          <fpage>301</fpage>
          -
          <lpage>331</lpage>
          . CELCT,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Potthast</surname>
            , Martin,
            <given-names>Matthias Hagen</given-names>
          </string-name>
          , Anna Beyer, Matthias Busse, Martin Tippmann,
          <source>Paolo Rosso, and Benno Stein "Overview of the 6th International Competition on Plagiarism Detection." CLEF</source>
          (Online Working Notes/Labs/Workshop).
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Shamsfard</surname>
          </string-name>
          , Mehrnoush.
          <article-title>"Challenges and open problems in Persian text processing</article-title>
          .
          <source>" Proceedings of LTC 11</source>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Sarabi</surname>
            , Zahra,
            <given-names>Hamidreza</given-names>
          </string-name>
          <string-name>
            <surname>Mahyar</surname>
            , and
            <given-names>Mojgan</given-names>
          </string-name>
          <string-name>
            <surname>Farhoodi</surname>
          </string-name>
          .
          <article-title>"ParsiPardaz: Persian Language Processing Toolkit." In Computer and Knowledge Engineering (ICCKE</article-title>
          ),
          <year>2013</year>
          3th International eConference on, pp.
          <fpage>73</fpage>
          -
          <lpage>79</lpage>
          . IEEE,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Seraji</surname>
            , Mojgan,
            <given-names>Beáta</given-names>
          </string-name>
          <string-name>
            <surname>Megyesi</surname>
            , and
            <given-names>Joakim</given-names>
          </string-name>
          <string-name>
            <surname>Nivre</surname>
          </string-name>
          .
          <article-title>"A basic language resource kit for Persian."</article-title>
          <source>In Eight International Conference on Language Resources and Evaluation (LREC</source>
          <year>2012</year>
          ),
          <fpage>23</fpage>
          -
          <lpage>25</lpage>
          May
          <year>2012</year>
          , Istanbul, Turkey, pp.
          <fpage>2245</fpage>
          -
          <lpage>2252</lpage>
          . European Language Resources Association,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>