<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>we
collected Persian papers from peer reviewed journals. We crawled
the websites of journals introduced in the System for Evaluation
of Scientific Journals2 (affiliated to Iran's Ministry of Science</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Mahak Samim: A Corpus of Persian Academic Texts for Evaluating Plagiarism Detection Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>CCS Concepts</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Morteza Rezaei Sharifabadi Computer Research Center of Islamic Sciences Tehran</institution>
          ,
          <addr-line>I. R.</addr-line>
          <country country="IR">Iran</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Seyed Ahmad Eftekhari Computer Research Center of Islamic Sciences Tehran</institution>
          ,
          <addr-line>I. R.</addr-line>
          <country country="IR">Iran</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we introduce Mahak Samim, a plagiarism detection corpus that consists of Persian academic texts in which plagiarism cases are embedded. This corpus, which can be used for evaluating plagiarism detection systems, consists of more than five thousand artificial plagiarism cases with various lengths and diverse degrees of obfuscation. The development process and the features of the corpus are described here. • Information systems ➝ Information retrieval ➝ Retrieval tasks and goals ➝ Near-duplicate and plagiarism detection. plagiarism detection; evaluation corpus; Persian; academic texts.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Plagiarism is defined as “copying or closely imitating the work of
another writer, composer, etc., without permission and with the
intention of passing the results off as original work” [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
Plagiarism detectors are software programs developed to detect
cases of such misconduct in documents. The PAN evaluation lab
series has provided a framework for evaluating plagiarism
detection systems. This framework relies on plagiarism corpora
which are basically collections of text that include cases of
plagiarism. Plagiarism detection systems receive the corpus texts
as input and their ability to detect the plagiarism cases embedded
in the texts are examined.
      </p>
      <p>Since plagiarism detectors are not entirely language independent,
there is a need for plagiarism corpora in various languages. In
recent years a couple of Persian plagiarism detection systems have
been developed. Proper evaluation of these systems is dependent
on reliable Persian plagiarism corpora.</p>
      <p>In this paper we introduce Mahak Samim1, a corpus suitable for
evaluating Persian plagiarism detectors. We first briefly review
previous works in this field and then we introduce our own
approach. The paper concludes with a summary and an outlook
for further work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. RELATED WORK</title>
      <p>
        Prior to PAN evaluation lab series, plagiarism corpora were rare.
The corpora used in PAN labs held in 2009 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], 2010 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and 2011
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] had basically the same structure. The documents used in these
1 Samim-Noor is a commercial plagiarism detection system
developed by the Computer Research Center of Islamic
Sciences. Mahak in Persian means “Touchstone”.
corpora were books from Project Gutenberg. The corpora were
used to evaluate both external and intrinsic plagiarism detection.
In external plagiarism detection suspicious documents are
checked against a collection of source documents, but in intrinsic
plagiarism suspicious documents are analyzed in isolation for
changes in writing style etc. Fifty percent of the documents were
used as source documents and fifty percent as suspicious
documents. The corpora contain plagiarism cases with different
lengths and various degrees of artificial and simulated
obfuscation. Artificial obfuscation includes techniques such as
automatically shuffling and replacing words and simulated
obfuscation was achieved through crowdsourcing the obfuscation
task. The major shortcoming of corpora presented in these years
was their relatively small size. The plagiarism detectors were
expected to include a stage of heuristic retrieval in which they
selected a group of candidate documents among the total
collection of source documents. However, since the size of the
corpora were not large enough, the systems skipped this stage. In
PAN 2012 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] this issue is addressed and a new approach is
adopted for developing the plagiarism corpus. For this purpose a
number of professional writers were asked to write articles –
containing plagiarism - on a set of topics. A one billion document
corpus resembling the web was used as the collection of source
documents. The writers compiled their articles by searching
through this huge collection. In PAN 2013 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and PAN 2014 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
expanded versions of the 2012 corpus were used.
      </p>
      <p>
        In PAN 2015 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] a task of corpus construction was introduced. In
this task, participants were asked to provide their own plagiarism
corpora. Eight plagiarism corpora were provided for this task
among which two included Persian documents. Khoshnavataher
et.al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] present a monolingual Persian corpus based on about
2100 Wikipedia articles with plagiarism cases obfuscated
artificially and intended for evaluation of extrinsic plagiarism
detection. Asghari et.al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] use Wikipedia documents and a
Persian-English sentence-aligned corpus to develop a bilingual
plagiarism detection corpus.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. CORPUS DEVELOPMENT</title>
    </sec>
    <sec id="sec-4">
      <title>3.1 Document Collection</title>
      <p>documents in each subject, as grouped by the System for
Evaluation of Scientific Journals, and Table 2 provides
information about the document lengths.</p>
    </sec>
    <sec id="sec-5">
      <title>3.4 Plagiarism case length</title>
      <p>Our corpus consists of a total of 5862 plagiarism cases with
lengths between 50 and 5000 words. Table 4 shows the statistics.
Long plagiarism cases may include more than one sentence.</p>
    </sec>
    <sec id="sec-6">
      <title>3.2 Source / suspicious documents</title>
      <p>In plagiarism corpora, the documents collection is usually split
into two main subgroups i.e. source documents and suspicious
documents. Source documents are documents from which parts of
text are selected as plagiarism cases. These parts are then inserted
inside the text of so-called suspicious documents. In other words,
suspicious documents are documents which include text used in
source documents. We follow PANs tradition of using half of the
documents as source documents and half as suspicious
documents. It is noteworthy that the subjects of the papers were
taken into consideration while dividing the collection into halves.
i.e. 50 percent of the papers in humanities were used as source
documents and 50 percent as suspicious documents etc.</p>
    </sec>
    <sec id="sec-7">
      <title>3.3 Plagiarism per document</title>
      <p>
        50 percent of the suspicious documents have no plagiarism cases.
As mentioned in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], the documents without plagiarism allow to
determine whether or not a detector can distinguish plagiarism
cases from overlaps that occur naturally between random
documents. Statistics of plagiarism per document in the rest of the
suspicious documents, i.e. 25 percent of the whole corpus, is
available in Table 3.
      </p>
      <sec id="sec-7-1">
        <title>Plagiarism Per Document</title>
      </sec>
      <sec id="sec-7-2">
        <title>Percent of Documents</title>
        <p>hardly (5%-20%)
medium (20%-50%)
much</p>
        <p>(50%-80%)
entirely (&gt;80%)
Short (50-150 words)
Medium (300-500 words)
Long (3000-5000 words)</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>3.5 Topic match</title>
      <p>The six general topic categories of the papers used in our corpus
were introduced in table 1. Fifty percent of the plagiarism cases
were made between papers with same topics (intra-topic cases)
and fifty percent between papers with different topics (inter-topic
cases).</p>
    </sec>
    <sec id="sec-9">
      <title>3.6 Obfuscation types</title>
      <p>
        In many cases, plagiarized texts are manipulated by those
committing plagiarism in order to avoid being detected by
plagiarism detection systems or human readers. Plagiarism
corpora developers use different techniques to include such
obfuscations in their plagiarism cases. An overview of different
types of obfuscation in our plagiarism cases is available in Table
5.
As shown in table 4, 40 percent of the plagiarism cases have no
obfuscation. As explained in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], since the writing style of the
original author is preserved in plagiarism cases without
obfuscation, these cases are especially appropriate for evaluating
intrinsic plagiarism detection. Random text operations are
operations such as adding, deleting and substituting words, which
are all done randomly. Semantic word variation, on the other
hand, is the random substitution of words with their synonyms.
We use the Comprehensive Dictionary of Persian Synonyms and
Antonyms3 as a resource for extracting synonyms. The terms “low
obfuscation” and “high obfuscation” mentioned in table 4 show
the degree of obfuscation i.e. how many words have been added,
deleted or substituted etc.
      </p>
    </sec>
    <sec id="sec-10">
      <title>4. SUMMARY AND FUTURE WORK</title>
      <p>As explained above, Mahak Samim is a plagiarism corpus which
can be used for evaluating both intrinsic and external plagiarism
detection systems. In order to preserve overall balance, many
factors – plagiarism per document, plagiarism case length, topic
match, obfuscation type, and obfuscation degree – were taken into
consideration while preparing each plagiarism case. The corpus
files are prepared according to the format of previous PAN
3 The plain-text version of this dictionary can be downloaded from
this link: http://dadegan.ir/catalog/D3911124a
corpora which include xml files that have information about the
starting point of the plagiarism in relevant source and suspicious
documents and the length of the plagiarism case.</p>
      <p>
        Plagiarism cases in our corpus are cases of “artificial plagiarism”.
Using “real plagiarism” cases in plagiarism corpora is problematic
due to ethical, legal, and financial issues [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. However, we may
enrich our corpus by adding cases of simulated plagiarism. Other
types of artificial obfuscation, such as POS-preserving word
shuffling could also be employed. The corpus may be easily
expanded with both academic papers and other types of
documents such as books, web articles, etc.
      </p>
      <p>
        This paper has been submitted to The PAN@FIRE2016 Shared
Task on Persian Plagiarism Detection and Text Alignment Corpus
Construction [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and the corpus is available through Peykaregan4
website.
      </p>
    </sec>
    <sec id="sec-11">
      <title>5. ACKNOWLEDGMENTS</title>
      <p>Special thanks to Dr. Martin Potthast for his valuable help and to
Dr. Mahdi Behnia and Mr. Amirhossein Rajabzadeh Assarha, our
colleagues in the Computer Research Center of Islamic Sciences,
for their support and their comments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Asghari</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khoshnava</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fatemi</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Faili</surname>
          </string-name>
          , H.,
          <year>2015</year>
          .
          <article-title>Developing Bilingual Plagiarism Detection Corpus Using Sentence Aligned Parallel Corpus</article-title>
          .
          <source>In Working Notes Papers of the CLEF</source>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Asghari</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohtaj</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fatemi</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faili</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <year>2016</year>
          .
          <article-title>Algorithms and Corpora for Persian Plagiarism Detection: Overview of PAN at FIRE 2016</article-title>
          . In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Khoshnavataher</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zarrabi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohtaj</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Asghari</surname>
          </string-name>
          , H.,
          <year>2015</year>
          .
          <article-title>Developing Monolingual Persian Corpus for Extrinsic Plagiarism Detection Using Artificial Obfuscation</article-title>
          .
          <source>In Working Notes Papers of the CLEF</source>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eiselt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>2009</year>
          .
          <article-title>Overview of the 1st international competition on plagiarism detection</article-title>
          .
          <source>In 3rd PAN Workshop</source>
          . Uncovering Plagiarism,
          <source>Authorship and Social Software Misuse.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eiselt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>2010</year>
          .
          <article-title>Overview of the 2nd international competition on plagiarism detection</article-title>
          .
          <source>In Notebook Papers of CLEF 2010 LABs and Workshops.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eiselt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón</surname>
            <given-names>Cedeño</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>L.A.</given-names>
            ,
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            and
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          ,
          <year>2011</year>
          .
          <article-title>Overview of the 3rd international competition on plagiarism detection</article-title>
          .
          <source>InCEUR Workshop Proceedings.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oberländer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tippmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2012</year>
          .
          <article-title>Overview of the 4th international competition on plagiarism detection</article-title>
          .
          <source>In Working Notes Papers of CLEF 2012 Evaluation Labs and Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beyer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Busse</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tippmann</surname>
            , Rosso,
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2014</year>
          .
          <article-title>Overview of the 6th international competition on plagiarism detection</article-title>
          .
          <source>Working Notes Papers of the CLEF 2014 Evaluation Labs, CEUR Workshop Proceedings.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tippmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2013</year>
          .
          <article-title>Overview of the 5th international competition on plagiarism detection</article-title>
          .
          <source>In CLEF Conference on Multilingual and Multimodal Information Access Evaluation</source>
          (pp.
          <fpage>301</fpage>
          -
          <lpage>331</lpage>
          ). CELCT.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Göring</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2015</year>
          .
          <article-title>Towards data submissions for shared tasks: first experiences for the task of text alignment</article-title>
          .
          <source>Working Notes Papers of the CLEF</source>
          <year>2015</year>
          , pp.
          <fpage>1613</fpage>
          -
          <lpage>0073</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>2010</year>
          .
          <article-title>An evaluation framework for plagiarism detection</article-title>
          .
          <source>In Proceedings of the 23rd international conference on computational linguistics: Posters</source>
          (pp.
          <fpage>997</fpage>
          -
          <lpage>1005</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Reitz</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <year>1996</year>
          . ODLIS:
          <article-title>Online dictionary for library and information science</article-title>
          .
          <source>Libraries Unlimited.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>