<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Algorithms and Corpora for Persian Plagiarism Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Habibollah Asghari</string-name>
          <email>habib.asghari@ictrc.ac.ir</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Salar Mohtaj</string-name>
          <email>salar.mohtaj@ictrc.ac.ir</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heshaam Faili</string-name>
          <email>hfaili@ut.ac.ir</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Rosso</string-name>
          <email>prosso@dsic.upv.es</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Omid Fatemi</string-name>
          <email>omid@fatemi.net</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Potthast</string-name>
          <email>martin.potthast@uni-weimar.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bauhaus-Universität Weimar</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ICT Research Institute, Academic Center for Education</institution>
          ,
          <addr-line>Culture and, Research (ACECR)</addr-line>
          ,
          <country country="IR">Iran</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>PRHLT Research Center, Universitat Politècnica de València</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>School of Electrical and Computer Engineering, College of Engineering, University of Tehran</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>School of Electrical and Computer, Engineering, College of Engineering, University of Tehran</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <abstract>
        <p>The task of plagiarism detection is to find passages of text-reuse in a suspicious document. This task is of increasing relevance, since scholars around the world take advantage of the fact that information about nearly any subject can be found on the World Wide Web by reusing existing text instead of writing their own. We organized the Persian PlagDet shared task at PAN 2016 in an effort to promote the comparative assessment of NLP techniques for plagiarism detection with a special focus on plagiarism that appears in a Persian text corpus. The goal of this shared task is to bring together researchers and practitioners around the exciting topic of plagiarism detection and text-reuse detection. We report on the outcome of the shared task, which divides into two subtasks: text alignment and corpus construction. In the first subtask, nine teams participated, whereas the best result achieved was a PlagDet score of 0.922. For the second subtask of corpus construction, five teams submitted a corpus, which were evaluated using the systems submitted for the first subtask. The results show that significant challenges remain in evaluating newly constructed corpora. •General and reference → General conference proceedings.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Plagiarism Detection</kwd>
        <kwd>Evaluation Framework</kwd>
        <kwd>TIRA Platform</kwd>
        <kwd>Shared Task</kwd>
        <kwd>Persian PlagDet</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>In recent years, a lot of research has been carried out concerning
text reuse and plagiarism detection for English. But the detection
of plagiarism in languages other than English has received
comparably little attention. Although there have been previous
developments on tools and algorithms to assist detecting text
reuse in Persian, little is known about their detection performance.
Therefore, to foster research and development on Persian
plagiarism detection, we have organized the first corresponding
competition, held in conjunction with the PAN evaluation lab at
FIRE 2016.</p>
      <p>
        We overview the detection approaches of nine participating
teams and evaluate their respective retrieval performance.
Participants were asked to submit their software to the TIRA
Evaluation-as-a-Service (EaaS) platform [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] instead of just
sending run outputs, rendering the shared task more reproducible.
The submitted pieces of software are maintained in executable
form so that they can be re-run against new corpora later on. To
demonstrate this possibility, we asked participants to also submit
evaluation corpora of their own design, which were examined
using the detection systems submitted by other participants.
In what follows, Section 2 reviews related work with respect to
shared tasks on plagiarism detection. Section 3 describes the main
steps of tasks. Section 4 describes the evaluation framework,
explaining the TIRA evaluation platform as well as the
construction of our training and test datasets alongside the
performance measures used. In Section 5, the evaluation results of
both the text alignment and the corpus construction subtasks are
reported.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. RELATED WORK</title>
      <p>This section reviews recent competitions and shared tasks on
plagiarism detection in English, Arabic and Persian.</p>
      <p>
        PAN. Potthast et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] first pointed out the lack of a
controlled evaluation environment and corresponding detection
quality measures to evaluate plagiarism detection systems as a
major obstacle to evaluating plagiarism detection approaches. To
overcome these shortcomings, they organized the first
international competition on plagiarism detection in 2009
featuring two subtasks: external plagiarism detection and intrinsic
plagiarism detection. An important by-product of this competition
was the first evaluation framework for plagiarism detection,
which consists of a large-scale plagiarism corpus and a detection
quality measure called as PlagDet [
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ].
      </p>
      <p>
        The PAN competition was continued in the next years,
improving the evaluation corpora with each iteration. As of 2012,
the competition was revamped in the form of two new subtasks:
source retrieval and text alignment. Moreover, at PAN 2015, for
the first time, participants were invited to submit their own
alignment corpora. Here, participants were asked to compile
corpora comprising artificial, simulated, or even real plagiarism,
formatted according to the data format established for the
previous shared tasks [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        AraPlagDet. AraPlagDet is the first international
competition on detecting plagiarism in Arabic documents. The
competition was held as a PAN shared task at FIRE 2015 and
included two sub-tasks corresponding to the first shared tasks at
PAN: external plagiarism detection and intrinsic plagiarism
detection [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The competition followed the formats used at PAN.
One of the main motivations of organizers for this shared task was
to raise awareness in the Arab world on the seriousness of
plagiarism, and, to promote the development of plagiarism
detection approaches that deal with the peculiarities of the Arabic
language, providing for an evaluation corpus that allows for
proper performance comparison between Arabic plagiarism
detectors.
      </p>
      <p>
        PlagDet Task at AAIC. The first competition on Persian
plagiarism detection was held as the 3rd AmirKabir Artificial
Intelligence Competition (AAIC) in 2015. The competition was
the first to plagiarism detection in the Persian language and led to
the release of the first plagiarism detection corpus in Persian [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
Like AraPlagDet, the PAN standard framework on evaluation and
corpus annotation has been used in this competition.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. TASK DESCRIPTION</title>
      <p>The shared task of Persian plagiarism detection divides into two
subtasks: text alignment and corpus construction.</p>
      <p>Text alignment is based on PAN evaluation framework to assess
the detection performance plagiarism detectors: given two
documents, the task is to determine all contiguous passages of
reused texts between them. Nine teams participated in this
subtask.</p>
      <p>The corpus construction subtask invited participants to submit
evaluation corpora of their own design for text alignment,
following the standard corpus format. Five corpora were
submitted to the competition. Their evaluation consisted of
evaluating the validity of annotations via analyzing corpus
statistics, such as the length distribution of the documents, the
length distribution of the plagiarized passages, and the ratio of
plagiarism per document. Moreover, we report on the
performance of the aforementioned nine plagiarism detectors in
detecting the plagiarism comprised within the submitted corpora.</p>
    </sec>
    <sec id="sec-4">
      <title>4. EVALUATION FRAMEWORK</title>
      <p>The text alignment subtask consists of identifying the exact
positions of reused text passages in a given pair of suspicious
document and source document. This section describes the
evaluation platform, corpus, and performance measure that were
used in this subtask. Moreover, the submitted detection
approaches and their respective evaluation results are presented.</p>
    </sec>
    <sec id="sec-5">
      <title>4.1 Evaluation Platform</title>
      <p>Establishing an evaluation framework for Persian plagiarism
detection was one of the primary goals of our competition,
consisting of a large-scale plagiarism detection corpus along with
performance measures. The framework may serve as a unified test
environment for future activities on Persian plagiarism detection
research.</p>
      <p>
        Due to the diverse development environments of participants,
it is preferable to set up a common platform that satisfies all their
requirements. We decided to use the TIRA experimentation
platform [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. TIRA provides for a set of features that facilitate the
reproducibility of our shared
organizational overhead [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]:
task
while
reducing
its





      </p>
      <p>TIRA provides every participant with a virtual machine
that allows for the convenient deployment and
execution of submitted software.</p>
      <p>Both Windows and Linux machines are available to
participants, whereas deployed software need only be
executable from a POSIX command line.</p>
      <p>TIRA offers a convenient web user interface that allows
participants to self-evaluate their software by
remotecontrolling its execution.</p>
      <p>TIRA allows for evaluating submitted software against
test datasets hosted at server side. Test datasets are
never visible to participants providing for a blind
evaluation, and also allowing for sensitive datasets to be
used for evaluation that cannot otherwise be shared
publicly.</p>
      <p>At the click of a button, the run output of given software
is evaluated against the ground truth of a given dataset.
Evaluation results are stored and made accessible on
TIRA web page as well as for download.</p>
      <p>
        TIRA is widely used as an Evaluation-as-a-Service platform for
experimenting information retrieval tasks [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In particular, the
evaluation platform was used in since the 4th international
competition on plagiarism detection at PAN 2012 [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], and now it
is a common platform for all of PAN shared tasks [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
    </sec>
    <sec id="sec-6">
      <title>4.2 Evaluation Corpus Construction</title>
      <p>
        In this section we describe the methodology for compiling the
Persian Plagdet evaluation corpus used for our shared task. The
corpus comprises cases of simulated, artificial, and real
plagiarism. In general, there are a number of reasons why
collecting only real plagiarism is not sufficient for evaluating
plagiarism detectors. First, collections of real plagiarism that have
been detected manually are usually skewed towards ease of
detection (i.e. the more difficult a plagiarism case is to be
detected, the less likely it will be detected after the fact). Second,
collecting real plagiarism is expensive and time consuming. Third,
a corpus comprising real plagiarism cases cannot be published due
to ethical and legal issues [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Because of these reasons, methods
to artificially create plagiarism, or to simulate plagiarism are often
employed to compile plagiarism corpora. These methods aim at
emulating humans who try to obfuscate their plagiarism by
paraphrasing reused portions of text. An artificial method for
compiling plagiarism corpora includes the use of automatic
paraphrasing technology to obfuscate plagiarized passages.
Simulated passages of plagiarized text are created manually using
human resources and crowdsourcing. Simulated methods yield
more realistic cases of plagiarism compared to artificial ones,
whereas artificial methods are cheaper in terms of both cost and
time and hence scalable.
      </p>
      <sec id="sec-6-1">
        <title>Simulated cases of plagiarism. To create simulated cases of</title>
        <p>plagiarism, a crowdsourcing approach has been used. For this
purpose, a dedicated crowdsourcing platform has been developed,
and a paraphrasing task was designed for crowd workers.
Paraphrased passages obtained via crowdsourcing were reviewed
by experts to ensure quality. All told, about 10% of the
crowdsourced paraphrases were rejected because of poor quality.
Table 1 gives an overview of the demographics of the crowd
workers recruited.</p>
      </sec>
      <sec id="sec-6-2">
        <title>Artificial cases of plagiarism. In addition to simulated</title>
        <p>
          plagiarism based on manual paraphrasing, a large number of
artificially created plagiarism has been constructed for the corpus.
As mentioned above, artificial plagiarism is cheaper and faster to
compile than simulated plagiarism. To create artificial plagiarism,
the previously proposed method of random obfuscation has been
used [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The method consists of random text operations (i.e.
word addition, deletion, shuffling), semantic word variation, and
POS-preserving word shuffling. A composition of these
operations has been used to create low and high degrees of
random obfuscation.
        </p>
        <p>As a result, after the obfuscation of passages extracted from a set
of source documents, the simulated and artificial cases of
plagiarism were inserted into a selection of suspicious documents.
Some key statistics of the plagiarism cases and the final corpus
are shown in the Tables 2 and 3.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>4.3 Performance Measures</title>
      <p>
        The PlagDet measure was used to evaluate the submitted
software. PlagDet is a weighted F-measure that combines
character level precision, recall, and granularity into one metric so
that plagiarism detection systems can be ranked [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The run
output of a given detector lists detected passages of allegedly
plagiarized text as character offsets and lengths. Detection
precision and recall are then computed as shown in Equations 1
and 2 below. In these equations, S is the set of the actual
plagiarism cases and R is the set of detected plagiarism cases:
(
(
)
)
| |
| |
∑
∑
{
|⋃
|⋃
(
| |
(
| |
)|
)|
( )
( )
( )
( )
The granularity measure assesses the capability of a detector to
detect a plagiarism case as a whole as opposed to in several
pieces. The granularity of a detector is defined as follows:
(
)
| |
∑
| |
where S denotes the set of plagiarism cases in the corpus, R
denotes the set of detections reported by a plagiarism detector,
S_R ⊆ S the cases detected by detections in R, and R_S ⊆ R
detections that detect cases in S. Finally, the PlagDet measure is a
combination of F1, the equally-weighted harmonic mean of
precision and recall, and granularity:
(
)
(
(
))
      </p>
    </sec>
    <sec id="sec-8">
      <title>5. SUBTASK 1: TEXT ALIGNMENT</title>
      <p>This section overviews the submitted software and reports on their
evaluation results.</p>
    </sec>
    <sec id="sec-9">
      <title>5.1 Survey of Detection Approaches</title>
      <p>Nine of 12 registered teams successfully submitted a software to
TIRA for the text alignment task. All of the nine participants
submitted working notes describing their approaches. In what
follows, we survey the approaches.</p>
      <p>
        Talebpour et al. [
        <xref ref-type="bibr" rid="ref24">23</xref>
        ] use -trie trees to index the source
documents after preprocessing. The preprocessing steps are text
tokenization, POS tagging, text cleansing, text normalization to
transform text characters into a unique and normal form, removal
of stop words and frequent words, and stemming. Moreover,
FarsNet (the Persian WordNet) [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] is used to find words’
synonyms and synsets. This may allow for detecting cases of
paraphrased plagiarism based on replacing words with their
synonyms. After preprocessing both documents, all of the words
of a source document and their exact positions are inserted into a
trie. After inserting all source documents into a -trie structure, the
suspicious document are iteratively analyzed, checking each word
one by one against the –trie to find potential sources.
      </p>
      <p>
        Minaei et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] employ n-grams as seed heuristic to find
primary matches between suspicious and source documents. Cases
of plagiarism without obfuscation and similar parts of paraphrased
text can be found this way. In order to detect cases of plagiarized
passages, matches closer than a specified threshold are merged.
Finally, to decrease false positive cases, detected cases shorter
than a pre-defined threshold are eliminated.
      </p>
      <p>
        Momtaz et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] use sentence boundaries to split source
and suspicious documents. After text normalization and removal
of stop words and punctuations, sentences of both documents are
turned into graphs, where words represent nodes and an edge is
established between each word and its four surrounding words.
Such graphs obtained from suspicious and source documents are
compared and their similarity computed, whereas sentences of
high similarity are labeled as plagiarism. Finally, to improve
granularity, sentences close to each other are merged to create
contiguous cases of detected plagiarism.
      </p>
      <p>
        Gillam et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] use an approach based on their previous
PAN efforts. The task of finding textual matching is undertaken
without direct using of the textual content. The proposed approach
produces a minimal representation of text by distinguishing
content and auxiliary words. Moreover it produces matchable
binary patterns directly from these dependent words on the
number of classes of interest. Although the approach act similar to
hashing functions, but no effort is taken to prevent collision.
Contrary, hash collision is encouraged over short distances, by
preventing reverse-engineering of the patterns, and uses the
number of coincident matches to indicate the extent of similarity.
      </p>
      <p>
        Mansoorizadeh et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and Ehsan et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] use sentence
boundaries to split source and suspicious documents like the
approach in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In both approaches, each sentence is represented
under the vector space model, using TF-IDF as weighting scheme.
Finally, sentences with cosine similarity greater than a pre-defined
threshold between corresponding vectors are considered as cases
of plagiarism. In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] a subsequent match merging stage improves
performance with respect to granularity. Moreover, overlapping
passages and extremely short passages are removed for the same
reason. The lack of such a merging stage in Mansoorizadeh et
al.’s [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] approach yields high granularity and therefore a poor
PlagDet score.
      </p>
      <p>
        Like most of the submitted software, Esteki et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] split
documents into sentences to detect plagiarism cases. After a
preprocessing phase, which includes normalization, stemming and
stop words removal, a Support Vector Machine (SVM) classifier
is used to separate “similar” sentences non-similar ones. The
Levenshtein distance, the Jaccard coefficient, and the Longest
Common Subsequence (LCS) are used as features extracted from
pairs of sentences. Moreover, synonyms are detected to increase
the likelihood of detecting paraphrased sentences.
      </p>
      <p>
        Gharavi et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] use a deep learning approach to represent
sentences of suspicious and source documents as vectors. For this
purpose, they use Word2Vec to extract words’ vectors and to
compute sentence vectors as average word vectors. The most
similar sentences between pairs of source document and
suspicious document are found using the cosine similarity, the
Jaccard coefficient, reporting them as plagiarism cases.
      </p>
      <p>
        Mashhadirajab et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] use the vector space model (VSM)
with TF-IDF weighting to create sentence vectors from source and
suspicious documents. To gain better results, they use an SVM
neural net to predict the obfuscation type in order to adjust the
required parameters. Moreover, to calculate the semantic
similarity between sentences, FarsNet [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] is used to extract
synsets of terms. Finally, within extension and filtering steps
similar sentences that are close to each other are merged while
passages that either overlap or are too short are removed.
      </p>
    </sec>
    <sec id="sec-10">
      <title>5.2 Evaluation Results</title>
      <p>
        Table 4 shows the overall performance and runtimes of the nine
submitted text alignment approaches. As can be seen, the
approach of Mashhadirajab [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] has achieved the highest PlagDet
score on the complete corpus and is hence ranked highest.
Regarding runtime, the submission of Gharavi [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and Minaei [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
are outstanding: they process the entire corpus in only 1:03 and
1:33 minutes, respectively. Table 5 shows the performance of the
submitted software dependent on obfuscation types in the corpus.
Although, due to the lack of true positives, no performance values
can be computed for the sub-corpus without plagiarism, at least
false positive detections for this sub-corpus influence the overall
performance of participants on the whole corpus [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Gharavi [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
is ranked first in detection performance with highest PlagDet for
“No obfuscation,” and Mashhadirajab [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] achieves best
performance for both “Artificial” and “Simulated” plagiarism.
Among all participants, Mashhadirajab achieves best recall across
all parts of the corpus, whereas Talebpour [
        <xref ref-type="bibr" rid="ref24">23</xref>
        ] and Gharavi [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
outperform it in precision.
      </p>
    </sec>
    <sec id="sec-11">
      <title>6. SUBTASK 2: CORPUS CONSTRUCTION</title>
      <p>This section overviews the five submitted text alignment corpora.
In the first subsection we will have a survey of submitted corpora
and will give a statistical overview of them. In the next subsection
the results of validation and evaluation on the submitted corpora
will be presented.</p>
    </sec>
    <sec id="sec-12">
      <title>6.1 Survey of Submitted Corpora</title>
      <p>
        All of the submitted corpora consist of Persian mono-lingual
plagiarism for the task of text alignment, except for
Mashhadirajab corpus [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] which also contains a set of
crosslingual English-Persian plagiarism cases. All of the corpora are
formatted in accordance with the PAN standard annotation format
for text alignment corpora. In particular, this includes two sets of
documents, namely source documents and suspicious documents,
where the latter are to be analyzed for plagiarism from any of the
source documents. The annotations of plagiarism cases are stored
separately from the text documents within XML documents for
each pair of suspicious and source documents. Therein, each
plagiarism case is annotated as follows:



      </p>
      <sec id="sec-12-1">
        <title>Start position and length of the source passage in the source document Start position and length of the suspicious passage in the suspicious document</title>
      </sec>
      <sec id="sec-12-2">
        <title>Obfuscation type (e.g., indicating to the way that a</title>
        <p>
          source passage has been paraphrased before being added
as suspicious passage to the suspicious documents)
6.1.1 Dataset Overview
Table 6 shows an overview of the submitted text alignment
corpora in terms of the corpus statistics also reported for our
corpus. Mashhadirajab corpus [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] is the biggest one in terms of
number of documents, whereas Abnar corpus contains the largest
number of plagiarism cases. Samim corpus [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] includes larger
documents compared to the other corpora, whereas a large volume
of small documents have been used for construction of the ICTRC
corpus. Samim corpus and the ICTRC corpus comprise the largest
and the smallest plagiarism case, respectively. A variety of
different obfuscation strategies have been employed. No
obfuscation (i.e., exact copy) and artificial obfuscation (random
text operations) are two common strategies.
        </p>
        <p>The length distributions of documents and plagiarized
passages are depicted in Figures 1 and 2. Here, the ICTRC corpus
contains stands out, containing the smallest documents and
plagiarized passages among all submitted corpora. Figure 3 shows
the distribution of the plagiarism ratio per suspicious document.
The ratio of plagiarism per suspicious documents in Samim
corpus is distributed more uniformly compared to the other
submitted corpora. In what follows, the documents used to
compile the corpora as well as the construction approaches are
discussed in detail.
6.1.2 Document Sources
The first step to compile a plagiarism detection corpus is choosing
the documents which will be used as the sets of source documents
and suspicious documents. Many plagiarism detection corpora
intend to simulate plagiarism in technical texts, so that Wikipedia
articles and scientific papers are often employed as source and
suspicious documents sources in these corpora. This also pertains
to the corpora submitted, which mainly employ journal articles
and Wikipedia articles. Wikipedia articles have been used as
resource to compiling the ICTRC corpus and Niknam corpus.
1 Mashhadirajab
2 Gharavi
3 Momtaz</p>
        <sec id="sec-12-2-1">
          <title>4 Minaei</title>
        </sec>
        <sec id="sec-12-2-2">
          <title>5 Esteki</title>
        </sec>
        <sec id="sec-12-2-3">
          <title>6 Talebpour</title>
        </sec>
        <sec id="sec-12-2-4">
          <title>7 Ehsan</title>
        </sec>
        <sec id="sec-12-2-5">
          <title>8 Gillam</title>
        </sec>
        <sec id="sec-12-2-6">
          <title>9 Mansourizadeh</title>
          <p>
            Niknam used 3000 documents larger than 4000 characters, and
ICTRC used about 6000 documents larger than 1500 characters.
Abnar used texts from a set of novels that were translated to
Persian. Despite the genre of books, the documents found in the
corpus are not as large as might be expected. Mashhadirajab [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ]
and Samim [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ] used scientific papers to compile their corpora.
Mashhadirajab used a combination of Wikipedia articles (40%),
articles from the Computer Society of Iran Computer Conference
(CSICC) (13%), theses available in online (13%) and Persian
open access articles (34%). Samim also collected Persian open
access papers from peer reviewed journals to compile their text
alignment corpus. The papers used include papers from the
humanities (57%), science (25%), veterinary science (10%) and
other related subjects (8%).
6.1.3 Obfuscation Synthesis
The second step in compiling a plagiarism detection corpus is to
obfuscate passages selected from source documents and then
insert them into suspicious documents. Obfuscating text passages
aims at emulating plagiarism cases whose authors try to conceal
the fact their plagiarized, making it more difficult for human
reviewers and plagiarism detection systems alike to identify the
plagiarized passages afterwards. As discussed above, creating
obfuscated plagiarism manually is laborious and expensive, so
that most participants resorted to automatic obfuscation methods.
It is remarkable that two of the corpora (the ones of
Mashhadirajab and ICTRC) comprise plagiarism that has been
manually created. Otherwise, a variety of different approaches
have been employed for obfuscation (see Table 6, rows
“Obfuscation type”). All of the submitted corpora also contain a
portion of plagiarized passages without any obfuscation to
simulate verbatim copying.
          </p>
          <p>Niknam employed a set of text operations consisting of addition,
deletion and shuffling of words, replacing words with their
synonyms and POS-preserving word replacement. Similar
obfuscation strategies have been used to compile Samim’s corpus.
It contains “Random Text Operations” and “Semantic Word
Variation” in addition to “No obfuscation.” In addition to these
obfuscation types, the authors of the ICTRC corpus used a
crowdsourcing platform for paraphrasing test passages. About 30
people of various ages, both genders, and different levels of
education have participated in the paraphrasing process. Abnar’s
corpus comprises obfuscation approaches such as replacing words
with synonyms, shuffling sentences, circular translation, and a
combination of the aforementioned ones. The circular translation
approach includes translating the text to an intermediate language
and then translating it back to the original one, hoping that the
resulting text will significantly differ from the original one while
maintaining its meaning. From a diversity point of view,
Mashhadirajab’s corpus contains the most variety in terms of
obfuscation. In addition to artificial and simulated cases, they
used summarizing cyclic translation and text manipulation
approaches to create cases of plagiarism. Moreover, the corpus
comprises also cross-lingual plagiarism where source documents
have been translated to Persian using manual and automatic
translation.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-13">
      <title>6.2 Corpus Validation</title>
      <p>In order to validate the submitted corpora, we analyzed them
quantitatively and qualitatively. For the latter, samples have been
drawn from each corpus and obfuscation type for manual review.
The review involved of validating the plagiarism annotations,
such as offsets and lengths of annotated plagiarism in both source
and suspicious documents. Moreover, the suspicious passage and
its corresponding source have been checked manually to observe
the impact of different obfuscation strategies as well as the level
of obfuscation. Altogether, no important issues have been found
among the studied samples during peer-review.</p>
      <p>In addition to manual review, we also analyzed the corpora
quantitatively: Figures 1 and 2 depict the length distributions of
the documents and the plagiarism cases in the corpora. Both
Abnar’s corpus and the ICTRC corpus have clear expected values,
whereas the other corpora are more evenly distributed. Figure 3
depicts the ratio of plagiarism per document, showing that the
ratios are quite unevenly distributed across corpora; Niknam’s
corpus and the ICTRC corpus comprise mostly suspicious
documents with a small ratio of plagiarism. Figures 4 and 5 show
the distribution of plagiarized passages in terms of where they
start within suspicious documents (i.e., their character offset), and
where they start within source documents. The distributions of
start offsets within suspicious documents are similar across all
corpora with a negative bias against offsets at the beginning of a
suspicious document (see Figure 4). The distributions are also
similar for the start offsets within source documents with one
notable exception: the source passages of Samim’s corpus have
almost always been chosen from the same offsets of source
documents which is a clear bias and may allow for trivial
detection.</p>
      <p>
        Finally, we analyzed the plagiarized passages in the submitted
corpora with regard to their similarity between source passage and
suspicious passage. The experiment consists of comparing source
passages with suspicious passages using 10 retrieval models. Each
model is an n-gram vector space model (VSM), where n ranges
from 1 to 10 words, employing stop word removal, TF-weighting
and the cosine similarity [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. For high-quality corpora, a pattern
similar to that of PAN corpora is expected.
      </p>
      <p>Since there are many obfuscate types to choose from, we only
compare a selection: the simulated plagiarism cases of
Mashhadirajab and ICTRC are compared to the PAN corpora
(Figure 6). Moreover, the artificial parts of all corpora are
compared to each other (Figure 7). Abnar’s corpus is omitted
since it lacks artificial obfuscation. Almost all of the corpora show
same patterns of similarity for different ranges of n, except the
Mashhadirajab’s corpus which has a higher range of similarity in
comparison others.</p>
    </sec>
    <sec id="sec-14">
      <title>6.3 Corpus Evaluation</title>
      <p>Exploiting the virtues of TIRA, our final experiment was to run
the nine submitted detection approaches on the five submitted
corpora, providing for a first impression on how difficult it is to
detect plagiarism within these corpora. Table 7 overviews the
results of this experiment. Unfortunately, not all submitted
approaches succeeded in processing all corpora. One reason was
scalability issues: since some of the submitted corpora are
significantly larger than our evaluation corpus, it seems
participants did not pay a lot of attention to scalability. The
approaches of Talebpour, Mashhadi, and Gillam failed to process
the corpora in time. The approaches of Momtaz and Esteki failed
to process some of the corpora at first, the results of the former are
only partially reliable to date, whereas the latter of which could be
fixed in time. This shows that submitting datasets to shared tasks
presents its own challenges. Participants will be invited to fix their
software to make it work on all corpora, so that further results
may become available after publication of this paper, e.g., on
TIRA’s web page. Considering the detection performance, it can
be seen that the PlagDet scores are generally lower compared to
our corpus, except for the ICTRC corpus, where the same
performance scores have been reached. This shows that the
submitted corpora present their own challenges, rendering them
more difficult, and presenting future researchers with new
opportunities for contributions.</p>
      <p>Given the results from all our experiments, the submitted
corpora are of reasonable quality. Although some of them are too
easy to be solved and comprise a biased sample of plagiarism
cases, the diversity of corpora ensures that future evaluations can
be done with confidence as long as all available datasets are
employed.</p>
    </sec>
    <sec id="sec-15">
      <title>7. CONCLUSION</title>
      <p>In conclusion, our shared task has attracted considerable attention
from the community of scientists working on plagiarism
detection. The shared task has served as a means to establish a
new state of the art in performance evaluation for Persian
plagiarism detection. Altogether six new evaluation corpora are
available now, and nine detection approaches have been evaluated
on them. The results show that Persian plagiarism detection is far
from being a solved problem. In addition, our contributions
broaden the scope of the text alignment task which has been
studied mostly for English until now. This may allow future work
on plagiarism detection approaches that work on both languages
simultaneously.</p>
    </sec>
    <sec id="sec-16">
      <title>8. ACKNOWLEDGMENTS</title>
      <p>This work has been funded by ICT Research Institute, ACECR,
under the partial support of Vice Presidency for Science and
Technology of Iran - Grant No. 1164331. The work of Paolo
Rosso has been partially funded by the SomEMBED MINECO
TIN2015-71147-C2-1-P research project and by the Generalitat
Valenciana under the grant ALMAMATER
(PrometeoII/2014/030). We would like to thank the participants of
the competition for their dedicated work. Our special thanks go to
the renowned experts who served on the organizing committee for
their contributions and devoted work to make this shared task
possible. We would like to thank Javad Rafiei and Khadijeh
Khoshnava for their help in construction of evaluation corpus. We
are also immensely grateful to Vahid Zarrabi for his comments
and valuable help along the way which greatly assisted this
challenging shared task.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Bensalem</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boukhalfa</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abouenour</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darwish</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Chikhi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Overview of the AraPlagDet PAN@ FIRE2015 Shared Task on Arabic Plagiarism Detection, CEUR-WS.org</article-title>
          , vol.
          <volume>1587</volume>
          , pp.
          <fpage>111</fpage>
          -
          <lpage>122</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Ehsan</surname>
            ,
            <given-names>N</given-names>
          </string-name>
          , Shakery,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>A Pairwise Document Analysis Approach for Monolingual Plagiarism Detection</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Esteki</surname>
            ,
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Safi</given-names>
            <surname>Esfahani</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>A Plagiarism Detection Approach Based on SVM for Persian Texts</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Gharavi</surname>
            ,
            <given-names>E</given-names>
          </string-name>
          , Bijari, k, Zahirnia,
          <string-name>
            <surname>K</surname>
          </string-name>
          , Veisi,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>A Deep Learning Approach to Persian Plagiarism Detection</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Gillam</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Vartapetiance</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <year>2016</year>
          . From English to Persian:
          <article-title>Conversion of Text Alignment for Plagiarism Detection</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burrows</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2012</year>
          ,
          <string-name>
            <surname>August.</surname>
          </string-name>
          <article-title>First experiences with TIRA for reproducible evaluation in information retrieval</article-title>
          .
          <source>In SIGIR</source>
          (Vol.
          <volume>12</volume>
          , pp.
          <fpage>52</fpage>
          -
          <lpage>55</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Burrows</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2012</year>
          ,
          <article-title>August. Ousting ivory tower research: towards a web framework for providing experiments as a service</article-title>
          .
          <source>In Proceedings of the 35th international ACM SIGIR conference on Research and development in information retrieval</source>
          (pp.
          <fpage>1125</fpage>
          -
          <lpage>1126</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burrows</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Hoppe</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>September. TIRA: Configuring, executing, and disseminating information retrieval experiments</article-title>
          .
          <source>In 2012 23rd International Workshop on Database and Expert Systems Applications</source>
          (pp.
          <fpage>151</fpage>
          -
          <lpage>155</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Hopfgartner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Müller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mercer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalpathy-Cramer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krithara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Balog</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <year>2015</year>
          ,
          <string-name>
            <surname>June.</surname>
          </string-name>
          <article-title>Report on the Evaluation-as-a-Service (EaaS) expert workshop</article-title>
          .
          <source>In ACM SIGIR Forum</source>
          (Vol.
          <volume>49</volume>
          , No.
          <issue>1</issue>
          , pp.
          <fpage>57</fpage>
          -
          <lpage>65</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Khoshnavataher</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zarrabi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohtaj</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Asghari</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Developing Monolingual Persian Corpus for Extrinsic Plagiarism Detection Using Artificial Obfuscation. Notebook for PAN at CLEF 2015</article-title>
          . CLEF (Working Notes).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Mansoorizadeh</surname>
            ,
            <given-names>M</given-names>
          </string-name>
          , Rahgooy,
          <string-name>
            <surname>T.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Persian Plagiarism Detection Using Sentence Correlations</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Mashhadirajab</surname>
            ,
            <given-names>F</given-names>
          </string-name>
          , Shamsfard,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>A Text Alignment Algorithm Based on Prediction of Obfuscation Types Using SVM Neural Network</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Mashhadirajab</surname>
            ,
            <given-names>F</given-names>
          </string-name>
          , Shamsfard,
          <string-name>
            <surname>M</surname>
          </string-name>
          , Adelkhah,
          <string-name>
            <surname>R</surname>
          </string-name>
          , Shafiee,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Saedi</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>A Text Alignment Corpus for Persian Plagiarism Detection</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Minaei</surname>
            ,
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Niknam</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>An n-gram based Method for Nearly Copy Detection in Plagiarism Systems</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Momtaz</surname>
            ,
            <given-names>M</given-names>
          </string-name>
          , Bijari,
          <string-name>
            <surname>K</surname>
          </string-name>
          , Salehi,
          <string-name>
            <surname>M</surname>
          </string-name>
          , Veisi,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Graphbased Approach to Text Alignment for Plagiarism Detection in Persian Documents</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eiselt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barron</surname>
            , Cedeno,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>2009</year>
          .
          <article-title>Overview of the 1st international competition on plagiarism detection</article-title>
          .
          <source>In 3rd PAN Workshop</source>
          . Uncovering Plagiarism,
          <source>Authorship and Social Software Misuse</source>
          (p.
          <fpage>1</fpage>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>2010</year>
          ,
          <string-name>
            <surname>August.</surname>
          </string-name>
          <article-title>An evaluation framework for plagiarism detection</article-title>
          .
          <source>In Proceedings of the 23rd international conference on computational linguistics: Posters</source>
          (pp.
          <fpage>997</fpage>
          -
          <lpage>1005</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Graßegger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oberländer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tippmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2012</year>
          .
          <article-title>Overview of the 4th International Competition on Plagiarism Detection</article-title>
          . In CLEF (Online Working Notes/Labs/Workshop).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2014</year>
          ,
          <article-title>September. Improving the Reproducibility of PAN's Shared Tasks</article-title>
          .
          <source>In International Conference of the Cross-Language Evaluation Forum for European Languages</source>
          (pp.
          <fpage>268</fpage>
          -
          <lpage>299</lpage>
          ). Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Göring</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2015</year>
          .
          <article-title>Towards data submissions for shared tasks: first experiences for the task of text alignment</article-title>
          .
          <source>Working Notes Papers of the CLEF</source>
          , pp.
          <fpage>1613</fpage>
          -
          <lpage>0073</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Rezaei</given-names>
            <surname>Sharifabadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Eftekhari</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. A.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Mahak Samim: A Corpus of Persian Academic Texts for Evaluating Plagiarism Detection Systems</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Shamsfard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Ontology for Persian</article-title>
          . WordNet conference.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Developing FarsNet</surname>
          </string-name>
          :
          <article-title>A lexical Proceedings of the 4th global</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Talebpour</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          , Shirzadi,
          <string-name>
            <given-names>M</given-names>
            ,
            <surname>Aminolroaya</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Plagiarism Detection based on a Novel Trie-based Approach</article-title>
          , In Working notes of FIRE 2016 -
          <article-title>Forum for Information Retrieval Evaluation, Kolkata</article-title>
          , India, December 7-
          <issue>10</issue>
          ,
          <year>2016</year>
          , CEUR Workshop Proceedings, CEUR-WS.org.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>