<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cross-Language Urdu-English (CLUE) Text Alignment Corpus</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, COMSATS Institute of Information Technology (Wah &amp; Lahore Campuses)</institution>
          ,
          <country country="PK">Pakistan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Israr Hanif, Rao Muhammad Adeel Nawab</institution>
          ,
          <addr-line>Affiffa Arbab, Huma Jamshed, Sara Riaz, and Ehsan Ullah Munir</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>Plagiarism is well known problem of the day. Easy access to print and electronic media and ready to use material made it easy to reuse the existing text in new document. The severity of the problem is much reduced in monolingual context by the automated and tailored effort made by the research community but the issue is yet not properly addressed in cross language (CL) text reuse. Any story or article written in any source language like Urdu is simply translated in target language like English and translator claims it as his own. Availability of standard and simulated resource address the issue and act as test bed for analyzing and implementing available plagiarism detection approaches. The research work is aimed at enriching the available cross- language corpus and on the other hand providing a benchmark corpus to Cross Language Plagiarism (CLP) domain.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Text reuse is the process of developing a new document using the data of existing
documents. Plagiarism is a most familiar type of text reuse. In general, plagiarism is
considered as reuse of thoughts, procedures, outcomes, or words without clearly showing
the original source. The size of text that is reused varies from case to case. In some
conditions authors use only phrases, sentences or passages to create new document while
in some conditions, word by word document is reused to create a new document. To
create a new document data can be collected from different source documents. In some
conditions entire document of original text is reused to create new document. Possible
ways to detect plagiarism are (1) Intrinsic Plagiarism Detection- indicating whether all
passages written by single author and (2) Extrinsic Plagiarism Detection- pointing all
sources from where passages are used to create the suspicious document [18].</p>
      <p>
        Plagiarism has crossed the language boundaries now like Urdu to English or any.
Translational technologies are giving new ways of plagiarism, known as cross language
plagiarism (CLP). In cross language plagiarism, source material is translated from one
language to another and then translated data is reused to develop a new document
without giving references of the original source. Generally such unattributed text reuse is
also labeled as plagiarism [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In this type of plagiarism only language change occurs
such as from Urdu to English or vice versa. That’s why cross language plagiarism is
also called translation plagiarism. Barron-Cedeno also defines CLP as a piece of text in
one language translated into a target language while keeping the content and semantics
same without referring the origin [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>Availability of ready to use data in different formats and in multiple languages on
internet is also boosting the case of CLP. Student assignments, and newspaper stories
and articles are hot domains for CLP as education and information has no barriers and
boundaries. CLP needs to develop a benchmark corpus having source and target
language document pairs to detect any level of plagiarism.</p>
      <p>
        Urdu is a language with more than 100 million native speakers1. Few corpuses
are developed for cross-language information retrieval (CLIR) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] but no serious effort
has been made to address CLP problem. English is an official and almost educational
language in indo-Pak region. This diversity raised the CLP issues with more potential
in this region especially in higher education sector [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Therefore developing an
UrduEnglish corpora for CLP detection is much needed area to be focused.
      </p>
      <p>This research is aimed at generating a standard corpus in Urdu-English language
pairs. The corpus will serve as base for CLP detection and analyzing multiple
evaluation techniques in context of performance. Three levels of plagiarism (Near Copy,
Light Revision, and Heavy Revision) enabled it to detect plagiarism at different levels.
Automated and manual effort to generate suspicious document made the corpus more
realistic and precise.</p>
      <p>The rest of the paper is organized as follows. Section 2 summarizes the related
work. Section 3 describes corpus generation process in detail. Analysis about corpus is
presented in section 4. Finally, section 5 concludes the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Generation of corpus using simulated and artificial approach as recommended by
Potthast et al. is in practice now [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Clough and Stevenson in 2011 created a short answer
corpus which contains plagiarized examples generated based on simulated format [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
Similar effort was made by stein et al. for PAN-PC-09 [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and by Potthast et al. for
PAN-PC-10 corpus [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>In spite of the fact that research community is addressing the plagiarism issue
potentially, it is majorly yet limited to monolingual aspect. The minor effort made in cross
lingual aspect of the problem is also limited to few European languages like Spanish
and German as source and English as suspicious in source-suspicious language pairs.</p>
      <p>
        Different cross lingual corpora like English-Spanish [19] and English-German
corpus [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and many others have been developed for detection and
analysis in this domain. New PAN@FIRE tasks like (CL!NSS) is an effort to trace
similar news stories across the languages [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In 2009, European Commissions office
for official publications (OPOCE) created a corpus for cross language research.
CrossLanguage Indian Text Reuse Competition corpus is a standard corpus in English-Hindi
language pair perspective [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Wikipedia articles were selected as source in computer
science and tourism with 112 documents as source and 276 suspicious documents for
different levels of plagiarized fragments. Along with corpus creation, applying
plagiarism detection approaches on newly created and already available corpuses is also in
practice. The JRC-Acquis Multilingual Parallel Corpus was used by Potthast et al. to
apply CLP detection approaches. 23,564 documents, extracted from legal documents of
European Union, incorporate the corpus [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Out of 22 languages in legal document
collection, only 5 including French, Germen, Polish, Dutch and Spanish was selected to
generate source-suspicious language pair with English language as source.
Comparable Wikipedia Corpus is another example of experimenting with similar approach. The
corpus contains 45,984 documents.
      </p>
      <p>
        Applying CLP detection approaches on multiple corpora have also been done by
Ceska et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Two corpuses JRC-EU and Fairy-tale Corpus were used for the
purpose. JRC-EU composed of 400 documents randomly extracted from legislation reports
of European Union. Out these 400 documents, 200 were in English as source and
remaining 200 were in Czech. Fairy-tale Corpus with 54 documents out of which 27 in
English and 27 in Czech translated from English, was the part of experiment.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Corpus Generation Process</title>
      <p>For the PAN 2015 Text Alignment task, we submitted a cross-language corpus
(UrduEnglish language pair) for evaluating the performance of CLP detection system. The
CLUE corpus contains simulated cases of plagiarism (source fragments are in English
and suspicious ones in English).
3.1</p>
      <p>Generation of Source-Suspicious Fragment Pairs
To generate source-suspicious fragment pairs, we collected source texts from two
domains: (1) computer science and (2) general essay topics. All the source fragments were
collected from Wikipedia (http://ur.wikipedia.org/wiki/urdu in footnote). It is likely that
the amount of text reused for plagiarize may vary from a phrase, sentence, paragraph to
entire document. Therefore, the source fragments were divided into three categories: (1)
small (less than 50 words), (2) medium (50-100 words) and (3) large (100-200) words.
Table 1 shows the distribution of source-suspicious fragment Paris.</p>
      <p>To generate simulated cases of plagiarism participants (volunteers), who were
university students (undergraduate and postgraduate) were asked to rewrite the source
fragment (in Urdu) to generate the plagiarized fragment (in English) using one of the three
methods.</p>
      <p>i. Near Copy: Participants were told to automatically translate the source fragment
to generate the plagiarized fragment.
ii. Light Revision: Using this approach, the plagiarized fragment was created in two
steps. In the first step source fragment (in Urdu) is automatically translated into
English. In the second step, the translated fragment is passed through an automatic
text rewriting tool called Article Rewriter1 to generate the plagiarized fragment
(i.e. light revision of the source fragment).
iii. Heavy Revision: Participants were instructed not to use the automatic machine
translation tools for generating heavy revision of the source text. Instead, they were
asked to manually translate the original source text in such a way that it looks like
a paraphrased copy of the source text.
The proposed corpus contains total 1000 documents (500 source documents (in Urdu)
and 500 suspicious documents (in English)). All the documents in the corpus are
collected from freely available online resources. A document in the corpus belongs to
the domain of computer science or general topics. Computer science topics (Total 50)
mainly includes: Free software, Open Source, Binary Numbers, Database
Normalization, Artificial intelligence, Robotics, Mobile Apps, Yahoo, MSN, Google, Whatsapp,
Android, twitter, Facebook, RUBY language, Gmail, Skype, Daily motion, HTML and
few others. General domain topics were also same in count and mainly include: Global
warming, Muhammad Iqbal, Capitalism, Bookselling, Mosque, Pakistan Air Force,
Two-Nation theory, Cricket, Fashion, Capitalism, Lahore Forte, Badshahi Masjid,
Globalization and few others. Out of 500 suspicious documents, 270 are plagiarized and
remaining 230 are non-plagiarized. Only one source-plagiarized fragment pair was
inserted into one source-suspicious document pair. Computer science source-plagiaries
fragment pairs were inserted into computer science source-suspicious documents and
similarly source-plagiarized fragment pairs on general topics were inserted into
sourcesuspicious document pairs which belonged to the domain of general topics.</p>
      <p>Out of 270 source-plagiarized fragment pairs, 180 are from Computer Science
domain and 90 from General topics domain.</p>
      <p>All the source-plagiarized fragment pairs were randomly inserted into source-
suspicious document pairs.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Analysis and Discussion</title>
      <p>The developed corpus is divided into source and suspicious documents. Although the
manual revision is done on each source fragment to generate its NC, LR and HR
version but the order of sentences was kept same. Manual revision was done to overcome
issues generated by automatic translation tools outcome. Providing Source (Urdu)
version to participant for generating its Heavy Revision (HR) made the plagiarized text
more realistic.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The paper describes the corpus creation process for detection of plagiarism in cross
language domain of Urdu-English pairs. The corpus can be used as benchmark or test
bed for upcoming tasks of performance evaluation among different plagiarism detection
techniques. In future we intend to increase the size of corpus.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Peer Review</title>
      <p>Following data sets were observed and most of the xml features including length and
offset of the fragment inserted in source and suspicious documents were found correct
in all data sets. A mismatch was also found in few cases due to newline and some
special characters. Dataset wise other findings are described as:
– cheema15-training-dataset-english</p>
      <p>Different folders are used to consider cases of plagiarism at undergrad, Master and
Ph. D levels. Fragments are inserted at character level at random places. Source to
suspicious ratio is on to one as single source fragment is used to make a document
suspicious. Obfuscation strategy is almost paraphrasing with good quality
Pair Entry / Example
suspicious-document0099source-document0391.xml
suspicious-document0259source-document0189.xml
suspicious-document0309source-document0321.xml
suspicious-document0386source-document0186.xml
suspicious-document0485source-document0447.xml</p>
      <p>Type /Artificial / Simulated
Simulated</p>
      <p>Quality of Plagiarism</p>
      <p>Well paraphrased
Simulated
Simulated
Simulated
Simulated</p>
      <p>Good
Well paraphrased
Well paraphrased
– palkovskii15-training-dataset-english</p>
      <p>Multilingual features although described but obfuscation is limited to English only.
Fragments are inserted at word level at random places in suspicious document.</p>
      <p>Source and suspicious documents are of large size and from general domain.
Pair Entry / Example
suspicious-document00021source-document02467.xml
suspicious-document00067source-document02563.xml
suspicious-document00081source-document03075.xml
suspicious-document00380source-document00270.xml
suspicious-document00407source-document02140.xml</p>
      <p>Type /Artificial / Simulated
Artificial</p>
      <p>Quality of Plagiarism</p>
      <p>NEAR COPY
Artificial
Artificial
translation-chain
translation-chain
– mohtaj15-training-dataset-english</p>
      <p>Multiple fragments are inserted in single document at random places. In most of
the cases 3 fragments are inserted at character level. Placement of fragments is at
random places in source and suspicious documents. Although in few cases
fragments in source and suspicious documents were found irrelevant but dataset is well
composed overall. Large sized documents from general domain are used.</p>
      <p>Type /Artificial / Simulated
Artificial</p>
      <p>Quality of Plagiarism
NEAR COPY
Pair Entry / Example
suspicious-document110926source-document307308.xml
suspicious-document179883source-document517709.xml
suspicious-document235057source-document534046.xml
suspicious-document102450source-document106487.xml
suspicious-document405184source-document26685.xml
suspicious-document105415source-document149775.xml
suspicious-document157936source-document198805.xml</p>
      <p>Artificial
Artificial
Artificial
Simulated
Artificial</p>
      <p>Simulated
– kong15-training-dataset-chinese</p>
      <p>Same text is used to suspect many documents. Small sized dataset with only 4
suspicious and 78 source documents. Suspicious text is inserted at consecutive
locations probably at character level. Both source and suspicious documents are in
Chinese but documents also have large English text in few cases. Quality of
plagiarism cannot be judged.
Poor
Good
Good
Good
– khoshnava15-training-dataset-persian</p>
      <p>A data set with 720 suspicious and 802 source documents. Almost one-to-one
source to suspicious ratio is there. Artificial type of plagiarism cases with no
obfuscation strategy mostly. Both source and suspicious documents are in Persian
therefore quality of plagiarism cannot be judged.
– Asghari15-training-dataset-english-persian</p>
      <p>Large data set with 15959 source and 5470 suspicious documents. Most of the
Plagiarism cases are artificially generated. Due to English to Persian nature quality
of plagiarism cannot be judged properly. Formation of dataset is fine.
– alvi15-training-dataset-english</p>
      <p>A data set with 70 source and 90 suspicious documents. Three types of
obfuscation strategies are used: character substitution, synonym replacement and human
retelling. One source fragment is used in different obfuscation strategies to suspect
the suspicious document. Insertion is at sentence level and almost near copy or
exact copy of source fragment is used in suspicious documents. There is some
difference in the source length, source offset, suspicious length and suspicious offset
because of new line character.</p>
      <p>Pair Entry / Example Type /Artificial / Simulated
suspicious-document00003.txt Retelling
source-document00002.txt
suspicious-document00043- Retelling
source-document00018.xml
suspicious-document00102- Retelling
source-document00040.xml
suspicious-document00128- Automatic
source-document00078.xml
suspicious-document00039- character-substitution
source-document00010.xml
suspicious-document00078- character-substitution
source-document00020.xml
suspicious-document00099- character-substitution
source-document00025.xml
Quality of Plagiarism
Good
Good
Good
Good
Well paraphrased
Well paraphrased
Well paraphrased</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>We are thankful to all volunteers for their valuable contribution in construction of this
corpus.
18. Stein, B., zu Eissen, S.M., Potthast, M.: Strategies for retrieving plagiarized documents. In:
Proceedings of the 30th annual international ACM SIGIR conference on Research and
development in information retrieval. pp. 825–826. ACM (2007)
19. Stein, B., Rosso, P., Stamatatos, E., Koppel, M., Agirre, E.: 3rd pan workshop on
uncovering plagiarism, authorship and social software misuse. In: 25th Annual Conference
of the Spanish Society for Natural Language Processing (SEPLN). pp. 1–77 (2009)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Barrón-Cedeno</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Devi</surname>
            ,
            <given-names>S.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stevenson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Pan@ fire: Overview of the cross-language! ndian text re-use detection competition</article-title>
          .
          <source>In: Multilingual Information Access in South Asian Languages</source>
          , pp.
          <fpage>59</fpage>
          -
          <lpage>70</lpage>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Barrón-Cedeno</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pinto</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>On cross-lingual plagiarism analysis using a statistical model</article-title>
          .
          <source>In: PAN</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ceska</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jezek</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Multilingual plagiarism detection</article-title>
          .
          <source>In: Artificial Intelligence: Methodology, Systems, and Applications</source>
          , pp.
          <fpage>83</fpage>
          -
          <lpage>92</lpage>
          . Springer (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stevenson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Developing a corpus of plagiarised short answers</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>45</volume>
          (
          <issue>1</issue>
          ),
          <fpage>5</fpage>
          -
          <lpage>24</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stevenson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Pan@ fire: Overview of the cross-language! ndian news story search (cl! nss) track</article-title>
          . In:
          <article-title>Forum for Information Retrieval Evaluation</article-title>
          ,
          <string-name>
            <surname>ISI</surname>
          </string-name>
          , Kolkata, India (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Judge</surname>
          </string-name>
          , G.:
          <article-title>Plagiarism: Bringing economics and education together (with a little help from it). Computers in Higher Education Economics Reviews (Virtual edition</article-title>
          )
          <volume>20</volume>
          ,
          <fpage>21</fpage>
          -
          <lpage>26</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Littman</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landauer</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          :
          <article-title>Automatic cross-language information retrieval using latent semantic indexing</article-title>
          .
          <source>In: Cross-language information retrieval</source>
          , pp.
          <fpage>51</fpage>
          -
          <lpage>62</lpage>
          . Springer (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Plagiarism: a misplaced emphasis</article-title>
          .
          <source>Journal of Information Ethics</source>
          <volume>3</volume>
          (
          <issue>2</issue>
          ),
          <fpage>36</fpage>
          -
          <lpage>47</lpage>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eiselt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the 2nd international competition on plagiarism detection</article-title>
          . In: CLEF (Notebook Papers/LABs/Workshops) (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eiselt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the 3rd international competition on plagiarism detection</article-title>
          .
          <source>In: Notebook Papers of CLEF 11 Labs and Workshops</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Cross-language plagiarism detection</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>45</volume>
          (
          <issue>1</issue>
          ),
          <fpage>45</fpage>
          -
          <lpage>62</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , OberlÂ´lander, A.,
          <string-name>
            <surname>Tippmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the 4th international competition on plagiarism detection</article-title>
          . In: CLEF (Online Working Notes/Labs/Workshop) (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Improving the reproducibility of panâA˘ Z´s shared tasks</article-title>
          .
          <source>In: Information Access Evaluation</source>
          . Multilinguality, Multimodality, and Interaction, pp.
          <fpage>268</fpage>
          -
          <lpage>299</lpage>
          . Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beyer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Busse</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tippmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the 6th international competition on plagiarism detection</article-title>
          . In: CLEF (Online Working Notes/Labs/Workshop) (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tippmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the 5th international competition on plagiarism detection</article-title>
          . In: CLEF (Online Working Notes/Labs/Workshop) (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>An evaluation framework for plagiarism detection</article-title>
          .
          <source>In: Proceedings of the 23rd international conference on computational linguistics: Posters</source>
          . pp.
          <fpage>997</fpage>
          -
          <lpage>1005</lpage>
          . Association for Computational Linguistics (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eiselt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the 1st International Competition on Plagiarism Detection</article-title>
          . In: Stein,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Koppel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Agirre</surname>
          </string-name>
          , E. (eds.) SEPLN 09 Workshop on Uncovering Plagiarism, Authorship, and
          <source>Social Software Misuse (PAN 09)</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          . CEUR-WS.
          <source>org (Sep</source>
          <year>2009</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-502
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>