<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Supervised Ranking for Plagiarism Source Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kyle Williamsy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hung-Hsuan Chenz</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>C. Lee Gilesy</string-name>
          <email>giles@ist.psu.edu</email>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>1021</fpage>
      <lpage>1026</lpage>
      <abstract>
        <p>Source retrieval involves making use of a search engine to retrieve candidate sources of plagiarism for a given suspicious document so that more accurate comparisons can be made. We describe a strategy for source retrieval that uses a supervised method to classify and rank search engine results as potential sources of plagiarism without retrieving the documents themselves. Evaluation shows the performance of our approach, which achieved the highest precision (0.57) and F1 score (0.47) in the 2014 PAN Source Retrieval task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        The advent of the Web has had many benefits in terms of access to information.
However, it has also made it increasingly easy for people to plagiarize. For instance, in a
study from 2002-2005, it was found that 36% of undergraduate college students
admitted to plagiarizing [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Given the prevalence of plagiarism, there has been significant
research and systems built for plagiarism detection [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>There is an inherent assumption in most plagiarism detection systems that potential
sources of plagiarism have already been identified that can be compared to a
suspicious document. For small collections of documents, it may be reasonable to perform
a comparison between the suspicious document and every document in the collection.
However, this is infeasible for large document collections and on the Web. The source
retrieval problem involves using a search engine to retrieve potential sources of
plagiarism for a given suspicious document by submitting queries to the search engine and
retrieving the search results.</p>
      <p>
        In this paper we describe our approach to the source retrieval task at PAN 2014.
The approach builds on our previous approach in 2013 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], but differs by including a
supervised search result ranking strategy.
Our approach builds on our previous approach to source retrieval [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], which achieved
the highest F1 score at PAN 2013 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The current approach involves four main steps:
1) query generation, 2) query submission, 3) result ranking and 4) candidate retrieval.
The key difference between our approach this year and our previous approach is in the
result ranking step (step 3). Previously, we used a simple method whereby we re-ranked
the results returned by the search engine based on the similarity of each result snippet
and the suspicious document. This year, we make use of a supervised method for result
ranking. Beyond that, the approach remains the same as in 2013. In the remainder of
this section, we describe the approach while focusing on the supervised result ranking.
2.1
      </p>
      <sec id="sec-1-1">
        <title>Query Generation</title>
        <p>
          To generate queries, the suspicious document is partitioned into paragraphs with each
paragraph consisting of 5 sentences as tagged by the Stanford Tagger [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Stop words
are then removed and each word is tagged with its part of speech (POS). Following
previous work [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], only words tagged as nouns, verbs and adjectives are retained
while all others are discarded. Queries are then created by concatenating the remaining
words to form queries consisting of 10 words each.
2.2
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>Query Submission</title>
        <p>Queries are submitted in batches consisting of 3 queries and the results returned by
each query are combined to form a single set of results. The intuition behind submitting
queries in batches is that the union of the results returned by the three queries is more
likely to contain a true positive than the results returned by a query individually. For
each query, a maximum of 3 results is returned resulting in a set of at most 9 results. The
intuition behind submitting 3 queries is that they likely capture sufficient information
about the paragraph.
2.3</p>
      </sec>
      <sec id="sec-1-3">
        <title>Result Ranking</title>
        <p>We assume that the order of search results as produced by the search engine does not
necessarily reflect the probability of the result being a source of plagiarism. We thus
infer a new ranking of the results based on a supervised method. This is a major
component of our approach and will be discussed in Section 3.
2.4</p>
      </sec>
      <sec id="sec-1-4">
        <title>Candidate Document Retrieval</title>
        <p>Having inferred a new ordering of the results, they are then retrieved in the new ranked
order. For each result retrieved, the PAN Oracle is consulted to determine if that result
is a source of plagiarism for the input suspicious document. If it is, then candidate
retrieval is stopped and the whole process is repeated on the next paragraph. If it is not,
then the next document is retrieved. Furthermore, a list is maintained of the URLs of
each document retrieved and used to prevent a URL from being retrieved more than
once since this does not improve precision or recall.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Supervised Result Ranking</title>
      <p>The main difference between our current approach and our approach in the previous
year is that now we make use of a supervised method. Specifically, we train a search
result classifier that classifies each search result as either being a potential source of
plagiarism or not. The classifier only makes use of features that are available at search
time and without retrieving a search result unless it is classified as a potential source
of plagiarism. Furthermore, an ordering of the positively classified search results is
produced based on the probabilities output by the classifier.
3.1</p>
      <sec id="sec-2-1">
        <title>Classifier</title>
        <p>
          We make use of a Linear Discriminant Analysis (LDA) classifier that tries to find a
linear combination of features for classification. We make use of the implementation of
the LDA classifier from the scikit-learn machine learning toolkit [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. The classifier
produces a binary prediction of whether a search result is a candidate source of plagiarism.
We sort the positively classified results by their probabilities of being positive as output
by the classifier. This probability essentially reflects the confidence of the classifier in
its prediction.
3.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Features</title>
        <p>
          For each search result that is returned by the search engine, we extract the following
features, which we use for classification. All of these features are available at search
result time and do not require the search result to be retrieved, which allows for
classification to be performed as the search results become available. Many of these features
are available from the ChatNoir search engine at search result time.
1. Readability. The readability of the result document as measured by the Flesh-Kincaid
grade level formula [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] (ChatNoir).
2. Weight. A weight assigned to the result by the search engine (ChatNoir).
3. Proximity. A proximity factor [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] (ChatNoir).
4. PageRank. The PageRank of the result (ChatNoir).
5. BM25. The BM25 score of the result (ChatNoir).
6. Sentences. The number of sentences in the result (ChatNoir).
7. Words. The number of words in the result (ChatNoir).
8. Characters. The number of characters in the result (ChatNoir).
9. Syllables. The number of syllables in the result (ChatNoir).
10. Rank. The rank of the result, i.e. the rank at which it appeared in the search results.
11. Document-snippet word 5-gram Intersection. The set of word 5-grams from the
suspicious document are extracted as well as the set of 5 grams from each search
result snippet, where the snippet is the small sample of text that appears under each
search result. A document-snippet 5-gram intersection score is then calculated as:
Sim(s; d) = S(s) \ S(d);
(1)
where s is the snippet, d is the suspicious document and S( ) is a set of 5-grams.
12. Snippet-document Cosine Similarity. The cosine similarity between the snippet
and the suspicious document, which is given by:
        </p>
        <p>Cosine(s; d) = cos( ) =</p>
        <p>
          Vs Vd
jjVsjjjjVdjj
;
(2)
where V is a term vector.
13. Title-document Cosine Similarity. The cosine similarity between the result title
and the suspicious document (Eq. 2).
14. Query-snippet Cosine Similarity. The cosine similarity between the query and
the snippet (Eq. 2).
15. Query-title Cosine Similarity. The cosine similarity between the query and the
result title (Eq. 2) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
16. Title length. The number of words in the result title.
17. Wikipedia source. Boolean value for whether or not the source was a Wikipedia
article (based on the existence of the word “Wikipedia” in the title).
18. #Nouns. Number of nouns in the title as tagged by the Stanford POS Tagger [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
19. #Verbs. Number of verbs in the title as tagged by the Stanford POS Tagger.
20. #Adjectives Number of adjectives in the title as tagged by the Stanford POS
Tagger.
3.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Training</title>
        <p>We collect training data for the classifier by running our source retrieval method (as
described above) over the training data provided as part of the PAN 2014 source retrieval
task. However, instead of classifying and ranking the search results returned by the
queries, we instead retrieve all of the results and consult the Oracle to determine if they
are sources of plagiarism. This provides labels for the set of search results.</p>
        <p>
          The training data was heavily imbalanced and skewed towards negative samples,
which made up around 70% of the training data. It is well known that most classifiers
expect an even class distribution in order for them to work well [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], thus oversampling
is used to even the class distribution. The oversampling is performed using the SMOTE
method, which creates synthetic examples of the minority class based on existing
samples [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. For each sample xi of the minority class, the k nearest neighbors are identified
and one of those nearest neighbors x^i is randomly selected. The difference between xi
and x^i is multiplied by a random number r 2 [0; 1], which is then added to xi to create
a new data point xnew:
xnew = xi + (x^i
xi)
r:
(3)
        </p>
        <p>The number of nearest neighbors considered is set to k = 3 and the minority class
is increased by 200%. Since nearest neighbors are randomly selected, we train 5
classifiers. The final classification of a search result is based on the majority vote of the 5
classifiers and the probability output for ranking is based on the average probabilities
produced by the 5 classifiers.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation</title>
      <p>
        Our approach achieved a relatively high precision of 0.57 and recall of 0.48. The
F1 score, which is the harmonic mean of precision and recall was 0.47. Overall, our
approach was very competitive. Both the precision and F1 score were the highest achieved
among all participants. The number of queries submitted was relatively high compared
to the other participants; however, as shown in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], queries are relatively cheap from
a bandwidth perspective though they do put additional strain on the search engine.
      </p>
      <p>
        The features used for classification are also of interest to gain a better understanding
of what features are important for source retrieval. While we do not discuss it here, a
complete discussion is presented in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions References</title>
      <p>We describe an approach to the source retrieval problem that makes used of supervised
ranking and classification of search results. Overall, the approach was very competitive
and achieved the highest precision and F1 score among all task participants.
1 The F1 score is computed by averaging the F1 score of each run rather than from the average
precision and recall.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beyer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Busse</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Recent Trends in Digital Text Forensics and Its Evaluation Plagiarism Detection, Author Identification, and Author Profiling</article-title>
          .
          <source>In: Information Access Evaluation</source>
          . Multilinguality, Multimodality, and Visualization. pp.
          <fpage>282</fpage>
          -
          <lpage>302</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Learning from Imbalanced Data</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>21</volume>
          (
          <issue>9</issue>
          ),
          <fpage>1263</fpage>
          -
          <lpage>1284</lpage>
          (
          <year>Sep 2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Jayapal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Similarity Overlap Metric and Greedy String Tiling at PAN 2012: Plagiarism Detection</article-title>
          . CLEF (Online Working Notes/Labs/Workshop) (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Joachims</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Optimizing search engines using clickthrough data</article-title>
          .
          <source>In: Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining - KDD '02</source>
          . pp.
          <fpage>133</fpage>
          -
          <lpage>142</lpage>
          . ACM Press, New York, New York, USA (Jul
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Unsupervised Approaches for Automatic Keyword Extraction Using Meeting Transcripts</article-title>
          .
          <source>In: Proceedings of Human Language Technologies</source>
          :
          <article-title>The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics</article-title>
          . pp.
          <fpage>620</fpage>
          -
          <lpage>628</lpage>
          . No.
          <string-name>
            <surname>June</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Maurer</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Media</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kappe</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaka</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Plagiarism - A Survey</surname>
          </string-name>
          .
          <source>Journal of Universal Computer Science</source>
          <volume>12</volume>
          (
          <issue>8</issue>
          ),
          <fpage>1050</fpage>
          -
          <lpage>1084</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Mccabe</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          :
          <article-title>Cheating among college and university students : A North American perspective</article-title>
          .
          <source>International Journal for Educational Integrity</source>
          <volume>1</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanderplas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Passos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cournapear</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brucher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perrot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duchesnay</surname>
          </string-name>
          , E.:
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          ,
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Graß</surname>
            <given-names>egger</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oberländer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tippmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <article-title>Barrón-cede no, A</article-title>
          .,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the 4th International Competition on Plagiarism Detection pp</article-title>
          .
          <fpage>17</fpage>
          -
          <lpage>20</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tippmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the 5th International Competition on Plagiarism Detection</article-title>
          . In:
          <article-title>CLEF 2013 Evaluation Labs</article-title>
          and Workshop Working Notes Papers (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Graß</surname>
            <given-names>egger</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tippmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Welsch</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>ChatNoir: A Search Engine for the ClueWeb09 Corpus</article-title>
          .
          <source>In: Proceedings of the35th International ACM Conference on Research and Development in Information Retrieval (SIGIR 12)</source>
          . p.
          <volume>1004</volume>
          (
          <year>Aug 2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Feature-rich part-of-speech tagging with a cyclic dependency network</article-title>
          .
          <source>In: Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology - NAACL '03</source>
          . vol.
          <volume>1</volume>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>180</lpage>
          (May
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choudhury</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giles</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Unsupervised Ranking for Plagiarism Source Retrieval - Notebook for PAN at CLEF 2013</article-title>
          .
          <article-title>In: CLEF 2013 Evaluation Labs</article-title>
          and Workshop Working Notes Papers (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giles</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          :
          <article-title>Classifying and Ranking Search Engine Results as Potential Sources of Plagiarism</article-title>
          .
          <source>In: ACM Symposium of Document Engineering</source>
          (
          <year>2014</year>
          ), (To appear)
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>